mariadb-columnstore-engine

mirror of https://github.com/mariadb-corporation/mariadb-columnstore-engine.git synced 2025-07-30 19:23:07 +03:00

Author	SHA1	Message	Date
Petr Vaněk	86159cc899	Migration for Boost 1.85 Boost 1.85 removed some deprecated code in filesystem module which is still used in columnstore: - The boost/filesystem/convenience.hpp was removed but columnstore does not use any functionality from that file except indirect includes. Therefore this include is removed or replaced with more general boost/filesystem.hpp. The convenience.hpp header file was deprecated in filesystem V3 introduced in Boost 1.46.0. - `normalize` method was removed and users are suggested to replace it with `lexically_normal` method, which was introduced in Boost 1.60.0. Original `normalize` call is preserved for backward compatibility with old Boost version, however`, `lexically_normal` method is preferably used with Boost 1.60.0 and newer. - The `copy_option` was removed in favor of `copy_options` (note the trailing 's'), but enum values were renamed. Namely, `fail_if_exists` is replaced with `none` and `overwrite_if_exists` is replaced with `overwrite_existing`. The `copy_options` was introduced in Boost 1.74.0. New form is used instead, but a backward compatibility layer for Boost 1.73.0 and older was introduced in boost_copy_options_compat.hpp file. This solution seems to be less awkward than using multiple #if #else #endif blocks in source code.	2025-02-23 02:52:52 +04:00
Daniel Black	9dedc57df6	feat(): MCOL-5881 set/getThreadName use FreeBSD API (#3381 ) Taken from FreeBSD ports, this uses the FreeBSD APIs rather than the Linux specific prctl to change and retreive the thread names. Co-authored-by: Bernard Spil <brnrd@FreeBSD.org>	2025-01-15 22:02:49 +00:00
drrtuy	baae1f66a5	* fix(compilation): `New warn` AKA `3919c541` backport * fix(compilation): `New warn` AKA `3919c541` backport * chore: Bump version to 6.4.8-2	2024-04-24 14:33:13 +03:00
Roman Nozdrin	fd9fe182d5	MCOL-5199 This patch solves the overal performance degradation introduced with a new way of char columns hashing in aggregation code The patch disables padding that forces hasher to calculate over the whole 2k buffer. This patch also moves hashing code into the common place where it belongs.	2022-08-22 13:39:45 +00:00
Roman Nozdrin	5f485f40ca	MCOL-5153 This patch replaces MDB collation aware hash function with the (#2487 ) exact functionality that does not use MDB hash function. This patch also takes a bit from Robin Hood hash map implementation forgotten that reduces hash function collision rate.	2022-08-04 16:22:11 -05:00
Roman Nozdrin	1fe39fff65	MCOL-5105 This patch raises pipe read operation timeout to 20 minutes to enable DMLProc to survive rollbacks on startup. The patch also fixes linter warnings in service.h and pipe.h.	2022-06-09 14:32:15 +00:00
Leonid Fedorov	b6b9a86dcf	Merge branch 'develop-6' into fix-arm-6	2022-02-11 19:10:03 +03:00
Leonid Fedorov	41ab33fa06	fix arm build withenum KIND	2022-02-11 13:27:14 +00:00
Leonid Fedorov	7c808317dc	clang format apply	2022-02-11 12:24:40 +00:00
David.Hall	509f005be7	Mcol 4841 dev6 Handle large joins without OOM (#2155 ) * MCOL-4846 dev-6 Handle large join results Use a loop to shrink the number of results reported per message to something manageable. * MCOL-4841 small changes requested by review * Add EXTRA threads to prioritythreadpool prioritythreadpool is configured at startup with a fixed number of threads available. This is to prevent thread thrashing. Since most of the time, BPP job steps are short lived, and a rescheduling mechanism exist if no threads are available, this works to keep cpu wastage to a minimum. However, if a query or queries consume all the threads in prioritythreadpool and then block (due to the consumer not consuming fast enough) we can run out of threads and no work will be done until some threads unblock. A new mechanism allows for EXTRA threads to be generated for the duration of the blocking action. These threads can act on new queries. When all blocking is completed, these threads will be released when idle. * MCOL-4841 dev6 Reconcile with changes in develop-6 * MCOL-4841 Some format corrections * MCOL-4841 dev clean up some things based on review * MCOL-4841 dev 6 ExeMgr Crashes after large join This commit fixes up memory accounting issues in ExeMgr * MCOL-4841 remove LDI change Opened MCOL-4968 to address the issue * MCOL-4841 Add fMaxBPPSendQueue to ResourceManager This causes the setting to be loaded at run time (requires restart to accept a change) BPPSendthread gets this in it's ctor Also rolled back changes to TupleHashJoinStep::smallRunnerFcn() that used a local variable to count locally allocated memory, then added it into the global counter at function's end. Not counting the memory globally caused conversion to UM only join way later than it should. This resulted in MCOL-4971. * MCOL-4841 make blockedThreads and extraThreads atomic Also restore previous scope of locks in bppsendthread. There is some small chance the new scope could be incorrect, and the performance boost is negligible. Better safe than sorry.	2022-02-09 21:38:32 +03:00
Roman Nozdrin	f7417c0b10	MCOL-4809 This patch adds support for float data types filtering and scanning vectorization	2022-02-03 16:02:37 +00:00
Roman Nozdrin	dafb9c5dae	This patch introduces support for scanning/filtering vectorized execution for numeric-based data types TEXT, CHAR, VARCHAR, FLOAT and DOUBLE are not yet supported by vectorized path This patch introduces an example for Google benchmarking suite to measure a perf diff b/w legacy scan/filtering code and the templated version	2021-10-28 16:47:18 +00:00
Leonid Fedorov	ef09342d47	MCOL-4839: Fix clang build (#2102 ) * Fix clang build * Extern C returned to plugin_instance Co-authored-by: Leonid Fedorov <l.fedorov@mail.corp.ru>	2021-08-23 15:58:56 -05:00
Alexander Barkov	9794f24369	MCOL-4801 Replace Row methods getStringLength() and getStringPointer() to getConstString()	2021-07-06 21:15:32 +04:00
Denis Khalikov	1d5f309b8f	MCOL-1205 Support queries with circular joins This patch adds support for queries with circular joins. Currently support added for inner joins only.	2021-07-02 18:37:07 +03:00
Denis Khalikov	c20015a7b2	MCOL-4713 Analyze table implementation.	2021-07-02 12:37:12 +03:00
Roman Nozdrin	d8cbc000e2	Merge pull request #2004 from drrtuy/MCOL-4759 MCOL-4759 Upmerge for MCOL-4564 code that implements hash merging fam…	2021-06-28 14:05:16 +03:00
Roman Nozdrin	8c360a1a27	MCOL-4759 Upmerge for MCOL-4564 code that implements hash merging family to reduce performance penalty using MDB hashing functions	2021-06-24 14:48:01 +00:00
Roman Nozdrin	2de4888899	Merge pull request #1990 from drrtuy/MCOL-4173_9 MCOL-4173 This patch adds support for wide-DECIMAL INNER, OUTER, SEMI…	2021-06-24 16:15:07 +03:00
Roman Nozdrin	bed0b7c6bc	MCOL-4173 This patch adds support for wide-DECIMAL INNER, OUTER, SEMI, functional JOINs based on top of TypelessData	2021-06-24 08:07:23 +00:00
Gagan Goel	7c8b502dc2	Fix regression in a query involving an aggregate function on a non-wide decimal column in the HAVING clause. In buildAggregateColumn(), if an aggregate function (such as avg) is applied on a non-wide decimal column, we were setting the precision of the resulting column as -1. This later down in the execution got converted to 255 as in some cases, precision is stored as uint8_t. The predicate operations on a DECIMAL column has logic that uses the wide Decimal::s128value field if precision > 18. This logic incorrectly used the Decimal::s128value instead of the correct value stored in the narrow Decimal::value field, since precision of the Decimal column was 255. The fix is to set the aggregate column precision to datatypes::INT64MAXPRECISION (18) in buildAggregateColumn() when the aggregate is applied on a non-wide decimal column. This commit also partially fixes -Wstrict-aliasing GCC warnings.	2021-06-22 11:11:34 +00:00
Alexey Antipovsky	475104e4d3	[MCOL-4709] Disk-based aggregation * Introduce multigeneration aggregation * Do not save unused part of RGDatas to disk * Add IO error explanation (strerror) * Reduce memory usage while aggregating * introduce in-memory generations to better memory utilization * Try to limit the qty of buckets at a low limit * Refactor disk aggregation a bit * pass calculated hash into RowAggregation * try to keep some RGData with free space in memory * do not dump more than half of rowgroups to disk if generations are allowed, instead start a new generation * for each thread shift the first processed bucket at each iteration, so the generations start more evenly * Unify temp data location * Explicitly create temp subdirectories whether disk aggregation/join are enabled or not	2021-06-06 16:09:15 +03:00
Alexander Barkov	9608533d92	MCOL-4734 Compilation failure: MariaDB-10.6 + ColumnStore-develop mcsconfig.h and my_config.h have the following pre-processor definitions: 1. Conflicting definitions coming from the standard cmake definitions: - PACKAGE - PACKAGE_BUGREPORT - PACKAGE_NAME - PACKAGE_STRING - PACKAGE_TARNAME - PACKAGE_VERSION - VERSION 2. Conflicting definitions of other kinds: - HAVE_STRTOLL - this is a dirt in MariaDB headers. Should be fixed in the server code. my_config.h erroneously performs "#define HAVE_STRTOLL" instead of "#define HAVE_STRTOLL 1". in some cases. The former is not CMake compatible style. The latter is. 3. Non-conflicting definitions: Otherwise, mcsconfig.h and my_config.h should be mutually compatible, because both are generated by cmake on the same host machine. So they should have exactly equal definitions like "HAVE_XXX", "SIZEOF_XXX", etc. Observations: - It's OK to include both mcsconfig.h and my_config.h providing that we suppress duplicate definition of the above conflicting types #1 and #2. - There is no a need to suppress duplicate definitions mentioned in #3, as they are compatible! - my_sys.h and m_ctype.h must always follow a CMake configuation header, either my_config.h or mcsconfig.h (or both). They must never be included without any preceeding configuration header. This change make sure that we resolve conflicts by: - either disallowing inclusion of mcsconfig.h and my_config.h at the same time - or by hiding conflicting definitions #1 and #2 (with their later restoring). - also, by making sure that my_sys.h and m_ctype.h always follow a CMake configuration file. Details: - idb_mysql.h can now only be included only after my_config.h An attempt to use idb_mysql.h with mcsconfig.h instead of my_config.h is caught by the "#error" preprocessor directive. - mariadb_my_sys.h can now be only included after mcsconfig.h. An attempt to use mariadb_my_sys.h without mcscofig.h (e.g. with my_config.h) is also caught by "#error". - collation.h now can now be included in two ways. It now has the following effective structure: #if defined(PREFER_MY_CONFIG_H) && defined(MY_CONFIG_H) // Remember current conflicting definitions on the preprocessor stack // Undefine current conflicting definitions #endif #include "mcsconfig.h" #include "m_ctype.h" #if defined(PREFER_MY_CONFIG_H) && defined(MY_CONFIG_H) # Restore conflicting definitions from the preprocessor stack #endif and can be included as follows: a. using only mcsconfig.h as a configuration header: // my_config.h must not be included so far #include "collation.h" b. using my_config.h as the first included configuration file: #define PREFER_MY_CONFIG_H // Force conflict resolution #include "my_config.h" // can be included directly or indirectly ... #include "collation.h" Other changes: - Adding helper header files utils/common/mcsconfig_conflicting_defs_remember.h utils/common/mcsconfig_conflicting_defs_restore.h utils/common/mcsconfig_conflicting_defs_undef.h to perform conflict resolution easier. - Removing `#include "collation.h"` from a number of files, as it's automatically included from rowgroup.h. - Removing redundant `#include "utils_utf8.h"`. This change is not directly related to the problem being fixed, but it's nice to remove redundant directives for both collation.h and utils_utf8.h from all the files that do not really need them. (this change could probably have gone as a separate commit) - Changing my_init() to MY_INIT(argv[0]) in the MCS services sources. After the fix of the complitation failure it appeared that ColumnStore services compiled with the debug build crash due to recent changes in safemalloc. The crash happened in strcmp() with `my_progname` as an argument (where my_progname is a mysys global variable). This problem should probably be fixed on the server side as well to avoid passing NULL. But, the majority of MariaDB executable programs also use MY_INIT(argv[0]) rather than my_init(). So let's make MCS do like the other programs do.	2021-05-25 12:34:36 +04:00
Alexander Barkov	765858bc5b	MCOL-4498 LIKE is not collation aware	2021-03-22 20:42:01 +04:00
benthompson15	afa88866bb	MCOL-4483: Fix and consolidate log files and cpimport logging.	2021-02-12 15:40:16 -06:00
Alexander Barkov	69da915160	MCOL-4531 New string-to-decimal conversion implementation This change fixes: MCOL-4462 CAST(varchar_expr AS DECIMAL(M,N)) returns a wrong result MCOL-4500 Bit functions processing throws internally trying to cast char into decimal representation MCOL-4532 CAST(AS DECIMAL) returns a garbage for large values Also, this change makes string-to-decimal conversion 5-10 times faster, depending on exact data. Performance implemenent is achieved by the fact that (unlike in the old implementation), the new version does not do any "string" object copying.	2021-02-09 13:02:27 +04:00
Roman Nozdrin	5fce19df0a	MCOL-4412 Introduce TypeHandler::getEmptyValueForType to return const ptr for an empty value WE changes for SQL DML and DDL operations Changes for bulk operations Changes for scanning operations Cleanup	2021-01-18 12:30:17 +00:00
Alexander Barkov	b08d719593	A cleanup for MCOL-4064 Make JOIN collation aware A non-JOIN condition like `WHERE c1=c2` (with c1 and c2 being columns of the same table) was not collation-aware yet after the main patches for MCOL-4064. Additionally fixing StrFilterCmd::compare*() to address this.	2020-12-08 16:43:07 +04:00
Alexander Barkov	c6158eee31	Part#1 MCOL-4064 Make JOIN collation aware Making field1=field2 collation aware for long CHAR/VARCHAR.	2020-12-04 08:41:26 +04:00
Alexander Barkov	52c5af054a	Part#2 MCOL-495 Make string comparison not case sensitive Fixing field='str' for short (non-Dict) CHAR and VARCHAR data types.	2020-12-04 08:40:29 +04:00
Alexander Barkov	0ff6a6ec20	Part#1 MCOL-495 Make string comparison not case sensitive Fixing field='str' for long (Dict) string data types.	2020-12-04 07:49:00 +04:00
Alexander Barkov	2ea73846b9	MCOL-4422 Remove mariadb.h and my_sys.h dependency from collation.h	2020-11-30 14:26:35 +04:00
Gagan Goel	995cadef2d	MCOL-641 Fix alter table add wide decimal column. This patch also removes CalpontSystemCatalog::BINARY and ddlpackage::DDL_BINARY that were added during the initial stages of the work on MCOL-641.	2020-11-20 19:49:54 -05:00
Roman Nozdrin	aa44bca473	A pack of fixes for compilation errors and warnings for all platforms Add libdatatypes.so into debian packaging	2020-11-19 10:21:45 +00:00
Roman Nozdrin	3eb26c0d4a	MCOL-4313 Introduced TSInt128 that is a storage class for int128 Removed uint128 from joblist/lbidlist.* Another toString() method for wide-decimal that is EMPTY/NULL aware Unified decimal processing in WF functions Fixed a potential issue in EqualCompData::operator() for wide-decimal processing Fixed some signedness warnings	2020-11-18 13:53:15 +00:00
Alexander Barkov	d5c6645ba1	Adding mcs_basic_types.h For now it consists of only: using int128_t = __int128; using uint128_t = unsigned __int128; All new privitive data types should go into this file in the future.	2020-11-18 13:53:15 +00:00
Alexander Barkov	129d5b5a0f	MCOL-4174 Review/refactor frontend/connector code	2020-11-18 13:53:15 +00:00
Roman Nozdrin	8de9764f84	MCOL-4172 Add support for wide-DECIMAL into statistical aggregate and regr_* UDAF functions The patch fixes wrong results returned when multiple UDAF exist in projection aggregate over wide decimal literals now works	2020-11-18 13:52:20 +00:00
David Hall	af80081c94	MCOL-4171 Some fixes	2020-11-18 13:52:20 +00:00
Roman Nozdrin	1588ebe439	MCOL-641 Clean up primitives code Add int128_t support into ByteStream Fixed UTs broken after collation patch	2020-11-18 13:52:19 +00:00
Gagan Goel	d3bc68b02f	MCOL-641 Refactor initial extent elimination support. This commit also adds support in TupleHashJoinStep::forwardCPData, although we currently do not support wide decimals as join keys. Row estimation to determine large-side of the join is also updated.	2020-11-18 13:52:19 +00:00
Gagan Goel	6aea838360	MCOL-641 Add support for functions (Part 2).	2020-11-18 13:51:55 +00:00
Roman Nozdrin	a7fcf39f2a	MCOL-641 Fixed group_concat for narrow-DECIMALs.	2020-11-18 13:51:26 +00:00
Gagan Goel	74b64eb4f1	MCOL-641 1. Add support for int128_t in ParsedColumnFilter. 2. Set Decimal precision in SimpleColumn::evaluate(). 3. Add support for int128_t in ConstantColumn. 4. Set IDB_Decimal::s128Value in buildDecimalColumn(). 5. Use width 16 as first if predicate for branching based on decimal width.	2020-11-18 13:47:45 +00:00
Roman Nozdrin	b09f3088ca	MCOL-641 Initial version of Math operations for wide decimal.	2020-11-18 13:47:44 +00:00
Gagan Goel	62d0c82d75	MCOL-641 1. Templatized convertValueNum() function. 2. Allocate int128_t buffers in batchprimitiveprocessor if a query involves wide decimal columns.	2020-11-18 13:47:44 +00:00
Roman Nozdrin	2e8e7d52c3	Renamed datatypes/decimal.* into csdecimal to avoid collision with MDB.	2020-11-18 13:47:44 +00:00
Gagan Goel	824615a55b	MCOL-641 Refactor empty value implementation in writeengine.	2020-11-18 13:47:44 +00:00
Roman Nozdrin	97ee1609b2	MCOL-641 Replaced NULL binary constants. DataConvert::decimalToString, toString, writeIntPart, writeFractionalPart are not templates anymore.	2020-11-18 13:47:44 +00:00
Gagan Goel	8f80c1dee6	MCOL-641 1. Implement int128 version of strtoll. 2. Templatize number_int_value. 3. Add test cases for strtoll128 and number_int_value for Decimal38.	2020-11-18 13:47:02 +00:00

1 2 3

114 Commits