mariadb-columnstore-engine

mirror of https://github.com/mariadb-corporation/mariadb-columnstore-engine.git synced 2025-07-30 19:23:07 +03:00

Author	SHA1	Message	Date
drrtuy	baae1f66a5	* fix(compilation): `New warn` AKA `3919c541` backport * fix(compilation): `New warn` AKA `3919c541` backport * chore: Bump version to 6.4.8-2	2024-04-24 14:33:13 +03:00
Alexey Antipovsky	69fd36847d	[MCOL-5213] Fix a rare IO error	2022-09-14 17:09:56 +03:00
Roman Nozdrin	a33597b073	MCOL-5198 This patch enables RowStorage to dump data on disk using startNewGeneration if there is 50 Megs left free	2022-08-23 21:30:55 +00:00
Alexey Antipovsky	1be82f859b	Randomly start a new generation if the free memory is less than 30%	2022-08-23 17:59:52 +00:00
Roman Nozdrin	fd9fe182d5	MCOL-5199 This patch solves the overal performance degradation introduced with a new way of char columns hashing in aggregation code The patch disables padding that forces hasher to calculate over the whole 2k buffer. This patch also moves hashing code into the common place where it belongs.	2022-08-22 13:39:45 +00:00
Alexey Antipovsky	30429a7f6c	Fix excessive memory consumption at the last stage of aggregation	2022-08-18 13:58:23 +03:00
Roman Nozdrin	5f485f40ca	MCOL-5153 This patch replaces MDB collation aware hash function with the (#2487 ) exact functionality that does not use MDB hash function. This patch also takes a bit from Robin Hood hash map implementation forgotten that reduces hash function collision rate.	2022-08-04 16:22:11 -05:00
Roman Nozdrin	eabca67c8d	MCOL-5153 This increases the size of the multiplier in the guarding check in RowAggStorage::increaseSize() so that it doesn't throw w/o a reason	2022-07-07 08:58:05 +00:00
Andrey Piskunov	b5d3663936	Renamed variables + removed server tests	2022-06-06 21:32:53 +03:00
Andrey Piskunov	eec85f1118	Welford algorithm for STD and VAR Naive algorithm for calculating STD and VAR is subject to catastrophic cancellation. A well-known Welford's algorithms is used instead.	2022-06-06 21:32:20 +03:00
Gagan Goel	1fc399451a	MCOL-4957 Fix performance slowdown for processing TIMESTAMP columns. Part 1: As part of MCOL-3776 to address synchronization issue while accessing the fTimeZone member of the Func class, mutex locks were added to the accessor and mutator methods. However, this slows down processing of TIMESTAMP columns in PrimProc significantly as all threads across all concurrently running queries would serialize on the mutex. This is because PrimProc only has a single global object for the functor class (class derived from Func in utils/funcexp/functor.h) for a given function name. To fix this problem: (1) We remove the fTimeZone as a member of the Func derived classes (hence removing the mutexes) and instead use the fOperationType member of the FunctionColumn class to propagate the timezone values down to the individual functor processing functions such as FunctionColumn::getStrVal(), FunctionColumn::getIntVal(), etc. (2) To achieve (1), a timezone member is added to the execplan::CalpontSystemCatalog::ColType class. Part 2: Several functors in the Funcexp code call dataconvert::gmtSecToMySQLTime() and dataconvert::mySQLTimeToGmtSec() functions for conversion between seconds since unix epoch and broken-down representation. These functions in turn call the C library function localtime_r() which currently has a known bug of holding a global lock via a call to __tz_convert. This significantly reduces performance in multi-threaded applications where multiple threads concurrently call localtime_r(). More details on the bug: https://sourceware.org/bugzilla/show_bug.cgi?id=16145 This bug in localtime_r() caused processing of the Functors in PrimProc to slowdown significantly since a query execution causes Functors code to be processed in a multi-threaded manner. As a fix, we remove the calls to localtime_r() from gmtSecToMySQLTime() and mySQLTimeToGmtSec() by performing the timezone-to-offset conversion (done in dataconvert::timeZoneToOffset()) during the execution plan creation in the plugin. Note that localtime_r() is only called when the time_zone system variable is set to "SYSTEM". This fix also required changing the timezone type from a std::string to a long across the system.	2022-02-11 19:03:32 -05:00
Leonid Fedorov	7c808317dc	clang format apply	2022-02-11 12:24:40 +00:00
David.Hall	509f005be7	Mcol 4841 dev6 Handle large joins without OOM (#2155 ) * MCOL-4846 dev-6 Handle large join results Use a loop to shrink the number of results reported per message to something manageable. * MCOL-4841 small changes requested by review * Add EXTRA threads to prioritythreadpool prioritythreadpool is configured at startup with a fixed number of threads available. This is to prevent thread thrashing. Since most of the time, BPP job steps are short lived, and a rescheduling mechanism exist if no threads are available, this works to keep cpu wastage to a minimum. However, if a query or queries consume all the threads in prioritythreadpool and then block (due to the consumer not consuming fast enough) we can run out of threads and no work will be done until some threads unblock. A new mechanism allows for EXTRA threads to be generated for the duration of the blocking action. These threads can act on new queries. When all blocking is completed, these threads will be released when idle. * MCOL-4841 dev6 Reconcile with changes in develop-6 * MCOL-4841 Some format corrections * MCOL-4841 dev clean up some things based on review * MCOL-4841 dev 6 ExeMgr Crashes after large join This commit fixes up memory accounting issues in ExeMgr * MCOL-4841 remove LDI change Opened MCOL-4968 to address the issue * MCOL-4841 Add fMaxBPPSendQueue to ResourceManager This causes the setting to be loaded at run time (requires restart to accept a change) BPPSendthread gets this in it's ctor Also rolled back changes to TupleHashJoinStep::smallRunnerFcn() that used a local variable to count locally allocated memory, then added it into the global counter at function's end. Not counting the memory globally caused conversion to UM only join way later than it should. This resulted in MCOL-4971. * MCOL-4841 make blockedThreads and extraThreads atomic Also restore previous scope of locks in bppsendthread. There is some small chance the new scope could be incorrect, and the performance boost is negligible. Better safe than sorry.	2022-02-09 21:38:32 +03:00
Denis Khalikov	19a41384bb	MCOL-4810 Add support for missed operation.	2021-10-25 15:04:00 +03:00
Roman Nozdrin	7847312448	MCOL-4876 This patch enables continues buffer to be used by ColumnCommand and aligns BPP::blockData (#2119 ) that in most cases was unaligned	2021-10-05 12:22:24 +03:00
Roman Nozdrin	0a9bbf8d67	Merge pull request #2064 from mariadb-AlexeyAntipovsky/MCOL-4829-dev6 [MCOL-4829] Compression for the temp disk-based aggregation files	2021-09-13 23:12:46 +03:00
Alexey Antipovsky	2328f4ef2a	[MCOL-4829] More accurate memory counting	2021-09-07 19:48:53 +03:00
Leonid Fedorov	4e62ae2058	Add ctest for google unittests	2021-09-07 15:19:25 +00:00
Alexey Antipovsky	bf1640be65	[MCOL-4829] Compression for the temp disk-based aggregation files	2021-09-02 19:31:38 +03:00
Roman Nozdrin	339ce8260b	Merge pull request #2103 from denis0x0D/MCOL-4810_dev_6 MCOL-4810 Redundant copying and wasting memory in PrimProc	2021-08-27 20:15:17 +03:00
Denis Khalikov	4a0150e93d	MCOL-4810 Redundant copying and wasting memory in PrimProc This patch eliminates a copying `long string`s into the bytestream.	2021-08-27 14:13:02 +03:00
Leonid Fedorov	ef09342d47	MCOL-4839: Fix clang build (#2102 ) * Fix clang build * Extern C returned to plugin_instance Co-authored-by: Leonid Fedorov <l.fedorov@mail.corp.ru>	2021-08-23 15:58:56 -05:00
Gagan Goel	3d557a2f1e	Merge pull request #2044 from dhall-MariaDB/MCOL-3738 MCOL-3738 COUNT(DISTINCT) with multiple parms	2021-07-12 07:34:56 -04:00
Leonid Fedorov	51a8ffcb6a	Fix sumavgoverflow.sql test	2021-07-09 22:41:28 +00:00
David Hall	76607be63a	MCOL-3738 COUNT(DISTINCT) with multiple parms Fixed regression Added a few more mtr tests	2021-07-09 09:07:03 -05:00
Leonid Fedorov	f81f743282	Replace underlying type for avg and sum for int types from long double to wide decimal	2021-07-08 17:04:43 +00:00
David Hall	1113470551	MCOL-4738 AVG gives wrong results with strict_aliasing A f fix that works with strict_aliasing	2021-07-07 13:08:32 -05:00
Alexander Barkov	8988253ff4	Merge pull request #2031 from mariadb-corporation/bar-develop-MCOL-4801 MCOL-4801 Replace Row methods getStringLength() and getStringPointer(…	2021-07-07 13:53:19 +04:00
David Hall	8332ab8974	MCOL-4738 AVG() returns a wrong result On AMD64 machines, the fpu is 80 bits. The unused bits must be masked for memcmp to work properly. For other archetectures, we don't want to mask those bits.	2021-07-06 19:50:00 -05:00
Alexander Barkov	9794f24369	MCOL-4801 Replace Row methods getStringLength() and getStringPointer() to getConstString()	2021-07-06 21:15:32 +04:00
Gagan Goel	8520f87237	MCOL-641 Cleanup.	2021-07-06 09:01:49 +00:00
Alexey Antipovsky	60495564b8	[MCOL-4709] Fix another UB in disk aggregation	2021-06-29 17:47:07 +03:00
Alexey Antipovsky	8a0b68f25e	[MCOL-4709] Fix UB in disk aggregation	2021-06-28 20:07:23 +03:00
Roman Nozdrin	8c360a1a27	MCOL-4759 Upmerge for MCOL-4564 code that implements hash merging family to reduce performance penalty using MDB hashing functions	2021-06-24 14:48:01 +00:00
Roman Nozdrin	bed0b7c6bc	MCOL-4173 This patch adds support for wide-DECIMAL INNER, OUTER, SEMI, functional JOINs based on top of TypelessData	2021-06-24 08:07:23 +00:00
Alexander Barkov	b3d6f62964	MCOL-4753 Performance problem in Typeless join	2021-06-10 09:26:26 +00:00
Alexey Antipovsky	0dedb7e628	Fix compilation warnings	2021-06-09 16:51:00 +03:00
Alexey Antipovsky	475104e4d3	[MCOL-4709] Disk-based aggregation * Introduce multigeneration aggregation * Do not save unused part of RGDatas to disk * Add IO error explanation (strerror) * Reduce memory usage while aggregating * introduce in-memory generations to better memory utilization * Try to limit the qty of buckets at a low limit * Refactor disk aggregation a bit * pass calculated hash into RowAggregation * try to keep some RGData with free space in memory * do not dump more than half of rowgroups to disk if generations are allowed, instead start a new generation * for each thread shift the first processed bucket at each iteration, so the generations start more evenly * Unify temp data location * Explicitly create temp subdirectories whether disk aggregation/join are enabled or not	2021-06-06 16:09:15 +03:00
Alexander Barkov	9608533d92	MCOL-4734 Compilation failure: MariaDB-10.6 + ColumnStore-develop mcsconfig.h and my_config.h have the following pre-processor definitions: 1. Conflicting definitions coming from the standard cmake definitions: - PACKAGE - PACKAGE_BUGREPORT - PACKAGE_NAME - PACKAGE_STRING - PACKAGE_TARNAME - PACKAGE_VERSION - VERSION 2. Conflicting definitions of other kinds: - HAVE_STRTOLL - this is a dirt in MariaDB headers. Should be fixed in the server code. my_config.h erroneously performs "#define HAVE_STRTOLL" instead of "#define HAVE_STRTOLL 1". in some cases. The former is not CMake compatible style. The latter is. 3. Non-conflicting definitions: Otherwise, mcsconfig.h and my_config.h should be mutually compatible, because both are generated by cmake on the same host machine. So they should have exactly equal definitions like "HAVE_XXX", "SIZEOF_XXX", etc. Observations: - It's OK to include both mcsconfig.h and my_config.h providing that we suppress duplicate definition of the above conflicting types #1 and #2. - There is no a need to suppress duplicate definitions mentioned in #3, as they are compatible! - my_sys.h and m_ctype.h must always follow a CMake configuation header, either my_config.h or mcsconfig.h (or both). They must never be included without any preceeding configuration header. This change make sure that we resolve conflicts by: - either disallowing inclusion of mcsconfig.h and my_config.h at the same time - or by hiding conflicting definitions #1 and #2 (with their later restoring). - also, by making sure that my_sys.h and m_ctype.h always follow a CMake configuration file. Details: - idb_mysql.h can now only be included only after my_config.h An attempt to use idb_mysql.h with mcsconfig.h instead of my_config.h is caught by the "#error" preprocessor directive. - mariadb_my_sys.h can now be only included after mcsconfig.h. An attempt to use mariadb_my_sys.h without mcscofig.h (e.g. with my_config.h) is also caught by "#error". - collation.h now can now be included in two ways. It now has the following effective structure: #if defined(PREFER_MY_CONFIG_H) && defined(MY_CONFIG_H) // Remember current conflicting definitions on the preprocessor stack // Undefine current conflicting definitions #endif #include "mcsconfig.h" #include "m_ctype.h" #if defined(PREFER_MY_CONFIG_H) && defined(MY_CONFIG_H) # Restore conflicting definitions from the preprocessor stack #endif and can be included as follows: a. using only mcsconfig.h as a configuration header: // my_config.h must not be included so far #include "collation.h" b. using my_config.h as the first included configuration file: #define PREFER_MY_CONFIG_H // Force conflict resolution #include "my_config.h" // can be included directly or indirectly ... #include "collation.h" Other changes: - Adding helper header files utils/common/mcsconfig_conflicting_defs_remember.h utils/common/mcsconfig_conflicting_defs_restore.h utils/common/mcsconfig_conflicting_defs_undef.h to perform conflict resolution easier. - Removing `#include "collation.h"` from a number of files, as it's automatically included from rowgroup.h. - Removing redundant `#include "utils_utf8.h"`. This change is not directly related to the problem being fixed, but it's nice to remove redundant directives for both collation.h and utils_utf8.h from all the files that do not really need them. (this change could probably have gone as a separate commit) - Changing my_init() to MY_INIT(argv[0]) in the MCS services sources. After the fix of the complitation failure it appeared that ColumnStore services compiled with the debug build crash due to recent changes in safemalloc. The crash happened in strcmp() with `my_progname` as an argument (where my_progname is a mysys global variable). This problem should probably be fixed on the server side as well to avoid passing NULL. But, the majority of MariaDB executable programs also use MY_INIT(argv[0]) rather than my_init(). So let's make MCS do like the other programs do.	2021-05-25 12:34:36 +04:00
Alexander Barkov	bd4cbb542d	MCOL-4721 CHAR(1) is not collation-aware for GROUP/DISTINCT	2021-05-18 16:14:53 +04:00
David Hall	f4e6939139	MCOL-4643 dev 5 reset valOut after processing UDAF After a UDAF result has been inserted in the output stream, the valOut object needs to be reset to empty in preparation for the next value. Failing to do so may cause what should be a NULL value to erroneously take the last value inserted.	2021-04-30 10:57:40 -05:00
Alexander Barkov	362bfcd15e	MCOL-4361 Replace pow(10.0, (double)scale) expressions with a static dictionary lookup.	2021-04-09 12:41:04 +04:00
Alexander Barkov	69911c2710	A joint patch for MCOL-4614, MCOL-4615, MCOL-4660 (decimal to string conversion) This patch fixes: - MCOL-4614 calShowPartitions() precision loss for huge narrow decimal - MCOL-4615 GROUP_CONCAT() precision loss for huge narrow decimal - MCOL-4660 Narow decimal to string conversion is inconsistent about zero integral Changes: - Implementing Row::getDecimalField() - Removing double arithmetic from the code printing DECIMAL values in TypeHandlerXDecimal::format64() and GroupConcator::outputRow(). Using Decimal::toString() instead. - Rewriting Decimal::toStringTSInt64(). The old implementation was wrong, too complex and slow (used unnecessary memmove, memcpy). An additional cleanup: - Removing the ENGINE=COLUMNSTORE clause from tests for MCOL-4532 and MCOL-4640 type_decimal.test is combinations-aware. It's run two times with default_storage_engine=MyISAM and default_storage_engine=COLUMNSTORE. So the CREATE TABLE statements should not specify the engine explicitly. - Adding --disable_warnings in the old fixed test. We needed to suppress warnings when the MyISAM combination is being run. Previously the table was erroneously created with ENGINE=COLUMNSTORE even with the MyISAM combination run. So warning were not generated.	2021-04-05 16:36:19 +04:00
David Hall	0eee6cfc62	MCOL-4643 reset valOut after UDAF evaluation	2021-03-26 16:09:15 -05:00
David Hall	13b7a794e4	MCOL-4620 Add charset to various RowGroup initializers Specifically to operator+=	2021-03-19 16:57:54 -05:00
Alexey Antipovsky	5080e1ae53	MCOL-4031 More accurate memory usage counting while sorting	2021-01-29 18:31:20 +03:00
Alexander Barkov	a687df48b9	MCOL-4065 DISTINCT is case sensitive This patch makes DISTINCT and GROUP BY collation aware.	2021-01-21 15:46:54 +04:00
Roman Nozdrin	5b9689ce55	MCOL-4478 MCS now rounds the last digits of an avg() result for wide-DECIMAL argument	2020-12-30 15:02:12 +00:00
Roman Nozdrin	5815c5c526	MCOL-4452 RowAggregationUMP2::doUDAF() now calls setUserData() using a correct UDAF context	2020-12-22 15:43:51 +00:00
Roman Nozdrin	494bde61e1	MCOL-4409 Moved static Decimal conversion methods into VDecimal class MCOL-4409 This patch combines VDecimal and Decimal and makes IDB_Decimal an alias for the result class MCOL-4409 More boilerplate reduction in Func_mod Removed couple TSInt128::toType() methods	2020-11-30 12:08:52 +00:00

1 2 3 4 5

205 Commits