mariadb-columnstore-engine

mirror of https://github.com/mariadb-corporation/mariadb-columnstore-engine.git synced 2025-12-20 01:42:27 +03:00

Author	SHA1	Message	Date
Leonid Fedorov	37fd915a08	Serg`s patch for develop-6 revised for develop https://github.com/mariadb-corporation/mariadb-columnstore-engine/pull/2614	2022-11-09 22:41:38 +00:00
Gagan Goel	f4e3022fbd	This commit fixes an incorrect predicate in the if condition (#2608 ) that checks for HWM extent in WE_DMLCommandProc::processBatchInsertHwm().	2022-11-08 14:51:42 -06:00
Leonid Fedorov	d2432f9bf6	get rid of pointers for 128 fields	2022-08-26 15:12:22 +00:00
mariadb-AndreyPiskunov	0863ecd279	Replace getBinaryField	2022-08-25 18:21:43 +03:00
Gagan Goel	cbfdae3481	MCOL-5021 Code changes based on review feedback.	2022-08-05 14:40:50 -04:00
Gagan Goel	1355237ca3	MCOL-5021 Some minor fixes.	2022-08-05 14:40:50 -04:00
Gagan Goel	9b6d3c3870	MCOL-5021 Add support for AUX column in the client code calling CalpontSystemCatalog::columnRIDs().	2022-08-05 14:40:49 -04:00
Gagan Goel	262cd5c501	MCOL-5021 Remove hard-coded values for data type, column width and compression type for the AUX column, and replace them with constants defined in the execplan namespace.	2022-08-05 14:40:49 -04:00
Gagan Goel	c8b6b154bf	MCOL-5021 Add an option in Columnstore.xml, fastdelete (disabled by default), which when enabled, indiscriminately invalidates all column extents and performs the actual DELETE only on the AUX column. The trade-off with this approach would now be that the first SELECT for certain query patterns (those containing a WHERE predicate) after the DELETE operation will slow down as the invalidated column extent would need to be scanned again to set the min/max values.	2022-08-05 14:40:49 -04:00
Gagan Goel	60eb0f86ec	MCOL-5021 non-AUX column files are opened in read-only mode during the DELETE operation. ColumnOp::readBlock() calls can cause writes to database files when the active chunk list in ChunkManager is full. Since non-AUX columns are read-only for the DELETE operation, we prevent writes of compressed chunks and header for these columns by passing an isReadOnly flag to CompFileData which indicates whether the column is read-only or read-write.	2022-08-05 14:40:49 -04:00
Gagan Goel	35a3a93964	MCOL-5021 For the DELETE operation, empty magic values are only written to database files for AUX column. Perform read-only operation for other columns in the table to update the Casual Partitioning information.	2022-08-05 14:40:49 -04:00
Gagan Goel	86df9a972c	MCOL-5021 Add prototype support for the AUX column in CREATE/DROP DDL commands, single and multi-value INSERTs, cpimport, and DELETE.	2022-08-05 14:40:49 -04:00
Denis Khalikov	fb1e23bb83	[MCOL-5106] Add support to work with StorageManager. This patch eliminates boost::filesystem from `mcsRebuildEM` tool. After this change we should be able to work with any filesystem even S3.	2022-07-28 16:47:34 +03:00
Leonid Fedorov	f5b2a6885f	MCOL-5013: Load Data from S3 into Columnstore Introduced UDF and stored prodecure. usage: set columnstore_s3_key='<s3_key>'; set columnstore_s3_secret='<s3_secret>'; set columnstore_s3_region='region'; and then use UDF select columnstore_dataload("<tablename>", "<filename>", "<bucket>", "<db_name>"); for UDF db_name can be ommited, then current connection db will be used or stored function call calpontsys.columnstore_load_from_s3("<tablename>", "<filename>", "<bucket>", "<db_name>");	2022-07-04 19:52:37 +03:00
Roman Nozdrin	4c26e4f960	MCOL-4912 This patch introduces Extent Map index to improve EM scaleability EM scaleability project has two parts: phase1 and phase2. This is phase1 that brings EM index to speed up(from O(n) down to the speed of boost::unordered_map) EM lookups looking for <dbroot, oid, partition> tuple to turn it into LBID, e.g. most bulk insertion meta info operations. The basis is boost::shared_managed_object where EMIndex is stored. Whilst it is not debug-friendly it allows to put a nested structs into shmem. EMIndex has 3 tiers. Top down description: vector of dbroots, map of oids to partition vectors, partition vectors that have EM indices. Separate EM methods now queries index before they do EM run. EMIndex has a separate shmem file with the fixed id MCS-shm-00060001.	2022-05-04 12:59:16 +00:00
Commander thrashdin	4bb3743110	Added word bytes	2022-04-20 16:04:14 +03:00
benthompson15	2ec502aaf3	MCOL-4576: remove S3 options from cpimport. (#2328 )	2022-04-01 09:18:34 -05:00
Leonid Fedorov	65252df4f6	C++20 fixes	2022-03-28 12:32:29 +00:00
Roman Nozdrin	2b4946f53a	Revert "MCOL-4576: remove S3 options from cpimport. (#2307 )" This reverts commit `14c4840d53`.	2022-03-23 12:33:23 +00:00
benthompson15	14c4840d53	MCOL-4576: remove S3 options from cpimport. (#2307 )	2022-03-21 09:54:39 -05:00
Serguey Zefirov	53b9a2a0f9	MCOL-4580 extent elimination for dictionary-based text/varchar types The idea is relatively simple - encode prefixes of collated strings as integers and use them to compute extents' ranges. Then we can eliminate extents with strings. The actual patch does have all the code there but miss one important step: we do not keep collation index, we keep charset index. Because of this, some of the tests in the bugfix suite fail and thus main functionality is turned off. The reason of this patch to be put into PR at all is that it contains changes that made CHAR/VARCHAR columns unsigned. This change is needed in vectorization work.	2022-03-02 23:53:39 +03:00
Leonid Fedorov	6b1c696991	chars are unsigned on arm y default	2022-02-17 23:18:27 +00:00
Leonid Fedorov	3919c541ac	New warnfixes (#2254 ) * Fix clang warnings * Remove vim tab guides * initialize variables * 'strncpy' output truncated before terminating nul copying as many bytes from a string as its length * Fix ISO C++17 does not allow 'register' storage class specifier for outdated bison * chars are unsigned on ARM, having if (ival < 0) always false * chars are unsigned by default on ARM and comparison with -1 if always true	2022-02-17 13:08:58 +03:00
Gagan Goel	973e5024d8	MCOL-4957 Fix performance slowdown for processing TIMESTAMP columns. Part 1: As part of MCOL-3776 to address synchronization issue while accessing the fTimeZone member of the Func class, mutex locks were added to the accessor and mutator methods. However, this slows down processing of TIMESTAMP columns in PrimProc significantly as all threads across all concurrently running queries would serialize on the mutex. This is because PrimProc only has a single global object for the functor class (class derived from Func in utils/funcexp/functor.h) for a given function name. To fix this problem: (1) We remove the fTimeZone as a member of the Func derived classes (hence removing the mutexes) and instead use the fOperationType member of the FunctionColumn class to propagate the timezone values down to the individual functor processing functions such as FunctionColumn::getStrVal(), FunctionColumn::getIntVal(), etc. (2) To achieve (1), a timezone member is added to the execplan::CalpontSystemCatalog::ColType class. Part 2: Several functors in the Funcexp code call dataconvert::gmtSecToMySQLTime() and dataconvert::mySQLTimeToGmtSec() functions for conversion between seconds since unix epoch and broken-down representation. These functions in turn call the C library function localtime_r() which currently has a known bug of holding a global lock via a call to __tz_convert. This significantly reduces performance in multi-threaded applications where multiple threads concurrently call localtime_r(). More details on the bug: https://sourceware.org/bugzilla/show_bug.cgi?id=16145 This bug in localtime_r() caused processing of the Functors in PrimProc to slowdown significantly since a query execution causes Functors code to be processed in a multi-threaded manner. As a fix, we remove the calls to localtime_r() from gmtSecToMySQLTime() and mySQLTimeToGmtSec() by performing the timezone-to-offset conversion (done in dataconvert::timeZoneToOffset()) during the execution plan creation in the plugin. Note that localtime_r() is only called when the time_zone system variable is set to "SYSTEM". This fix also required changing the timezone type from a std::string to a long across the system.	2022-02-14 14:12:27 -05:00
Leonid Fedorov	04752ec546	clang format apply	2022-01-21 16:43:49 +00:00
Leonid Fedorov	6b6411229f	build fixes	2022-01-21 16:34:04 +00:00
Leonid Fedorov	01f3ceb437	replace header guards with #pragma once	2022-01-21 15:24:58 +00:00
Roman Nozdrin	af36f9940f	This patch introduces support for scanning/filtering vectorized execution for numeric-based data types TEXT, CHAR, VARCHAR, FLOAT and DOUBLE are not yet supported by vectorized path This patch introduces an example for Google benchmarking suite to measure a perf diff b/w legacy scan/filtering code and the templated version	2021-12-10 10:30:00 +00:00
Vicențiu Ciorbaru	342f71e7fb	Fix compilation failure on aarch64 with gcc 10.3 Ubuntu 21.04 According to C++ spec: If an inline function or variable (since C++17) with external linkage is defined differently in different translation units, the behavior is undefined. The undefined behaviour causes link errors for cpimport binary. /usr/bin/ld: /tmp/cpimport.bin.av067N.ltrans0.ltrans.o:(.data.rel.ro+0x6c8): undefined reference to `WriteEngine::ColumnOp::isEmptyRow(unsigned long, unsigned char const, int)' The isEmptyRow method is defined as inline in the cpp file and not inline in the header file. As the method is not used as part of an external API by any of the callers, nor is it subclassed, mark it as a normal (non virtual) inline member function.	2021-11-04 10:04:08 +00:00
Leonid Fedorov	56d8a33f0b	filesystem ambiguaty	2021-10-29 14:57:11 +00:00
Roman Nozdrin	550d1cf1c4	MCOL-4858 This patch fixes HWM comparison for 16 columns and reduces boilerplate code for other column widths	2021-09-06 17:10:33 +00:00
Leonid Fedorov	5c5f103f98	MCOL-4839: Fix clang build (#2100 ) * Fix clang build * Extern C returned to plugin_instance Co-authored-by: Leonid Fedorov <l.fedorov@mail.corp.ru>	2021-08-23 10:45:10 -05:00
Leonid Fedorov	202885ae37	obviously wrong bytesTx assignment (#2086 )	2021-08-18 11:50:10 -05:00
Leonid Fedorov	37ad263a7a	targetDbroot should be assined, not compared (#2087 )	2021-08-18 11:49:27 -05:00
Sergey Zefirov	6eaee180f3	MCOL-4779 Keep correct ranges during DML for short char columns	2021-07-12 14:07:39 +03:00
Sergey Zefirov	9e0851e4cf	MCOL-4766 ROLLBACK kept ranges changed inside rolled back transaction Now ROLLBACK drops ranges to INVALID state which makes engine to rescan blocks and discover correct ranges.	2021-07-07 18:16:56 +03:00
Roman Nozdrin	866dc25729	Merge pull request #1842 from denis0x0D/MCOL-987_LZ MCOL-987 LZ4 compression support.	2021-07-07 13:13:18 +03:00
Denis Khalikov	cc1c3629c5	MCOL-987 Add LZ4 compression. * Adds CompressInterfaceLZ4 which uses LZ4 API for compress/uncompress. * Adds CMake machinery to search LZ4 on running host. * All methods which use static data and do not modify any internal data - become `static`, so we can use them without creation of the specific object. This is possible, because the header specification has not been modified. We still use 2 sections in header, first one with file meta data, the second one with pointers for compressed chunks. * Methods `compress`, `uncompress`, `maxCompressedSize`, `getUncompressedSize` - become pure virtual, so we can override them for the other compression algos. * Adds method `getChunkMagicNumber`, so we can verify chunk magic number for each compression algo. * Renames "s/IDBCompressInterface/CompressInterface/g" according to requirement.	2021-07-06 18:04:37 +03:00
Gagan Goel	8520f87237	MCOL-641 Cleanup.	2021-07-06 09:01:49 +00:00
David.Hall	237cad347f	MCOL-4758 Limit LONGTEXT and LONGBLOB to 16MB (#1995 ) MCOL-4758 Limit LONGTEXT and LONGBLOB to 16MB Also add the original test case from MCOL-3879.	2021-07-05 02:09:41 -04:00
Alexander Barkov	b3d6f62964	MCOL-4753 Performance problem in Typeless join	2021-06-10 09:26:26 +00:00
Roman Nozdrin	c6d0b46bc6	Merge pull request #1984 from denis0x0D/MCOL-4685_fix_warn MCOL-4685: Fix GCC warnings.	2021-06-10 11:39:10 +03:00
Alexey Antipovsky	0dedb7e628	Fix compilation warnings	2021-06-09 16:51:00 +03:00
Denis Khalikov	2cd024145c	MCOL-4685: Fix GCC warnings. This patch fixes GCC warnings: 1.'const long int' and 'long unsigned int' [-Werror=sign-compare] 2. unused variable 'cf' [-Werror=unused-variable]	2021-06-09 14:33:32 +03:00
Roman Nozdrin	3c33a816c3	Merge pull request #1981 from denis0x0D/MCOL-4685_fix_rebase MCOL-4685 Fix bug after rebase and add comments.	2021-06-08 20:45:05 +03:00
Roman Nozdrin	7a152c6a19	Merge pull request #1944 from mariadb-AlexeyAntipovsky/MCOL-563-dev [MCOL-4709] Disk-based aggregation	2021-06-08 20:42:58 +03:00
Denis Khalikov	caa02a383a	MCOL-4685 Fix bug after rebase and add comments.	2021-06-07 13:00:11 +03:00
Alexey Antipovsky	475104e4d3	[MCOL-4709] Disk-based aggregation * Introduce multigeneration aggregation * Do not save unused part of RGDatas to disk * Add IO error explanation (strerror) * Reduce memory usage while aggregating * introduce in-memory generations to better memory utilization * Try to limit the qty of buckets at a low limit * Refactor disk aggregation a bit * pass calculated hash into RowAggregation * try to keep some RGData with free space in memory * do not dump more than half of rowgroups to disk if generations are allowed, instead start a new generation * for each thread shift the first processed bucket at each iteration, so the generations start more evenly * Unify temp data location * Explicitly create temp subdirectories whether disk aggregation/join are enabled or not	2021-06-06 16:09:15 +03:00
Denis Khalikov	606194e6e4	MCOL-4685: Eliminate some irrelevant settings (uncompressed data and extents per file). This patch: 1. Removes the option to declare uncompressed columns (set columnstore_compression_type = 0). 2. Ignores [COMMENT '[compression=0] option at table or column level (no error messages, just disregard). 3. Removes the option to set more than 2 extents per file (ExtentsPreSegmentFile). 4. Updates rebuildEM tool to support up to 10 dictionary extent per dictionary segment file. 5. Adds check for `DBRootStorageType` for rebuildEM tool. 6. Renamed rebuildEM to mcsRebuildEM.	2021-06-03 14:44:33 +03:00
Sergey Zefirov	04d5b55c37	MCOL-4652 Fixes for wide-decimal support in bulk insert operations Previously cpimport didn't send wide min/max-es talking to BRM	2021-05-31 18:40:57 +00:00

1 2 3 4 5 ...

409 Commits