mariadb-columnstore-engine

mirror of https://github.com/mariadb-corporation/mariadb-columnstore-engine.git synced 2025-10-27 08:55:33 +03:00

Author	SHA1	Message	Date
Sergei Golubchik	0f9d2d978d	cpimport.bin doesn't really need libboost_program_options linking with unused libraries creates a difference in dependencies between --no-as-needed (gcc default) and --as-needed (default on Fedora rpmbuild) builds.	2023-06-29 18:39:35 -04:00
Gagan Goel	101aa079ab	Fix resource leak in DDLProc/DMLProc/PrimProc/WriteengineServer processes. As part of the charset support, a call to MY_INIT() was added at the initialization of the above processes. This call initializes the MySQL thread environment required by the charset library. However, the accompanying my_end() call required to terminate this thread environment was not added at the termination of these process, hence leaking resources. As a fix, we move the MY_INIT() calls to the Child() functions of these services and also add the missing my_end() call.	2023-06-23 20:10:10 +00:00
Leonid Fedorov	030144127e	Remove boost shared array [develop 23.02] (#2812 ) * remove boost/shared_array include * replace boost::shared_array<T> to std::shared_ptr<T[]>	2023-04-17 20:56:09 +03:00
Leonid Fedorov	2f153184c3	Fixes of bugs from ASAN warnings, part one (#2796 )	2023-03-30 18:29:04 +03:00
Leonid Fedorov	56f2346083	Remove windows ifdefs	2023-03-02 15:59:42 +00:00
Gagan Goel	006b92bba2	Revert "This commit fixes an incorrect predicate in the if condition (#2608 )" This reverts commit `f4e3022fbd`. The commit apparently caused MCOL-5318 and MCOL-5319 which involve the internal ColumnStore batch insert mechanism passing through the SQL layer. The code block involved in this change is a predicate checking for the HWM extent in WriteEngineServer at the end of the batch insert. This is done in WE_DMLCommandProc::processBatchInsertHwm(). The original predicate check in this function for the HWM extent is restored until further investigation.	2023-02-02 08:07:18 -05:00
Gagan Goel	ad59ed5402	MCOL-5367 Fix a bug introduced in MCOL-5021 (AUX column implementation). In the implementation of MCOL-5021, an assert was added in `WE_DMLCommandProc::processBatchInsertHwm()` that assumed the `WriteEngine::TableMetaData` cache is uniform across the cluster. However, this assumption is incorrect. This bug caused undefined behaviour in ColumnStore resulting in bugs such as MCOL-5367. In MCOL-5367, in a multi-node ColumnStore cluster, an INSERT ... SELECT in a transaction with system variable `columnstore_use_import_for_batchinsert=OFF/ON` did not show inserted records when a SELECT query was issued. Assuming a 3-node cluster setup, DMLProc only sends a given batch of records to be inserted to one of the 3 nodes, and not all nodes. As a result, the `WriteEngine::TableMetaData` cache is only populated for that one node and is not uniform across the cluster, causing the assert to fail. As a fix, we simply remove this assert as it is redundant and should not have been added in the first place.	2023-01-16 05:54:44 -05:00
Denis Khalikov	d61780cab1	MCOL-5263 Add support to ROLLBACK when PP were restarted. DMLProc starts ROLLBACK when SELECT part of UPDATE fails b/c EM facility in PP were restarted. Unfortunately this ROLLBACK stuck if EM/PP are not yet available. DMLProc must have a t/o with re-try doing ROLLBACK.	2022-12-13 16:18:53 +03:00
Leonid Fedorov	b936ed8b2e	Fix some GCC-12 Build errors	2022-11-22 03:28:17 +03:00
Leonid Fedorov	37fd915a08	Serg`s patch for develop-6 revised for develop https://github.com/mariadb-corporation/mariadb-columnstore-engine/pull/2614	2022-11-09 22:41:38 +00:00
Gagan Goel	f4e3022fbd	This commit fixes an incorrect predicate in the if condition (#2608 ) that checks for HWM extent in WE_DMLCommandProc::processBatchInsertHwm().	2022-11-08 14:51:42 -06:00
Leonid Fedorov	d2432f9bf6	get rid of pointers for 128 fields	2022-08-26 15:12:22 +00:00
mariadb-AndreyPiskunov	0863ecd279	Replace getBinaryField	2022-08-25 18:21:43 +03:00
Gagan Goel	cbfdae3481	MCOL-5021 Code changes based on review feedback.	2022-08-05 14:40:50 -04:00
Gagan Goel	1355237ca3	MCOL-5021 Some minor fixes.	2022-08-05 14:40:50 -04:00
Gagan Goel	9b6d3c3870	MCOL-5021 Add support for AUX column in the client code calling CalpontSystemCatalog::columnRIDs().	2022-08-05 14:40:49 -04:00
Gagan Goel	262cd5c501	MCOL-5021 Remove hard-coded values for data type, column width and compression type for the AUX column, and replace them with constants defined in the execplan namespace.	2022-08-05 14:40:49 -04:00
Gagan Goel	c8b6b154bf	MCOL-5021 Add an option in Columnstore.xml, fastdelete (disabled by default), which when enabled, indiscriminately invalidates all column extents and performs the actual DELETE only on the AUX column. The trade-off with this approach would now be that the first SELECT for certain query patterns (those containing a WHERE predicate) after the DELETE operation will slow down as the invalidated column extent would need to be scanned again to set the min/max values.	2022-08-05 14:40:49 -04:00
Gagan Goel	60eb0f86ec	MCOL-5021 non-AUX column files are opened in read-only mode during the DELETE operation. ColumnOp::readBlock() calls can cause writes to database files when the active chunk list in ChunkManager is full. Since non-AUX columns are read-only for the DELETE operation, we prevent writes of compressed chunks and header for these columns by passing an isReadOnly flag to CompFileData which indicates whether the column is read-only or read-write.	2022-08-05 14:40:49 -04:00
Gagan Goel	35a3a93964	MCOL-5021 For the DELETE operation, empty magic values are only written to database files for AUX column. Perform read-only operation for other columns in the table to update the Casual Partitioning information.	2022-08-05 14:40:49 -04:00
Gagan Goel	86df9a972c	MCOL-5021 Add prototype support for the AUX column in CREATE/DROP DDL commands, single and multi-value INSERTs, cpimport, and DELETE.	2022-08-05 14:40:49 -04:00
Denis Khalikov	fb1e23bb83	[MCOL-5106] Add support to work with StorageManager. This patch eliminates boost::filesystem from `mcsRebuildEM` tool. After this change we should be able to work with any filesystem even S3.	2022-07-28 16:47:34 +03:00
Leonid Fedorov	f5b2a6885f	MCOL-5013: Load Data from S3 into Columnstore Introduced UDF and stored prodecure. usage: set columnstore_s3_key='<s3_key>'; set columnstore_s3_secret='<s3_secret>'; set columnstore_s3_region='region'; and then use UDF select columnstore_dataload("<tablename>", "<filename>", "<bucket>", "<db_name>"); for UDF db_name can be ommited, then current connection db will be used or stored function call calpontsys.columnstore_load_from_s3("<tablename>", "<filename>", "<bucket>", "<db_name>");	2022-07-04 19:52:37 +03:00
Roman Nozdrin	4c26e4f960	MCOL-4912 This patch introduces Extent Map index to improve EM scaleability EM scaleability project has two parts: phase1 and phase2. This is phase1 that brings EM index to speed up(from O(n) down to the speed of boost::unordered_map) EM lookups looking for <dbroot, oid, partition> tuple to turn it into LBID, e.g. most bulk insertion meta info operations. The basis is boost::shared_managed_object where EMIndex is stored. Whilst it is not debug-friendly it allows to put a nested structs into shmem. EMIndex has 3 tiers. Top down description: vector of dbroots, map of oids to partition vectors, partition vectors that have EM indices. Separate EM methods now queries index before they do EM run. EMIndex has a separate shmem file with the fixed id MCS-shm-00060001.	2022-05-04 12:59:16 +00:00
Commander thrashdin	4bb3743110	Added word bytes	2022-04-20 16:04:14 +03:00
benthompson15	2ec502aaf3	MCOL-4576: remove S3 options from cpimport. (#2328 )	2022-04-01 09:18:34 -05:00
Leonid Fedorov	65252df4f6	C++20 fixes	2022-03-28 12:32:29 +00:00
Roman Nozdrin	2b4946f53a	Revert "MCOL-4576: remove S3 options from cpimport. (#2307 )" This reverts commit `14c4840d53`.	2022-03-23 12:33:23 +00:00
benthompson15	14c4840d53	MCOL-4576: remove S3 options from cpimport. (#2307 )	2022-03-21 09:54:39 -05:00
Serguey Zefirov	53b9a2a0f9	MCOL-4580 extent elimination for dictionary-based text/varchar types The idea is relatively simple - encode prefixes of collated strings as integers and use them to compute extents' ranges. Then we can eliminate extents with strings. The actual patch does have all the code there but miss one important step: we do not keep collation index, we keep charset index. Because of this, some of the tests in the bugfix suite fail and thus main functionality is turned off. The reason of this patch to be put into PR at all is that it contains changes that made CHAR/VARCHAR columns unsigned. This change is needed in vectorization work.	2022-03-02 23:53:39 +03:00
Leonid Fedorov	6b1c696991	chars are unsigned on arm y default	2022-02-17 23:18:27 +00:00
Leonid Fedorov	3919c541ac	New warnfixes (#2254 ) * Fix clang warnings * Remove vim tab guides * initialize variables * 'strncpy' output truncated before terminating nul copying as many bytes from a string as its length * Fix ISO C++17 does not allow 'register' storage class specifier for outdated bison * chars are unsigned on ARM, having if (ival < 0) always false * chars are unsigned by default on ARM and comparison with -1 if always true	2022-02-17 13:08:58 +03:00
Gagan Goel	973e5024d8	MCOL-4957 Fix performance slowdown for processing TIMESTAMP columns. Part 1: As part of MCOL-3776 to address synchronization issue while accessing the fTimeZone member of the Func class, mutex locks were added to the accessor and mutator methods. However, this slows down processing of TIMESTAMP columns in PrimProc significantly as all threads across all concurrently running queries would serialize on the mutex. This is because PrimProc only has a single global object for the functor class (class derived from Func in utils/funcexp/functor.h) for a given function name. To fix this problem: (1) We remove the fTimeZone as a member of the Func derived classes (hence removing the mutexes) and instead use the fOperationType member of the FunctionColumn class to propagate the timezone values down to the individual functor processing functions such as FunctionColumn::getStrVal(), FunctionColumn::getIntVal(), etc. (2) To achieve (1), a timezone member is added to the execplan::CalpontSystemCatalog::ColType class. Part 2: Several functors in the Funcexp code call dataconvert::gmtSecToMySQLTime() and dataconvert::mySQLTimeToGmtSec() functions for conversion between seconds since unix epoch and broken-down representation. These functions in turn call the C library function localtime_r() which currently has a known bug of holding a global lock via a call to __tz_convert. This significantly reduces performance in multi-threaded applications where multiple threads concurrently call localtime_r(). More details on the bug: https://sourceware.org/bugzilla/show_bug.cgi?id=16145 This bug in localtime_r() caused processing of the Functors in PrimProc to slowdown significantly since a query execution causes Functors code to be processed in a multi-threaded manner. As a fix, we remove the calls to localtime_r() from gmtSecToMySQLTime() and mySQLTimeToGmtSec() by performing the timezone-to-offset conversion (done in dataconvert::timeZoneToOffset()) during the execution plan creation in the plugin. Note that localtime_r() is only called when the time_zone system variable is set to "SYSTEM". This fix also required changing the timezone type from a std::string to a long across the system.	2022-02-14 14:12:27 -05:00
Leonid Fedorov	04752ec546	clang format apply	2022-01-21 16:43:49 +00:00
Leonid Fedorov	6b6411229f	build fixes	2022-01-21 16:34:04 +00:00
Leonid Fedorov	01f3ceb437	replace header guards with #pragma once	2022-01-21 15:24:58 +00:00
Roman Nozdrin	af36f9940f	This patch introduces support for scanning/filtering vectorized execution for numeric-based data types TEXT, CHAR, VARCHAR, FLOAT and DOUBLE are not yet supported by vectorized path This patch introduces an example for Google benchmarking suite to measure a perf diff b/w legacy scan/filtering code and the templated version	2021-12-10 10:30:00 +00:00
Vicențiu Ciorbaru	342f71e7fb	Fix compilation failure on aarch64 with gcc 10.3 Ubuntu 21.04 According to C++ spec: If an inline function or variable (since C++17) with external linkage is defined differently in different translation units, the behavior is undefined. The undefined behaviour causes link errors for cpimport binary. /usr/bin/ld: /tmp/cpimport.bin.av067N.ltrans0.ltrans.o:(.data.rel.ro+0x6c8): undefined reference to `WriteEngine::ColumnOp::isEmptyRow(unsigned long, unsigned char const, int)' The isEmptyRow method is defined as inline in the cpp file and not inline in the header file. As the method is not used as part of an external API by any of the callers, nor is it subclassed, mark it as a normal (non virtual) inline member function.	2021-11-04 10:04:08 +00:00
Leonid Fedorov	56d8a33f0b	filesystem ambiguaty	2021-10-29 14:57:11 +00:00
Roman Nozdrin	550d1cf1c4	MCOL-4858 This patch fixes HWM comparison for 16 columns and reduces boilerplate code for other column widths	2021-09-06 17:10:33 +00:00
Leonid Fedorov	5c5f103f98	MCOL-4839: Fix clang build (#2100 ) * Fix clang build * Extern C returned to plugin_instance Co-authored-by: Leonid Fedorov <l.fedorov@mail.corp.ru>	2021-08-23 10:45:10 -05:00
Leonid Fedorov	202885ae37	obviously wrong bytesTx assignment (#2086 )	2021-08-18 11:50:10 -05:00
Leonid Fedorov	37ad263a7a	targetDbroot should be assined, not compared (#2087 )	2021-08-18 11:49:27 -05:00
Sergey Zefirov	6eaee180f3	MCOL-4779 Keep correct ranges during DML for short char columns	2021-07-12 14:07:39 +03:00
Sergey Zefirov	9e0851e4cf	MCOL-4766 ROLLBACK kept ranges changed inside rolled back transaction Now ROLLBACK drops ranges to INVALID state which makes engine to rescan blocks and discover correct ranges.	2021-07-07 18:16:56 +03:00
Roman Nozdrin	866dc25729	Merge pull request #1842 from denis0x0D/MCOL-987_LZ MCOL-987 LZ4 compression support.	2021-07-07 13:13:18 +03:00
Denis Khalikov	cc1c3629c5	MCOL-987 Add LZ4 compression. * Adds CompressInterfaceLZ4 which uses LZ4 API for compress/uncompress. * Adds CMake machinery to search LZ4 on running host. * All methods which use static data and do not modify any internal data - become `static`, so we can use them without creation of the specific object. This is possible, because the header specification has not been modified. We still use 2 sections in header, first one with file meta data, the second one with pointers for compressed chunks. * Methods `compress`, `uncompress`, `maxCompressedSize`, `getUncompressedSize` - become pure virtual, so we can override them for the other compression algos. * Adds method `getChunkMagicNumber`, so we can verify chunk magic number for each compression algo. * Renames "s/IDBCompressInterface/CompressInterface/g" according to requirement.	2021-07-06 18:04:37 +03:00
Gagan Goel	8520f87237	MCOL-641 Cleanup.	2021-07-06 09:01:49 +00:00
David.Hall	237cad347f	MCOL-4758 Limit LONGTEXT and LONGBLOB to 16MB (#1995 ) MCOL-4758 Limit LONGTEXT and LONGBLOB to 16MB Also add the original test case from MCOL-3879.	2021-07-05 02:09:41 -04:00
Alexander Barkov	b3d6f62964	MCOL-4753 Performance problem in Typeless join	2021-06-10 09:26:26 +00:00

1 2 3 4 5 ...

368 Commits