glibc

lib/glibc

mirror of https://sourceware.org/git/glibc.git synced 2025-10-12 19:04:54 +03:00

Author	SHA1	Message	Date
Wilco Dijkstra	ce2f26a22e	AArch64: Remove PTR_ARG/SIZE_ARG defines This series removes various ILP32 defines that are now no longer needed. Remove PTR_ARG/SIZE_ARG. Reviewed-by: Adhemerval Zanella <adhemerval.zanella@linaro.org>	2025-02-24 14:15:15 +00:00
Wilco Dijkstra	be0cfd848d	stdlib: Add single-threaded fast path to rand() Improve performance of rand() and __random() by adding a single-threaded fast path. Bench-random-lock shows about 5x speedup on Neoverse V1. Reviewed-by: Adhemerval Zanella <adhemerval.zanella@linaro.org>	2025-02-24 14:13:43 +00:00
Stefan Liebler	4734d0f8ad	Increase the amount of data tested in stdio-common/tst-fwrite-pipe.c The number of iterations and the length of the string are not high enough on some systems causing the test to return false-positives. Testcase stdio-common/tst-fwrite-bz29459.c was fixed in the same way in `1b6f868625` (Increase the amount of data tested in stdio-common/tst-fwrite-bz29459.c, 2025-02-14) Testcases stdio-common/tst-fwrite-bz29459.c and stdio-common/tst-fwrite-pipe.c were introcued in `596a61cf6b` (libio: Start to return errors when flushing fwrite's buffer [BZ #29459], 2025-01-28)	2025-02-24 14:45:55 +01:00
Frédéric Bérat	8a46bf41e5	posix: Rewrite cpuset tests Rewriting the cpuset macros test to cover more use cases and port the tests to the new test infrastructure. The use cases include bad actor access attempts, before and after the CPU set structure. Reviewed-by: Tulio Magno Quites Machado Filho <tuliom@redhat.com>	2025-02-24 14:19:51 +01:00
Frédéric Bérat	fa53723cdb	support: Add support_next_to_fault_before support function Refactor the support_next_to_fault and add the support_next_to_fault_before method returns a buffer with a protected page before it, to be able to test buffer underflow accesses. Reviewed-by: Tulio Magno Quites Machado Filho <tuliom@redhat.com>	2025-02-24 14:19:36 +01:00
koraynilay	29803ed3ce	math: Fix `unknown type name '__float128'` for clang 3.4 to 3.8.1 (bug 32694) When compiling a program that includes <bits/floatn.h> using a clang version between 3.4 (included) and 3.8.1 (included), clang will fail with `unknown type name '__float128'; did you mean '__cfloat128'?`. This changes fixes the clang prerequirements macro call in floatn.h to check for clang 3.9 instead of 3.4, since support for __float128 was actually enabled in 3.9 by: commit 50f29e06a1b6a38f0bba9360cbff72c82d46cdd4 Author: Nemanja Ivanovic <nemanja.i.ibm@gmail.com> Date: Wed Apr 13 09:49:45 2016 +0000 Enable support for __float128 in Clang This fixes bug 32694. Signed-off-by: koraynilay <koray.fra@gmail.com> Reviewed-by: H.J. Lu <hjl.tools@gmail.com>	2025-02-23 11:47:11 +08:00
Michael Jeanson	689a62a421	nptl: clear the whole rseq area before registration Due to the extensible nature of the rseq area we can't explictly initialize fields that are not part of the ABI yet. It was agreed with upstream that all new fields will be documented as zero initialized by userspace. Future kernels configured with CONFIG_DEBUG_RSEQ will validate the content of all fields during registration. Replace the explicit field initialization with a memset of the whole rseq area which will cover fields as they are added to future kernels. Signed-off-by: Michael Jeanson <mjeanson@efficios.com> Reviewed-by: Florian Weimer <fweimer@redhat.com>	2025-02-21 22:21:25 +00:00
Yury Khrustalev	41f6684557	aarch64: Add GCS test with signal handler Test that when we return from a function that enabled GCS at runtime we get SIGSEGV. Also test that ucontext contains GCS block with the GCS pointer. Reviewed-by: Adhemerval Zanella <adhemerval.zanella@linaro.org>	2025-02-21 16:23:44 +00:00
Yury Khrustalev	15afd01e80	aarch64: Add GCS tests for dlopen Reviewed-by: Adhemerval Zanella <adhemerval.zanella@linaro.org>	2025-02-21 16:10:44 +00:00
Yury Khrustalev	57ee1deb1f	aarch64: Add GCS tests for transitive dependencies Reviewed-by: Adhemerval Zanella <adhemerval.zanella@linaro.org>	2025-02-21 16:09:06 +00:00
Yury Khrustalev	82decb59bc	aarch64: Add tests for Guarded Control Stack These tests validate that GCS tunable works as expected depending on the GCS markings in the test binaries. Tests validate both static and dynamically linked binaries. These new tests are AArch64 specific. Moreover, they are included only if linker supports the "-z gcs=<value>" option. If built, these tests will run on systems with and without HWCAP_GCS. In the latter case the tests will be reported as UNSUPPORTED. Reviewed-by: Adhemerval Zanella <adhemerval.zanella@linaro.org>	2025-02-21 16:08:00 +00:00
Yury Khrustalev	c05086d904	aarch64: Add configure checks for GCS support - Add check that linker supports -z gcs=... - Add checks that main and test compiler support -mbranch-protection=gcs Reviewed-by: Adhemerval Zanella <adhemerval.zanella@linaro.org>	2025-02-21 16:06:44 +00:00
Carlos O'Donell	6d24313e4a	manual: Mark setlogmask as AS-unsafe and AC-unsafe. This fixes the check-safety.sh failure with commit `ad9c4c5361`, and correctly marks the function AS-unsafe and AC-unsafe due to the use of the non-recursive lock. Tested on x86_64 without regressions. Reviewed-by: Frédéric Bérat <fberat@redhat.com>	2025-02-20 11:55:21 -05:00
Wilco Dijkstra	163b1bbb76	AArch64: Add SVE memset Add SVE memset based on the generic memset with predicated load for sizes < 16. Unaligned memsets of 128-1024 are improved by ~20% on average by using aligned stores for the last 64 bytes. Performance of random memset benchmark improves by ~2% on Neoverse V1. Reviewed-by: Yury Khrustalev <yury.khrustalev@arm.com>	2025-02-20 15:31:50 +00:00
H.J. Lu	5a4573be6f	x86 (__HAVE_FLOAT128): Defined to 0 for Intel SYCL compiler [BZ #32723 ] Intel compiler always defines __INTEL_LLVM_COMPILER. When SYCL is enabled by -fsycl, it also defines SYCL_LANGUAGE_VERSION. Since Intel SYCL compiler doesn't support _Float128: https://github.com/intel/llvm/issues/16903 define __HAVE_FLOAT128 to 0 for Intel SYCL compiler. This fixes BZ #32723. Signed-off-by: H.J. Lu <hjl.tools@gmail.com> Reviewed-by: Sam James <sam@gentoo.org>	2025-02-20 08:42:37 +08:00
Carlos O'Donell	ad9c4c5361	manual: Document setlogmask as MT-safe. setlogmask(3) was made MT-safe in glibc-2.33 with the fix for bug 26100. Reviewed-by: Florian Weimer <fweimer@redhat.com>	2025-02-19 16:22:54 -05:00
Adhemerval Zanella	0242c9f9e6	math: Consolidate acosf and asinf internal tables The libm size improvement built with gcc-14, "--enable-stack-protector=strong --enable-bind-now=yes --enable-fortify-source=2": Before: 582292 844 12 583148 8e5ec aarch64-linux-gnu/math/libm.so 975133 1076 12 976221 ee55d x86_64-linux-gnu/math/libm.so 1203586 5608 368 1209562 1274da powerpc64le-linux-gnu/math/libm.so After: 581972 844 12 582828 8e4ac aarch64-linux-gnu/math/libm.so 974941 1076 12 976029 ee49d x86_64-linux-gnu/math/libm.so 1203394 5608 368 1209370 12741a powerpc64le-linux-gnu/math/libm.so Reviewed-by: Andreas K. Huettel <dilfridge@gentoo.org>	2025-02-17 10:09:09 -03:00
Adhemerval Zanella	1faccf388a	math: Consolidate acospif and asinpif internal tables The libm size improvement built with gcc-14, "--enable-stack-protector=strong --enable-bind-now=yes --enable-fortify-source=2": Before: text data bss dec hex filename 583444 844 12 584300 8ea6c aarch64-linux-gnu/math/libm.so 976349 1076 12 977437 eea1d x86_64-linux-gnu/math/libm.so 1204738 5608 368 1210714 12795a powerpc64le-linux-gnu/math/libm.so After: 582292 844 12 583148 8e5ec aarch64-linux-gnu/math/libm.so 975133 1076 12 976221 ee55d x86_64-linux-gnu/math/libm.so 1203586 5608 368 1209562 1274da powerpc64le-linux-gnu/math/libm.so Reviewed-by: Andreas K. Huettel <dilfridge@gentoo.org>	2025-02-17 10:09:09 -03:00
Adhemerval Zanella	246e52574d	math: Consolidate cospif and sinpif internal tables The libm size improvement built with gcc-14, "--enable-stack-protector=strong --enable-bind-now=yes --enable-fortify-source=2": Before: text data bss dec hex filename 584500 844 12 585356 8ee8c aarch64-linux-gnu/math/libm.so 977341 1076 12 978429 eedfd x86_64-linux-gnu/math/libm.so 1205762 5608 368 1211738 127d5a powerpc64le-linux-gnu/math/libm.so After: text data bss dec hex filename 583444 844 12 584300 8ea6c aarch64-linux-gnu/math/libm.so 976349 1076 12 977437 eea1d x86_64-linux-gnu/math/libm.so 1204738 5608 368 1210714 12795a powerpc64le-linux-gnu/math/libm.so Reviewed-by: Andreas K. Huettel <dilfridge@gentoo.org>	2025-02-17 10:09:09 -03:00
gfleury	4afbc1aa2e	htl: don't export __pthread_default_rwlockattr anymore. since now all symbloy that use it are in libc Message-ID: <20250216145434.7089-11-gfleury@disroot.org>	2025-02-16 23:43:04 +01:00
gfleury	6f6732c1c4	htl: move pthread_rwlock_init into libc. Signed-off-by: gfleury <gfleury@disroot.org> Message-ID: <20250216145434.7089-10-gfleury@disroot.org>	2025-02-16 23:43:03 +01:00
gfleury	d3ef1b56aa	htl: move pthread_rwlock_destroy into libc. Signed-off-by: gfleury <gfleury@disroot.org> Message-ID: <20250216145434.7089-9-gfleury@disroot.org>	2025-02-16 23:42:38 +01:00
gfleury	25650ef6b9	htl: move pthread_rwlock_{rdlock, timedrdlock, timedwrlock, wrlock, clockrdlock, clockwrlock} into libc. Signed-off-by: gfleury <gfleury@disroot.org> Message-ID: <20250216145434.7089-8-gfleury@disroot.org>	2025-02-16 23:08:54 +01:00
gfleury	119798a7b1	htl: move pthread_rwlock_unlock into libc. Signed-off-by: gfleury <gfleury@disroot.org> Message-ID: <20250216145434.7089-7-gfleury@disroot.org>	2025-02-16 23:08:54 +01:00
gfleury	18accc19b9	htl: move pthread_rwlock_tryrdlock, pthread_rwlock_trywrlock into libc. Signed-off-by: gfleury <gfleury@disroot.org> Message-ID: <20250216145434.7089-6-gfleury@disroot.org>	2025-02-16 22:59:34 +01:00
gfleury	4b25413df5	htl: move pthread_rwlockattr_getpshared, pthread_rwlockattr_setpshared into libc. Signed-off-by: gfleury <gfleury@disroot.org> Message-ID: <20250216145434.7089-5-gfleury@disroot.org>	2025-02-16 22:59:25 +01:00
gfleury	cd2d31ed58	htl: move pthread_rwlockattr_destroy into libc. Signed-off-by: gfleury <gfleury@disroot.org> Message-ID: <20250216145434.7089-4-gfleury@disroot.org>	2025-02-16 22:59:16 +01:00
gfleury	e618b671cd	htl: move pthread_rwlockattr_init into libc. Signed-off-by: gfleury <gfleury@disroot.org> Message-ID: <20250216145434.7089-3-gfleury@disroot.org>	2025-02-16 22:59:07 +01:00
gfleury	8f842ce13e	htl: move __pthread_default_rwlockattr into libc. Signed-off-by: gfleury <gfleury@disroot.org> Message-ID: <20250216145434.7089-2-gfleury@disroot.org>	2025-02-16 22:59:00 +01:00
Aurelien Jarno	60f2d6be65	Fix tst-aarch64-pkey to handle ENOSPC as not supported The syscall pkey_alloc can return ENOSPC to indicate either that all keys are in use or that the system runs in a mode in which memory protection keys are disabled. In such case the test should not fail and just return unsupported. This matches the behaviour of the generic tst-pkey. Reviewed-by: Adhemerval Zanella <adhemerval.zanella@linaro.org> Reviewed-by: Florian Weimer <fweimer@redhat.com>	2025-02-15 11:08:43 +01:00
Tulio Magno Quites Machado Filho	1b6f868625	Increase the amount of data tested in stdio-common/tst-fwrite-bz29459.c The number of iterations and the length of the string are not high enough on some systems causing the test to return false-positives. Fixes: `596a61cf6b` (libio: Start to return errors when flushing fwrite's buffer [BZ #29459], 2025-01-28) Reported-by: Florian Weimer <fweimer@redhat.com>	2025-02-14 15:46:38 -03:00
Florian Weimer	aa3d7bd529	elf: Keep using minimal malloc after early DTV resize (bug 32412) If an auditor loads many TLS-using modules during startup, it is possible to trigger DTV resizing. Previously, the DTV was marked as allocated by the main malloc afterwards, even if the minimal malloc was still in use. With this change, _dl_resize_dtv marks the resized DTV as allocated with the minimal malloc. The new test reuses TLS-using modules from other auditing tests. Reviewed-by: DJ Delorie <dj@redhat.com>	2025-02-13 21:56:52 +01:00
Tulio Magno Quites Machado Filho	88f7ef881d	libio: Initialize _total_written for all kinds of streams Move the initialization code to a general place instead of keeping it specific to file-backed streams. Fixes: `596a61cf6b` (libio: Start to return errors when flushing fwrite's buffer [BZ #29459], 2025-01-28) Reported-by: Florian Weimer <fweimer@redhat.com> Reviewed-by: Arjun Shankar <arjun@redhat.com>	2025-02-13 16:34:54 -03:00
Ben Kallus	d10176c0ff	malloc: Add size check when moving fastbin->tcache By overwriting a forward link in a fastbin chunk that is subsequently moved into the tcache, it's possible to get malloc to return an arbitrary address [0]. When a chunk is fetched from a fastbin, its size is checked against the expected chunk size for that fastbin (see malloc.c:3991). This patch adds a similar check for chunks being moved from a fastbin to tcache, which renders obsolete the exploitation technique described above. Now updated to use __glibc_unlikely instead of __builtin_expect, as requested. [0]: https://github.com/shellphish/how2heap/blob/master/glibc_2.39/fastbin_reverse_into_tcache.c Signed-off-by: Ben Kallus <benjamin.p.kallus.gr@dartmouth.edu> Reviewed-by: Adhemerval Zanella <adhemerval.zanella@linaro.org>	2025-02-13 16:31:28 -03:00
Tobias Stoeckmann	6a3cb6b1bd	nss: Improve network number parsers (bz 32573, 32575) Make sure that numbers never overflow uint32_t in inet_network to properly validate octets encountered in IPv4 addresses. Avoid malloca in NSS networks file code because /etc/networks lines can be arbitrarily long. Instead of handcrafting the input for inet_network by adding ".0" octets if they are missing, just left shift the result. Also, do not accept invalid entries, but ignore the line instead. Reviewed-by: Adhemerval Zanella <adhemerval.zanella@linaro.org> Signed-off-by: Tobias Stoeckmann <tobias@stoeckmann.org>	2025-02-13 16:31:28 -03:00
Carlos O'Donell	991febc2f4	nptl: Remove unused __g_refs comment. In the block comment for __pthread_cond_wait_common we mention __g_refs, but the implementation no longer uses group references.	2025-02-13 14:25:25 -05:00
Siddhesh Poyarekar	a30374e4ce	advisories: Fix up GLIBC-SA-2025-0001 Add ref for the test case as well as backports. Signed-off-by: Siddhesh Poyarekar <siddhesh@sourceware.org>	2025-02-13 14:09:44 -05:00
Yat Long Poon	95e807209b	AArch64: Improve codegen for SVE powf Improve memory access with indexed/unpredicated instructions. Eliminate register spills. Speedup on Neoverse V1: 3%. Reviewed-by: Wilco Dijkstra <Wilco.Dijkstra@arm.com>	2025-02-13 18:16:54 +00:00
Yat Long Poon	0b195651db	AArch64: Improve codegen for SVE pow Move constants to struct. Improve memory access with indexed/unpredicated instructions. Eliminate register spills. Speedup on Neoverse V1: 24%. Reviewed-by: Wilco Dijkstra <Wilco.Dijkstra@arm.com>	2025-02-13 18:16:54 +00:00
Yat Long Poon	f5ff34cb3c	AArch64: Improve codegen for SVE erfcf Reduce number of MOV/MOVPRFXs and use unpredicated FMUL. Replace MUL with LSL. Speedup on Neoverse V1: 6%. Reviewed-by: Wilco Dijkstra <Wilco.Dijkstra@arm.com>	2025-02-13 18:16:54 +00:00
Luna Lamb	c0ff447edf	Aarch64: Improve codegen in SVE exp and users, and update expf_inline Use unpredicted muls, and improve memory access. 7%, 3% and 1% improvement in throughput microbenchmark on Neoverse V1, for exp, exp2 and cosh respectively. Reviewed-by: Wilco Dijkstra <Wilco.Dijkstra@arm.com>	2025-02-13 18:16:54 +00:00
Luna Lamb	8f0e7fe61e	Aarch64: Improve codegen in SVE asinh Use unpredicated muls, use lanewise mla's and improve memory access. 1% regression in throughput microbenchmark on Neoverse V1. Reviewed-by: Wilco Dijkstra <Wilco.Dijkstra@arm.com>	2025-02-13 18:16:54 +00:00
Wilco Dijkstra	5afaf99edb	math: Improve layout of exp/exp10 data GCC aligns global data to 16 bytes if their size is >= 16 bytes. This patch changes the exp_data struct slightly so that the fields are better aligned and without gaps. As a result on targets that support them, more load-pair instructions are used in exp. Exp10 is improved by moving invlog10_2N later so that neglog10_2hiN and neglog10_2loN can be loaded using load-pair. The exp benchmark improves 2.5%, "144bits" by 7.2%, "768bits" by 12.7% on Neoverse V2. Exp10 improves by 1.5%. Reviewed-by: Adhemerval Zanella <adhemerval.zanella@linaro.org>	2025-02-13 18:16:54 +00:00
Siddhesh Poyarekar	cdb9ba8419	assert: Add test for CVE-2025-0395 Use the __progname symbol to override the program name to induce the failure that CVE-2025-0395 describes. This is related to BZ #32582 Signed-off-by: Siddhesh Poyarekar <siddhesh@sourceware.org> Reviewed-by: Adhemerval Zanella <adhemerval.zanella@linaro.org>	2025-02-13 12:33:27 -05:00
Adhemerval Zanella	b81252c4b9	math: Consolidate coshf and sinhf internal tables The libm size improvement built with "--enable-stack-protector=strong --enable-bind-now=yes --enable-fortify-source=2": Before: text data bss dec hex filename 585192 860 12 586064 8f150 aarch64-linux-gnu/math/libm.so 960775 1068 12 961855 ead3f x86_64-linux-gnu/math/libm.so 1189174 5544 368 1195086 123c4e powerpc64le-linux-gnu/math/libm.so After: text data bss dec hex filename 584952 860 12 585824 8f060 aarch64-linux-gnu/math/libm.so 960615 1068 12 961695 eac9f x86_64-linux-gnu/math/libm.so 1189078 5544 368 1194990 123bee powerpc64le-linux-gnu/math/libm.so The are small code changes for x86_64 and powerpc64le, which do not affect performance; but on aarch64 with gcc-14 I see a slight better code generation due the usage of ldq for floating point constant loading. Reviewed-by: Andreas K. Huettel <dilfridge@gentoo.org>	2025-02-12 16:31:57 -03:00
Adhemerval Zanella	994007ff29	math: Consolidate acoshf and asinhf internal tables The libm size improvement built with "--enable-stack-protector=strong --enable-bind-now=yes --enable-fortify-source=2": Before: text data bss dec hex filename 587304 860 12 588176 8f990 aarch64-linux-gnu-master/math/libm.so 962855 1068 12 963935 eb55f x86_64-linux-gnu-master/math/libm.so 1191222 5544 368 1197134 12444e powerpc64le-linux-gnu-master/math/libm.so After: text data bss dec hex filename 585192 860 12 586064 8f150 aarch64-linux-gnu/math/libm.so 960775 1068 12 961855 ead3f x86_64-linux-gnu/math/libm.so 1189174 5544 368 1195086 123c4e powerpc64le-linux-gnu/math/libm.so The are small code changes for x86_64 and powerpc64le, which do not affect performance; but on aarch64 with gcc-14 I see a slight better code generation due the usage of ldq for floating point constant loading. Reviewed-by: Andreas K. Huettel <dilfridge@gentoo.org>	2025-02-12 16:31:57 -03:00
Adhemerval Zanella	8f170dc819	math: Use tanpif from CORE-MATH The CORE-MATH implementation is correctly rounded (for any rounding mode) and shows better performance to the generic tanpif. The code was adapted to glibc style and to use the definition of math_config.h (to handle errno, overflow, and underflow). Benchtest on x64_64 (Ryzen 9 5900X, gcc 14.2.1), aarch64 (Neoverse-N1, gcc 13.3.1), and powerpc (POWER10, gcc 13.2.1): latency master patched improvement x86_64 85.1683 47.7990 43.88% x86_64v2 76.8219 41.4679 46.02% x86_64v3 73.7775 37.7734 48.80% aarch64 (Neoverse) 35.4514 18.0742 49.02% power8 22.7604 10.1054 55.60% power10 22.1358 9.9553 55.03% reciprocal-throughput master patched improvement x86_64 41.0174 19.4718 52.53% x86_64v2 34.8565 11.3761 67.36% x86_64v3 34.0325 9.6989 71.50% aarch64 (Neoverse) 25.4349 9.2017 63.82% power8 13.8626 3.8486 72.24% power10 11.7933 3.6420 69.12% Reviewed-by: DJ Delorie <dj@redhat.com>	2025-02-12 16:31:57 -03:00
Adhemerval Zanella	de2fca9fe2	math: Use sinpif from CORE-MATH The CORE-MATH implementation is correctly rounded (for any rounding mode) and shows better performance to the generic sinpif. The code was adapted to glibc style and to use the definition of math_config.h (to handle errno, overflow, and underflow). Benchtest on x64_64 (Ryzen 9 5900X, gcc 14.2.1), aarch64 (Neoverse-N1, gcc 13.3.1), and powerpc (POWER10, gcc 13.2.1): latency master patched improvement x86_64 47.5710 38.4455 19.18% x86_64v2 46.8828 40.7563 13.07% x86_64v3 44.0034 34.1497 22.39% aarch64 (Neoverse) 19.2493 14.1968 26.25% power8 23.5312 16.3854 30.37% power10 22.6485 10.2888 54.57% reciprocal-throughput master patched improvement x86_64 21.8858 11.6717 46.67% x86_64v2 22.0620 11.9853 45.67% x86_64v3 21.5653 11.3291 47.47% aarch64 (Neoverse) 13.0615 6.5499 49.85% power8 16.2030 6.9580 57.06% power10 12.8911 4.2858 66.75% Reviewed-by: DJ Delorie <dj@redhat.com>	2025-02-12 16:31:57 -03:00
Adhemerval Zanella	be85208b9f	math: Use cospif from CORE-MATH The CORE-MATH implementation is correctly rounded (for any rounding mode) and shows better performance to the generic cospif. The code was adapted to glibc style and to use the definition of math_config.h (to handle errno, overflow, and underflow). Benchtest on x64_64 (Ryzen 9 5900X, gcc 14.2.1), aarch64 (Neoverse-N1, gcc 13.3.1), and powerpc (POWER10, gcc 13.2.1): latency master patched improvement x86_64 47.4679 38.4157 19.07% x86_64v2 46.9686 38.3329 18.39% x86_64v3 43.8929 31.8510 27.43% aarch64 (Neoverse) 18.8867 13.2089 30.06% power8 22.9435 7.8023 65.99% power10 15.4472 7.77505 49.67% reciprocal-throughput master patched improvement x86_64 20.9518 11.4991 45.12% x86_64v2 19.8699 10.5921 46.69% x86_64v3 19.3475 9.3998 51.42% aarch64 (Neoverse) 12.5767 6.2158 50.58% power8 15.0566 3.2654 78.31% power10 9.2866 3.1147 66.46% Reviewed-by: DJ Delorie <dj@redhat.com>	2025-02-12 16:31:57 -03:00
Adhemerval Zanella	95a01ea955	math: Use atanpif from CORE-MATH The CORE-MATH implementation is correctly rounded (for any rounding mode) and shows better performance to the generic atanpif. The code was adapted to glibc style and to use the definition of math_config.h (to handle errno, overflow, and underflow). Benchtest on x64_64 (Ryzen 9 5900X, gcc 14.2.1), aarch64 (Neoverse-N1, gcc 13.3.1), and powerpc (POWER10, gcc 13.2.1): latency master patched improvement x86_64 66.3296 52.7558 20.46% x86_64v2 66.0429 51.4007 22.17% x86_64v3 60.6294 48.7876 19.53% aarch64 (Neoverse) 24.3163 20.9110 14.00% power8 16.5766 13.3620 19.39% power10 16.5115 13.4072 18.80% reciprocal-throughput master patched improvement x86_64 30.8599 16.0866 47.87% x86_64v2 29.2286 15.4688 47.08% x86_64v3 23.0960 12.8510 44.36% aarch64 (Neoverse) 15.4619 10.6752 30.96% power8 7.9200 5.2483 33.73% power10 6.8539 4.6262 32.50% Reviewed-by: DJ Delorie <dj@redhat.com>	2025-02-12 16:31:57 -03:00

1 2 3 4 5 ...

42160 Commits