diff options
author | H.J. Lu <hjl.tools@gmail.com> | 2018-01-08 08:04:26 -0800 |
---|---|---|
committer | H.J. Lu <hjl.tools@gmail.com> | 2018-01-08 08:04:40 -0800 |
commit | c70e4e9c9efff9df4c847dd7cfd81bae674219ab (patch) | |
tree | 46cbbfb74a8c03e933fc4245c66559def374b1a8 /ChangeLog | |
parent | 579396ee082565ab5f42ff166a264891223b7b82 (diff) | |
download | glibc-c70e4e9c9efff9df4c847dd7cfd81bae674219ab.zip glibc-c70e4e9c9efff9df4c847dd7cfd81bae674219ab.tar.gz glibc-c70e4e9c9efff9df4c847dd7cfd81bae674219ab.tar.bz2 |
x86-64: Add sincosf with vector FMA
Since the x86-64 assembly version of sincosf is higly optimized with
vector instructions, there isn't much room for improvement. However
s_sincosf.c written in C with vector math and intrinsics can be
optimized by GCC with FMA.
On Skylake, bench-sincosf reports performance improvement:
Assembly FMA improvement
max 104.042 101.008 3%
min 9.426 8.586 10%
mean 20.6209 18.2238 13%
* sysdeps/x86_64/fpu/multiarch/Makefile (libm-sysdep_routines):
Add s_sincosf-sse2 and s_sincosf-fma.
(CFLAGS-s_sincosf-fma.c): New.
* sysdeps/x86_64/fpu/multiarch/s_sincosf-fma.c: New file.
* sysdeps/x86_64/fpu/multiarch/s_sincosf-sse2.S: Likewise.
* sysdeps/x86_64/fpu/multiarch/s_sincosf.c: Likewise.
* sysdeps/x86_64/fpu/s_sincosf.S: Don't add alias if
__sincosf is defined.
Diffstat (limited to 'ChangeLog')
-rw-r--r-- | ChangeLog | 11 |
1 files changed, 11 insertions, 0 deletions
@@ -1,3 +1,14 @@ +2018-01-08 H.J. Lu <hongjiu.lu@intel.com> + + * sysdeps/x86_64/fpu/multiarch/Makefile (libm-sysdep_routines): + Add s_sincosf-sse2 and s_sincosf-fma. + (CFLAGS-s_sincosf-fma.c): New. + * sysdeps/x86_64/fpu/multiarch/s_sincosf-fma.c: New file. + * sysdeps/x86_64/fpu/multiarch/s_sincosf-sse2.S: Likewise. + * sysdeps/x86_64/fpu/multiarch/s_sincosf.c: Likewise. + * sysdeps/x86_64/fpu/s_sincosf.S: Don't add alias if + __sincosf is defined. + 2018-01-08 Florian Weimer <fweimer@redhat.com> * nptl/tst-thread-exit-clobber.cc: New file. |