https://gcc.gnu.org/bugzilla/show_bug.cgi?id=117487
--- Comment #2 from GCC Commits <cvs-commit at gcc dot gnu.org> --- The releases/gcc-16 branch has been updated by Michael Meissner <[email protected]>: https://gcc.gnu.org/g:8828492b3eb1f847791a38963e57a5afd77366a0 commit r16-9605-g8828492b3eb1f847791a38963e57a5afd77366a0 Author: Michael Meissner <[email protected]> Date: Thu Aug 27 12:22:59 2026 -0400 Backport PR target/117487 I was answering an email from a co-worker and I pointed him to work I had done for the Power8 era that optimizes the 32-bit float math library in Glibc. In doing so, I discovered with the Power9 and later computers, this optimization is no longer taking place. The glibc 32-bit floating point math functions have code that looks like: union u { float f; uint32_t u32; }; float math_foo (float x, unsigned int mask) { union u arg; float x2; arg.f = x; arg.u32 &= mask; x2 = arg.f; /* ... */ } On power8 with the optimization it generates: xscvdpspn 0,1 sldi 9,4,32 mtvsrd 32,9 xxland 1,0,32 xscvspdpn 1,1 I.e., it converts the SFmode to the memory format (instead of the DFmode that is used within the register), converts the mask so that it is in the vector register in the upper 32-bits, and does a XXLAND (i.e. there is only one direct move from GPR to vector register). Then after doing this, it converts the upper 32-bits back to DFmode. If the XSCVSPDN instruction took the value in the normal 32-bit scalar in a vector register, we wouldn't have needed the SLDI of the mask. On power9/power10/power11 it currently generates: xscvdpspn 0,1 mfvsrwz 2,0 and 2,2,4 mtvsrws 1,2 xscvspdpn 1,1 blr I.e convert to SFmode representation, move the value to a GPR, do an AND operation, move the 32-bit value with a splat, and then convert it back to DFmode format. With this patch, it now generates: xscvdpspn 0,1 mtvsrwz 32,2 xxland 32,0,32 xxspltw 1,32,1 xscvspdpn 1,1 blr I.e. convert to SFmode representation, move the mask to the vector register, do the operation using XXLAND. Splat the value to get the value in the correct location, and then convert back to DFmode. This patch was back ported from the changes to the master branch on 2026-08-07. 2026-08-27 Michael Meissner <[email protected]> gcc/ PR target/117487 * config/rs6000/vsx.md (SFmode logical peephoole): Update comments in the original code that supports power8. (SFBOOL2_*): New constants. (power9/power10 define_peephol2): Add a new define_peephole2 to optimize float and logical operations on power9/power10/power11 similar to the optimiztion that is done on power8. gcc/testsuite/ PR target/117487 * gcc.target/powerpc/pr117487.c: New test.
