https://gcc.gnu.org/bugzilla/show_bug.cgi?id=126028

--- Comment #10 from rguenther at suse dot de <rguenther at suse dot de> ---
On Thu, 27 Aug 2026, Reshma.Roy at amd dot com wrote:

> https://gcc.gnu.org/bugzilla/show_bug.cgi?id=126028
> 
> --- Comment #9 from Reshma Roy <Reshma.Roy at amd dot com> ---
> (In reply to Richard Biener from comment #8)
> > The testcase is now vectorized.  What's missing is more optimal handling of
> > 
> > 
> >   pj_x_0 = p_1[jprev0].position.x;  pj_x_1 = p_1[jprev1].position.x;
> >   pj_y_0 = p_1[jprev0].position.y;  pj_y_1 = p_1[jprev1].position.y;
> > ..
> >   for (;;)
> > ...
> >      pj_x_0 = p_1[jprev0].position.x;  pj_x_1 = p_1[jprev1].position.x;
> >      pj_y_0 = p_1[jprev0].position.y;  pj_y_1 = p_1[jprev1].position.y;
> > 
> > where the three vectors {pj_x_1, pj_x_0}, {pj_y_1, pj_y_0} and {pj_z_1,
> > pj_z_0 }
> > are constructed from scalar loads (because the vector SLP involves two
> > distinct load groups).  It's not clear whether two V2DF loads plus
> > shuffling would be more efficient.  I have a heuristic patch that moves
> > the CTORs to the latch and the preheader instead of constructing from the
> > PHI which removes an IV.
> 
> The code of interest in the test case as well as the namd benchmark is
> vectorized in the latest gcc trunk. With the fix the loop vectorizes, i.e.,
> packed vsubpd/vmulpd/vfmadd132pd instead of the fully scalar *sd codegen. 
> 
> I created a standalone microbenchmark of this loop and there is a runtime
> improvement with this vectorization. But the namd benchmark does not show any
> measurable uplift. The remaining codegen caveat as you mentioned is that 
> {pj_x,
> pj_y, pj_z} are still built from scalar loads + vunpcklpd CTORs rather than 
> two
> V2DF loads plus a shuffle. Following your suggestion I hand-edited the kernel
> to (with the understanding that this is what you intended):
> 
> vmovupd (%rax), %xmm4          
> vmovupd (%rdx), %xmm8          
> vpunpckhqdq %xmm8, %xmm4, %xmm5  
> vpunpcklqdq %xmm8, %xmm4, %xmm4  
> vmovsd 16(%rax), %xmm0
> vmovhpd 16(%rdx), %xmm0, %xmm0   
> 
> This binary was faster than the current generated binary. So, I think that two
> V2DF loads plus shuffling is the better form here, it would be great if that
> heuristic patch landed.

I'll note the heuristic patch does not produce V2DF loads + shuffles
but only duplicates the construction from scalars to the loop
preheader and the latch.  Creating V2DF loads + shuffles is more a
complicated task.

Reply via email to