On Fri, 31 Jul 2026 16:53:45 +0200
Morten Brørup <[email protected]> wrote:

> +TO: x86 maintainers, ThunderX maintainers
> 
> > From: Stephen Hemminger [mailto:[email protected]]
> > Sent: Wednesday, 29 July 2026 19.54
> > 
> > The generic and C11 lock-free stack implementations differ only in
> > memory ordering. The generic version uses a full barrier where its
> > own comments state an acquire fence is sufficient, and seq_cst for
> > all length counter operations.
> > 
> > Only x86 and ThunderX still used the generic version. On x86 the
> > switch removes a locked add per CAS attempt in push and pop; TSO
> > provides the acquire semantics. On ThunderX the pop fence weakens
> > from dmb ish to dmb ishld and the push fence goes away. Unlike the
> > ring, no platform selected the generic stack for measured
> > performance reasons.
> > 
> > Remove it and use the C11 implementation everywhere.  
> 
> The lack of measured performance difference documentation is not a valid 
> reason to remove the generic version!
> 
> It would be reasonable to assume that x86 (and ThunderX) use the generic 
> version for non-insignificant performance reasons.
> 
> If there is no performance difference, I agree with this patch. Otherwise not.
> This could be verified by providing the missing measurements.
> 

The reason is that smp_mb's need to go away.
The other option is revise to replace with thread fence, at which
point it becomes the same as generic code.

Reply via email to