On Wed, Aug 12, 2026 at 01:47:56PM +0000, Morten Brørup wrote:
> Added a new high-performance lock-free "pile", using the Stack API.
> The pile behaves roughly like a stack, but is not strictly LIFO.
> 
> The pile is optimized for pushing/popping bulks of objects, which
> it does significantly faster than the lock-free stack.
> 
> Pushing/popping a number of objects not divisible by the compile time
> configurable bulk size is handled gracefully, but not as fast as
> complete bulks.
> 
> Performance examples, stack_pile_perf_autotest vs. stack_lf_autotest:
> 
> On a single core, pushing/popping 1 or 8 objects is similar speed.
> On a single core, pushing/popping 32 objects is 2x faster.
> On a single core, pushing/popping 512 objects is 10x faster.
> 
> On four cores, pushing/popping 1, 8 or 32 objects is slightly faster.
> On four cores, pushing/popping 512 objects is 4x faster.
> 
> Signed-off-by: Morten Brørup <[email protected]>
> ---
>  app/test/test_stack.c               |  71 +++++-
>  app/test/test_stack_perf.c          |  15 +-
>  config/rte_config.h                 |   3 +
>  doc/guides/prog_guide/stack_lib.rst |  67 +++++-
>  lib/mempool/rte_mempool.h           |   2 +-
>  lib/stack/meson.build               |   3 +-
>  lib/stack/rte_stack.c               |  18 +-
>  lib/stack/rte_stack.h               |  79 +++++++
>  lib/stack/rte_stack_lf.h            |   1 +
>  lib/stack/rte_stack_pile.c          |  35 +++
>  lib/stack/rte_stack_pile.h          | 334 ++++++++++++++++++++++++++++
>  11 files changed, 609 insertions(+), 19 deletions(-)
>  create mode 100644 lib/stack/rte_stack_pile.c
>  create mode 100644 lib/stack/rte_stack_pile.h
> 
Looking at this a little closer, and thinking some more, here are some of
my further thoughts/ideas on this:

* For most cases using a mempool, I can't see having non-strict LIFO
  behaviour being an issue, and there is nothing in the mempool API that
  makes any ordering guarantees about what buffers get given by get/put,
  and in fact we can't make any guarantees because of the fact of multiple
  cores doing allocs and frees. Therefore, for the mempool driver, I
  believe one pre-emptible implementation is enough, so therefore the pile
  mempool driver should just replace the current LF one.
* For apps which may want to use the stack structs directly, not through a
  mempool, I can see that having defined ordering behaviour may be
  beneficial. However, if multiple cores are involved, then we can never
  guarantee ordering, I believe, so I'm not sure its worth trying to
  enforce strict LIFO for such cases. [If you need the same elements back
  in the correct order from a core, then use a regular stack without sharing
  it].
* IF we decide that we really, really want LIFO ordering across multiple
  cores - despite the likely random ordering of allocs/frees between those
  cores, I still think that this implementation should replace the LF
  stack. If we reverse the order of elements on enqueue (or dequeue) then
  we should be closer to correct LIFO ordering - and fully lifo if allocs
  and frees are based on multiples of the burst size.

I suppose for me the main question to be resolved is - do we have scenarios
where we a) have multi-core operation on the stacks and b) absolutely must
have strict LIFO ordering? In the absense of that, I'd very much be in
favour of replacing the existing LF implementation completely with this
one.

/Bruce

Reply via email to