costin commented on PR #16623:
URL: https://github.com/apache/lucene/pull/16623#issuecomment-6020646300

   Grazie @salvatore-campagna for the in-depth review, I've addressed the 
comments inline.
   
   Below the updated benchmark
   
   ### Benchmark
   
   AMD EPYC 7R32 (c5a.2xlarge), JDK 25. 2 forks, 4 iters x 2s. Baseline is the 
merge base with the same benchmark classes.
   
   #### DenseLiveDocs:
   
   | deletionRate | size | before (us) | after (us) | speedup |
   |---|---|---:|---:|---:|
   | 0.1% | 100k | 50.7 | 1.53 | 33.1x |
   | 0.1% | 1M | 481.1 | 15.4 | 31.3x |
   | 1% | 100k | 58.8 | 1.53 | 38.3x |
   | 1% | 1M | 610.1 | 15.5 | 39.4x |
   | 10% | 100k | 124.1 | 1.54 | 80.7x |
   | 10% | 1M | 1244.2 | 15.6 | 79.8x |
   
   #### SparseLiveDocs:
   
   | deletionRate | size | before (us) | after (us) | speedup |
   |---|---|---:|---:|---:|
   | 0.1% | 100k | 77.2 | 1.27 | 60.9x |
   | 0.1% | 1M | 722.0 | 12.7 | 56.8x |
   | 1% | 100k | 112.5 | 2.73 | 41.2x |
   | 1% | 1M | 1142.0 | 27.6 | 41.4x |
   | 10% | 100k | 220.6 | 4.32 | 51.0x |
   | 10% | 1M | 2173.8 | 43.5 | 50.0x |
   
   #### Generic Bits (unchanged, within noise):
   
   | deletionRate | size | before (us) | after (us) | speedup |
   |---|---|---:|---:|---:|
   | 0.1% | 100k | 52.1 | 56.6 | 0.92x |
   | 0.1% | 1M | 479.6 | 481.1 | 1.00x |
   | 1% | 100k | 61.0 | 56.2 | 1.09x |
   | 1% | 1M | 552.4 | 550.7 | 1.00x |
   | 10% | 100k | 122.6 | 123.7 | 0.99x |
   | 10% | 1M | 1225.6 | 1223.6 | 1.00x |
   
   #### SoftDeletesReaderBenchmark
   
   | hardDeleteRate | LiveDocs | softDeleteRate | numDocs | before (us) | after 
(us) | speedup |
   |---|---|---|---|---:|---:|---:|
   | 0.5% | sparse | 1% | 200k | 446.3 | 17.4 | 25.7x |
   | 0.5% | sparse | 5% | 200k | 494.8 | 67.2 | 7.4x |
   | 5% | dense | 1% | 200k | 149.6 | 15.6 | 9.6x |
   | 5% | dense | 5% | 200k | 383.0 | 63.2 | 6.1x |
   | 0.5% | sparse | 1% | 1M | 2043.9 | 89.7 | 22.8x |
   | 0.5% | sparse | 5% | 1M | 1879.9 | 335.9 | 5.6x |
   | 5% | dense | 1% | 1M | 1192.2 | 78.6 | 15.2x |
   | 5% | dense | 5% | 1M | 969.4 | 315.7 | 3.1x |
   
   Dense is a bit slower than the earlier clone() version (about 15 vs 10 us at 
1M), but it gets the length right, and it is still 30-80x faster than main. 
Sparse is faster than before at higher deletion rates thanks to the word-wise 
applyMask from #16593. 
   In the reader benchmark the gain shrinks as the soft delete rate grows, 
since applying the soft deletes takes a bigger share of the wrap cost however, 
while the 1M baseline numbers are noisy (±150-960 us), every case improves.


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to