costin commented on PR #16623: URL: https://github.com/apache/lucene/pull/16623#issuecomment-6020646300
Grazie @salvatore-campagna for the in-depth review, I've addressed the comments inline. Below the updated benchmark ### Benchmark AMD EPYC 7R32 (c5a.2xlarge), JDK 25. 2 forks, 4 iters x 2s. Baseline is the merge base with the same benchmark classes. #### DenseLiveDocs: | deletionRate | size | before (us) | after (us) | speedup | |---|---|---:|---:|---:| | 0.1% | 100k | 50.7 | 1.53 | 33.1x | | 0.1% | 1M | 481.1 | 15.4 | 31.3x | | 1% | 100k | 58.8 | 1.53 | 38.3x | | 1% | 1M | 610.1 | 15.5 | 39.4x | | 10% | 100k | 124.1 | 1.54 | 80.7x | | 10% | 1M | 1244.2 | 15.6 | 79.8x | #### SparseLiveDocs: | deletionRate | size | before (us) | after (us) | speedup | |---|---|---:|---:|---:| | 0.1% | 100k | 77.2 | 1.27 | 60.9x | | 0.1% | 1M | 722.0 | 12.7 | 56.8x | | 1% | 100k | 112.5 | 2.73 | 41.2x | | 1% | 1M | 1142.0 | 27.6 | 41.4x | | 10% | 100k | 220.6 | 4.32 | 51.0x | | 10% | 1M | 2173.8 | 43.5 | 50.0x | #### Generic Bits (unchanged, within noise): | deletionRate | size | before (us) | after (us) | speedup | |---|---|---:|---:|---:| | 0.1% | 100k | 52.1 | 56.6 | 0.92x | | 0.1% | 1M | 479.6 | 481.1 | 1.00x | | 1% | 100k | 61.0 | 56.2 | 1.09x | | 1% | 1M | 552.4 | 550.7 | 1.00x | | 10% | 100k | 122.6 | 123.7 | 0.99x | | 10% | 1M | 1225.6 | 1223.6 | 1.00x | #### SoftDeletesReaderBenchmark | hardDeleteRate | LiveDocs | softDeleteRate | numDocs | before (us) | after (us) | speedup | |---|---|---|---|---:|---:|---:| | 0.5% | sparse | 1% | 200k | 446.3 | 17.4 | 25.7x | | 0.5% | sparse | 5% | 200k | 494.8 | 67.2 | 7.4x | | 5% | dense | 1% | 200k | 149.6 | 15.6 | 9.6x | | 5% | dense | 5% | 200k | 383.0 | 63.2 | 6.1x | | 0.5% | sparse | 1% | 1M | 2043.9 | 89.7 | 22.8x | | 0.5% | sparse | 5% | 1M | 1879.9 | 335.9 | 5.6x | | 5% | dense | 1% | 1M | 1192.2 | 78.6 | 15.2x | | 5% | dense | 5% | 1M | 969.4 | 315.7 | 3.1x | Dense is a bit slower than the earlier clone() version (about 15 vs 10 us at 1M), but it gets the length right, and it is still 30-80x faster than main. Sparse is faster than before at higher deletion rates thanks to the word-wise applyMask from #16593. In the reader benchmark the gain shrinks as the soft delete rate grows, since applying the soft deletes takes a bigger share of the wrap cost however, while the 1M baseline numbers are noisy (±150-960 us), every case improves. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
