Pranshu-S commented on PR #16710: URL: https://github.com/apache/lucene/pull/16710#issuecomment-5933595912
Sharing the benchmark runs - baseline here is only raw float16 and candidate here is quantized float16 ndoc: 10000, nquery1000 ``` ===================================================================== COMPARISON (candidate vs baseline; delta/pct = candidate - baseline) metrics are mean±stddev over iterations (baseline n=5, candidate n=5) ===================================================================== metric baseline(n=5) candidate(n=5) delta pct ----------------- --------------- --------------- -------- ------ recall 0.998±0.000 0.977±0.000 -0.021 -2.1% latency(ms) 3.409±0.050 0.608±0.011 -2.801 -82.2% netCPU 3.403±0.050 0.602±0.011 -2.801 -82.3% avgCpuCount 0.998±0.000 0.990±0.000 -0.008 -0.8% nDoc 10000 10000 searchType KNN KNN topK 10 10 fanout 100 100 resultSimilarity N/A N/A decay N/A N/A resultCount 10.000 10.000 maxConn 32 32 beamWidth 200 200 quantized 7 bits 7 bits visited 2247.600±4.159 2250.600±4.879 +3.000 +0.1% index(s) 4.140±0.106 5.044±0.123 +0.904 +21.8% index_docs/s 2416.156±62.095 1984.008±50.093 -432.148 -17.9% merge(s) 0.000±0.000 0.000±0.000 +0.000 force_merge(s) 39.856±0.348 6.816±0.111 -33.040 -82.9% num_segments 1.000±0.000 1.000±0.000 +0.000 +0.0% index_size(MB) 20.188±0.004 30.102±0.004 +9.914 +49.1% filterStrategy null null filterSelectivity N/A N/A overSample 1.000 1.000 vec_disk(MB) 29.449±0.000 29.449±0.000 +0.000 +0.0% vec_RAM(MB) 9.918±0.000 9.918±0.000 +0.000 +0.0% bp-reorder false false indexType HNSW HNSW rerank no no ``` ndoc100000, nquery10000 ``` ===================================================================== COMPARISON (candidate vs baseline; delta/pct = candidate - baseline) metrics are mean±stddev over iterations (baseline n=5, candidate n=5) ===================================================================== metric baseline(n=5) candidate(n=5) delta pct ----------------- --------------- --------------- -------- ------ recall 0.993±0.000 0.967±0.000 -0.026 -2.6% latency(ms) 5.045±0.038 1.170±0.014 -3.875 -76.8% netCPU 5.044±0.038 1.169±0.014 -3.875 -76.8% avgCpuCount 1.000±0.000 0.999±0.000 -0.001 -0.1% nDoc 100000 100000 searchType KNN KNN topK 10 10 fanout 100 100 resultSimilarity N/A N/A decay N/A N/A resultCount 10.000 10.000 maxConn 32 32 beamWidth 200 200 quantized 7 bits 7 bits visited 2984.400±2.966 2983.200±1.304 -1.200 -0.0% index(s) 74.166±1.458 77.218±1.173 +3.052 +4.1% index_docs/s 1348.744±26.648 1295.276±19.401 -53.468 -4.0% merge(s) 647.172±1.826 82.688±0.860 -564.484 -87.2% force_merge(s) 0.000±0.000 36.008±0.238 +36.008 num_segments 1.000±0.000 1.000±0.000 +0.000 +0.0% index_size(MB) 204.712±0.008 303.892±0.004 +99.180 +48.4% filterStrategy null null filterSelectivity N/A N/A overSample 1.000 1.000 vec_disk(MB) 294.495±0.000 294.495±0.000 +0.000 +0.0% vec_RAM(MB) 99.182±0.000 99.182±0.000 +0.000 +0.0% bp-reorder false false indexType HNSW HNSW rerank no no ``` -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
