Pranshu-S commented on PR #16706:
URL: https://github.com/apache/lucene/pull/16706#issuecomment-5896330909

   Here's search latency with the dedup HNSW format on Cohere v3 1M, where the 
ordinals pack at 20 bits. That width hits the unaligned `DirectPackedReader20` 
path you linked, and every scored vector goes through one 
`fieldOrdToGroupOrd.get`.
   
   | filterStrategy | | recall | visited | latency (ms) | netCPU (ms) |
   |---|---|---:|---:|---:|---:|
   | index-time-filter | baseline | 0.988 | 8131 | 1.885 | 1.875 |
   | index-time-filter | candidate | 0.988 | 8135 | 1.873 | 1.863 |
   | query-time-pre-filter | baseline | 0.973 | 16776 | 4.533 | 4.486 |
   | query-time-pre-filter | candidate | 0.973 | 16788 | 4.459 | 4.420 |
   
   (1024d float32, 10K queries, `topK=100`, `fanout=100`, `maxConn=64`, 
`beamWidth=250`, force-merged to 1 segment, `filterSelectivity=0.5`, medians of 
3 interleaved search-only runs)
   
   No measurable regression, even with query-time pre-filtering, which does ~2x 
the lookups per query. At 1024d the extra shifts/masks are likely hidden behind 
the vector scoring and memory access, and the smaller map (2.5 MB vs 4 MB per 
1M docs) may be offsetting them a bit. 
   
   Most likely, the cost would matter most for low-dim or quantized vectors, 
where scoring is cheaper?


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to