kaivalnp commented on PR #16706:
URL: https://github.com/apache/lucene/pull/16706#issuecomment-5894838399

   Thanks @Pranshu-S, the first few rows are too extreme (up to 65536 unique 
vectors across 1M docs) -- I'd focus on the last two for a more realistic 
measure.
   
   The percentage saving in "hot" memory would be higher with quantization: a 
1-bit quantized vector of 1024 dimensions takes up 1024 bits + the HNSW 
neighbor list takes up something of the same order (`maxConn` integer node IDs 
stored as [group 
vints](https://github.com/apache/lucene/blob/020802ff07053f50ed215244a90f3696cef498a3/lucene/core/src/java/org/apache/lucene/codecs/lucene99/Lucene99HnswVectorsReader.java#L604-L607)).
 The savings from using fewer bits for each entry of `fieldToGroupOrd` is 8 
bits per doc for a segment with 1M docs, which is still <1%
   
   It also [adds a few 
cycles](https://github.com/apache/lucene/blob/020802ff07053f50ed215244a90f3696cef498a3/lucene/core/src/java/org/apache/lucene/util/packed/DirectReader.java#L320-L322)
 on the hot path compared to the [32-bit 
lookup](https://github.com/apache/lucene/blob/020802ff07053f50ed215244a90f3696cef498a3/lucene/core/src/java/org/apache/lucene/util/packed/DirectReader.java#L381)
 -- could you also share search latency from your benchmark?


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to