kaivalnp commented on PR #16706: URL: https://github.com/apache/lucene/pull/16706#issuecomment-5894838399
Thanks @Pranshu-S, the first few rows are too extreme (up to 65536 unique vectors across 1M docs) -- I'd focus on the last two for a more realistic measure. The percentage saving in "hot" memory would be higher with quantization: a 1-bit quantized vector of 1024 dimensions takes up 1024 bits + the HNSW neighbor list takes up something of the same order (`maxConn` integer node IDs stored as [group vints](https://github.com/apache/lucene/blob/020802ff07053f50ed215244a90f3696cef498a3/lucene/core/src/java/org/apache/lucene/codecs/lucene99/Lucene99HnswVectorsReader.java#L604-L607)). The savings from using fewer bits for each entry of `fieldToGroupOrd` is 8 bits per doc for a segment with 1M docs, which is still <1% It also [adds a few cycles](https://github.com/apache/lucene/blob/020802ff07053f50ed215244a90f3696cef498a3/lucene/core/src/java/org/apache/lucene/util/packed/DirectReader.java#L320-L322) on the hot path compared to the [32-bit lookup](https://github.com/apache/lucene/blob/020802ff07053f50ed215244a90f3696cef498a3/lucene/core/src/java/org/apache/lucene/util/packed/DirectReader.java#L381) -- could you also share search latency from your benchmark? -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
