Pranshu-S commented on PR #16731: URL: https://github.com/apache/lucene/pull/16731#issuecomment-5980010785
Ran some benchmark suits, I was able to find meaningful results only when there close to 70-80% deduplication WITHIN A FIELD which is very unlikely in any dataset. Even in those cases, it would make sense to just run exact KNN on the dedup than build a Hnsw Graph. What might be good next points to explore would be IVF and tiered search similar to Vamana? -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
