Pranshu-S commented on PR #16706:
URL: https://github.com/apache/lucene/pull/16706#issuecomment-5894146388

   ### Benchmarks Results 
   
   I compared this PR using `luceneutil` with `DedupHnswVectorsFormat`.
   
   **Vectors:** Cohere v3 Wikipedia, 1024d, float32, 1M docs.
   
   **Methodology:**
   - **Unique vectors:** each doc file has 1M docs built from a fixed set of 
unique Cohere vectors (1, 256, 4,096 or 65,536), shuffled. The last rows use 
the real Cohere 1M.
   - **Fields:**
     - "1 field": one vector field.
     - "2 fields": luceneutil's `index-time-filter` at 0.5, which adds a second 
field with the same vectors for 499,739 docs. Both fields share one dedup group.
   - **Indexing:** 8 threads, then force-merge to 1 segment. No HNSW graph is 
built, so file sizes isolate the ordinal-map change.
   
   I also added metrics around actual `.vdd` and `.vdm` file size. Capture the 
saving here:
   
   | unique vectors | bits | fields | baseline (MB) | candidate (MB) | index 
size | `.vdd` saved (bytes) |
   |---:|---:|---:|---:|---:|---:|---:|
   | 1 | 1 | 1 | 7.70 | 4.01 | **-47.9%** | 3,875,000 |
   | 1 | 1 | 2 | 10.45 | 4.91 | **-53.0%** | 5,811,488 |
   | 256 | 8 | 1 | 8.70 | 5.84 | **-32.9%** | 3,000,000 |
   | 256 | 8 | 2 | 11.45 | 7.16 | **-37.5%** | 4,499,217 |
   | 4,096 | 12 | 1 | 23.70 | 21.32 | **-10.0%** | 2,499,999 |
   | 4,096 | 12 | 2 | 26.45 | 22.88 | **-13.5%** | 3,749,346 |
   | 65,535 | 16 | 1 | 263.70 | 261.79 | -0.72% | 2,000,000 |
   | 65,535 | 16 | 2 | 266.45 | 263.58 | -1.08% | 2,999,478 |
   | 999,987 (Cohere 1M) | 20 | 1 | 3,913.90 | 3,912.47 | -0.04% | 1,499,998 |
   | 999,987 (Cohere 1M) | 20 | 2 | 3,916.65 | 3,914.50 | -0.05% | 2,249,602 |
   
   ### Result
   - The savings are around `entries × (32 - bits) / 8`, give or take a few 
bytes of alignment. `.vdm` grows by 4 bytes per field.
   - The gain is large when ordinals dominate the index. With ≤ 4,096 unique 
vectors the index shrinks 10–53%. With mostly unique 1024d vectors it saves 
~1.5 MB per 1M-doc field. Each extra field sharing the group adds 
proportionally more.
   
   cc @kaivalnp 


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to