jimczi commented on PR #16656:
URL: https://github.com/apache/lucene/pull/16656#issuecomment-5748849380

   Thanks for this. 
   The conclusion that the time is lost waiting rather than computing looks 
right.
   I also think the structure is right: queue the candidates, submit them 
together, then score them. 
   My point is that Lucene already has a hook for this: 
`IndexInput#prefetch(offset, length)`, “start this read, I’ll consume it 
later.” So the three phases don’t need a new SPI.
   
   `RescoreTopNQuery` never calls it today. That means the “stock mmap” result 
is really “mmap at queue depth 1.” It’s slow because nothing tells the kernel 
about the next candidate, not because it’s mmap.
   
   I’d also revisit the two reasons for routing `.vec` around the page cache. 
Lucene already applies `MADV_RANDOM` to that file through 
`DataAccessHint.RANDOM` in `Lucene99FlatVectorsReader`, and it addresses both:
   
   - **Read amplification.** `VM_RAND_READ` disables readahead in the mmap 
fault path. Both `do_sync_mmap_readahead` and `do_async_mmap_readahead` return 
early, so a fault reads only the pages it needs. That makes the 27.6 KB read 
per 4 KB vector surprising. I’d double-check that `--enable-native-access` was 
enabled for that run. Without it, `madvise` never runs and 
`MemorySegmentIndexInput#prefetch` always returns false.
   
   - **Eviction of the graph and codes.** The mmap fault path doesn’t mark 
folios as accessed. A page faulted once under a random mapping goes onto the 
inactive list and gets reclaimed first. The graph and quantized codes are 
touched on every query, so they stay active. As long as the rerank working set 
is manageable, the page cache already acts like the bounded rerank buffer 
you’re building by hand. It would also be worth testing with MGLRU enabled and 
disabled. MGLRU marks mapped folios active on fault, and it’s disabled by 
default on 6.1, so your current runs use classic LRU.
   
   This is why mmap + `MADV_RANDOM` + batched `WILLNEED` is hard to beat here. 
There’s already a lot of relevant work:
   
   - **#16044** started from the same observation: mmap page-fault storms in 
cgroup-limited containers with indices larger than RAM. It proposed a native 
`pread` Directory and grew into a broad study across three platforms, four 
cache regimes, 1–16 threads, mmap `NORMAL`, `MADV_RANDOM`, batched `WILLNEED`, 
FFI `pread`, `FileChannel`, and `pread` + `O_DIRECT`.
   
     Batched prefetch won every memory-pressured case. The biggest margin was 
exactly where you are: one thread, cold, and file ≫ RAM.
   
     | 16 KB random reads, cold, file > RAM, NVMe (ops/ms) | T01 | T08 | T16 |
     |---|---|---|---|
     | `pread` | 0.58 | 4.02 | 4.65 |
     | mmap, no prefetch (`MADV_RANDOM`) | 0.15 | 1.10 | 1.94 |
     | **mmap + batched prefetch** | **4.19** | 5.62 | 5.98 |
   
     That T01 result is around 67k IOPS / 1.25 GB/s from one thread, which is 
what `fio` gets from that device at iodepth 16. It’s a different box from your 
g6, so it’s only indicative, but it’s the same order as your 67.3k without a 
ring or native dependency.
   
   - **#16279** is the JMH harness behind it — basically `fio` in Java over 
Lucene’s store primitives. It’s the easiest way to compare prefetch with your 
O_DIRECT results on the same hardware without touching the search path.
   
   - **#16145** and **#14156** cover the weak spot in mmap prefetch. The 
power-of-two backoff in `MemorySegmentIndexInput#prefetch` suppresses `madvise` 
when the index barely fits in RAM, exactly when it matters most. That’s a live 
tuning issue in one method.
   
   Here’s how we do it in Elasticsearch using APIs Lucene already has. The ring 
sits outside the leaf loop, so a candidate in segment N+1 can be prefetched 
while segment N is still being scored. The window isn’t tied to one input, so 
there’s no per-file ceiling:
   
   ```java
   PrefetchRing ring = new PrefetchRing(WINDOW);   // WINDOW = 100 for us; doc 
ids, never vectors
   
   for (LeafReaderContext leaf : reader.leaves()) {
     FloatVectorValues values = leaf.reader().getFloatVectorValues(field);
     Scorer inner = weight.scorer(leaf);
     if (values == null || inner == null) continue;
   
     VectorScorer rescorer = values.rescorer(queryVector);      // full 
precision
     KnnVectorValues.DocIndexIterator vectorIter = values.iterator();
     DocIdSetIterator conj =
         ConjunctionUtils.intersectIterators(List.of(vectorIter, 
inner.iterator()));
   
     for (int doc = conj.nextDoc(); doc != NO_MORE_DOCS; doc = conj.nextDoc()) {
       values.prefetch(vectorIter.index());                     // fire and 
forget
       if (ring.isFull()) {
         // oldest entry was prefetched WINDOW candidates ago; scores via 
VectorScorer.Bulk
         ring.advance(scoreOldest(ring, buffer, results));
       }
       ring.append(doc, leaf.docBase, rescorer);
     }
   }
   while (ring.size() > 0) ring.advance(scoreOldest(ring, buffer, results));
   ```
   
   Heap stays flat because the ring holds `WINDOW` doc IDs instead of count × 
dim floats across a barrier. Outstanding submissions from one thread provide 
the depth. There’s no read pool and no context-switch cost.
   
   The current limitation is that `values.prefetch(ord)` doesn’t work. 
`KnnVectorValues#prefetch` only accepts an array and returns early below two 
ords. Also, `Lucene104ScalarQuantizedVectorsReader.ScalarQuantizedVectorValues` 
(what `getFloatVectorValues()` returns for a quantized field) forwards 
`vectorValue` and `rescorer`, but not prefetching. So it silently does nothing 
on the BBQ rerank path.
   
   We hit the same issue in Elasticsearch and fixed it in our wrapper. The 
upstream change is small: add `prefetch(int ord)`, make the array version loop 
over it, and forward it from the wrapper like the other methods. I’m happy to 
open that PR so you have something to build on.
   
   This doesn’t rule out O_DIRECT. It should also be able to rescore 
efficiently, but it doesn’t need the new SPI either. The input can submit 
asynchronously during `prefetch` and reap during `readFloats`. Inputs from the 
same directory can share a ring. Submissions can also accumulate until the 
first dependent read forces a flush, which is what `VectorBatch#execute` 
expresses.
   
   The query loop stays the same and the store chooses the mechanism.
   
   Would you be up for measuring the prefetch path before adding the API, on 
the same box and at the same operating point? I still expect mmap to be better, 
but if O_DIRECT clearly wins, that’s a real result. The next step would be your 
Directory behind `prefetch`, not a new interface.


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to