jimczi commented on PR #16656:
URL: https://github.com/apache/lucene/pull/16656#issuecomment-5748849380
Thanks for this.
The conclusion that the time is lost waiting rather than computing looks
right.
I also think the structure is right: queue the candidates, submit them
together, then score them.
My point is that Lucene already has a hook for this:
`IndexInput#prefetch(offset, length)`, “start this read, I’ll consume it
later.” So the three phases don’t need a new SPI.
`RescoreTopNQuery` never calls it today. That means the “stock mmap” result
is really “mmap at queue depth 1.” It’s slow because nothing tells the kernel
about the next candidate, not because it’s mmap.
I’d also revisit the two reasons for routing `.vec` around the page cache.
Lucene already applies `MADV_RANDOM` to that file through
`DataAccessHint.RANDOM` in `Lucene99FlatVectorsReader`, and it addresses both:
- **Read amplification.** `VM_RAND_READ` disables readahead in the mmap
fault path. Both `do_sync_mmap_readahead` and `do_async_mmap_readahead` return
early, so a fault reads only the pages it needs. That makes the 27.6 KB read
per 4 KB vector surprising. I’d double-check that `--enable-native-access` was
enabled for that run. Without it, `madvise` never runs and
`MemorySegmentIndexInput#prefetch` always returns false.
- **Eviction of the graph and codes.** The mmap fault path doesn’t mark
folios as accessed. A page faulted once under a random mapping goes onto the
inactive list and gets reclaimed first. The graph and quantized codes are
touched on every query, so they stay active. As long as the rerank working set
is manageable, the page cache already acts like the bounded rerank buffer
you’re building by hand. It would also be worth testing with MGLRU enabled and
disabled. MGLRU marks mapped folios active on fault, and it’s disabled by
default on 6.1, so your current runs use classic LRU.
This is why mmap + `MADV_RANDOM` + batched `WILLNEED` is hard to beat here.
There’s already a lot of relevant work:
- **#16044** started from the same observation: mmap page-fault storms in
cgroup-limited containers with indices larger than RAM. It proposed a native
`pread` Directory and grew into a broad study across three platforms, four
cache regimes, 1–16 threads, mmap `NORMAL`, `MADV_RANDOM`, batched `WILLNEED`,
FFI `pread`, `FileChannel`, and `pread` + `O_DIRECT`.
Batched prefetch won every memory-pressured case. The biggest margin was
exactly where you are: one thread, cold, and file ≫ RAM.
| 16 KB random reads, cold, file > RAM, NVMe (ops/ms) | T01 | T08 | T16 |
|---|---|---|---|
| `pread` | 0.58 | 4.02 | 4.65 |
| mmap, no prefetch (`MADV_RANDOM`) | 0.15 | 1.10 | 1.94 |
| **mmap + batched prefetch** | **4.19** | 5.62 | 5.98 |
That T01 result is around 67k IOPS / 1.25 GB/s from one thread, which is
what `fio` gets from that device at iodepth 16. It’s a different box from your
g6, so it’s only indicative, but it’s the same order as your 67.3k without a
ring or native dependency.
- **#16279** is the JMH harness behind it — basically `fio` in Java over
Lucene’s store primitives. It’s the easiest way to compare prefetch with your
O_DIRECT results on the same hardware without touching the search path.
- **#16145** and **#14156** cover the weak spot in mmap prefetch. The
power-of-two backoff in `MemorySegmentIndexInput#prefetch` suppresses `madvise`
when the index barely fits in RAM, exactly when it matters most. That’s a live
tuning issue in one method.
Here’s how we do it in Elasticsearch using APIs Lucene already has. The ring
sits outside the leaf loop, so a candidate in segment N+1 can be prefetched
while segment N is still being scored. The window isn’t tied to one input, so
there’s no per-file ceiling:
```java
PrefetchRing ring = new PrefetchRing(WINDOW); // WINDOW = 100 for us; doc
ids, never vectors
for (LeafReaderContext leaf : reader.leaves()) {
FloatVectorValues values = leaf.reader().getFloatVectorValues(field);
Scorer inner = weight.scorer(leaf);
if (values == null || inner == null) continue;
VectorScorer rescorer = values.rescorer(queryVector); // full
precision
KnnVectorValues.DocIndexIterator vectorIter = values.iterator();
DocIdSetIterator conj =
ConjunctionUtils.intersectIterators(List.of(vectorIter,
inner.iterator()));
for (int doc = conj.nextDoc(); doc != NO_MORE_DOCS; doc = conj.nextDoc()) {
values.prefetch(vectorIter.index()); // fire and
forget
if (ring.isFull()) {
// oldest entry was prefetched WINDOW candidates ago; scores via
VectorScorer.Bulk
ring.advance(scoreOldest(ring, buffer, results));
}
ring.append(doc, leaf.docBase, rescorer);
}
}
while (ring.size() > 0) ring.advance(scoreOldest(ring, buffer, results));
```
Heap stays flat because the ring holds `WINDOW` doc IDs instead of count ×
dim floats across a barrier. Outstanding submissions from one thread provide
the depth. There’s no read pool and no context-switch cost.
The current limitation is that `values.prefetch(ord)` doesn’t work.
`KnnVectorValues#prefetch` only accepts an array and returns early below two
ords. Also, `Lucene104ScalarQuantizedVectorsReader.ScalarQuantizedVectorValues`
(what `getFloatVectorValues()` returns for a quantized field) forwards
`vectorValue` and `rescorer`, but not prefetching. So it silently does nothing
on the BBQ rerank path.
We hit the same issue in Elasticsearch and fixed it in our wrapper. The
upstream change is small: add `prefetch(int ord)`, make the array version loop
over it, and forward it from the wrapper like the other methods. I’m happy to
open that PR so you have something to build on.
This doesn’t rule out O_DIRECT. It should also be able to rescore
efficiently, but it doesn’t need the new SPI either. The input can submit
asynchronously during `prefetch` and reap during `readFloats`. Inputs from the
same directory can share a ring. Submissions can also accumulate until the
first dependent read forces a flush, which is what `VectorBatch#execute`
expresses.
The query loop stays the same and the store chooses the mechanism.
Would you be up for measuring the prefetch path before adding the API, on
the same box and at the same operating point? I still expect mmap to be better,
but if O_DIRECT clearly wins, that’s a real result. The next step would be your
Directory behind `prefetch`, not a new interface.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]