goankur commented on PR #16656:
URL: https://github.com/apache/lucene/pull/16656#issuecomment-5806091673
> @jimczi
>
> A more natural API for prefetch would be `int prefetch(int ord, int
size)` to prefetch contiguous
> vectors (size=1 for a single one) and then the multiple ords one just
call the contiguous prefetch
> N times.
Done, and opened as #16705. `prefetch(int ord, int count)` is the primitive;
the array form
coalesces ordinals that are consecutive in the array and calls it once per
run. Two notes on the shape:
- I named the second parameter `count` rather than `size`, since `size`
collides with `size()` meaning "number of vectors in the field."
- I read the `int` return as "the number of vectors a prefetch was
actually issued for, 0 if none," mirroring `IndexInput#prefetch`'s boolean. Let
me know if you meant something else.
The dedup codecs in sandbox needed their own override: they map field ords
to group ords, so a contiguous run of field ords is not contiguous on disk.
Without it they would inherit the no-op default, the same silent break as the
quantized wrapper.
Closing this one. The `VectorBatch`/`ParallelVectorReadable` SPI is not
justified — prefetch on the existing API wins in both the single-stream and
concurrent regimes, as measured in #16705. Executor-parallel prefetch issuance
will follow as a separate PR on top of it.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]