kaivalnp commented on code in PR #16030:
URL: https://github.com/apache/lucene/pull/16030#discussion_r3980521525


##########
lucene/core/src/java/org/apache/lucene/codecs/lucene104/Lucene104ScalarQuantizedVectorsFormat.java:
##########
@@ -85,19 +89,27 @@
  *   <li><b>vlong</b> the length of the vector data in the .veq file
  *   <li><b>vint</b> the number of vectors
  *   <li><b>vint</b> the wire number for ScalarEncoding
- *   <li><b>[float]</b> the centroid
- *   <li><b>float</b> the centroid square magnitude
+ *   <li><b>[float]</b> the centroid (omitted when the metadata version 
indicates data-blind mode)
+ *   <li><b>float</b> the centroid square magnitude (omitted when the metadata 
version indicates
+ *       data-blind mode)
  *   <li>The sparse vector information, if required, mapping vector ordinal to 
doc ID
  * </ul>
  *
+ * <p>{@code enableCentering} manifests in the version: when true we write 
version 0 and when false

Review Comment:
   Interesting use of the version for this value -- I assume it is to keep the 
change backwards compatible with the `Lucene104` format by **not** storing an 
additional value on disk?
   
   It works because `enableCentering` is a boolean that maps well to the two 
versions, but it would've been tricky if there was more information to add.
   
   I guess this is slightly unconventional, because the writer generally writes 
with only the latest version. I can't think of immediate downsides though, 
except if there's a future minor version that (say) does centering + something 
else, then functions like `isDataBlind` get convoluted -- but this is 
far-fetched.
   
   I'd personally prefer creating a new format (`Lucene106*`) to store the 
value explicitly, but the current state looks clean to me -- so I'll leave this 
up to you.
   
   Perhaps _if_ we create a new format, we should end the metadata with a 
configurable block for future params?



##########
lucene/core/src/java/org/apache/lucene/codecs/lucene104/Lucene104ScalarQuantizedVectorsWriter.java:
##########
@@ -320,39 +346,129 @@ private QuantizedByteVectorValues 
mergedQuantizedVectorValues(
     return new QuantizedFloatVectorValues(vectorValues, quantizer, encoding, 
centroid);
   }
 
+  /**
+   * Returns a view that quantizes a single segment's float vectors against 
{@code centroid} using
+   * this writer's encoding, without consulting any quantized bytes the 
segment may already store.
+   */
+  private QuantizedFloatVectorValues quantizeFromFloats(
+      KnnVectorsReader reader, FieldInfo fieldInfo, float[] centroid) throws 
IOException {
+    OptimizedScalarQuantizer quantizer =
+        new OptimizedScalarQuantizer(fieldInfo.getVectorSimilarityFunction());
+    FloatVectorValues vectorValues =
+        fieldInfo.getVectorEncoding() == VectorEncoding.FLOAT16
+            ? new 
Float16AsFloatVectorValues(reader.getFloat16VectorValues(fieldInfo.name))
+            : reader.getFloatVectorValues(fieldInfo.name);
+    if (fieldInfo.getVectorSimilarityFunction() == COSINE) {
+      vectorValues = new NormalizedFloatVectorValues(vectorValues);
+    }
+    return new QuantizedFloatVectorValues(vectorValues, quantizer, encoding, 
centroid);
+  }
+
   @Override
   public void mergeOneFlatVectorField(FieldInfo fieldInfo, MergeState 
mergeState)
       throws IOException {
-    // Don't need access to the random vectors, we can just use the merged
-    rawVectorDelegate.mergeOneFlatVectorField(fieldInfo, mergeState);
     if (fieldInfo.getVectorEncoding().isFloatingPoint() == false) {
+      rawVectorDelegate.mergeOneFlatVectorField(fieldInfo, mergeState);
       return;
     }
-    final float[] centroid;
+    if (enableCentering) {
+      mergeOneFlatVectorFieldCentered(fieldInfo, mergeState);
+    } else {
+      mergeOneFlatVectorFieldDataBlind(fieldInfo, mergeState);
+    }
+  }
+
+  private void mergeOneFlatVectorFieldCentered(FieldInfo fieldInfo, MergeState 
mergeState)
+      throws IOException {
+    // Don't need access to the random vectors, we can just use the merged
+    rawVectorDelegate.mergeOneFlatVectorField(fieldInfo, mergeState);
     final float[] mergedCentroid = new float[fieldInfo.getVectorDimension()];
     int vectorCount = mergeAndRecalculateCentroids(mergeState, fieldInfo, 
mergedCentroid);
-    centroid = mergedCentroid;
     if (segmentWriteState.infoStream.isEnabled(QUANTIZED_VECTOR_COMPONENT)) {
       segmentWriteState.infoStream.message(
           QUANTIZED_VECTOR_COMPONENT, "Vectors' count:" + vectorCount);
     }
     QuantizedByteVectorValues quantizedVectorValues =
-        mergedQuantizedVectorValues(fieldInfo, mergeState, centroid);
+        mergedQuantizedVectorValues(fieldInfo, mergeState, mergedCentroid);
     long vectorDataOffset = vectorData.alignFilePointer(Float.BYTES);
     DocsWithFieldSet docsWithField = writeVectorData(vectorData, 
quantizedVectorValues);
     long vectorDataLength = vectorData.getFilePointer() - vectorDataOffset;
     float centroidDp =
-        docsWithField.cardinality() > 0 ? VectorUtil.dotProduct(centroid, 
centroid) : 0;
+        docsWithField.cardinality() > 0 ? 
VectorUtil.dotProduct(mergedCentroid, mergedCentroid) : 0;
     writeMeta(
         fieldInfo,
         segmentWriteState.segmentInfo.maxDoc(),
         vectorDataOffset,
         vectorDataLength,
-        centroid,
+        mergedCentroid,
         centroidDp,
         docsWithField);
   }
 
+  private void mergeOneFlatVectorFieldDataBlind(FieldInfo fieldInfo, 
MergeState mergeState)
+      throws IOException {
+    float[] zeroCentroid = new float[fieldInfo.getVectorDimension()];
+    // Build one merged view where, per contributing segment, either its 
existing quantized bytes
+    // are passed through or its float vectors are quantized fresh. Inputs 
already quantized to
+    // {@code encoding} against a zero centroid (data-blind segments) are 
copied directly; they are
+    // never dequantized and re-quantized, which would only add loss. Segments 
with raw floats are
+    // quantized fresh, as their stored bytes live in a different (centered) 
quantization space.
+    List<QuantizedByteVectorValuesSub> subs = new ArrayList<>();
+    for (int i = 0; i < mergeState.knnVectorsReaders.length; i++) {
+      KnnVectorsReader reader = mergeState.knnVectorsReaders[i];
+      if (reader == null) {
+        continue;
+      }
+      QuantizedByteVectorValues values;
+      if (hasRawVectorValues(reader, fieldInfo)) {
+        // Segment stored full-precision floats; quantize them against the 
zero centroid.
+        values = quantizeFromFloats(reader, fieldInfo, zeroCentroid);
+      } else {
+        QuantizedByteVectorValues qvv = getQuantizedVectorValues(reader, 
fieldInfo.name);
+        if (qvv == null || qvv.size() == 0) {
+          continue;
+        }
+        if (qvv.getScalarEncoding() != encoding) {
+          // Re-quantization from raw floats would be required, which is not 
possible when raw
+          // floats were never written.
+          throw new IllegalStateException(

Review Comment:
   Nice!
   
   IMO on similar lines: we should disallow de-quantizing vectors while merging 
for the centered version too (i.e. in `mergeOneFlatVectorFieldCentered`)?
   
   As it stands today, one might be able to create a data-blind index, merge it 
to a centered one, and convert it back to data-blind with another quantization 
level to bypass this check?



##########
lucene/core/src/java/org/apache/lucene/codecs/lucene104/Lucene104ScalarQuantizedVectorsWriter.java:
##########
@@ -456,8 +610,10 @@ static int mergeAndRecalculateCentroids(
       float[] centroid = getCentroid(knnVectorsReader, fieldInfo.name);
       totalVectorCount += vectorCount;
       // If there aren't centroids, or previously clustered with more than one 
cluster
-      // or if there are deleted docs, we must recalculate the centroid
-      if (centroid == null || mergeState.liveDocs[i] != null) {
+      // or if there are deleted docs, we must recalculate the centroid. An 
all-zero centroid
+      // indicates a data-blind segment (no centering was done); it can't be 
combined with the
+      // others, so recompute from the (possibly dequantized) vectors.
+      if (centroid == null || isAllZero(centroid) || mergeState.liveDocs[i] != 
null) {

Review Comment:
   `isAllZero(centroid)` does not necessarily point to a data-blind segment, 
should `isDataBlind` be checked instead?
   
   If we do this, we can also remove `isAllZero` entirely? (the prior usage in 
`mergeOneFlatVectorFieldDataBlind` can be replaced with 
`Arrays.equals(centroid, zeroCentroid)`, which might be faster too)



##########
lucene/core/src/java/org/apache/lucene/codecs/lucene104/OffHeapScalarQuantizedFloatVectorValues.java:
##########
@@ -93,7 +93,7 @@ abstract class OffHeapScalarQuantizedFloatVectorValues 
extends FloatVectorValues
     this.byteBuffer = ByteBuffer.allocate(docPackedLength);
     this.vectorValue = new float[dimension];
     this.byteValue = byteBuffer.array();
-    this.unpackedByteVectorValue = new byte[dimension];
+    this.unpackedByteVectorValue = new 
byte[encoding.getDiscreteDimensions(dimension)];

Review Comment:
   I didn't get why this had to change?
   
   Is it required independently of this PR?



##########
lucene/core/src/java/org/apache/lucene/codecs/lucene104/Lucene104ScalarQuantizedVectorsFormat.java:
##########
@@ -44,6 +44,10 @@
  *       quantized vectors in the index.
  *   <li>Transforming the half-byte quantized query vectors in such a way that 
the comparison with
  *       single bit vectors can be done with bit arithmetic.
+ *   <li>Data blind mode: vectors are quantized without centering and float 
vectors are discarded.

Review Comment:
   I wonder if "disable centering" and "drop float vectors" should be separate 
options? (the former for faster merges, the latter for smaller indexes). No 
strong opinions though.
   
   If we want to return "reconstructed" vectors, can we update the comment to 
include that the float vectors are approximate too? (in addition to the 
distance estimates)



##########
lucene/CHANGES.txt:
##########
@@ -393,6 +393,10 @@ New Features
   translates each term into a SpanTermQuery within a {@link SpanOrQuery}, but 
retains only the most frequent terms so it
   will not overflow the boolean max clause count. (Alan Woodward, Nishant 
Mehta, David Smiley)
 
+* GITHUB#16029: Scalar quantization option to disable centering and writing of 
float vectors. This
+  reduces vector storage costs by 4x or more but also reduces quantization 
accuracy.
+  (Trevor McCulloch)

Review Comment:
   Should we also specify that going data-blind is a one-way door? (i.e. you 
can only merge data-blind segments of the same encoding, and re-encoding to 
another level is not possible)



##########
lucene/core/src/java/org/apache/lucene/codecs/lucene104/Lucene104ScalarQuantizedVectorsWriter.java:
##########
@@ -320,39 +346,129 @@ private QuantizedByteVectorValues 
mergedQuantizedVectorValues(
     return new QuantizedFloatVectorValues(vectorValues, quantizer, encoding, 
centroid);
   }
 
+  /**
+   * Returns a view that quantizes a single segment's float vectors against 
{@code centroid} using
+   * this writer's encoding, without consulting any quantized bytes the 
segment may already store.
+   */
+  private QuantizedFloatVectorValues quantizeFromFloats(
+      KnnVectorsReader reader, FieldInfo fieldInfo, float[] centroid) throws 
IOException {
+    OptimizedScalarQuantizer quantizer =
+        new OptimizedScalarQuantizer(fieldInfo.getVectorSimilarityFunction());
+    FloatVectorValues vectorValues =
+        fieldInfo.getVectorEncoding() == VectorEncoding.FLOAT16
+            ? new 
Float16AsFloatVectorValues(reader.getFloat16VectorValues(fieldInfo.name))
+            : reader.getFloatVectorValues(fieldInfo.name);
+    if (fieldInfo.getVectorSimilarityFunction() == COSINE) {
+      vectorValues = new NormalizedFloatVectorValues(vectorValues);
+    }
+    return new QuantizedFloatVectorValues(vectorValues, quantizer, encoding, 
centroid);
+  }
+
   @Override
   public void mergeOneFlatVectorField(FieldInfo fieldInfo, MergeState 
mergeState)
       throws IOException {
-    // Don't need access to the random vectors, we can just use the merged
-    rawVectorDelegate.mergeOneFlatVectorField(fieldInfo, mergeState);
     if (fieldInfo.getVectorEncoding().isFloatingPoint() == false) {
+      rawVectorDelegate.mergeOneFlatVectorField(fieldInfo, mergeState);
       return;
     }
-    final float[] centroid;
+    if (enableCentering) {
+      mergeOneFlatVectorFieldCentered(fieldInfo, mergeState);
+    } else {
+      mergeOneFlatVectorFieldDataBlind(fieldInfo, mergeState);
+    }
+  }
+
+  private void mergeOneFlatVectorFieldCentered(FieldInfo fieldInfo, MergeState 
mergeState)
+      throws IOException {
+    // Don't need access to the random vectors, we can just use the merged
+    rawVectorDelegate.mergeOneFlatVectorField(fieldInfo, mergeState);
     final float[] mergedCentroid = new float[fieldInfo.getVectorDimension()];
     int vectorCount = mergeAndRecalculateCentroids(mergeState, fieldInfo, 
mergedCentroid);
-    centroid = mergedCentroid;
     if (segmentWriteState.infoStream.isEnabled(QUANTIZED_VECTOR_COMPONENT)) {
       segmentWriteState.infoStream.message(
           QUANTIZED_VECTOR_COMPONENT, "Vectors' count:" + vectorCount);
     }
     QuantizedByteVectorValues quantizedVectorValues =
-        mergedQuantizedVectorValues(fieldInfo, mergeState, centroid);
+        mergedQuantizedVectorValues(fieldInfo, mergeState, mergedCentroid);
     long vectorDataOffset = vectorData.alignFilePointer(Float.BYTES);
     DocsWithFieldSet docsWithField = writeVectorData(vectorData, 
quantizedVectorValues);
     long vectorDataLength = vectorData.getFilePointer() - vectorDataOffset;
     float centroidDp =
-        docsWithField.cardinality() > 0 ? VectorUtil.dotProduct(centroid, 
centroid) : 0;
+        docsWithField.cardinality() > 0 ? 
VectorUtil.dotProduct(mergedCentroid, mergedCentroid) : 0;
     writeMeta(
         fieldInfo,
         segmentWriteState.segmentInfo.maxDoc(),
         vectorDataOffset,
         vectorDataLength,
-        centroid,
+        mergedCentroid,
         centroidDp,
         docsWithField);
   }
 
+  private void mergeOneFlatVectorFieldDataBlind(FieldInfo fieldInfo, 
MergeState mergeState)
+      throws IOException {
+    float[] zeroCentroid = new float[fieldInfo.getVectorDimension()];
+    // Build one merged view where, per contributing segment, either its 
existing quantized bytes
+    // are passed through or its float vectors are quantized fresh. Inputs 
already quantized to
+    // {@code encoding} against a zero centroid (data-blind segments) are 
copied directly; they are
+    // never dequantized and re-quantized, which would only add loss. Segments 
with raw floats are
+    // quantized fresh, as their stored bytes live in a different (centered) 
quantization space.
+    List<QuantizedByteVectorValuesSub> subs = new ArrayList<>();
+    for (int i = 0; i < mergeState.knnVectorsReaders.length; i++) {
+      KnnVectorsReader reader = mergeState.knnVectorsReaders[i];
+      if (reader == null) {
+        continue;
+      }
+      QuantizedByteVectorValues values;
+      if (hasRawVectorValues(reader, fieldInfo)) {
+        // Segment stored full-precision floats; quantize them against the 
zero centroid.
+        values = quantizeFromFloats(reader, fieldInfo, zeroCentroid);
+      } else {
+        QuantizedByteVectorValues qvv = getQuantizedVectorValues(reader, 
fieldInfo.name);
+        if (qvv == null || qvv.size() == 0) {
+          continue;
+        }
+        if (qvv.getScalarEncoding() != encoding) {
+          // Re-quantization from raw floats would be required, which is not 
possible when raw
+          // floats were never written.
+          throw new IllegalStateException(
+              "Cannot merge field \""
+                  + fieldInfo.name
+                  + "\" from data-blind segment with encoding "
+                  + qvv.getScalarEncoding()
+                  + " into data-blind format with encoding "
+                  + encoding
+                  + ": re-quantization requires raw float vectors");
+        }
+        float[] centroid = getCentroid(reader, fieldInfo.name);
+        if (centroid != null && isAllZero(centroid)) {
+          // Quantized-only segment whose bytes already match the output 
format (encoding and zero

Review Comment:
   Nice!



##########
lucene/core/src/java/org/apache/lucene/codecs/lucene104/Lucene104ScalarQuantizedVectorsWriter.java:
##########
@@ -740,6 +936,148 @@ long quantizationOverheadBytesUsed() {
     }
   }
 
+  /** In-memory storage for fp32 vectors used in data-blind mode; nothing is 
written to disk. */
+  private static class InMemoryFloatFieldWriter extends 
FlatFieldVectorsWriter<float[]> {

Review Comment:
   This is largely the same as the fp16 class, should we use generics?
   
   One step further: this looks almost identical to the [default `Lucene99` 
flat field 
writer](https://github.com/apache/lucene/blob/372b1db88a5614ba3f0a927ab05e858e5087782e/lucene/core/src/java/org/apache/lucene/codecs/lucene99/Lucene99FlatVectorsWriter.java#L420),
 perhaps we can extract that out into a package-private 
`InMemoryFlatFieldVectorsWriter` class and re-use in both places?



##########
lucene/core/src/java/org/apache/lucene/codecs/lucene104/Lucene104ScalarQuantizedVectorsWriter.java:
##########
@@ -420,6 +536,44 @@ static float[] getCentroid(KnnVectorsReader vectorsReader, 
String fieldName) {
     return null;
   }
 
+  static QuantizedByteVectorValues getQuantizedVectorValues(
+      KnnVectorsReader vectorsReader, String fieldName) throws IOException {
+    vectorsReader = vectorsReader.unwrapReaderForField(fieldName);
+    if (vectorsReader instanceof Lucene104ScalarQuantizedVectorsReader reader) 
{
+      return reader.getQuantizedVectorValues(fieldName);
+    }
+    return null;
+  }
+
+  /**
+   * Returns whether the segment stores full-precision vectors for this field, 
or null when the

Review Comment:
   nit: `null` -> `false`?



##########
lucene/core/src/java/org/apache/lucene/codecs/lucene104/Lucene104ScalarQuantizedVectorsWriter.java:
##########
@@ -320,39 +346,129 @@ private QuantizedByteVectorValues 
mergedQuantizedVectorValues(
     return new QuantizedFloatVectorValues(vectorValues, quantizer, encoding, 
centroid);
   }
 
+  /**
+   * Returns a view that quantizes a single segment's float vectors against 
{@code centroid} using
+   * this writer's encoding, without consulting any quantized bytes the 
segment may already store.
+   */
+  private QuantizedFloatVectorValues quantizeFromFloats(
+      KnnVectorsReader reader, FieldInfo fieldInfo, float[] centroid) throws 
IOException {
+    OptimizedScalarQuantizer quantizer =
+        new OptimizedScalarQuantizer(fieldInfo.getVectorSimilarityFunction());
+    FloatVectorValues vectorValues =
+        fieldInfo.getVectorEncoding() == VectorEncoding.FLOAT16
+            ? new 
Float16AsFloatVectorValues(reader.getFloat16VectorValues(fieldInfo.name))
+            : reader.getFloatVectorValues(fieldInfo.name);
+    if (fieldInfo.getVectorSimilarityFunction() == COSINE) {
+      vectorValues = new NormalizedFloatVectorValues(vectorValues);
+    }
+    return new QuantizedFloatVectorValues(vectorValues, quantizer, encoding, 
centroid);
+  }
+
   @Override
   public void mergeOneFlatVectorField(FieldInfo fieldInfo, MergeState 
mergeState)
       throws IOException {
-    // Don't need access to the random vectors, we can just use the merged
-    rawVectorDelegate.mergeOneFlatVectorField(fieldInfo, mergeState);
     if (fieldInfo.getVectorEncoding().isFloatingPoint() == false) {
+      rawVectorDelegate.mergeOneFlatVectorField(fieldInfo, mergeState);
       return;
     }
-    final float[] centroid;
+    if (enableCentering) {
+      mergeOneFlatVectorFieldCentered(fieldInfo, mergeState);
+    } else {
+      mergeOneFlatVectorFieldDataBlind(fieldInfo, mergeState);
+    }
+  }
+
+  private void mergeOneFlatVectorFieldCentered(FieldInfo fieldInfo, MergeState 
mergeState)
+      throws IOException {
+    // Don't need access to the random vectors, we can just use the merged
+    rawVectorDelegate.mergeOneFlatVectorField(fieldInfo, mergeState);
     final float[] mergedCentroid = new float[fieldInfo.getVectorDimension()];
     int vectorCount = mergeAndRecalculateCentroids(mergeState, fieldInfo, 
mergedCentroid);
-    centroid = mergedCentroid;
     if (segmentWriteState.infoStream.isEnabled(QUANTIZED_VECTOR_COMPONENT)) {
       segmentWriteState.infoStream.message(
           QUANTIZED_VECTOR_COMPONENT, "Vectors' count:" + vectorCount);
     }
     QuantizedByteVectorValues quantizedVectorValues =
-        mergedQuantizedVectorValues(fieldInfo, mergeState, centroid);
+        mergedQuantizedVectorValues(fieldInfo, mergeState, mergedCentroid);
     long vectorDataOffset = vectorData.alignFilePointer(Float.BYTES);
     DocsWithFieldSet docsWithField = writeVectorData(vectorData, 
quantizedVectorValues);
     long vectorDataLength = vectorData.getFilePointer() - vectorDataOffset;
     float centroidDp =
-        docsWithField.cardinality() > 0 ? VectorUtil.dotProduct(centroid, 
centroid) : 0;
+        docsWithField.cardinality() > 0 ? 
VectorUtil.dotProduct(mergedCentroid, mergedCentroid) : 0;
     writeMeta(
         fieldInfo,
         segmentWriteState.segmentInfo.maxDoc(),
         vectorDataOffset,
         vectorDataLength,
-        centroid,
+        mergedCentroid,
         centroidDp,
         docsWithField);
   }
 
+  private void mergeOneFlatVectorFieldDataBlind(FieldInfo fieldInfo, 
MergeState mergeState)
+      throws IOException {
+    float[] zeroCentroid = new float[fieldInfo.getVectorDimension()];
+    // Build one merged view where, per contributing segment, either its 
existing quantized bytes
+    // are passed through or its float vectors are quantized fresh. Inputs 
already quantized to
+    // {@code encoding} against a zero centroid (data-blind segments) are 
copied directly; they are
+    // never dequantized and re-quantized, which would only add loss. Segments 
with raw floats are
+    // quantized fresh, as their stored bytes live in a different (centered) 
quantization space.
+    List<QuantizedByteVectorValuesSub> subs = new ArrayList<>();
+    for (int i = 0; i < mergeState.knnVectorsReaders.length; i++) {
+      KnnVectorsReader reader = mergeState.knnVectorsReaders[i];
+      if (reader == null) {
+        continue;
+      }
+      QuantizedByteVectorValues values;
+      if (hasRawVectorValues(reader, fieldInfo)) {
+        // Segment stored full-precision floats; quantize them against the 
zero centroid.
+        values = quantizeFromFloats(reader, fieldInfo, zeroCentroid);
+      } else {
+        QuantizedByteVectorValues qvv = getQuantizedVectorValues(reader, 
fieldInfo.name);
+        if (qvv == null || qvv.size() == 0) {
+          continue;
+        }
+        if (qvv.getScalarEncoding() != encoding) {
+          // Re-quantization from raw floats would be required, which is not 
possible when raw
+          // floats were never written.
+          throw new IllegalStateException(
+              "Cannot merge field \""
+                  + fieldInfo.name
+                  + "\" from data-blind segment with encoding "
+                  + qvv.getScalarEncoding()
+                  + " into data-blind format with encoding "
+                  + encoding
+                  + ": re-quantization requires raw float vectors");
+        }
+        float[] centroid = getCentroid(reader, fieldInfo.name);
+        if (centroid != null && isAllZero(centroid)) {
+          // Quantized-only segment whose bytes already match the output 
format (encoding and zero
+          // centroid): copy them directly.
+          values = qvv;
+        } else {
+          // Bytes were produced against a non-zero (or unknown) centroid, so 
they cannot be passed
+          // through into the zero-centroid output; re-quantize from floats.
+          values = quantizeFromFloats(reader, fieldInfo, zeroCentroid);
+        }
+      }
+      subs.add(new QuantizedByteVectorValuesSub(mergeState.docMaps[i], 
values));
+    }
+    long vectorDataOffset = vectorData.alignFilePointer(Float.BYTES);
+    MergedQuantizedByteVectorValues mergedQBVV =
+        MergedQuantizedByteVectorValues.merge(mergeState, zeroCentroid, 
encoding, subs);
+    DocsWithFieldSet docsWithField = writeVectorData(vectorData, mergedQBVV);
+    long vectorDataLength = vectorData.getFilePointer() - vectorDataOffset;
+    // centroidDp is 0 (zero centroid); the data-blind metadata omits it and 
the centroid.
+    writeMeta(
+        fieldInfo,
+        segmentWriteState.segmentInfo.maxDoc(),
+        vectorDataOffset,
+        vectorDataLength,
+        zeroCentroid,

Review Comment:
   nit: `zeroCentroid` -> `null` for clarity / ensuring that it's not used?



-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to