osscm opened a new pull request, #18187:
URL: https://github.com/apache/iceberg/pull/18187

   Reference implementation of the SCALAR index type defined in #16961.
   
   Supersedes #17426, which the stale bot auto-closed for 30 days of inactivity 
(explicitly not a
   judgement on the merits) before this branch's Spark integration existed. 
Continuing that work here
   with a full Spark integration added since.
   
   ## What is included
   
   Spec (format/index.md from #16961) + Core/API/Data/Spark layers:
   
   Layer 0 — Data model: IndexMetadata, IndexSnapshot interfaces and JSON 
parsers
   Layer 1 — Storage: IndexCatalog, IndexMetadataIO, InMemoryIndexCatalog
   Layer 2 — Tracking file: TrackingFileWriter + TrackingFileReader (Avro, same 
as Iceberg manifests)
   Layer 3 — Build: HashTransform, LeafFileMetadata, ScalarIndexCommitter, 
LeafFileEntry/Writer/Reader
   Layer 4 — Spark integration:
     - `CALL system.build_scalar_index(table, columns, transform, options)` 
populates an index from a live table
     - Read-path: SparkScanBuilder resolves equality-predicate matches against 
an existing index
     - FileScanTaskFilteringScan enforces file-level pruning at scan-planning 
time (not just advisory) --
       self-verifying: falls back to the unfiltered file set if resolved paths 
don't match real
       candidates, so a bug here can only miss an optimization, never return a 
wrong result
   
   ## Test status
   
   Core/API/Data: unit tests from earlier layers (index metadata round-trip, 
tracking/leaf file
   read-write, in-memory catalog, end-to-end build+commit+lookup wiring).
   Spark: TestBuildScalarIndexProcedure, TestScalarIndexScanPruning, and 
TestFileScanTaskFilteringScan
   green; full spark-extensions module regression run (2722 tests) confirms no 
unrelated breakage.
   
   ## Known gaps / explicit non-goals for this pass
   
   - Pruning is file-level only, not exact-row/position-level (tracked as a 
follow-up)
   - When the index confirms a key is absent, the scan is NOT pruned to zero 
files -- deliberately
     deferred since a bug there would silently return wrong (empty) results, 
not just miss an
     optimization
   - Scoped to plain SELECT batch scans; 
incremental-append/changelog/merge-on-read/copy-on-write
     scans are untouched
   - Index catalog registration is JVM-session-scoped only 
(InMemoryIndexCatalog/SparkIndexCatalogs) --
     not yet durable across restarts
   - `build_scalar_index` only supports full rebuilds -- no 
incremental/append-only update path yet.
     A related primary-key-index PoC for equality-delete-to-position-delete 
conversion found full-leaf
     rebuild maintenance falls behind realistic checkpoint budgets under 
CDC-style churn at large key
     counts, while an append-only update stays within budget -- the same risk 
likely applies here. An
     append-only (MOR-style) update path is the natural follow-up.
   - Trino integration not started (planned as a separate follow-up phase)
   - Marked Draft -- looking for early feedback on direction, not requesting a 
merge review yet
   
   ## Builds on
   
   - Design doc: 
https://docs.google.com/document/d/1N6a2IOzC6Qsqv7NBqHKesees4N6WF49YUSIX2FrF7S0
   - #16961
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to