ParyshevSergey opened a new pull request, #18272:
URL: https://github.com/apache/iceberg/pull/18272

   INT96 timestamps can be decoded in the wrong unit for nanosecond columns, 
causing dictionary pruning to discard matching rows. Day conversion can 
saturate and arithmetic can wrap, corrupting values near the long boundaries. 
This change decodes INT96 in the requested unit using checked arithmetic: 
MICROS truncates sub-microsecond digits, NANOS preserves them, and both accept 
their full long range and reject overflow. Generic readers return LocalDateTime 
for timestamps without a timezone, fixing residual filtering and equality 
deletes. Failed dictionary materialization releases its temporary Arrow vector. 
Spark coverage uses its microsecond representation; Flink continues to reject 
INT96.
   
   Related: [#12266](https://github.com/apache/iceberg/issues/12266).
   
   Test plan:
   
   - Verify both precisions and timezone variants, long boundaries, negative 
values, truncation, nulls, nested fields and container reuse using raw PLAIN, 
dictionary and mixed-encoding fixtures.
   - Assert surviving row IDs after timestamp residuals and equality deletes 
with an unselected timestamp key, including INT96 delete files. The 24 new 
integration cases pass; 12 failed with ClassCastException before the timezone 
fix.
   - Run full Parquet, Arrow and Data suites; targeted INT96 tests for both 
Hadoop InputFormat APIs, Spark 3.5/4.0/4.1/4.2 and Flink 1.20/2.1/2.2/2.3; and 
the existing Spark-generated INT96 regression. Latest local results: 1,805 
passed, 94 skipped, 0 failures/errors.
   - Spotless, Parquet Revapi and git diff --check pass.
   
   ---
   **AI Disclosure**
   - Model: GPT-6
   - Platform/Tool: Codex GPT-6 Astra
   - Human Oversight: reviewed
   - Prompt Summary: Fix INT96 timestamp units, overflow handling and generic 
timezone representation while retaining microsecond truncation. Production 
changes, fixtures, regression tests and PR text were generated with AI 
assistance.


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to