eneskeles opened a new issue, #2133:
URL: https://github.com/apache/iceberg-go/issues/2133
### Apache Iceberg version
main (development)
### Please describe the bug 🐞
Came across this one while working on #2111.
A column that is missing from an old data file should read as NULL. For
`NotNaN` and multi-valued `NotIn`, those rows get a different answer than rows
with a stored NULL.
Setup: append ids 1-4, add an optional `x double`, then append id 5 (`x =
1.0`), id 6 (`x = NaN`) and id 7 (`x = NULL`).
| scan filter | ids returned | expected |
|---|---|---|
| `NotNaN(x)` | `5, 7` | `1, 2, 3, 4, 5, 7` |
| `NOT(IsNaN(x))` | `1, 2, 3, 4, 5, 7` | same |
| `NotIn(x, 1, 2)` | `6, 7` | `1, 2, 3, 4, 6, 7` |
| `NOT(In(x, 1, 2))` | `1, 2, 3, 4, 6, 7` | same |
The stored NULL (id 7) matches, but the rows of the old file (ids 1-4)
don't. Same on v2 and v3.
### Why
`scanTranslator.VisitBound` in `visitors.go` handles a column that is not in
the file schema and has no initial default like this:
```go
if pred.Op() == OpIsNull {
return AlwaysTrue{}
}
return AlwaysFalse{}
```
That is right for most predicates, but `NotNaN` is true for NULL, and so is
`NotIn` with more than one value.
### Expected
Rows from a file without the column should match the same way as rows with a
stored NULL. `NotNaN` on a missing column should become `AlwaysTrue`.
The `NotIn` part depends on #<issue 1>: if `NotIn` stops matching NULL,
`AlwaysFalse` is already correct for it and only `NotNaN` needs fixing.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]