quinncheong commented on issue #4079:
URL: 
https://github.com/apache/iceberg-python/issues/4079#issuecomment-6008441752

   Two details from reproducing the first two items on 0.12.0 and on `main` 
(068aae5), in case they help scope the work. Line numbers are from `main`.
   
   **Item 1 (unquoted form) also needs the emit side.** Reading is only half of 
it: PyIceberg writes the quoted form, which is not the spec form 
(`"geometry(srid:4326)"`, `"geography(srid:4326, spherical)"` in the JSON 
serialization table of format/spec.md):
   
   ```python
   from pyiceberg.schema import Schema
   from pyiceberg.types import GeographyType, GeometryType, NestedField
   
   for t in ["geometry('srid:4326')", "geometry(srid:4326)", 
"geometry(EPSG:7415)", "geography(srid:4326, spherical)"]:
       try:
           print(f"{t!r:36} -> {NestedField(1, 'g', t, 
required=False).field_type!r}")
       except Exception as e:
           print(f"{t!r:36} -> {type(e).__name__}: {str(e).splitlines()[0]}")
   print("emitted:", GeometryType("srid:4326").model_dump_json(), 
GeographyType("srid:4326", "planar").model_dump_json())
   
Schema.model_validate_json('{"type":"struct","schema-id":0,"fields":[{"id":1,"name":"geom","required":false,"type":"geometry(srid:4326)"}]}')
   ```
   
   ```
   "geometry('srid:4326')"              -> GeometryType(crs='srid:4326')
   'geometry(srid:4326)'                -> ValidationError: Could not parse 
geometry(srid:4326) into a GeometryType
   'geometry(EPSG:7415)'                -> ValidationError: Could not parse 
geometry(EPSG:7415) into a GeometryType
   'geography(srid:4326, spherical)'    -> ValidationError: Could not parse 
geography(srid:4326, spherical) into a GeographyType
   emitted: "geometry('srid:4326')" "geography('srid:4326', 'planar')"
   pyiceberg.exceptions.ValidationError: Could not parse geometry(srid:4326) 
into a GeometryType
   ```
   
   Java's `GEOMETRY_PARAMETERS` pattern (`Types.java` lines 67-71 on 
apache/iceberg main) does not strip quotes, so it reads PyIceberg's output as 
CRS `'srid:4326'` with the quotes included. Java writes `geometry(%s)` / 
`geography(%s, %s)` (lines 632 and 709). This matches the split @moomindani 
proposed on #3530: read side there, emit side as a follow-up.
   
   **Item 2 is a hard error, not just plain binary.** Because the file column 
resolves to `BinaryType`, `_cast_if_needed` calls `promote(BinaryType, 
GeometryType)`, and that raises (pyiceberg/io/pyarrow.py line 2036, 
pyiceberg/schema.py line 1691). It happens on write as well as read. Without 
geoarrow-pyarrow, writing a binary or large_binary WKB column through 
PyIceberg's own writer fails like this:
   
   ```python
   import tempfile, uuid
   import pyarrow as pa
   from pyiceberg.io.pyarrow import PyArrowFileIO, _dataframe_to_data_files
   from pyiceberg.partitioning import UNPARTITIONED_PARTITION_SPEC
   from pyiceberg.schema import Schema
   from pyiceberg.table.metadata import new_table_metadata
   from pyiceberg.table.sorting import UNSORTED_SORT_ORDER
   from pyiceberg.types import GeometryType, LongType, NestedField
   
   schema = Schema(NestedField(1, "id", LongType(), required=False), 
NestedField(2, "geom", GeometryType(), required=False))
   meta = new_table_metadata(schema, UNPARTITIONED_PARTITION_SPEC, 
UNSORTED_SORT_ORDER, "file://" + tempfile.mkdtemp(),
                             properties={"format-version": "3"})
   df = pa.table({"id": pa.array([1], pa.int64()),
                  "geom": 
pa.array([bytes.fromhex("0101000000000000000000f03f0000000000000040")], 
pa.large_binary())})
   list(_dataframe_to_data_files(table_metadata=meta, df=df, 
io=PyArrowFileIO(), write_uuid=uuid.uuid4()))
   ```
   
   ```
     File "pyiceberg/io/pyarrow.py", line 1937, in _to_requested_schema
     File "pyiceberg/io/pyarrow.py", line 2036, in _cast_if_needed
     File "pyiceberg/schema.py", line 1691, in _
   pyiceberg.exceptions.ResolveError: Cannot promote an binary to geometry
   ```
   
   Scanning a Parquet file that stores WKB as a plain BYTE_ARRAY column with 
the matching field id fails with the same `ResolveError` in 
`ArrowScan.to_table`. With geoarrow-pyarrow installed, a `geoarrow.wkb` input 
column is rejected earlier with `UnsupportedPyArrowTypeException: Column 'geom' 
has an unsupported type: extension<geoarrow.wkb<WkbType>>`. So as far as I can 
tell, no input form round-trips today. That differs from 
mkdocs/docs/geospatial.md, which says that without geoarrow-pyarrow "geometry 
and geography are written as binary in Parquet while the Iceberg schema still 
preserves the spatial type". An end-to-end write/read test for each of the 
three column forms would catch this. The existing geo tests in 
tests/io/test_pyarrow.py only cover the type conversion.
   
   I have a small patch for item 1, read and emit side. It makes the quotes 
optional in `GEOMETRY_REGEX` / `GEOGRAPHY_REGEX`, keeps accepting the quoted 
form so older type strings still parse, emits the unquoted form, and adds 
parametrized tests to tests/test_types.py (308 pass; the new tests fail on 
main). It does not yet do the case-insensitive matching. I'm happy to open it 
as the emit-side follow-up once #3530 settles the read side, or to leave it to 
@moomindani if that's already in progress.
   
   This report was prepared with an AI coding agent (Claude Code) and 
reproduced before posting by me.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to