peterxcli opened a new issue, #5436: URL: https://github.com/apache/datafusion-comet/issues/5436
### What is the problem the feature request solves? Comet keeps native columnar-to-row conversion on Spark fallback when a schema contains `VariantType`. This prevents a native scan or projection that carries Variant from feeding row-based Spark operators through Comet's native C2R path. Spark does not encode Variant in `UnsafeRow` as an ordinary nested Struct. [`VariantVal`](https://github.com/apache/spark/blob/v4.1.3/common/unsafe/src/main/java/org/apache/spark/unsafe/types/VariantVal.java#L28-L46) defines a dedicated payload, and [`UnsafeWriter.write(VariantVal)`](https://github.com/apache/spark/blob/v4.1.3/sql/catalyst/src/main/java/org/apache/spark/sql/catalyst/expressions/codegen/UnsafeWriter.java#L164-L182) writes `4-byte value length + value bytes + metadata bytes`. Comet's current [`supportsSchema`](https://github.com/peterxcli/datafusion-comet/blob/c355fefd9c0b7e96d86523a7214bb2cdd47e1a55/spark/src/main/scala/org/apache/spark/sql/comet/CometNativeColumnarToRowExec.scala#L230-L259) does not admit Variant, while the native [C2R initialization path](https://github.com/peterxcli/datafusion-comet/blob/c355fefd9c0b7e96d86523a7214bb2cdd47e1a55/native/core/src/execution/jni_api.rs#L1018-L1034) currently derives only the physical Arrow datatype. Treating the `[value, metadata]` storage as a normal Struct would produce the wrong UnsafeRow representation. ### Describe the potential solution - Preserve explicit Variant identity when initializing native C2R. - Add a top-level Variant writer that consumes Comet's canonical `[value: Binary, metadata: Binary]` array and emits Spark's exact UnsafeRow Variant payload. - Preserve the distinction between SQL NULL and Variant JSON null, and validate malformed child/null combinations rather than panicking. - Keep Variant out of generic Struct handling. Focused parity tests should cover objects, arrays/scalars, Variant JSON null, SQL NULL, nullable parents, empty batches, and fields before/after Variant. The resulting row should expose Spark `VariantType` and round-trip through `UnsafeRow.getVariant` with the same value and metadata bytes. ### Additional context Related work: #5407 establishes the native columnar `[value, metadata]` representation; #5425 tracks Variant-valued native expression output. This issue is only C2R transport. Nested Variant, shuffle/spill, Python transport, writes, and general Variant expressions remain separate. Spark 3.x behavior must remain unchanged. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
