Describe the bug
With the native from_json path enabled (spark.comet.expression.JsonToStructs.allowIncompatible=true), a schema that repeats a field name fails the task once the result reaches the JVM. Spark accepts such a schema and fills the last field with that name, so from_json('{"a":1}', 'a INT, a INT') returns {null, 1}. Comet's native projection builds the same struct, and to_json over it matches Spark, but importing the result on the JVM fails because Java Arrow keys struct children by name:
java.lang.IllegalStateException: ArrowArray struct has 2 children (expected 1)
at org.apache.arrow.util.Preconditions.checkState(Preconditions.java:562)
at org.apache.arrow.c.ArrayImporter.doImport(ArrayImporter.java:92)
at org.apache.arrow.c.ArrayImporter.importArray(ArrayImporter.java:68)
at org.apache.arrow.c.ArrowImporter.importVector(ArrowImporter.java:62)
at org.apache.comet.vector.NativeUtil.importVector(NativeUtil.scala:264)
at org.apache.comet.vector.NativeUtil.getNextBatch(NativeUtil.scala:211)
at org.apache.comet.CometExecIterator.getNextBatch(CometExecIterator.scala:236)
By default from_json runs through the codegen dispatcher, which declines structs with repeated field names (#5766), so only the opt-in native path is affected.
Steps to reproduce
Reproduced on main at fef94f6cd with Spark 4.1.3:
SET spark.comet.expression.JsonToStructs.allowIncompatible=true;
CREATE TABLE tj (j STRING) USING parquet;
INSERT INTO tj VALUES ('{"a":1}'), ('{"a":2,"b":3}'), (NULL);
SELECT from_json(j, 'a INT, a INT') FROM tj; -- IllegalStateException
SELECT to_json(from_json(j, 'a INT, a INT')) FROM tj; -- matches Spark
SELECT from_json(j, 'a INT, b INT') FROM tj; -- matches Spark
The plan is CometProject over CometNativeScan.
Expected behavior
Either match Spark or stay off the native path. CometJsonToStructs.isSupportedSchema checks the field types but not the names. If it rejected a struct with repeated field names at any depth, these schemas would go to the dispatcher, which already declines them, and the query would fall back to Spark.
Additional context
This is the same Java Arrow limitation as #6591 (struct casts) and #6251 (arrays_zip). Found while checking whether anything in #5603 was worth keeping.
Describe the bug
With the native
from_jsonpath enabled (spark.comet.expression.JsonToStructs.allowIncompatible=true), a schema that repeats a field name fails the task once the result reaches the JVM. Spark accepts such a schema and fills the last field with that name, sofrom_json('{"a":1}', 'a INT, a INT')returns{null, 1}. Comet's native projection builds the same struct, andto_jsonover it matches Spark, but importing the result on the JVM fails because Java Arrow keys struct children by name:By default
from_jsonruns through the codegen dispatcher, which declines structs with repeated field names (#5766), so only the opt-in native path is affected.Steps to reproduce
Reproduced on
mainatfef94f6cdwith Spark 4.1.3:The plan is
CometProjectoverCometNativeScan.Expected behavior
Either match Spark or stay off the native path.
CometJsonToStructs.isSupportedSchemachecks the field types but not the names. If it rejected a struct with repeated field names at any depth, these schemas would go to the dispatcher, which already declines them, and the query would fall back to Spark.Additional context
This is the same Java Arrow limitation as #6591 (struct casts) and #6251 (
arrays_zip). Found while checking whether anything in #5603 was worth keeping.