Skip to content

Native from_json with a schema that repeats a field name fails with "ArrowArray struct has 2 children (expected 1)" #6592

Description

@andygrove

Describe the bug

With the native from_json path enabled (spark.comet.expression.JsonToStructs.allowIncompatible=true), a schema that repeats a field name fails the task once the result reaches the JVM. Spark accepts such a schema and fills the last field with that name, so from_json('{"a":1}', 'a INT, a INT') returns {null, 1}. Comet's native projection builds the same struct, and to_json over it matches Spark, but importing the result on the JVM fails because Java Arrow keys struct children by name:

java.lang.IllegalStateException: ArrowArray struct has 2 children (expected 1)
	at org.apache.arrow.util.Preconditions.checkState(Preconditions.java:562)
	at org.apache.arrow.c.ArrayImporter.doImport(ArrayImporter.java:92)
	at org.apache.arrow.c.ArrayImporter.importArray(ArrayImporter.java:68)
	at org.apache.arrow.c.ArrowImporter.importVector(ArrowImporter.java:62)
	at org.apache.comet.vector.NativeUtil.importVector(NativeUtil.scala:264)
	at org.apache.comet.vector.NativeUtil.getNextBatch(NativeUtil.scala:211)
	at org.apache.comet.CometExecIterator.getNextBatch(CometExecIterator.scala:236)

By default from_json runs through the codegen dispatcher, which declines structs with repeated field names (#5766), so only the opt-in native path is affected.

Steps to reproduce

Reproduced on main at fef94f6cd with Spark 4.1.3:

SET spark.comet.expression.JsonToStructs.allowIncompatible=true;

CREATE TABLE tj (j STRING) USING parquet;
INSERT INTO tj VALUES ('{"a":1}'), ('{"a":2,"b":3}'), (NULL);

SELECT from_json(j, 'a INT, a INT') FROM tj;            -- IllegalStateException
SELECT to_json(from_json(j, 'a INT, a INT')) FROM tj;   -- matches Spark
SELECT from_json(j, 'a INT, b INT') FROM tj;            -- matches Spark

The plan is CometProject over CometNativeScan.

Expected behavior

Either match Spark or stay off the native path. CometJsonToStructs.isSupportedSchema checks the field types but not the names. If it rejected a struct with repeated field names at any depth, these schemas would go to the dispatcher, which already declines them, and the query would fall back to Spark.

Additional context

This is the same Java Arrow limitation as #6591 (struct casts) and #6251 (arrays_zip). Found while checking whether anything in #5603 was worth keeping.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    area:expressionsExpression evaluationarea:ffiArrow FFI / JNI boundarybugSomething isn't workinggood first issueGood for newcomerspriority:mediumFunctional bugs, performance regressions, broken features

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions