Skip to content

[flink][spark] Support composite BTree global index creation - #10345

Merged
JingsongLi merged 1 commit into
apache:masterfrom
JingsongLi:codex/composite-btree-engines
Oct 3, 2026
Merged

JingsongLi merged 1 commit into
apache:masterfrom
JingsongLi:codex/composite-btree-engines

Conversation

@JingsongLi

Copy link
Copy Markdown
Contributor

Purpose

Allow Flink and Spark to build composite BTree global indexes for Data Evolution tables through sys.create_global_index with an ordered, comma-separated column list.

This follows the storage and equality-query support merged in #10338 and #10339.

  • Pass every key column to the index scanner, writer and indexer.
  • Construct typed tuple keys in both engines and preserve key order during sorting.
  • Keep independently configured single-column indexes and composite definitions separate.
  • Add integration coverage for nullable components, reader/eager equality queries, incremental builds, updates to either component, unchanged rebuilds and precise index removal.
  • Document the current full-key equality contract and build lifecycle.

Python support remains deferred. Prefix, range and IN query optimization will follow separately.

Tests

All scoped tests pass without the fast-build profile:

Profile Runtime Scope Result
Flink 1 JDK 8 Composite/single BTree, Bitmap, MultiValue integration tests; topology and procedure unit tests 15 passed
Flink 2 JDK 11 Same Flink scope after rebuilding the common module 15 passed
Spark 3 JDK 8 / Scala 2.12 Composite and existing create-index procedure suites; topology and procedure Java unit tests 9 Scala + 15 Java passed
Spark 4 JDK 17 / Scala 2.13 Same Spark scope after a clean dependency build 9 Scala + 15 Java passed

Checkstyle, Spotless, Maven Enforcer, Apache RAT and git diff --check also pass.

@leaves12138 leaves12138 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed 08c3521. No blocking issues found.

Checked the ordered field propagation, tuple materialization and sorting in both engines, coexistence with independent single-column definitions, and the incremental build / component-update / drop lifecycle.

Local validation on JDK 8, without fast-build:

  • Flink 1: all 15 scoped integration/topology/procedure tests passed.
  • Spark 3 / Scala 2.12: all 9 Scala procedure tests and 15 Java topology/procedure tests passed.
  • An additional temporary Flink integration test passed with 6,000 rows across two partitions, a three-column key in non-schema order, object reuse enabled, and a 64 KiB sort buffer. It covered partition-filtered construction followed by incremental construction of the remaining partition and equality lookups in both eager and reader modes.

Full GitHub CI still has pending jobs at review time; Flink 2 and Spark 4 were not rerun locally.

@JingsongLi
JingsongLi merged commit 8063701 into apache:master Oct 3, 2026
15 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants