Skip to content

[core] Support composite BTree global indexes for data evolution tables - #10327

Closed
JingsongLi wants to merge 4 commits into
apache:masterfrom
JingsongLi:codex/composite-btree-index
Closed

JingsongLi wants to merge 4 commits into
apache:masterfrom
JingsongLi:codex/composite-btree-index

Conversation

@JingsongLi

@JingsongLi JingsongLi commented Oct 1, 2026 •

Copy link
Copy Markdown
Contributor

Split plan

This PR remains the complete implementation reference. The contribution is being submitted in dependency order:

  1. Common tuple storage, point lookups and ordered-field-list factory API: [core] Add composite BTree key storage and point lookups #10338 (merged).
  2. Complete-key equality matching/pruning, core builds/query/coverage, and Python compatibility: [core] Support composite BTree equality queries for data evolution tables #10339 (23 files).
  3. Spark/Flink SQL build integration, equality end-to-end tests and documentation.
  4. Composite prefix/range/IN/NULL queries, key-only filtering, budgets and index selection.

The first PR merged with the ordered List<DataField> API and exact tuple-arity validation. The second PR is based on that merged master and passed 283 Java and 92 Python tests, including an isolated mutation proving longest-complete-key selection. The original source branch remains preserved; arity validation, API simplification and the added selection regression are recorded for the final-series equivalence check.

Purpose

Support multi-column BTree global indexes for data evolution tables. Queries such as category = ? AND item_number = ? currently read each single-column posting list before intersecting them, which is expensive when one condition matches many row IDs. A composite index reads the posting list for the complete tuple directly.

  • Add typed, length-delimited tuple keys with lexicographic comparison, null handling, full-key bloom filtering and min/max pruning. Reuse the existing BTree file versions and retain single-column behavior.
  • Build composite indexes in Spark and Flink with index_column => 'category,item_number'. Carry all key columns through sorting, metadata, incremental builds and refresh after component updates.
  • Plan equality/IN/NULL prefixes and an optional following range, independently of predicate order. Evaluate later indexed-column conditions on tuple keys before decoding row-ID postings. Choose indexes using useful leading bounds, key filters, coverage, probe counts and selected file bytes. Allow composite and single-column indexes to coexist.
  • Track coverage for the selected index path so full/detail searches scan uncovered rows without borrowing coverage from a different index. Preserve coverage when key metadata prunes every index file.
  • Skip composite index files in PyPaimon's scalar reader, allow distinct BTree definitions to coexist in Python manifests, and document the supported behavior.

Composite queries support complete point lookups, leading prefixes, equality/IN/IS NULL prefixes followed by >, >=, <, <=, BETWEEN, or IS NOT NULL, and supported filters on later key columns. IN expansion is bounded at 256 intervals per lookup. Prefix/range scans obey the selected-file scan budget; budget rejection preserves available scalar alternatives. Complete point queries may evaluate in readers; ranges and mixed/scalar fallback paths resolve runtime support during planning. No skip scan or composite LIKE interval construction is included. Vector and full-text pre-filters continue to use single-column indexes or ordinary data fallback. PyPaimon does not build or read composite indexes. No production-scale benchmark has been run.

Tests

Previous equality-path verification passed without the fast-build profile:

mvn -pl paimon-core -am -DfailIfNoTests=false -DwildcardSuites=none \
  -Dtest=CompositeBTreeIndexTest,CompositeBTreeTableTest,BtreeGlobalIndexTableTest,GlobalIndexQueryTest,GlobalIndexEvaluatorTest,SortedGlobalIndexScannerTest,SortedGlobalIndexWriterTest,IndexManifestFileHandlerTest,DataEvolutionBatchScanTest,BitmapGlobalIndexTableTest,MultiValueGlobalIndexTableTest,BTreeIndexWriterCloseTest,BTreePostingListTest,VectorSearchRowFilterExactnessTest,FullTextSearchBuilderTest test

mvn -pl paimon-flink/paimon-flink-common,paimon-spark/paimon-spark-ut -am \
  -Pflink1 -Pspark3 -DfailIfNoTests=false \
  -Dtest=SortedGlobalIndexITCase#testCompositeBTreeIndex,SortedIndexTopoBuilderTest,CreateGlobalIndexProcedureTest \
  -DwildcardSuites=org.apache.paimon.spark.procedure.CompositeBTreeIndexProcedureTest test

PYTHONPATH=paimon-python python3 -m unittest pypaimon.tests.global_index_scalar_search_mode_test pypaimon.tests.index_manifest_write_test
  • 248 common/core tests, 12 Flink tests, 15 Spark Java tests, two Spark Scala integration tests and 16 Python tests passed.
  • New regressions cover two- and three-column lookups, reversed predicate order, OR branches, contradictions, nulls, negative integers, embedded delimiters, mutable input rows, BTree versions 1/2, partial coverage in both directions, serialized reader-side plans, incremental builds and refresh after either component changes.
  • The single-column index files are removed in a joint-query regression, including redundant null checks and same-key prefix/range residuals, to prove that composite lookups do not read their posting lists.
  • Maven Checkstyle, Spotless, RAT and Enforcer checks passed; git diff --check passed.

Review follow-up

Five review/fix rounds and a final confirmation pass cover query planning, storage and metadata, Spark/Flink build lifecycle, Python interoperability, and advanced-search compatibility.

  • Preserve runtime unsupported-predicate fallback and use the coverage of supported index groups. Keep mixed composite/scalar planning consistent across splits and the public scanner APIs.
  • Ignore index definitions with dropped schema fields in both query selection and coverage accounting.
  • Keep vector/full-text scalar evaluation aligned with its per-column coverage until composite pre-filters are integrated.
  • Keep unsupported residuals in the data filter; evaluate supported conditions on selected tuple keys before expanding row-ID postings to avoid large scalar postings.
  • Shard Flink tuple build tasks by row range, preserve Spark nullable-extra-fields API compatibility, and allow distinct BTree definitions to coexist in Python manifests.
  • Regressions reproduce the original failures, including mutation checks for ANN/raw-vector and full-text coverage.

CI fixes

Reproduced the JDK 8/11 type-validation failure and the five shared Python failures from the failed CI run.

  • Construct tuple serializers explicitly in composite BTree paths. Restore the generic scalar serializer's rejection of physical ROW types, including ARRAY in multi-value indexes, and add a regression test for BTree/bitmap ROW fields.
  • Skip extra-field metadata only for composite BTree files in PyPaimon. Preserve other scalar index types and update multi-field scalar fixtures to use bitmap indexes while retaining coverage, padding and reader-selection assertions.
  • Final local verification passed: 396 common/core Java tests without fast-build, 92 focused Python tests, Checkstyle, Spotless, RAT, Enforcer and flake8.
mvn -pl paimon-core -am -DwildcardSuites=none -DfailIfNoTests=false \
  -Dtest=MultiValueBitmapIndexReaderTest,LazyFilteredBitmapIndexReaderTest,BTreeIndexReaderTest,LazyFilteredBTreeIndexReaderTest,CompositeBTreeIndexTest,SortedFileMetaSelectorTest,CompositeBTreeTableTest,MultiValueGlobalIndexTableTest,BitmapGlobalIndexTableTest,SortedGlobalIndexScannerTest,SortedGlobalIndexWriterTest test

PYTHONPATH=paimon-python python3 -m unittest pypaimon.tests.vector_search_filter_test pypaimon.tests.global_index_scalar_search_mode_test pypaimon.tests.index_manifest_write_test

Prefix and range follow-up

  • Add virtual prefix boundaries for SST seek, preserving the existing on-disk tuple encoding and BTree versions 1/2. Inclusive/strict ranges, integer extremes and nullable/non-nullable keys are covered.
  • Bound IN Cartesian expansion at 256 intervals. Check key-only residuals before posting decode; a corrupt rejected posting fixture proves the decode is skipped.
  • Prune files using tuple min/max, preserve unknown endpoints conservatively, and apply the scan budget to selected file bytes per row-range group. Complete point probes and metadata-proved empty results remain supported at a zero budget.
  • Select candidates using leading bounds, key filters, row coverage, point probes, interval count and file bytes. This is a metadata heuristic, not a statistics-based cost optimizer.
  • Review reproduced a regression where rejecting a composite range also lost scalar equality pruning. Exclude budget-ineligible candidates before consuming predicates; resolve unsupported scalar alternatives before deferred split pruning.
  • Independent storage/query reviews found no remaining P0-P2 findings. A JDK 8 probe checked 6,274 plans over nullable tuples; four isolated semantic mutations were caught by the added tests.
  • Final Spark/Flink verification: one Flink integration test and two Spark Scala integration tests passed, including SQL equality, prefix, range, IN, IS NULL and IS NOT NULL under both query-in-reader settings. Checkstyle, Spotless, RAT and Enforcer passed without fast-build.
  • Final common/core verification: 608 tests passed, including shared SST seek, scalar readers, full/detail/fast coverage, serialized query plans, vector/full-text isolation and fallback paths. Checkstyle, Spotless, RAT and Enforcer passed without fast-build.
mvn -pl paimon-core -am -DwildcardSuites=none -DfailIfNoTests=false \
  -Dtest=CompositeBTreeIndexTest,CompositeBTreeTableTest,GlobalIndexQueryTest,BTreeIndexReaderTest,LazyFilteredBTreeIndexReaderTest,BlockIteratorTest,SstFileReaderCacheTest,GlobalIndexEvaluatorTest,SortedFileMetaSelectorTest,BtreeGlobalIndexTableTest,DataEvolutionBatchScanTest,IndexQuerySplitTest,VectorSearchRowFilterExactnessTest,FullTextSearchBuilderTest test

mvn -pl paimon-flink/paimon-flink-common,paimon-spark/paimon-spark-ut -am \
  -Pflink1,spark3 -DfailIfNoTests=false \
  '-Dtest=SortedGlobalIndexITCase#testCompositeBTreeIndex' \
  -DwildcardSuites=org.apache.paimon.spark.procedure.CompositeBTreeIndexProcedureTest test

@JingsongLi
JingsongLi force-pushed the codex/composite-btree-index branch from 72c674e to 86af449 Compare October 1, 2026 15:15
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant