Skip to content

[core] Support composite BTree equality queries for data evolution tables - #10339

Merged
JingsongLi merged 13 commits into
apache:masterfrom
JingsongLi:codex/composite-btree-equality
Oct 2, 2026
Merged

JingsongLi merged 13 commits into
apache:masterfrom
JingsongLi:codex/composite-btree-equality

Conversation

@JingsongLi

@JingsongLi JingsongLi commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

Purpose

Part 2 of the composite BTree series, following merged #10338 and split from #10327. Connect complete-key equality lookups to data evolution index construction, query planning and coverage accounting.

  • Match a conjunction of equalities against every indexed field, independent of predicate order, and construct the point key in declared index order. Prefer the longest fully matched composite definition before reading single-column postings.
  • Prune tuple files using typed min/max metadata. NULL or contradictory equalities yield no candidates; incomplete keys and missing metadata remain conservative.
  • Build ordered tuple keys through the core sorted-index writer/scanner, preserve ordered definitions in manifests, and refresh when either indexed component changes.
  • Track the row ranges covered by the index paths that actually contributed, including ranges whose files were safely pruned by metadata. Preserve unsupported-reader fallback and scan uncovered rows in full/detail modes.
  • Keep composite files out of scalar TopN and vector/full-text prefilters. Allow scalar and distinct composite BTree definitions to coexist in Java manifests.
  • Reuse the ordered List<DataField> factory APIs merged in [core] Add composite BTree key storage and point lookups #10338. Replace the custom tuple encoding with RowCompactedSerializer behind KeySerializer; canonicalize RowKind and floating NaN values, use one serializer per thread, and remove the extra tuple-arity checks.

SQL topology integration in Spark/Flink follows in part 3; this PR includes only the mechanical migration of existing single-column topology calls to the field-list API. Composite prefix/range/IN/IS NULL bounds and scan-budget/index-selection extensions follow in part 4. This PR replaces the unreleased composite key encoding from #10338; composite indexes built with that implementation must be rebuilt. It introduces no new table options.

Tests

  • 283 Java tests passed on JDK 8, with Checkstyle, Spotless and Enforcer enabled (fast-build was not used for final verification).
  • The root RAT licensing check and git diff --check passed.
  • Cover two/three-column keys, reversed predicate order, nested OR branches, contradictory equalities, NULL components, typed tuple pruning, metadata gaps, serialized reader-side plans, partial coverage in both directions, full/detail/fast modes, dropped columns, incremental builds and refresh of either component.
  • Remove scalar index data files while retaining their manifest entries to prove joint equality queries avoid those postings. A separate test builds matching two- and three-column indexes, removes the shorter index's data files, and proves eager/deferred queries choose the longest complete key.
  • An isolated semantic mutation reversing the longest-key preference failed that regression with the expected missing-shorter-index-file error. Production source was restored and the original regression passed again.
# Prepare the bundled codegen plugin in a fresh checkout.
mvn -pl paimon-codegen-loader -am -DskipTests package

mvn -pl paimon-core -am -DwildcardSuites=none -DfailIfNoTests=false \
  -Dtest=CompositeBTreePredicateTest,CompositeBTreeIndexTest,CompositeBTreeTableTest,GlobalIndexQueryTest,GlobalIndexEvaluatorTest,SortedFileMetaSelectorTest,SortedGlobalIndexScannerTest,SortedGlobalIndexWriterTest,IndexManifestFileHandlerTest,DataEvolutionBatchScanTest,BtreeGlobalIndexTableTest,BitmapGlobalIndexTableTest,MultiValueGlobalIndexTableTest,VectorSearchRowFilterExactnessTest,FullTextSearchBuilderTest,IndexQuerySplitTest test

Split verification

Reconstructed from the reviewed equality snapshot 4e34273c09 in a new worktree based on master df43423696. The complete source branch at c738a41321 was preserved. The merged field-list API remains in place; the serializer reuse and removal of extra arity checks are explicit follow-up changes requested during review; no composite range-planning or engine topology changes are included.

Independent dependency, query/coverage, and storage reviews found no remaining P0–P2 issues. The longest-key selection regression is recorded as an explicit test addition for the final-series tree-equivalence check.

Serializer follow-up

  • Removed the per-component length codec and delegated serialization/deserialization to RowCompactedSerializer. Typed object comparison remains in the KeySerializer adapter.

  • Canonicalize RowKind and floating NaN values in a temporary row to keep equality and Bloom hashing consistent without changing the source row or the shared row serializer.

  • Added compacted encoding compatibility, RowKind, concurrent round-trip, NaN payload and Bloom point-lookup regressions. The encoding and NaN regressions were observed failing before their respective fixes.

  • Final follow-up verification: 64 Java tests passed on JDK 8 across CompositeBTreeIndexTest, CompositeBTreePredicateTest, BTreeThreadSafetyTest, BTreeBloomFilterTest, SortedFileMetaSelectorTest, CompositeBTreeTableTest, SortedGlobalIndexWriterTest and SortedGlobalIndexScannerTest, with Checkstyle, Spotless and Enforcer enabled. Independent review found no remaining P0-P2 issues.

Key ownership follow-up

  • Remove the composite-specific copying and serializer checks from BTreeIndexWriter, restoring its original implementation. No KeySerializer.copy() API is added.

  • Composite deserialization follows the same ownership contract as string deserialization: copy the input bytes and return a new row with fields backed by that independent allocation. The existing RowCompactedSerializer adapter already satisfies this contract.

  • Verify that consecutive deserialization calls and modifications to input buffers leave previously returned keys intact. Use independently deserialized keys for posting/local-range tests and assert the retained min/max metadata for both BTree file versions.

  • Final verification: 65 Java tests passed on JDK 8, with Checkstyle, Spotless and Enforcer enabled. Independent review confirmed returned rows/field storage are independent and found no remaining P0-P2 issues.

Composite predicate entry point

  • Replace visitCompositeEqual(List<Object>) with visitComposite(Predicate) and remove the old API. Query planning passes its composite predicate directly to the reader.

  • Retain the ordered index fields in the BTree reader. Match complete equalities by field name/type and construct the point key in index order, independent of the predicate schema and conjunction order.

  • Keep the current complete-equality scope: NULL/contradictory equalities return a supported empty result; incomplete keys and currently unsupported range/IN/OR predicates return an empty Optional so callers can fall back. Single-column readers decline composite predicates. Future range/prefix support can extend this same entry point.

  • Update point-lookup/local-range tests to use predicates; cover field-order changes, contradictions, unsupported conditions and single-column isolation.

  • Final verification: 108 Java tests passed on JDK 8, covering direct reader behavior, eager/deferred query execution, coverage/fallback and sorted-index construction, with Checkstyle, Spotless and Enforcer enabled. Independent review found no remaining P0-P2 issues.

Evaluation constructor simplification

  • Remove the internal two-argument Evaluation constructor and update all six scalar/refinement call sites to pass null coverage explicitly. Composite queries continue passing their concrete coverage.
  • Preserve the distinction between unspecified (null) and explicitly empty coverage; scanner fallback behavior is unchanged.
  • Final verification: 93 Java tests passed on JDK 8 across GlobalIndexEvaluatorTest, GlobalIndexQueryTest and SortedGlobalIndexScannerTest, with Checkstyle, Spotless and Enforcer enabled.

Sorted builder field-list API

  • Remove withIndexField(String) from SortedGlobalIndexScanner and SortedGlobalIndexWriter. Migrate all 28 calls across core tests and Flink/Spark topology code to withIndexFields(Collections.singletonList(...)).
  • Existing single-column topology behavior is preserved; composite SQL topology support remains in part 3.
  • Final verification: 52 Java tests passed on JDK 8, including core sorted scanner/writer and composite-table tests plus Flink and Spark topology tests, with the flink1/spark3 profiles and Checkstyle, Spotless and Enforcer enabled. Independent review found no remaining P0-P2 issues; no old method definitions or calls remain.
mvn -pl paimon-core,paimon-flink/paimon-flink-common,paimon-spark/paimon-spark-common \
  -am -Pspark3,flink1 -DwildcardSuites=none -DfailIfNoTests=false \
  -Dtest=SortedGlobalIndexScannerTest,SortedGlobalIndexWriterTest,CompositeBTreeTableTest,SortedIndexTopoBuilderTest,IndexQuerySourceTest test

Python scope

  • Remove all Python production and test changes from this PR. All five affected Python files are restored byte-for-byte to the PR base.
  • Verify that paimon-python has no diff against the base and that the non-Python diff is identical to the previously verified version.

Per-field coverage rule

  • Remove the BTree-specific condition from DataEvolutionGlobalIndexCoverage. Every definition with non-empty extraFieldIds is excluded from per-field coverage, including its primary field. Only independent single-field definitions contribute; null and empty extra-field arrays remain valid single-field definitions.
  • Composite query paths continue using their explicitly selected coverage. FULL/DETAIL searches scan data wherever dedicated single-field coverage is missing; FAST mode retains its existing behavior.
  • Add FULL/DETAIL regressions for BTree and another index type, independently covering primary/extra columns, null/empty metadata, disjoint scalar ranges and explicit query coverage. The two non-BTree cases failed against the prior implementation with missing fallback ranges.
  • Update the vector regression to assert the new fallback ranges in FULL/DETAIL and exact results across FAST/FULL/DETAIL; retain reader input checks.
  • Final verification: 196 Java tests passed on JDK 8 across CompositeBTreeTableTest, GlobalIndexQueryTest, BtreeGlobalIndexTableTest, FullTextSearchBuilderTest, VectorSearchBuilderTest and VectorSearchRowFilterExactnessTest, with Checkstyle, Spotless and Enforcer enabled. Independent review found no remaining P0-P2 issues.

Scanner and query single-field selection

  • Replace Scanner's isCompositeBTree with index-type-independent isMultiFieldIndex based on non-empty extraFieldIds. Scalar prefilters and scalar reader construction exclude all multi-field definitions. Retain the BTree capability constraint for TopN.
  • Apply the same rule to GlobalIndexQuery scalar planning: only independent single-field definitions can serve scalar leaves, including in-reader execution and residuals next to a complete composite condition. The actual BTree composite matching remains in query planning.
  • Add missing-file regressions for BTree and another multi-field type, single-field planner coverage for primary/extra columns, and a mixed composite lookup with an unsupported multi-field residual. Before their respective fixes, the non-BTree scanner tried to load its factory, the non-BTree scalar planner incorrectly produced a plan, and mixed evaluation tried to open the residual definition.
  • Preserve metadata planning for unknown single-field algorithms, explicitly contributed field ids and composite coverage. No Python changes are included.
  • Final verification: 201 Java tests passed on JDK 8 across query planning, composite and scalar BTree tables, full-text search, vector search and exact row filtering, with Checkstyle, Spotless and Enforcer enabled. Independent review's Query-selection finding was fixed and re-reviewed CLEAN.

Generic manifest definition identity

  • Remove the BTree-specific coexistence condition from IndexManifestFileHandler. Index types are grouped by the existing outer layer; compare each definition's complete ordered field-id list within a type. Different definitions may coexist, while overlapping ranges for the same definition require deleting the prior file first.
  • Treat null and empty extra-field arrays as the same single-field definition. Non-overlapping ranges remain valid, and the existing source-meta rules are unchanged.
  • Add tests for three index types covering independent single/pair/triple/reordered definitions, overlapping replacement rejection, delete-before-replace and null/empty metadata equivalence. Five of six new parameterized cases failed against the prior implementation.
  • Final verification: 41 Java tests passed on JDK 8 across IndexManifestFileHandlerTest and CompositeBTreeTableTest, with Checkstyle, Spotless and Enforcer enabled. Independent review of the Java diff found no remaining P0-P2 issues.
  • Python manifest identity alignment remains a separate follow-up under the requested exclusion of Python changes from this PR. Java/Python coexistence rules are not yet aligned.

Reader best effort follow-up

  • Remove supportsPredicate, requiresGlobalEvaluation, and the special eager fallback from deferred index planning. Readers determine which predicates can use each index at runtime.
  • Distinguish metadata-proven no matches from unavailable index filtering. A metadata miss can prune a split only when its coverage includes the complete requested row range.
  • Retain candidates outside each supported leaf's coverage before combining AND/OR. AND can narrow candidates using other supported predicates; an unavailable OR branch retains the eligible rows for data filtering.
  • FAST limits reader candidates to rows covered by any selected query index. Missing or unsupported conditions within this domain remain residual filters, so partial coverage can retain additional correct matches compared with eager FAST. FULL/DETAIL retain all actual split ranges where index filtering is unavailable.
  • Cover partial coverage across separate and coalesced data splits, AND/OR, unsupported predicates, zero scan budgets, and alternative algorithms. Split serialization, pinned index files, and failure on missing required index files remain unchanged; Python changes remain excluded.

Final follow-up verification: 206 Java tests passed on JDK 8, with Checkstyle, Spotless and Enforcer enabled. Independent static review found no production P0–P2 findings.

mvn -pl paimon-core -am -DfailIfNoTests=false -DwildcardSuites=none \
  -Dtest=CompositeBTreeTableTest,IndexQuerySplitTest,GlobalIndexQueryTest,DataEvolutionBatchScanTest,BtreeGlobalIndexTableTest,BitmapGlobalIndexTableTest,GlobalIndexEvaluatorTest,CompositeBTreePredicateTest,CompositeBTreeIndexTest,SortedFileMetaSelectorTest test

CI follow-up: vector-search coverage expectations

Synchronize the mixed vector-search integration test with the single-field coverage rule: multi-field definitions do not remove rows from the FULL data fallback, and their matching index files remain attached to that fallback. Keep the exact final Top-K result assertion.

Verification: the original failure reproduced locally before the correction. All 21 vector-search procedure tests passed with Flink 1 on JDK 8, and all 21 passed with Flink 2 on JDK 11, with Checkstyle, Spotless and Enforcer enabled. The shared Flink module was rebuilt from clean output for the second profile.

@leaves12138 leaves12138 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed c2cb701. No blocking findings within the documented Java/common+core complete-equality scope.

Validation on JDK 8:

  • 321 focused tests passed across composite matching/storage, eager and deferred queries, evaluated coverage, manifest identities, incremental builds, scalar/search prefilters and split serialization.
  • An additional 328 existing scalar BTree reader/writer regression tests passed. Both suites ran without fast-build, with the normal validation checks enabled.
  • Added an isolated seeded differential test: 80 nested AND/OR predicates across FULL/DETAIL and eager/reader execution (320 comparisons), with differently covered scalar/composite definitions and unindexed tail rows. All results matched direct row predicate evaluation.

The separation between metadata-proven misses, actual evaluated coverage and unsupported-reader fallback is consistent in the reviewed paths. Scope boundaries remain important: composite files written by the earlier unreleased encoding must be rebuilt; Python multi-index compatibility and composite SQL/prefix/range integration remain follow-ups, not capabilities validated by this approval. The full remote CI matrix is still in progress.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants