[core] Add composite BTree key storage and point lookups - #10338
Merged
JingsongLi merged 2 commits intoOct 2, 2026
Merged
Conversation
leaves12138
approved these changes
Oct 2, 2026
leaves12138
left a comment
Contributor
There was a problem hiding this comment.
Reviewed 496e445 within the stated Part 1 scope: common-layer composite key storage and complete-key point lookups. No blocking findings.
Validation on JDK 8:
- Re-ran the documented five-class regression suite without fast-build: 339 tests passed, with Checkstyle, Spotless, RAT and Enforcer enabled.
- Ran four additional isolated test cases covering 600 six-component keys, typed ordering and round trips, reused input rows, three index files with small blocks, BTree versions 1/2, Bloom enabled/disabled, LZ4, duplicate postings, null-containing tuples, missing keys and split-local row-ID filtering. All passed.
The additive reader API and conservative factory-level pruning are appropriate for this storage-only stage. SQL construction, composite predicate planning and prefix/range support remain explicitly outside this PR. The full remote CI matrix is still in progress at review time.
This was referenced Oct 2, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Purpose
Part 1 of the composite BTree index series, split from #10327. Add the common-layer storage primitives for ordered, multi-column BTree keys and complete-key point lookups.
List<DataField>throughout BTree construction, global-index factory creation/file selection, and index-writer creation. Remove scalar and primary/extra-fields compatibility overloads; migrate all implementations and callers, including Spark, Flink, vector, full-text, Lumina and ESLib.GlobalIndexReader.visitCompositeEqualwith a default unsupported result and a BTree implementation using the existing point lookup, Bloom filter and row-ID filtering.The existing BTree file versions and posting encodings are reused. This PR is independently buildable and covers common storage and reader APIs, together with the mechanical factory-interface migration required by all consumers. Table query planning and SQL construction follow in the next parts.
Series
Tests
JDK 8 verification, with final checks run without
fast-build:git diff --check.The composite cases cover BTree versions 1/2, Bloom filters, LZ4 compression, reused mutable rows, duplicate-key postings, missing keys, split-local row ranges, typed ordering, NULLs, negative integers, embedded delimiters, malformed tuple arity, empty field-list rejection, and single-column factory restrictions.
A fresh worktree needs the bundled codegen plugin prepared with
mvn -pl paimon-codegen-loader -am -DskipTests packagebefore tests that generate projections or comparators.Split verification
The original implementation was captured at
c738a41321; its branch was preserved. The initial storage patch was reconstructed in an isolated worktree and rebased onto current master with identical patch contents. The subsequent ordered-field-list API migration follows the requested simplification and is recorded for the later parts.An independent dependency review identified the tuple-arity validation as an explicit correctness addition, included with regression assertions. An independent review of the API migration also verified field order and scalar/vector/full-text behavior; two missed engine call sites were corrected. The full-series tree-equivalence check remains for the final part, with the arity validation and API simplification recorded as intended changes.