Skip to content

[core] Add composite BTree key storage and point lookups - #10338

Merged
JingsongLi merged 2 commits into
apache:masterfrom
JingsongLi:codex/composite-btree-storage
Oct 2, 2026
Merged

JingsongLi merged 2 commits into
apache:masterfrom
JingsongLi:codex/composite-btree-storage

Conversation

@JingsongLi

@JingsongLi JingsongLi commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

Purpose

Part 1 of the composite BTree index series, split from #10327. Add the common-layer storage primitives for ordered, multi-column BTree keys and complete-key point lookups.

  • Encode typed tuple components with length delimiters and NULL markers; compare keys lexicographically in declared field order.
  • Use one ordered List<DataField> throughout BTree construction, global-index factory creation/file selection, and index-writer creation. Remove scalar and primary/extra-fields compatibility overloads; migrate all implementations and callers, including Spark, Flink, vector, full-text, Lumina and ESLib.
  • Copy tuple keys retained by the writer so reused input rows cannot mutate posting-list grouping or min/max metadata.
  • Add GlobalIndexReader.visitCompositeEqual with a default unsupported result and a BTree implementation using the existing point lookup, Bloom filter and row-ID filtering.
  • Require the exact tuple arity when serializing keys or querying complete tuples. SQL equality with a NULL literal returns an empty result.
  • Retain composite files conservatively in factory-level scalar metadata selection until composite predicate planning is introduced.

The existing BTree file versions and posting encodings are reused. This PR is independently buildable and covers common storage and reader APIs, together with the mechanical factory-interface migration required by all consumers. Table query planning and SQL construction follow in the next parts.

Series

  1. This PR: common tuple storage, complete-key point lookup and the ordered-field-list factory API migration.
  2. Common equality matcher/pruning/evaluated coverage; core multi-column build, equality query/coverage; Python manifest and scalar-reader compatibility.
  3. Spark/Flink SQL build integration, end-to-end equality tests and documentation.
  4. Composite prefix/range/IN/NULL queries, key-only filtering, scan budgets and index selection.

Tests

JDK 8 verification, with final checks run without fast-build:

  • 367 common-layer tests passed: composite BTree (5), scalar BTree readers/writer close (328), sorted metadata pruning (7), FM (22), and multivalue bitmap (5).
  • 222 core tests passed: global-index query planning (42), schema validation (70), primary-key sorted index building (7), vector search (41), full-text search (50), and vector row-filter exactness (12).
  • 21 factory tests passed across native vector, native full-text, and Lumina.
  • Flink 1 and Spark 3 common modules: production and test-source compilation.
  • ESLib: production and test-source compilation on JDK 11, as required by that module.
  • Checkstyle, Spotless and Enforcer checks; git diff --check.

The composite cases cover BTree versions 1/2, Bloom filters, LZ4 compression, reused mutable rows, duplicate-key postings, missing keys, split-local row ranges, typed ordering, NULLs, negative integers, embedded delimiters, malformed tuple arity, empty field-list rejection, and single-column factory restrictions.

A fresh worktree needs the bundled codegen plugin prepared with mvn -pl paimon-codegen-loader -am -DskipTests package before tests that generate projections or comparators.

Split verification

The original implementation was captured at c738a41321; its branch was preserved. The initial storage patch was reconstructed in an isolated worktree and rebased onto current master with identical patch contents. The subsequent ordered-field-list API migration follows the requested simplification and is recorded for the later parts.

An independent dependency review identified the tuple-arity validation as an explicit correctness addition, included with regression assertions. An independent review of the API migration also verified field order and scalar/vector/full-text behavior; two missed engine call sites were corrected. The full-series tree-equivalence check remains for the final part, with the arity validation and API simplification recorded as intended changes.

@leaves12138 leaves12138 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed 496e445 within the stated Part 1 scope: common-layer composite key storage and complete-key point lookups. No blocking findings.

Validation on JDK 8:

  • Re-ran the documented five-class regression suite without fast-build: 339 tests passed, with Checkstyle, Spotless, RAT and Enforcer enabled.
  • Ran four additional isolated test cases covering 600 six-component keys, typed ordering and round trips, reused input rows, three index files with small blocks, BTree versions 1/2, Bloom enabled/disabled, LZ4, duplicate postings, null-containing tuples, missing keys and split-local row-ID filtering. All passed.

The additive reader API and conservative factory-level pruning are appropriate for this storage-only stage. SQL construction, composite predicate planning and prefix/range support remain explicitly outside this PR. The full remote CI matrix is still in progress at review time.

@JingsongLi
JingsongLi merged commit df43423 into apache:master Oct 2, 2026
25 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants