Skip to content

Validate cloud object snapshots and range responses - #613

Open
quinnj wants to merge 2 commits into
core-rewritefrom
fix/cloud-source-snapshots
Open

quinnj wants to merge 2 commits into
core-rewritefrom
fix/cloud-source-snapshots

Conversation

@quinnj

@quinnj quinnj commented Oct 2, 2026 •

Copy link
Copy Markdown
Member

Arrow cloud reads could combine different versions of an object or accept bytes from the wrong range. In a reproduced scan, selecting column a = 1:30000 returned column b's -7 values: the server returned an equally sized slice, and the extension discarded the Content-Range that identified it.

Pin a strong ETag before reading, refreshing missing metadata once and checking the known size. Retain each range response and require HTTP 206, exact range bounds and total size, the pinned ETag, and the requested byte count. Request the identity representation, disable HTTP decompression, and reject encoded responses. Whole-object ranges follow the same checks; empty reads send no request. Complete object handles need no extra metadata request, and overwritten objects still surface HTTP 412.

This targets core-rewrite and depends on #609. The request path is verified against released CloudStore 1.6 and 1.8, with no runtime dependency or compatibility changes. Validation follows normal HTTP body buffering; it does not add a hard network receive limit.

Validation:

  • Reproduced mixed-version reads against signed MinIO and Azurite before the metadata fix, and reproduced the wrong-column scan through both provider clients on unchanged code.
  • The original wrong-column reproduction now passes 8/8 checks, preserving valid reads and rejecting the wrong range.
  • All 145 cloud checks pass on Julia 1.10.12 / CloudStore 1.6.0 / HTTP 1.11 and Julia 1.13.1 / CloudStore 1.8.0 / HTTP 2.8, including signed MinIO and Azurite Table/Stream scans, whole-object reads, and replacement checks.
  • Protocol controls cover missing, malformed, oversized, or inconsistent range metadata; missing, changed, or weak response ETags; body-length mismatches; and gzip responses. Valid chunked responses and case-insensitive range units remain supported.
  • The full Arrow suite passes on Julia 1.13.1 with four threads, including Aqua and the scan acceptance checks. Strict documentation build, JuliaFormatter, and git diff --check also pass.
  • Independent review found no actionable issue in this extension change. A separate response-level probe passed 53/53 checks on each dependency environment.

The research, implementation, and tests were AI-driven with OpenAI Codex, including an independent agent review. The broader Arrow rewrite in #609 still needs upstream review.

Co-authored by Codex

quinnj added 2 commits October 1, 2026 23:49
Resolve missing CloudStore metadata once and require a strong ETag before reading ranges. Check that the refreshed size agrees with the supplied handle, so same-size replacement fails with HTTP 412 instead of mixing versions.

Generated-by: OpenAI Codex
Retain provider responses and reject mismatched ranges, object versions, content encodings, and byte counts before Arrow interprets the bytes. Keep the released CloudStore request paths and disable HTTP decompression so returned bytes retain their stored offsets.

Cover equal-length wrong-column responses and protocol faults through both provider clients, and exercise the signed S3 and Azure Table and Stream paths.

Generated-by: OpenAI Codex
@quinnj quinnj changed the title Pin cloud source metadata before range reads Validate cloud object snapshots and range responses Oct 2, 2026

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant