Conversation
Co-Authored-By: Codex <codex@openai.com>
Co-Authored-By: Codex <codex@openai.com>
Co-Authored-By: Codex <codex@openai.com>
Retry interrupted mapping release at both constructor and finalizer ownership boundaries. Check munmap failures before publishing the mapping as released. Co-Authored-By: Codex <codex@openai.com>
Record pins directly in their root and keep stable registry keys across schema and array publication. Make private and multi-root cleanup retry interrupted work without rereading freed C structs. Co-Authored-By: Codex <codex@openai.com>
Publish the aggregate counter, node state, and public release callback as one rollbackable transaction. Transfer the claim under the registry lock so a post-commit exception cannot touch reaped control memory. Co-Authored-By: Codex <codex@openai.com>
Defer interruption across allocation and ledger registration. Free an unregistered allocation exactly once when the handoff fails. Co-Authored-By: Codex <codex@openai.com>
Defer interruption from mmap through owner construction, retry claim rollback, and make close-claim rollback no-escape. Add focused failures after mmap and during rollback. Co-Authored-By: Codex <codex@openai.com>
Carry a Julia-side commit token through the C callback retry loop. Once a root can be reaped, later exceptions return without reading or calling its freed raw pointer. Also clear local malloc ownership inside retryable cleanup so deferred interruption cannot repeat a completed free. Co-Authored-By: Codex <codex@openai.com>
Make guard, close, mmap, and manual-finalizer rollback handoffs retry interruption without stealing a later owner claim. Co-Authored-By: Codex <codex@openai.com>
Keep one persistent moved-array struct for producer callbacks and retry C void releases until they publish release=NULL. Install schema cleanup before owner construction and route schema-finally failures through moved-owner cleanup. Co-Authored-By: Codex <codex@openai.com>
Keep the public pull frame responsible for its claim and speculative index. Roll back a batch advance before releasing the single-puller gate on interruption. Co-Authored-By: Codex <codex@openai.com>
Keep the constructed region reachable behind a cleanup handler until finalizer registration and the public constructor return handoff complete. Co-Authored-By: Codex <codex@openai.com>
Do not let a second interruption escape while the constructor is still the only owner of a release callback. Co-Authored-By: Codex <codex@openai.com>
Document six lifecycle and ownership findings, their dispositions, scope decisions, and final validation. Update the README review index through round ten. Co-Authored-By: Codex <codex@openai.com>
State the committed ownership and cursor handoffs that roll back or retain a cleanup owner. Bound the remaining guarantee to Julia safepoints, safe retries, finalizer-backed owners, and consumer-released C exports. Co-Authored-By: Codex <codex@openai.com>
Record the clean material review, the interruption-class judgment, scope decisions, assumptions, and final validation. Update the README review index through round eleven. Co-Authored-By: Codex <codex@openai.com>
Per-reader codec contexts (created lazily, closed on every readstream exit path — no global pools), spec-exact per-buffer Int64 prefix handling with the -1 stored-raw sentinel, declared sizes bounded before any allocation and charged to a decode-side budget, exact declared/actual size matching, and each decompressed buffer in its own exact-sized owned region. Acceptance: 2.x-written lz4 and zstd streams (compressed dictionary batches included) decode through Core; hostile and understated prefixes are clean ValidationErrors located via the framer itself. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…Action Trim-compile groundwork: release behavior is a closed set in the real system (munmap, one C callback, notify/rendezvous observers), so encode it as data on one concrete struct executed by a static _run_release!. Fault injection and exactly-once observation become counters on the action rather than injected closures; the C-data gate arms a free-only action at construction (interruption before the move reclaims only our malloc'd copy) and upgrades to call+verify+free after the move commits, with the producer's null-the-release conformance check driven by a data offset. Hook kwargs are where-parameterized so every call site specializes statically. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Asynchronous interruption (SIGINT/InterruptException, task cancellation) is now explicitly out of contract, matching ecosystem practice — the pervasive disable_sigint/retry scaffolding accreted across review rounds is deleted wholesale (guard/close protocol, finalizers, mmap claim, C-data handoffs, IPC cursor), along with its fault-injection hooks and tests. Ordinary exception safety (error paths clean up, release exactly-once) remains tested. A formal revisit is planned on Julia 1.14's structured cancellation. All Threads.Atomic boxes are replaced with @atomic struct fields (ReleaseCounter, MapClaim, local test gates). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
… runs The trim gate (core/test/trim_compile_tests.jl + trim_entrypoint.jl, matching the JSON/HTTP/Reseau/StructUtils harness convention) compiles the Core workload with --trim=safe to a ~2.2MB binary that runs to exit 0, with zero verifier errors AND zero warnings. Design rules applied: closed-set @inline isa ladders over the descriptor registry (layoutspec_of/_value_of/_materialize_of/typeequal/descriptorname/ _validate_descriptor_of), literal load widths instead of runtime DataTypes, CAS loops instead of atomic RMW (Core.modifyfield! is unimplemented in the trim verifier), Ptr{Cvoid} @cfunction finalizers (Base's generic finalizer is @nospecialize'd), concrete boundary containers (struct scalars are Vector{Pair{String,Any}} always; the NamedTuple surface is facade work per report §14.2), unrolled decimal limb math (captured-reassigned closure locals box), and primitive-form file IO in the workload (Base's write/open conveniences splat; mktempdir's cleanup registry parks the trimmed scheduler). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Use exact-size native decoder destinations, enforce V5 compression semantics, and share one reader-wide allocation budget. Add adversarial coverage for frame completeness, feature negotiation, empty buffers, context cleanup, and aggregate limits. Co-Authored-By: Codex <codex@openai.com>
Free C-call wrapper storage when conformance checks fail and roll back ownership when initial finalizer registration fails. Add exactly-once, finalizer, and two-owner MapClaim race regressions. Co-Authored-By: Codex <codex@openai.com>
Delete unused fault hooks, retry wrappers, injection seams, and dead helpers left by the interruption-policy change. Keep real ordinary-error rollback tests and make void release failures return LIVE for a later explicit call. Co-Authored-By: Codex <codex@openai.com>
Use literal temporal load widths, inline the descriptor ladder, and keep test-only release primitives behind the module namespace. Update the trim workload and tests for the narrower export surface. Co-Authored-By: Codex <codex@openai.com>
Remove the obsolete interruption guarantee and document the actual struct, compression, allocation-budget, and C callback behavior. Co-Authored-By: Codex <codex@openai.com>
Co-Authored-By: Codex <codex@openai.com>
Co-Authored-By: Codex <codex@openai.com>
Co-Authored-By: Codex <codex@openai.com>
…ection The lock-free CAS state word, generation packing, seq_cst handshake reasoning, and every yield() spin are gone: region state and guard count are plain Ints under one Threads.Condition, waiters use wait/notify (timed waits via a Timer that notifies at the deadline), and the release action runs outside the lock so a blocking action cannot deadlock closers or acquirers. timeout_ms=0 never waits, which is what finalizers use. mmapregion now maps through the Mmap stdlib (cross-platform) and wraps the array as the region's GC anchor: forceclose! invalidates views and drops the anchor, with unmapping owned by the stdlib at collection. This deletes the hand-rolled POSIX ccalls, MapClaim, MunmapRelease, the two-owner release race machinery, and their tests wholesale. Trim gate unchanged: 0 verifier errors, 0 warnings, binary runs. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
Very cool! One question: could ArrowCore.jl be a real standalone, dependency-free package, similar to how ArrowStrings.jl is split out already? From reading the code, it looks like ArrowCore.jl could contain what it has at the moment in this branch, plus cdata.jl, and the uncompressed IPC read/write path (with the vendored FlatBuffers/Meta layer). And the Tables.jl facade, ArrowTypes lowering, and the compression codecs could then stay in Arrow.jl on top of it. As far as I can tell this split already exists at the code level: ArrowCore only imports Mmap, cdata.jl only uses ArrowCore, and neither ipc_read.jl nor ipc_write.jl references Tables at all — the codecs only appear at the buffer (de)compression step. So ArrowCore.jl would depend on Base and stdlibs only: no binary deps (Lz4_jll/Zstd_jll), no Tables.jl. The two places where the package boundary doesn't quite match the code today seem small: EnumX (but Meta is generated, so the generator could emit self-contained enum modules instead), and compression (which could be a plain dispatch seam — ArrowCore.jl defines compress/decompress generics and errors helpfully on compressed buffers, Arrow.jl keeps its hard codec deps and implements them; so using Arrow behaves exactly as it does now. The reason I'm asking: we could use this in the Julia VS Code extension. The REPL process could use ArrowCore.jl and ArrowStrings.jl to hand large tables to the extension's interactive table viewer — write an uncompressed IPC file, and the TypeScript side reads it with Arrow JS, with the file-format footer giving the viewer random access to just the visible row range. That would be really nice! But the way we ship code that loads into the user's REPL process rules out anything with a binary dependency. And we also couldn't use anything that pulls in Tables.jl, because that would pin the Tables.jl version users see in their REPL to whatever we ship. So this could only work if there is a core package without those deps. (I'd also have some uses for this in Queryverse, but the VS Code one probably the more interesting one) And, as all ideas these days, heavily Claude helped ;) |
|
I think it's close to that, but currently we still have the new ArrowStrings.jl package which provides the native offset + string view representation. We want that as a stand-alone package because I want to be able to use it as a dep for CSV.jl to do direct parse into that representation, I.e. we can parse a csv file directly into a zero-copy arrow representation. So not quite zero-dependency. I'm also not sure just ArrowCore.jl would be quite enough to have arrow JS be able to read easily. Currently we have the CData & IPC layers on top of the core buffers to actually fulfill those interfaces. I do have another project/data format I've been working on that would be a much cleaner/simpler binary table format with cross-language implementations. I've been slowly iterating on it, but it could potentially be a really good fit for this. |
|
Yeah, I also think it is super close. I think the ArrowStrings.jl would not be a problem, i.e. if ArrowCore.jl were to depend on ArrowStrings.jl that would work (we use a trick for that kind of situation a lot in the extension). The main thing is really that the compression stuff would have to not be in there, because we can't handle the binary dependencies that are pulled in. And yes, the cdata and ipc layers would have to move into ArrowCore.jl as well. Claude thinks that would be easy, but who knows ;) |
(cherry picked from commit 64ff77fb9ba290e7db08a2afb6f745ae6ef8067a)
Run constructor checks before a path sink can be truncated. Attempt owned sink cleanup after failed finalization and preserve the first error when construction, the callback, or finalization also fails. Cover existing and missing paths, borrowed IO, all constructor forms, valid output parity, and deterministic sink failures in regression tests. (cherry picked from commit 8c0dc762cf536e2fc6fd0eb13088b3c17940dcca)
Serialize only valid Utf8View and BinaryView content. Clear null and inline padding bytes, drop unused buffers, and retain sharing for repeated or overlapping ranges. Keep source storage and dictionary values unchanged. Cover selections, compressed file and stream output, borrowed entries, and independent range-union checks. (cherry picked from commit b21316c65e36e13d11ee4980743e4ddd0ade9c24)
Prove that valid entries cover each copied buffer prefix, then copy the prefix once. Keep entry and validity sanitation and fall back to range compaction for gaps or unused buffers. Cache validity lookup and load each entry once. (cherry picked from commit 8491d44e40e1884608dec33f5887405478aac6dd)
Accept undeclared stream replacements and concatenate validated dictionary deltas for stream, file, and range readers. Retain earlier batch snapshots and preserve physical nested and Union layouts. Cover PyArrow fixtures, resource limits, malformed updates, and V4 messages in V5 file footers. Fixes #610
Record the original upstream commit lineage after the rebased import. Keep the current JuliaCN Arrow 3 and Flight tree unchanged.
Durations.jl 1.4 ships ZonedTimestamp{P,Z}: a UTC Timestamp{P} whose
8 bytes are exactly an Arrow timestamp column's storage. The facade now
owns native timestamp mappings.
- A zone-declared timestamp column reads as ZonedTimestamp{P,Z} at every
unit, with no TimeZones.jl requirement; zone-naive micro/nanosecond
columns read as Timestamp{P} instead of raw Int64. Zone-naive
second/millisecond columns keep their DateTime read. An empty timezone
string reads zone-naive, per the Arrow spec.
- Fresh Timestamp and ZonedTimestamp columns write zero-conversion (one
zone and one resolution per column; mixed zones refuse with an
astimezone hint, and mixed resolutions promote to the finer unit).
- Scan literals lower zone-aware. Both new types are total,
order-preserving bijections with their storage, so every comparison
operator pushes down when the literal converts exactly; the zone-naive
and zone-declared literal domains never cross, and integer literals no
longer lower against timestamp columns (their public comparison is
false).
- ArrowTimeZonesExt keeps the fresh-ZonedDateTime write route and lowers
ZonedDateTime filter literals at every unit; its read hooks are gone,
so a zoned column's element type no longer depends on loaded packages.
Durations compat floor is 1.4.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Second and millisecond timestamp columns now materialize as
Timestamp{Second} and Timestamp{Millisecond} instead of DateTime,
matching the micro/nanosecond units: one rule, every unit, and the
conversion is a per-element reinterpret that cannot wrap or truncate.
Timestamp compares equal to DateTime at the same instant, so a written
DateTime column reads back as an equal Timestamp{Millisecond} column.
Dropping the DateTime read removes its Int64 wrap-aliasing, so every
comparison operator now lowers on every timestamp column (millisecond
was equality-only and second never lowered), and extreme second counts
no longer alias the epoch through the multiply-then-wrap conversion.
Date64 keeps its DateTime read and its equality-only lowering.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The facade's public value domain is now identical at every depth: convertible leaves (dates, times, zone-naive and zone-declared timestamps, durations, intervals, shared decimals) materialize as their public Julia values inside struct, list, map, and union rows and under run-end-encoded and dictionary wrappers — exactly as at the top level — and the writer accepts those public values back at every depth, fresh and retained. docs/dev/DESIGN-nested-public-values.md is the contract. - One recursive public row builder in the facade (`_publicvalue`); union branches convert at build time while the type id is in hand. - The two-pass ArrowTypes boundary splits labeled subtrees (raw, for restoration through the ArrowTypes contract) from label-free subtrees (public) via `underlabel`/`rawdomain` threading; a label with no registered target converts its raw leaves on the way into public rows. - `_declaredeltype(f, converted)` declares the public domain at every depth, so empty, all-missing, and populated columns agree. - Writer symmetry: retained write-back lowers nested public values through the same public-to-storage authority as the top level, and fresh composite children build through `_constructnativepart`; values that cannot lower exactly refuse. - Scan literal lowering is transparent through run-end encoding, like dictionary encoding; composite filters stay public. - The reader allocation budget covers the dynamic conversion work: per-value reserves at each conversion seam (facade converted leaf, raw-domain restoration, shared-decimal postconvert, routed-walker leaf estimate, registered fromarrow lifts split by storage shape, string-copy and isbits-payload reserves that scale with value size), allocation-free memo-cache hits and boxed-children caches in every per-row loop, an unboxed wide-decimal conversion barrier, and a nestedvalues testset pinning charge floors, refusals, and warmed allocation under the charge. Fixes twenty-four findings from twelve adversarial Codex review rounds (the twelfth returned CLEAN): union branch fidelity, unknown-label domain agreement, double conversion under wrappers, declared-type mismatches, and allocation-cap overshoot in every dynamic conversion path, including several budget gaps that predate this change on the registered-extension paths. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Resume append from the final dictionary updates, including replacements and deltas after the last record batch. Cover both file and IO append paths and verify older batch snapshots stay unchanged. AI disclosure: This work was prepared with assistance from OpenAI Codex.
Use job-scoped OIDC for repository-owned CI and nightly uploads. Preserve the public-fork tokenless path and all existing coverage files, matrices, and thresholds. AI disclosure: This work was prepared with assistance from OpenAI Codex.
Wait for one complete CI coverage run, with 16 Arrow and one ArrowTypes report. This also supports main commits and branches with no duplicate pull-request run; coverage percentage gates and CI waiting remain unchanged. AI disclosure: This work was prepared with assistance from OpenAI Codex.
Codecov Report❌ Patch coverage is
Additional details and impacted files@@ Coverage Diff @@
## main #609 +/- ##
==========================================
+ Coverage 87.43% 91.93% +4.49%
==========================================
Files 26 27 +1
Lines 3288 12367 +9079
==========================================
+ Hits 2875 11369 +8494
- Misses 413 998 +585 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
|
The trailing-dictionary append repair and coverage authentication changes are fully validated at The rewrite remains draft and unmerged. The scheduled nightly authentication change has not been exercised by a fresh scheduled run. AI disclosure: This work was prepared with assistance from OpenAI Codex. |
Summary
Arrow.appendfrom the final dictionary state even when an update follows the last record batch; fixesArrow.Tableuses last dictionary for all batches in streaming decode, resulting in incorrect reads #610Tables.Scanonce, keep storage and public-domain plans separate, preserve stable empty-result schemas, and support sparse byte-range reads with footer statisticsDurations.Timestamp{P}at the column unit (a writtenDateTimecolumn reads back as an instant-equalTimestamp{Millisecond}) and every timezone-declared timestamp asDurations.ZonedTimestamp{P,Z}(no TimeZones.jl requirement), both writing back zero-conversion with every comparison operator pushing down; preserve wire descriptors on rewriteCurrent validation
Latest head:
0ffdf1436664228629cbe619249011a415d216f7.Appending to a valid IPC stream with a trailing dictionary replacement or delta previously either wrote incorrect values or rejected a matching dictionary. The reader now retains the final dictionary map; append resumes that map while earlier record batches keep their original snapshots. Public file and IO regressions cover replacements, deltas, and uncompressed/LZ4/Zstd streams.
All 55 hosted checks pass at
0ffdf1436664228629cbe619249011a415d216f7: both CI runs and the conformance run completed successfully. All 34 native coverage uploads were accepted for apache/arrow-julia at that exact commit, and all 34 reports finished processing, with aggregate coverage at 91.93%. Prior green job logs show coverage uploads were rejected withToken required because branch is protected. Uploading jobs now use scoped OIDC for repository-owned runs and fail on upload errors; public forks retain tokenless upload support. Coverage percentage thresholds and processing paths are unchanged. The notification count now matches one 17-report CI run, so main commits and branches without a duplicate PR run can report coverage; all 34 uploads from this PR's push and PR runs were independently verified. The scheduled nightly upload path receives the same authentication repair; scheduled execution is not part of the local validation above.Dependency
JuliaData/Tables.jl#380 is merged and released as Tables.jl 1.14.0. Commit dd8c272 removes the temporary source and CI pins, raises the Tables.jl compatibility floor to 1.14, and resolves Tables from the General registry everywhere (package, CI workflows, conformance image, docs).
Breaking changes
This is the Arrow.jl 3.0 rewrite. Important changes include:
Arrow.TableandArrow.Streamreturn materialized Julia vectors instead of lazy Arrow vector wrappers.Arrow.writeis eager and whole-buffer.Arrow.Writerand IPC-stream append are reimplemented with fixed-schema semantics; the curried write andtobuffercompatibility forms remain. Obsolete writer tuning keywords are ignored with warnings.Arrow.write(io, table)writes file format by default; usefile=falsefor stream format.ArrowTypesstays exported only as an Arrow 2.x compatibility exception. Packages that define mappings should depend on andimport ArrowTypesdirectly.Arrow.ArrowTypesremains as a qualified compatibility binding.Arrow.Fresh declared heterogeneous Julia Union columns are written as canonical dense Arrow Unions. A retained Union rewrite fails closed when materialization has lost its original route. Nested retained dictionary pools follow the same fail-closed rule.
See
CHANGELOG.mdanddocs/src/migration.mdfor the complete list and migration guidance.Release preparation
Timestamp/ZonedTimestamp) resolve from General; the bundled ArrowStrings package and its release steps are removed3.0.0-DEVpending the rewrite merge and Apache release processCo-authored by Codex
AI disclosure: This work was prepared with assistance from OpenAI Codex.
🤖 Generated with Claude Code