DDIR: registered functions as corgi host kernels (and float math) - #906
Merged
Merged
Conversation
…mpares, fint) Adds the F64 operations a procedural-generation layer needs, spelled like the existing `fadd`/`fneg`: `f` + the Rust `f64` method name. - F64 -> F64: fabs, fsqrt, fexp, fln, ffloor, fceil, fround, fsin, fcos, ftan (`UnOp::F64Fn`), each the Rust method of the same name. - F64 -> Int: fint(x), Rust's `x as i64` (truncate toward zero, NaN is 0, saturating). - Binary: fpow (powf), fpowi (powi, Int exponent), fmin/fmax (skip a NaN operand, else total order, so signed zeros are deterministic), and the IEEE comparisons feq fne flt fle fgt fge returning Int 0/1. The generic `== < ...` were already right for negative F64 values: the payload is an order-preserving encoding, so they are `f64::total_cmp`. They differ from IEEE only on -0.0 vs 0.0 and NaN, which is what the f-comparisons are for. Corgi backend: fabs, fmin/fmax and the IEEE comparisons have columnar kernels built from existing corgi ops. The rest have no corgi kernel, so a term using one compiles to `Kernel::Rows`: the columns are untranscoded, `ir::eval` runs a row at a time, and the result is transcoded back. The output shape is still corgi's typing, of the term with each row-only op replaced by a same-shaped columnar stand-in. This applies to map, filter, flatmap, enter_at, and join projections. tests/programs/f64_math.ddp exercises every op, including negatives, both zeros, infinities and NaN, in a map, filter, flatmap and join, and is in the corgi-vs-vec gate. The tour gains a small float example. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
An embedding program registers a pure Rust function with `ir::register` (name, argument shapes, result shape, body), and any program parsed afterwards calls it like a builtin: `grow($0[0])`. The parser resolves a name that is not a builtin against the registry (a registered name may not shadow a builtin), producing `Term::Call`. - Row backend: `ir::eval` evaluates the arguments and calls the body. - Corgi backend: a call has no columnar kernel, so a term containing one runs through the row-at-a-time fallback (`Kernel::Rows`, from the float work). It is typed as a tuple of its arguments' stand-ins (so they still typecheck) projected to a literal of the declared result shape. The contract is purity: the same arguments must give the same value, since the dataflow re-evaluates terms to retract what they produced. tests/programs/registered.ddp calls three registered functions (a list-valued `grow`, a float-valued `blend`, a nested-shape `describe`) in a map, a filter, a flatmap, a join projection and a refinement fixpoint, and is in the corgi-vs-vec gate; the test also checks the fixpoint's contents. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
transcode/untranscode (the conversions around Kernel::Rows) cloned every value they read. They now take their input by value and move it: transcode_owned builds columns from owned rows, and untranscode moves out of the columns. Worldgen DDIR tour, one worker: 16.7 -> 12.6 s. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Term::Call compiles to NumOp::Host (corgi #38) over the tuple of its arguments, so the term around a call stays columnar instead of falling back to Kernel::Rows. A function's columnar body comes from ir::register_kernel; otherwise its row body runs behind RowKernel, which converts only the call's arguments and result. kernel_of keeps one Arc per function, so every call site shares one kernel and CSE can merge equal calls. Argument shapes are checked when the program is typed at install, by the host op. Pins corgi at the merge of #38. Worldgen DDIR tour at 4 workers: 8.6 -> 5.0 s with its six hottest rules as columnar kernels; a 1 km move 263 -> 60 ms. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…; review tidy - A call with no arguments passes its kernel a Unit over the anchor (ir::call_input): an empty product carries no row count, and corgi (#38) rejects it as a kernel input. register_kernel and RowKernel use the same input shape. - registered.ddp adds a zero-argument call (seven) and a function with a columnar kernel (double, registered with register_kernel); the gate checks both against the vec backend at 1-4 workers and over serializing channels. - Arc::clone for kernel handles; float-lowering locals renamed so no binding shadows an unrelated one. Clippy warnings for interactive: 165 -> 164. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The F64 functions that are not compositions of corgi ops (fsqrt, fexp, fln, ffloor, fceil, fround, fsin, fcos, ftan, fint, fpow, fpowi) are now built-in host kernels over the float column: decode each total-order key, apply the function as ir::eval does, encode. One kernel per function (CSE merges equal calls). With no row-only op left, Kernel, row_only, columnar_stand_in, literal_of_shape and Kernel::Rows go; lowering returns a Graph<NumOp> again, and the backend and join call eval_graph as on master-next. Review fixes: - register drops any kernel of the name being replaced (the row adapter held the old body; a register_kernel kernel could have the old shapes). - The row backend checks each call's arguments against the declared shapes (Value::has_shape), so a mis-shaped call fails on both backends. - Keywords (min, map, key, count, ...) cannot be registered: the lexer's table is one function, keyword(), and is_builtin consults it. - transcode asserts a tuple has exactly its shape's fields (was >=). - RowKernel's unreachable arm is gone; Term::Call and the float docs updated. Tests: every F64 op agrees with ir::eval on all special-value pairs (NaNs of both signs); a composite term over non-NaN pairs (corgi's float Add returns the positive NaN where ir::eval keeps an operand's sign, as fadd alone did before); re-registration; argument shapes on both backends. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
frankmcsherry
marked this pull request as ready for review
September 28, 2026 23:42
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Rust functions an embedding program registers by name, callable from DDIR terms, and run as corgi host kernels (frankmcsherry/WIP#38) on the corgi backend. Float math comes along because the registered-function work was built on it and the programs that use one use the other.
Commits
fsqrt,fexp,fln, rounding, trig,fpow/fpowi,fmin/fmax, IEEE compares,fint, over theF64newtype, in both backends. (In this commit the functions other thanfabs, the compares andfmin/fmaxran a row at a time on corgi; commit 6 makes them columnar.)ir::register(Function { name, args, result, body }); the parser resolves a non-builtin name against the registry (a registered name may not shadow a builtin) toTerm::Call. The row backend calls the body. The contract is purity.transcode/untranscodetake their input by value.NumOp::Host: the term around a call stays columnar. A function's columnar body comes fromir::register_kernel; otherwiseRowKernelruns the row body, converting only the call's arguments and result.kernel_ofkeeps oneArcper function, so CSE can merge equal calls. Argument shapes are checked when the program is typed at install, by the host op. Pins corgi at the merge of Implement indices for Traces #38.Unitover the anchor (corgi rejects an empty product as a kernel input, since it carries no row count); tests for a zero-argument call and a columnar kernel;Arc::clone; float-lowering locals renamed so nothing shadows an unrelated binding.ir::evaldoes, encode), one per function. With nothing row-only left,Kernel,row_only,columnar_stand_in,literal_of_shapeandKernel::Rowsare deleted and lowering returns aGraph<NumOp>again. Review fixes:registerdrops any kernel of the name it replaces; the row backend checks call arguments against the declared shapes (so both backends reject a mis-shaped call); keywords such asmincannot be registered (one keyword table, consulted byis_builtin);transcoderequires a tuple's exact field count.Tests
tests/programs/registered.ddpcalls five registered functions (list-valued, float-valued, nested-shape, zero-argument, and one with a columnar kernel) in a map, a filter, a flatmap, a join projection and a refinement fixpoint. The corgi-vs-vec gate runs it at 1–4 workers and over serializing channels.ir::evalon all pairs of special values (signed zeros, infinities, NaNs of both signs); a composite term over non-NaN pairs. One pre-existing difference, noted in the test: corgi's floatAddreturns the positive NaN whereir::evalkeeps a NaN operand's sign.mincannot be registered.cargo test --release -p interactive: 110 passed, 5 ignored. Clippy warnings forinteractive: 165 → 164.Float functions, before and after commit 6 (a projection
($0 ; f($0))over 2^20 F64 rows, lowered and run on one core, median of 7, ns per row):fsqrt152 → 1.9,fexp155 → 3.1,fln157 → 3.3,ffloor155 → 1.7,fsin158 → 3.6,fint130 → 1.1,fpow178 → 5.6; a mixed termtuple(fadd(fsqrt($0), $0), flt($0, $0))264 → 13.fabs(already columnar) 1.5 → 1.7, unchanged. The old cost was the row fallback converting the whole term's input to rows and back, not the math.Measured on worldgen's DDIR program (33 registered rules, 6 as columnar kernels; 27/27 tour steps agree with the native engine):
Not in this PR (follow-ups): columnar exports and export taps; a check that join key shapes agree (a mismatch silently yields nothing today); running
corgi::cseon lowered graphs; resolving a call's function once at lowering instead of a registry lookup per row on the row backend; whether the process-wide registry should become one passed with the program.🤖 Generated with Claude Code