Skip to content

DDIR: registered functions as corgi host kernels (and float math) - #906

Merged
frankmcsherry merged 6 commits into
master-nextfrom
ddir-host-calls
Sep 28, 2026
Merged

frankmcsherry merged 6 commits into
master-nextfrom
ddir-host-calls

Conversation

@frankmcsherry

@frankmcsherry frankmcsherry commented Sep 27, 2026 •

Copy link
Copy Markdown
Member

Rust functions an embedding program registers by name, callable from DDIR terms, and run as corgi host kernels (frankmcsherry/WIP#38) on the corgi backend. Float math comes along because the registered-function work was built on it and the programs that use one use the other.

Commits

  1. Floating-point math: fsqrt, fexp, fln, rounding, trig, fpow/fpowi, fmin/fmax, IEEE compares, fint, over the F64 newtype, in both backends. (In this commit the functions other than fabs, the compares and fmin/fmax ran a row at a time on corgi; commit 6 makes them columnar.)
  2. Registered functions: ir::register(Function { name, args, result, body }); the parser resolves a non-builtin name against the registry (a registered name may not shadow a builtin) to Term::Call. The row backend calls the body. The contract is purity.
  3. Move, don't clone, in the corgi row fallback: transcode/untranscode take their input by value.
  4. Calls lower to NumOp::Host: the term around a call stays columnar. A function's columnar body comes from ir::register_kernel; otherwise RowKernel runs the row body, converting only the call's arguments and result. kernel_of keeps one Arc per function, so CSE can merge equal calls. Argument shapes are checked when the program is typed at install, by the host op. Pins corgi at the merge of Implement indices for Traces #38.
  5. Review follow-ups: a no-argument call passes a Unit over the anchor (corgi rejects an empty product as a kernel input, since it carries no row count); tests for a zero-argument call and a columnar kernel; Arc::clone; float-lowering locals renamed so nothing shadows an unrelated binding.
  6. Float functions as host kernels; the row fallback goes. The F64 functions that are not compositions of corgi ops are built-in host kernels over the float column (decode the order key, apply as ir::eval does, encode), one per function. With nothing row-only left, Kernel, row_only, columnar_stand_in, literal_of_shape and Kernel::Rows are deleted and lowering returns a Graph<NumOp> again. Review fixes: register drops any kernel of the name it replaces; the row backend checks call arguments against the declared shapes (so both backends reject a mis-shaped call); keywords such as min cannot be registered (one keyword table, consulted by is_builtin); transcode requires a tuple's exact field count.

Tests

  • tests/programs/registered.ddp calls five registered functions (list-valued, float-valued, nested-shape, zero-argument, and one with a columnar kernel) in a map, a filter, a flatmap, a join projection and a refinement fixpoint. The corgi-vs-vec gate runs it at 1–4 workers and over serializing channels.
  • Every F64 op agrees with ir::eval on all pairs of special values (signed zeros, infinities, NaNs of both signs); a composite term over non-NaN pairs. One pre-existing difference, noted in the test: corgi's float Add returns the positive NaN where ir::eval keeps a NaN operand's sign.
  • Re-registration drops the old kernel; a mis-shaped call fails on both backends; min cannot be registered.
  • cargo test --release -p interactive: 110 passed, 5 ignored. Clippy warnings for interactive: 165 → 164.

Float functions, before and after commit 6 (a projection ($0 ; f($0)) over 2^20 F64 rows, lowered and run on one core, median of 7, ns per row): fsqrt 152 → 1.9, fexp 155 → 3.1, fln 157 → 3.3, ffloor 155 → 1.7, fsin 158 → 3.6, fint 130 → 1.1, fpow 178 → 5.6; a mixed term tuple(fadd(fsqrt($0), $0), flt($0, $0)) 264 → 13. fabs (already columnar) 1.5 → 1.7, unchanged. The old cost was the row fallback converting the whole term's input to rows and back, not the math.

Measured on worldgen's DDIR program (33 registered rules, 6 as columnar kernels; 27/27 tour steps agree with the native engine):

  • One worker, tour: 16.7 s → 12.6 s from commit 3.
  • Four workers, tour: 8.6 s → 5.0 s, and a 1 km move 263 → 60 ms, from commit 4 with the columnar rules.

Not in this PR (follow-ups): columnar exports and export taps; a check that join key shapes agree (a mismatch silently yields nothing today); running corgi::cse on lowered graphs; resolving a call's function once at lowering instead of a registry lookup per row on the row backend; whether the process-wide registry should become one passed with the program.

🤖 Generated with Claude Code

frankmcsherry and others added 6 commits September 27, 2026 09:10
…mpares, fint)

Adds the F64 operations a procedural-generation layer needs, spelled like the
existing `fadd`/`fneg`: `f` + the Rust `f64` method name.

- F64 -> F64: fabs, fsqrt, fexp, fln, ffloor, fceil, fround, fsin, fcos, ftan
  (`UnOp::F64Fn`), each the Rust method of the same name.
- F64 -> Int: fint(x), Rust's `x as i64` (truncate toward zero, NaN is 0,
  saturating).
- Binary: fpow (powf), fpowi (powi, Int exponent), fmin/fmax (skip a NaN
  operand, else total order, so signed zeros are deterministic), and the IEEE
  comparisons feq fne flt fle fgt fge returning Int 0/1.

The generic `== < ...` were already right for negative F64 values: the payload
is an order-preserving encoding, so they are `f64::total_cmp`. They differ from
IEEE only on -0.0 vs 0.0 and NaN, which is what the f-comparisons are for.

Corgi backend: fabs, fmin/fmax and the IEEE comparisons have columnar kernels
built from existing corgi ops. The rest have no corgi kernel, so a term using
one compiles to `Kernel::Rows`: the columns are untranscoded, `ir::eval` runs a
row at a time, and the result is transcoded back. The output shape is still
corgi's typing, of the term with each row-only op replaced by a same-shaped
columnar stand-in. This applies to map, filter, flatmap, enter_at, and join
projections.

tests/programs/f64_math.ddp exercises every op, including negatives, both
zeros, infinities and NaN, in a map, filter, flatmap and join, and is in the
corgi-vs-vec gate. The tour gains a small float example.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
An embedding program registers a pure Rust function with `ir::register`
(name, argument shapes, result shape, body), and any program parsed afterwards
calls it like a builtin: `grow($0[0])`. The parser resolves a name that is not
a builtin against the registry (a registered name may not shadow a builtin),
producing `Term::Call`.

- Row backend: `ir::eval` evaluates the arguments and calls the body.
- Corgi backend: a call has no columnar kernel, so a term containing one runs
  through the row-at-a-time fallback (`Kernel::Rows`, from the float work). It
  is typed as a tuple of its arguments' stand-ins (so they still typecheck)
  projected to a literal of the declared result shape.

The contract is purity: the same arguments must give the same value, since the
dataflow re-evaluates terms to retract what they produced.

tests/programs/registered.ddp calls three registered functions (a list-valued
`grow`, a float-valued `blend`, a nested-shape `describe`) in a map, a filter, a
flatmap, a join projection and a refinement fixpoint, and is in the corgi-vs-vec
gate; the test also checks the fixpoint's contents.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
transcode/untranscode (the conversions around Kernel::Rows) cloned every value
they read. They now take their input by value and move it: transcode_owned
builds columns from owned rows, and untranscode moves out of the columns.
Worldgen DDIR tour, one worker: 16.7 -> 12.6 s.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Term::Call compiles to NumOp::Host (corgi #38) over the tuple of its arguments,
so the term around a call stays columnar instead of falling back to
Kernel::Rows. A function's columnar body comes from ir::register_kernel;
otherwise its row body runs behind RowKernel, which converts only the call's
arguments and result. kernel_of keeps one Arc per function, so every call site
shares one kernel and CSE can merge equal calls. Argument shapes are checked
when the program is typed at install, by the host op. Pins corgi at the merge
of #38.

Worldgen DDIR tour at 4 workers: 8.6 -> 5.0 s with its six hottest rules as
columnar kernels; a 1 km move 263 -> 60 ms.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…; review tidy

- A call with no arguments passes its kernel a Unit over the anchor
  (ir::call_input): an empty product carries no row count, and corgi (#38)
  rejects it as a kernel input. register_kernel and RowKernel use the same
  input shape.
- registered.ddp adds a zero-argument call (seven) and a function with a
  columnar kernel (double, registered with register_kernel); the gate checks
  both against the vec backend at 1-4 workers and over serializing channels.
- Arc::clone for kernel handles; float-lowering locals renamed so no binding
  shadows an unrelated one. Clippy warnings for interactive: 165 -> 164.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The F64 functions that are not compositions of corgi ops (fsqrt, fexp, fln,
ffloor, fceil, fround, fsin, fcos, ftan, fint, fpow, fpowi) are now built-in
host kernels over the float column: decode each total-order key, apply the
function as ir::eval does, encode. One kernel per function (CSE merges equal
calls). With no row-only op left, Kernel, row_only, columnar_stand_in,
literal_of_shape and Kernel::Rows go; lowering returns a Graph<NumOp> again, and
the backend and join call eval_graph as on master-next.

Review fixes:
- register drops any kernel of the name being replaced (the row adapter held
  the old body; a register_kernel kernel could have the old shapes).
- The row backend checks each call's arguments against the declared shapes
  (Value::has_shape), so a mis-shaped call fails on both backends.
- Keywords (min, map, key, count, ...) cannot be registered: the lexer's table
  is one function, keyword(), and is_builtin consults it.
- transcode asserts a tuple has exactly its shape's fields (was >=).
- RowKernel's unreachable arm is gone; Term::Call and the float docs updated.

Tests: every F64 op agrees with ir::eval on all special-value pairs (NaNs of
both signs); a composite term over non-NaN pairs (corgi's float Add returns the
positive NaN where ir::eval keeps an operand's sign, as fadd alone did before);
re-registration; argument shapes on both backends.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@frankmcsherry
frankmcsherry marked this pull request as ready for review September 28, 2026 23:42
@frankmcsherry
frankmcsherry merged commit d1a80d9 into master-next Sep 28, 2026
6 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant