Skip to content

d1-omni-600M on Core ML: runtime, model store, moderation firehose demo - #27

Open
Alex-Wengg wants to merge 1 commit into
mainfrom
feat/d1-omni-moderation
Open

Alex-Wengg wants to merge 1 commit into
mainfrom
feat/d1-omni-moderation

Conversation

@Alex-Wengg

Copy link
Copy Markdown
Member

Adds Liquid AI's d1-omni-600M (text path) on Core ML, plus a comment-moderation demo that runs the Neural Engine and GPU at the same time.

  • D1OmniManager + D1OmniTokenizer: upstream prompt encoding, calibration temperatures, batched calls; one multifunction package (8 shapes, shared weights, 729 MB).
  • D1OmniModelStore: pins FluidInference/d1-omni-600m-coreml @ 8ae382b1 with SHA-256 checks.
  • ModerationDemo / ModerationCheck: 5,000 Civil Comments test comments (CC0); 0/5000 flags differ from PyTorch, 95.4% agreement with human labels, ~240–300 comments/s with Neural Engine + GPU on an M5 Pro.

The Neural Engine's fused silu is ~1.4% off and flipped decisions vs PyTorch; the package uses the tanh form of SiLU instead (details in the commit message). macOS 15 / iOS 18. Weights are under the LFM Open License v1.0.

🤖 Generated with Claude Code

D1OmniManager runs the text path of Liquid AI's d1-omni-600M (LFM2.5 encoder
trunk + decision head) from one multifunction Core ML package
(L{64,128,256}_K2_B{1,8}, L{128,256}_K8_B1; weights shared, 729 MB):
upstream prompt encoding, byte-level BPE tokenizer (D1OmniTokenizer),
calibration temperatures, batched calls. D1OmniModelStore pins
FluidInference/d1-omni-600m-coreml @ 8ae382b1 with SHA-256 checks.

The Neural Engine's fused silu is ~1.4% off; across 16 SwiGLU MLPs that
flipped 41/492 Snake decisions vs PyTorch. The package computes SiLU as
x*0.5*(1+tanh(x/2)) (same function): 2/492 flips, max dp 0.017. 1,140 of
1,150 ops run on the Neural Engine.

ModerationDemo streams 5,000 Civil Comments test comments (CC0, clear
labels) through a Neural Engine worker (1 per call) and a GPU worker (8 per
call) sharing one queue. ModerationCheck: 0/5000 flags differ from
PyTorch, 95.4% agreement with human labels; ~100/s Neural Engine only,
~180-215/s GPU only, ~240-300/s together on an M5 Pro.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant