Skip to content
 
 

Repository files navigation

FreeToken

| Download | Paper | Developer Slack | Community Discord | Community WeChat |

Key Additions in this Fork:

  1. TurboQuant (tq4) 4-Bit KV Cache Compression:

    • True 4-bit nibble-packed KV cache storage (0.5 B/elem, 4.00× memory savings vs fp16).
    • Orthogonal Walsh-Hadamard Transform (WHT) + 16-level Lloyd-Max optimal scalar quantizer with custom CUDA JIT store/dequant kernels (quant_store.cu / quant_dequant.cu).
    • Bit-matched reconstruction with MAE < 0.006 vs torch reference oracle.
    • Enabled via --kv-quant tq4.
  2. Multi-Shard GGUF, Qwen3 / Qwen3.5 / Qwen3.6 Families, & Every GGML Quant Type:

    • Supports reading multi-shard GGUF checkpoints (split-00001-of-000NN) across directories or direct shard paths.
    • Adds qwen3moe, qwen35moe, and qwen35 dense architectures.
    • Support for all 21 GGML quant types (including standard, K-quants, and I-quants).

Verified on my hardware

RTX 4060 Laptop, 8GB VRAM, 64GB system RAM, routed experts offloaded to host memory.

Model Arch Quant Result
Ornith-1.5-35B-A3B qwen35moe IQ3_S 8/8 factual, 47 to 50 tok/s
Ornith-1.5-35B-A3B qwen35moe IQ3_XXS 6/6 factual, 50 to 52 tok/s
Ornith-1.5-35B-A3B split into 3 shards qwen35moe IQ3_S 6/6 factual, 44 to 46 tok/s
Qwen3-30B-A3B qwen3moe IQ4_XS 6/6 factual, 45 to 47 tok/s

A 35B model with 3B active, in 16GB, decoding at 50 tok/s on a laptop GPU with 8GB of VRAM. For reference, llama.cpp on the same file on CPU does 11 tok/s.

Supported by architecture

These carry a general.architecture this fork now handles. I have not run all of them, so this is "the loader covers it", not "I benchmarked it". Bank column is from reading each file's tensor table.

Model Arch Expert banks Notes
Qwen3-235B-A22B qwen3moe check per quant split, 3 to 10 shards
Qwen3.5-122B-A10B qwen35moe uniform at IQ3_S should load as is
Qwen3.6-35B-A3B qwen35moe mixed at IQ3_S needs a --pure quant
Ornith-1.0-35B-AEON qwen35moe mixed at Q4_K_M needs a --pure quant
Qwen3.8-27B qwen35 dense none Q4_K_M, test in progress
Qwen3.6-27B qwen35 dense none Q4_K_M
Qwen3.5-27B qwen35 dense none Q4_K_M
Qwen3.5-9B qwen35 dense none Q4_K_M
Qwen3.5-4B qwen35 dense none Q4_K_M

Dense models have no expert banks, so the uniform-bank rule below does not apply to them and ordinary _M quants are fine.

More of my quants at huggingface.co/vcruz305.

Quant types

All 21 ggml types are now described on the Python side. csrc/gguf/ already dispatched 19 of them, but the tables only listed 6, so K-quants and I-quants were unreachable for every architecture rather than just this one.

Family Types Prefill Decode
Standard Q4_0, Q4_1, Q5_0, Q5_1, Q8_0 MMQ MMVQ
K-quants Q2_K, Q3_K, Q4_K, Q5_K, Q6_K MMQ MMVQ
I-quants IQ1_S, IQ1_M, IQ2_XXS, IQ2_XS, IQ2_S, IQ3_XXS, IQ3_S, IQ4_NL, IQ4_XS dequant plus matmul MMVQ

I-quants have no MMQ kernel upstream, so prefill falls back to ggml_dequantize and a torch matmul. That branch already existed but was unreachable, because _MMQ was tested before _DEQUANT and the two sets were identical.

What else changed

  • Fixed a silent data corruption bug. None of the five switch (type) blocks in gguf_kernel.cu had a default:, and the output tensor is allocated with torch::empty, so an unsupported quant type returned uninitialized memory instead of raising. ggml_moe_get_block_size returned 0. That one is worth having on its own, independent of the rest of this fork.
  • Documented the four transforms llama.cpp's converter applies that a loader has to undo: ssm_a holds A rather than A_log, the (1+w) norm shift is already folded in, V heads are stored tiled rather than grouped when there are fewer K heads than V heads, and merged projections are not uniformly typed. None of these are visible to shape, dtype or byte identity checks. The model loads, runs at full speed, and produces fluent nonsense.
  • Tests derive the block sizes from ggml-common.h and extract the case labels from gguf_kernel.cu at runtime, so the Python tables cannot drift from the kernels without a test failing.

Known limits

  • MoE expert banks have to use one ggml type across every layer. The GPU slot pool is a single allocation and moe_vec.cuh indexes it as expert * nrows * (ncols / qk) with no padding allowance, so two row strides in one pool would read every block at the wrong offset. llama.cpp's _M and _XXS levels raise the precision of the first few layers' ffn_down_exps, which trips this. Those refuse to load with an error naming the layers. Quantize with llama-quantize --pure to get one type throughout. Dense models are unaffected.
  • TP=1 only, same as the existing Gemma-4 GGUF path.
  • The NextN/MTP block is dropped, so no speculative decoding.
  • Expert bank loading is serial, so first load of a 16GB checkpoint takes about a minute. Converting to FTW with ft checkpoint avoids paying it every time.
  • Tested with the flashinfer attention backend. The triton fallback is unexercised here.
  • First token argmax matches llama.cpp CPU on 2 of 6 single token prompts. Every disagreement is a plausible near tie, and this runs CUDA W4A8 MMVQ against CPU AVX, so I do not think exact agreement is reachable. Stating the number rather than implying it is bit exact.

Unlock datacenter-class intelligence on the hardware you already own — Run 290B+ frontier MoE models locally on your gaming PC at blistering interactive speeds.

About

FreeToken is an edge-native Mixture-of-Experts (MoE) serving engine designed for running frontier-scale open-weight models on personal and consumer hardware. It treats heterogeneous edge resources—GPUs, CPUs, host memory, and interconnects—as a unified, elastic inference platform. Its core features include:

  • Fast Edge-Native Runtime: Provides efficient MoE serving with bandwidth-adaptive CPU–GPU co-execution ($q^\star$ policy), full-layer double-buffered prefill streaming, global LRU expert caching, graph-compatible execution, and the FTW fast weight format.
  • Semantic-Aware Caching: Features semantic anchor checkpoints for recurrent state and KV caches, allowing agentic context edits (e.g., tool calls, thinking blocks) to avoid redundant context recomputation.
  • Elastic Memory Management: Supports dynamic, runtime VRAM re-allocation between expert caches and KV memory without engine restarts or weight reloading.
  • Broad MoE & Ecosystem Support: Supports frontier open-weight MoE models (e.g., DeepSeek-V4-Flash, Qwen3.6-35B-A3B, GLM-5.2) across various parameter scales and quantization formats (e.g., MXFP4, NVFP4, FP8, BF16), with Anthropic/OpenAI-compatible APIs for seamless integration with real-world coding and tool-calling agents (e.g., Codex, Claude Code, OpenCode, OpenClaw, DeepSeek Harness).
  • Diverse Consumer Hardware: Scales across consumer laptops, gaming desktops, and workstation GPUs, with native support for NVIDIA RTX 30, RTX 40, and RTX 50 series GPUs.

Getting Started

Desktop app

Download FreeToken for Windows or Linux at flashml.ai. It sets the engine up for you and gives you a GUI for running models, chatting, and tuning the engine.

FreeToken Desktop

CLI

Install FreeToken with uv (recommended) or pip:

uv pip install "freetoken[accel]"

Or build from source:

git clone https://github.com/FlashML-org/FreeToken.git && cd FreeToken
uv venv && source .venv/bin/activate
uv pip install -e ".[accel]"

For More details:

Citation

If you use FreeToken for your research, please cite our paper:

@article{yang2026freetoken,
  title={FreeToken: Efficient Edge-Native MoE Serving with Bandwidth-Adaptive Execution},
  author={Yang, Shuo and Fan, Xiaoze and Pan, Melissa and Xi, Haocheng and Wang, Zhe and Sun, Shanlin and Keutzer, Kurt and Han, Song and Zaharia, Matei and Xu, Chenfeng and Stoica, Ion},
  journal={arXiv preprint arXiv:2608.16157},
  year={2026}
}

Acknowledgment

FreeToken was deeply inspired by mini-sglang, and learned the design and reused code from the following projects: SGLang, vLLM, FlashInfer, flash-linear-attention, LightLLM and llama.cpp.

License

Apache License 2.0.

About

No description, website, or topics provided.

Resources

Contributing

Stars

3 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages