Lattice and trellis-coded quantization for LLM weights — 2–8 bpw, mixed-precision, fused CUDA tensor-core inference on vLLM + HF
-
Updated
Sep 6, 2026 - Python
Lattice and trellis-coded quantization for LLM weights — 2–8 bpw, mixed-precision, fused CUDA tensor-core inference on vLLM + HF
💎 LTS Industrial Standard: PDF/Word optimization with scientific Trellis Mimic engine. Features Turbo Parallel processing & Global Camouflage (Ricoh/Fujitsu/Canon 2025 profiles). Embedded Python 3.12, zero-install, no admin needed. Ultimate document privacy for Windows LTSC/Enterprise.
To associate your repository with the trellis-quantization topic, visit your repo's landing page and select "manage topics."