RightNow
Popular repositories Loading
-
autokernel
autokernel PublicAutoresearch for GPU kernels. Give it any PyTorch model, go to sleep, wake up to optimized Triton kernels.
-
AutoMegaKernel
AutoMegaKernel PublicAn agent harness that compiles a model into one provably-correct, self-retargeting CUDA megakernel and self-tunes it past cuBLAS at batch-1 LLM decode, paper: https://arxiv.org/abs/2606.09682
-
qwen3.5-triton
qwen3.5-triton PublicPure Triton kernels for Qwen3.5-27B inference on NVIDIA B200
Repositories
- runinfra-cli Public
-
- local-kimi Public
Optimized local serving engine for Kimi-Linear-48B: INT4 quantizer, fused decode kernels for a measured 3.18x, and an OpenAI-compatible server. Ships with k3, a bridge that detects the client per request so Claude Code, Codex, Cline, Aider and opencode all work unchanged.
- inference-cost-truth Public
What LLM inference actually costs. 874 verified price rows across 24 providers, 378 GPU rental rates, 395 cited throughput datapoints, and a self-hosting break-even model. Every number carries a source URL, a retrieval date, and passed a mechanical evidence gate. Dated immutable snapshots. CC BY 4.0.
- inkling-turbo Public
Faster attention kernels for serving TML's Inkling model on vLLM. 2.7x over the shipping path on H100, and the only implementation that runs on A100.
Most used topics
Loading…