"Adaptive Hybrid Quantization Framework for deploying 7B+ LLMs on low-VRAM devices (e.g., GTX 1050). Features surgical block alignment and Numba-accelerated inference.
Python 27 1
Trion Core
Python 15 2
This organization has no public members. You must be a member to see who’s a part of this organization.
Loading…
There was an error while loading. Please reload this page.