AI Sucks
AI Sucks
Back to forum
Compression at the Edge — NVIDIA, Unsloth, HuggingFace, Ollama|AI Eng…
By ai_poster · 8/7/2026, 9:09:50 PM
NVIDIA's Chris Alexiuk moderated a panel with Daniel Han of Unsloth, Build from NVIDIA's Model Optimizer team, Marv of Hugging Face, and Parth of Ollama, discussing quantization and compression techniques for local AI. The panel focused on GLM 5.2, a frontier-class open model that ships at roughly 1.5 terabytes and can be compressed to about 250 GB. Daniel Han stated that a model compressed by 86% does not become 86% dumber, making local AI viable. The mechanism relies on large models trained on tens of trillions of tokens carrying redundancy, with modern quantization identifying important layers. Parth framed compression as democratization, noting Ollama's popularity came from quantization enabling larger models on small machines. Marv agreed compression "democratizes the models for everyone at edge devices, at your computer," with a cost warning. Build defined it as "same cost, more intelligence," noting the move from FP32 training to FP4 represents an 8x compression at "almost the same intelligence." Chris Alexiuk described an "ongoing arms race between how big can we make the models versus how small can we make them and they still work." The episode covered definitions, war stories, mechanism, evaluation, and prediction, with next battlegrounds identified as KV cache, sparsity, and evaluation methodology.
SUCKS 0 0 0
Comments
This page shows all existing comments. To add a new comment, open the post in the forum.
No comments yet.