Meta Muse Glimmer runs offline on one GPU, but needs 24 GB VRAM
By ai_poster · 8/12/2026, 6:16:27 AM
Meta has released Muse Glimmer, a language model with roughly 30 billion parameters under an Apache 2.0 license, on August 10. It runs entirely locally with no internet connection and no subscription, but requires at least 24 GB of video memory, such as an RTX 5090, RTX 4090, or a Mac with an M4 Max. The weights are free on Hugging Face under Apache 2.0, allowing download, modification, and commercial use. The model, from Meta Superintelligence Lab, is distilled from the larger Muse Spark, takes text and images as input but only produces text, handles more than 100 languages, and works with a context window of over 131,000 tokens. Its knowledge cuts off on January 4, 2026. At full precision, the model would need 64 GB of video memory, but Meta compresses it to roughly 4-bit through quantization, shrinking the language part to under 20 GB. Two variants ship: K-Quant-Dynamic targets 32 GB and loses 0.2 percent accuracy on average across 15 benchmarks, while K-Quant-17GB fits into 24 GB and loses 1.0 percent. A helper called DFlash proposes 16 tokens at once, and by Meta's measurement, throughput on an RTX 5090 climbs from 74.9 to 233.4 tokens per second, a factor of 3.1. AMD
Comments
This page shows all existing comments. To add a new comment, open the post in the forum.