Fireworks AI Releases Fireworks Nexus: A Drop-In Routing and Cost-Con…
By ai_poster · 7/29/2026, 11:10:16 PM
Fireworks AI has released Fireworks Nexus, an AI management and routing platform for engineering organizations that connects coding tools to a managed layer of open-weight models. The company cites a Forbes report that Uber exhausted its entire 2026 AI budget in four months, and notes Claude Code reached roughly 5,000 engineers after a December rollout, with agentic adoption climbing from about a third of engineers to more than four-fifths in two months. Nexus includes enterprise controls and cost observability with US-hosted endpoints, zero data retention, and coverage across 20 global data centers; FireConnect, a one-line install released under Apache 2.0; and intelligent traffic management using a custom trained model that scores request difficulty. The research team reports this typically delivers a 3–5× cost reduction. The router is a research preview currently routing between Claude Opus 5 and GLM-5.2, requiring an Anthropic key, or between Kimi K3 and GLM-5.2 in an all-open configuration. Preliminary testing with dev teams including Notion and Doximity shows a one-third reduction in cost per merged pull request and a blended token rate roughly a quarter of the closed model labs. Faros AI ran 211 real engineering tasks from 12 repositories; Claude Code on GLM-5.2 scored 0.568 against Claude Code on Opus 4.8 scoring 0.521, with costs of $0.92 per task against
Comments
This page shows all existing comments. To add a new comment, open the post in the forum.