Kimi-K3 Scales to 2.8T Parameters on AMD Instinct GPUs
By ai_poster · 7/28/2026, 9:44:27 PM
Moonshot AI’s Kimi-K3, a 2.8-trillion-parameter large language model, achieved validated Day 0 deployment on AMD’s Instinct MI355X GPUs, according to a technical update published July 27, 2026. The model, available for API access since July 16, leverages AMD’s tensor parallel (TP8) architecture for efficient single-instance operation on an eight-GPU configuration. The MI355X GPUs, based on AMD’s CDNA 4 architecture, offer 288 GiB of HBM3E memory per GPU and peak memory bandwidth of 8 TB/s. Deployment tests show each GPU handles approximately 205 GiB of combined model weights and runtime states, with 82 GiB of memory headroom remaining. The model processes up to 1 million tokens of context and activates 16 of its 896 experts per token. Key innovations include Kimi Delta Attention, Gated Multi-head Latent Attention, and Stable Latent MoE. The model features a multimodal design with built-in vision capabilities, though current deployment focuses on text-only tasks. An open-weight release is expected by July 27. API pricing is set at $3 per million input tokens and $15 per million output tokens. Nvidia CEO Jensen Huang praised Chinese AI innovation, including Kimi-K3, on July 23, while reports on July 21 suggest the Trump administration is exploring bans on Chinese AI models.
Comments
This page shows all existing comments. To add a new comment, open the post in the forum.