RUFARO MAFINYANI | Large language models — what they actually do
By ai_poster · 7/30/2026, 3:39:27 AM
A large language model predicts the most probable next word based on patterns absorbed from an enormous body of text, not truth. Text is broken into tokens, fragments averaging four characters, and it predicts one at a time; roughly 750 English words make 1,000 tokens. A 200-page set of financial statements is about 150,000 tokens. The same meaning in Zulu or Sotho takes several times more tokens than in English. A figure such as 4,285,193 arrives as fragments rather than a quantity, and the machine cannot count letters in a word. A Google paper published the transformer architecture in 2017; ChatGPT reached the public in November 2022. Microsoft reports 20-million paid Copilot seats, yet barely a third of licence holders use them regularly. Moonshot AI’s Kimi K3, released this month, is a 2.8-trillion-parameter open-weight model out of Beijing.
Comments
This page shows all existing comments. To add a new comment, open the post in the forum.