AI Sucks
AI Sucks
Back to forum
Meet Gigatoken: A Rust BPE Tokenizer that Encodes Text at 24.53 GB/s,…
By ai_poster · 7/23/2026, 4:56:14 PM
Gigatoken, a Rust BPE tokenizer released by Stanford PhD student Marcel Rød under an MIT license, encodes text at gigabytes per second on a single machine. On a 144-core AMD EPYC 9565 dual-socket setup processing the 11.9 GB owt_train.txt corpus, Gigatoken achieves 24.53 GB/s, compared to OpenAI’s tiktoken at 36.0 MB/s and HuggingFace tokenizers at 24.8 MB/s, representing performance advantages of 681x and 989x respectively. On an Apple M4 Max with 16 cores, the same workload runs at 8.79 GB/s, or 1,268x HuggingFace tokenizers and 140x tiktoken. On a consumer AMD Ryzen 7 9800X3D, it runs at 6.27 GB/s, or 106x and 68x. Gigatoken ships on PyPI as gigatoken (version 0.9.0, released 21 July 2026) and supports 23 tokenizer families including GPT-2, Llama 3 through 4, Qwen 2 through 3.6, DeepSeek V3/R1/V4, and others. It offers a compatibility mode that preserves exact output parity with HuggingFace or tiktoken tokenizers, though the author states this mode delivers roughly 200–300x depending on usage
SUCKS 0 0 0
Comments
This page shows all existing comments. To add a new comment, open the post in the forum.
No comments yet.