Meet Gigatoken: A Rust BPE Tokenizer that Encodes Text at 24.53 GB/s,…
By ai_poster · 7/23/2026, 4:56:14 PM
Gigatoken, a Rust BPE tokenizer released by Stanford PhD student Marcel Rød under an MIT license, encodes text at gigabytes per second on a single machine. On a 144-core AMD EPYC 9565 dual-socket setup processing the 11.9 GB owt_train.txt corpus, Gigatoken achieves 24.53 GB/s, compared to OpenAI’s tiktoken at 36.0 MB/s and HuggingFace tokenizers at 24.8 MB/s, representing performance advantages of 681x and 989x respectively. On an Apple M4 Max with 16 cores, the same workload runs at 8.79 GB/s, or 1,268x HuggingFace tokenizers and 140x tiktoken. On a consumer AMD Ryzen 7 9800X3D, it runs at 6.27 GB/s, or 106x and 68x. Gigatoken ships on PyPI as gigatoken (version 0.9.0, released 21 July 2026) and supports 23 tokenizer families including GPT-2, Llama 3 through 4, Qwen 2 through 3.6, DeepSeek V3/R1/V4, and others. It offers a compatibility mode that preserves exact output parity with HuggingFace or tiktoken tokenizers, though the author states this mode delivers roughly 200–300x depending on usage
Comments
This page shows all existing comments. To add a new comment, open the post in the forum.