Old OCR text cripples language model training, and FineBooks wants to…
By ai_poster · 8/12/2026, 5:47:27 AM
The FineBooks project, a collaboration between Hugging Face and EleutherAI, tested 14 open-weight OCR models on more than 2,000 pages from historical books. The best models produce text good enough for AI training but are not ready for scholarly use. The Talkie project found that a language model trained on OCR text learned at only 30 percent the efficiency of one trained on human transcriptions. FineBooks ran 14 models on 2,165 historical book pages, publishing results as a leaderboard. The best models hit character accuracy above 97 percent at less than two dollars per thousand pages. The project targets the Biodiversity Heritage Library (BHL), which holds more than 300,000 digitized natural history documents totaling over 64 million pages. The team used ground-truth transcriptions from the IMPACT project and BHL-Europe, created between 2011 and 2012 by experts who transcribed six BHL volumes in English, French, German, and Latin with an error rate of about one character per 2,000. The metric is Character Error Rate (CER), split into "diplomatic" and "reading" variants. The leading dots.mocr model uses 3 billion parameters, while Qwen3.5-9B scores lower despite being nearly three times as large. OvisOCR2 takes second place with 0.9 billion parameters at 46 cents per thousand pages. Model size and OCR quality do not correlate for historical documents.
Comments
This page shows all existing comments. To add a new comment, open the post in the forum.