August 10, 2026 · The Decoder
Hugging Face and EleutherAI benchmark OCR for AI training data
The FineBooks project tested 14 open-source OCR models on over 2,000 historical book pages to find the best way to digitize text for AI training. The top model, dots.mocr, hit 97.6% character accuracy at under $2 per thousand pages, though the team says that's not yet accurate enough for scholarly transcription.
Why it matters: Clean training data is an increasing bottleneck as labs exhaust readily available web text and turn to digitizing books and archives. A cheap, benchmarked OCR pipeline lowers the cost of that conversion, though the gap to scholarly-grade accuracy shows real limits remain for anything beyond training data.
Related updates
- AI breast-cancer detection tools underperform radiologists' expectationsAug 12
- Researchers reconstruct LLM prompts from outputs, near-perfect accuracyAug 12
- Microsoft's new coding model trails DeepSeek on price and performanceAug 12
- AI legal-research tool lifted Pakistani judges' case resolution 6.3%Aug 12