---
title: "Hugging Face and EleutherAI benchmark OCR for AI training data"
url: https://www.parallelquant.com/posts/hugging-face-and-eleutherai-benchmark-ocr-for-ai-training-data-707484
source_name: "The Decoder"
source_url: https://the-decoder.com/old-ocr-text-cripples-language-model-training-and-finebooks-wants-to-fix-that-at-scale/
published: 2026-08-10T18:20:39.000Z
topics: ["research", "open source"]
publisher: "Parallel Quant"
---

# Hugging Face and EleutherAI benchmark OCR for AI training data

*2026-08-10 · Source: [The Decoder](https://the-decoder.com/old-ocr-text-cripples-language-model-training-and-finebooks-wants-to-fix-that-at-scale/)*

The FineBooks project tested 14 open-source OCR models on over 2,000 historical book pages to find the best way to digitize text for AI training. The top model, dots.mocr, hit 97.6% character accuracy at under $2 per thousand pages, though the team says that's not yet accurate enough for scholarly transcription.

**Why it matters:** Clean training data is an increasing bottleneck as labs exhaust readily available web text and turn to digitizing books and archives. A cheap, benchmarked OCR pipeline lowers the cost of that conversion, though the gap to scholarly-grade accuracy shows real limits remain for anything beyond training data.

**Topics:** research, open source

---
Read the original: https://the-decoder.com/old-ocr-text-cripples-language-model-training-and-finebooks-wants-to-fix-that-at-scale/
Canonical: https://www.parallelquant.com/posts/hugging-face-and-eleutherai-benchmark-ocr-for-ai-training-data-707484
Published by Parallel Quant — https://www.parallelquant.com
