NSP
no shit papers
← Back to all papers
Step 1: Pre-print

Cost-Aware Human-LLM Collaboration for post-OCR Corrections in Swiss Historical Newspapers

Download PDF
Authors: iD Stergios Konstantinidis (0000-0003-1620-5871) ✉️ Corresponding iD Hayman Lotfy (0009-0001-1768-3675) • Uploaded July 24, 2026 DOI: 10.5555/nsp.505ac130
Step 1 Citation Goal: 5 Citations to reach "Accepted / Published" 0 / 5
Your pre-print is accessible to everyone for free, forever. Once it receives 5 citations, it is officially promoted to Step 2.

Abstract

OCR transcription errors in historical archives often hinder digital search and retrieval. While Large Language Models (LLMs) can correct many of these errors, applying them indiscriminately is costly and may negatively affect already-clean text. We propose a three-tier collaboration framework that routes each text segment to one of: (1) No Correction, (2) LLM Correction, or (3) Human Correction. We introduce a regression-guided routing approach that prioritizes segments by predicted CER improvement, paired with a safeguard layer that detects harmful LLM corrections and routes uncertain segments to human review. With only <5% of the corpus reviewed by human experts, our safeguard achieves a 14% relative reduction over the All-LLM baseline, and substantially outperforms standard confidence-based approaches. By dynamically routing degraded segments to humans and fixable errors to the LLM, the collaborative framework outperforms either corrector in isolation.

Citations Received (0)

No citations received yet on NSP.

Cite This Paper

Select one of your existing papers on NSP to add a citation to this paper:

No other published papers available to cite from. Upload another paper first.

Report Academic Fraud / Dishonesty