[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83330-en":3,"doc-seo-83330-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83330,1374391974585,"Genevieve","https://ap-avatar.wpscdn.com/davatar_276721f389ce27ea32af1340a28f341c",8,"Research & Report","ICDAR 2026 HIPE-OCRepair Competition on LLM-Assisted OCR Post-Correction for Historical Documents","HIPE-OCRepair-2026 presents results from an ICDAR competition focused on LLM-assisted OCR post-correction for historical documents. It addresses persistent OCR error issues in large digital heritage collections, where large-scale re-digitization is impractical. The shared task evaluates modern OCR correction systems and provides a reproducible multilingual evaluation framework grounded in a harmonized HIPE-OCRepair dataset spanning English, French, and German (17th–20th century).","arXiv :2607 .08 143v 1 [ cs .CL] 9 Jul 2026  \nICDAR 2026 HIPE-OCRepair Competition on LLM-Assisted OCR Post-Correction for Historical Documents  \nMaud Ehrmann 1[0000−0001−9900−2193], Emanuela Boros 1[0000−0001−6299−9452], Juri Opitz2[0000−0001−6892−4574], Andrianos Michail2[0009−0004−1025−7851], Florian Wagner2 , and Simon Clematide2[0000−0003−1365−0662]  \n1 École Polytechnique Fédérale de Lausanne (EPFL), Switzerland  \n2 University of Zurich, Switzerland  \nAbstract. We present the results of HIPE-OCRepair-2026, an ICDAR competition on LLM-assisted OCR post-correction of historical documents. OCR post-correction remains a long-standing challenge in digital heritage: large-scale collections of digitized documents are affected by legacy OCR errors, while re-digitization at scale remains impractical.  \nLarge language models (LLMs) offers a major opportunity to revisit this challenge, yet their effectiveness across languages, document types, and noise conditions—and their tendency to hallucinate—remains insufficiently understood.  \nHIPE-OCRepair-2026 pursues two objectives: (i) to evaluate the capabilities of modern OCR post-correction systems, and (ii) to provide a reproducible evaluation framework anchored in the HIPE-OCRepair-2026 dataset, a harmonized multilingual resource consolidating existing and newly curated historical datasets. Participants were tasked with correcting noisy OCR transcripts from historical newspapers and printed works in English, French, and German (17th–20th century), working at the level of coherent transcription units (paragraphs or articles) without access to source images. The evaluation adopts a retrieval-oriented rather than diplomatic scoring approach, reflecting the practical use case of search and access over digitized collections.  \nFour teams submitted systems ranging from zero-shot prompting to continued pre-training and fine-tuning, offering insights into the merits of different adaptation strategies. Results show that modern LLM-assisted systems can significantly improve OCR quality, but performance varies across datasets, languages, and noise levels. Overcorrection on low-noise inputs emerges as a recurring challenge, highlighting the importance of evaluation beyond character error reduction. The dataset, scorer, and evaluation pipeline are publicly released to support future research.  \nKeywords: OCR post-correction · historical documents · large language models · shared task · multilingual benchmark · digital heritage  \n2 Ehrmann et al.  \n1 Introduction  \nThe large-scale digitization of historical collections has made millions of newspaper pages, books, and archival records searchable and available to researchers worldwide [17,3] . Yet the quality of the resulting text is uneven. Many collections were processed years or decades ago with OCR systems whose performance was limited by the technology of the time, and even modern engines continue to struggle with the specific challenges posed by historical documents: degraded paper and ink, non-standard typography, complex layouts, and multilingual content all contribute to recognition failures that remain difficult to avoid [14,22] . The result is a systematic gap between what digitization has made available and what downstream applications—from keyword search and information retrieval to named entity recognition and large-scale corpus analysis—require to function reliably [8,29,12,6,9] .  \nOCR post-correction, the automatic improvement of existing transcripts without reprocessing source images, is a long-standing response to this challenge [19] . Despite decades of research spanning rule-based methods, statistical models, and neural sequence-to-sequence approaches [1,21], the problem remains far from solved. Performance is highly sensitive to language, historical period, document type and noise characteristics, and robust generalization across heterogeneous collections has remained elusive [27] . At the same time, the scale of the accumulate","cbCaihkMuS3BSJkk","https://ap.wps.com/l/cbCaihkMuS3BSJkk","pdf",468747,3,1,17,"English","en",105,"# Abstract\n# Introduction\n## Digital heritage digitization and OCR quality gap\n## OCR post-correction as a long-standing challenge\n## LLMs for noisy-to-clean transformation and correction","[{\"question\":\"What problem does HIPE-OCRepair-2026 address?\",\"answer\":\"It addresses OCR post-correction for historical documents, where legacy digitization leaves systematic OCR errors and re-digitization at scale is impractical.\"},{\"question\":\"What are the main objectives of the competition?\",\"answer\":\"The competition evaluates the capabilities of modern OCR post-correction systems and provides a reproducible evaluation framework anchored in the HIPE-OCRepair-2026 dataset.\"},{\"question\":\"How is system performance evaluated in the shared task?\",\"answer\":\"Evaluation uses a retrieval-oriented rather than diplomatic scoring approach, reflecting search and access use cases over digitized collections.\"}]",1784186761,43,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"icdar-2026-hipe-ocrepair-competition-on-llm-assisted-ocr-post-correction-for-historical-documents","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/icdar-2026-hipe-ocrepair-competition-on-llm-assisted-ocr-post-correction-for-historical-documents/83330/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does HIPE-OCRepair-2026 address?","Question",{"text":75,"@type":76},"It addresses OCR post-correction for historical documents, where legacy digitization leaves systematic OCR errors and re-digitization at scale is impractical.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What are the main objectives of the competition?",{"text":80,"@type":76},"The competition evaluates the capabilities of modern OCR post-correction systems and provides a reproducible evaluation framework anchored in the HIPE-OCRepair-2026 dataset.",{"name":82,"@type":73,"acceptedAnswer":83},"How is system performance evaluated in the shared task?",{"text":84,"@type":76},"Evaluation uses a retrieval-oriented rather than diplomatic scoring approach, reflecting search and access use cases over digitized collections.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]