[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83721-en":3,"doc-seo-83721-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83721,4398048949847,"Eliana","https://ap-avatar.wpscdn.com/avatar/400002536579ef2da7f?_k=1778318612642679267",8,"Research & Report","Improving Access to Historical Archives with Real-time RAG-based Systems","Digitized historical archives are large, heterogeneous cultural heritage repositories, yet access is constrained by noisy OCR text and rigid keyword-based retrieval that reduces answer quality. An end-to-end archival processing and retrieval framework integrates large language models into the pipeline. It adds an LLM-based OCR refinement module and a semantic retrieval plus cross-encoder reranking pipeline for natural-language question answering via retrieval-augmented generation. Experiments on 500,000 Swiss newspaper segments (1762–2001) show CER/WER reductions up to 44.52%/60.95% and a 31.9% NDCG@10 improvement, with gains in correctness and context relevance.","Improving Access to Historical Archives with Real-time  \nRAG-based Systems  \narXiv :2607 .03440v 1 [ cs .IR] 3 Jul 2026  \nStergios Konstantinidis University of Lausanne Lausanne, Switzerland [stergios@unil.ch](stergios@unil.ch)[ ](stergios@unil.ch)Corresponding author Faruk Zahiragic Swiss Federal Institute of Technology (EPFL) Lausanne, Switzerland [faruk.zahiragic@epfl.ch](faruk.zahiragic@epfl.ch)  \nHayman Lotfy University of Lausanne Lausanne, Switzerland [hayman.lotfy@unil.ch](hayman.lotfy@unil.ch)  \nMin-Yen Kan National University of Singapore (NUS) [knmnyn@nus.edu.sg](knmnyn@nus.edu.sg)  \nAlexis Erne University of Lausanne Lausanne, Switzerland [alexis.erne@ik.me](alexis.erne@ik.me)  \nMichalis Vlachos University of Lausanne Lausanne, Switzerland michalis.vlachos@unil.ch  \nAbstract  \nDigitized historical archives are large, heterogeneous cultural heritage repositories, but access methods for such archives face challenges such as noisy optical character recognition (OCR) output and rigid keyword-based retrieval, which limit retrieval quality. In this work, we present an end-to-end archival processing and retrieval framework that integrates large language models (LLMs) into the archival pipeline. Our system introduces two core components: (i) an LLM-based OCR refinement module that improves text quality, and (ii) a semantic retrieval and cross-encoder reranking pipeline supporting natural-language question answering via retrieval-augmented generation (RAG) . Our evaluations are done on a historical archival dataset of 500,000 Swiss newspaper segments spanning over three centuries (1762–2001) . Experiments are conducted across 384 natural-language test queries. Our results highlight that LLM refinements reduce OCR errors by up to 44.52%(CER) and 60.95%(WER) . More importantly, this is accompanied by downstream information retrieval improvements. Compared to traditional keyword baselines, our reranking pipeline increases NDCG@10 by 31.9%(from 65.99% to 87.05%) and achieves statistically significant gains in both answer correctness and context relevance. These results demonstrate that integrating LLMs with established document processing and retrieval pipelines can elevate digital libraries from static repositories to interactive, semantically searchable archival systems.  \nKeywords: large language models; digital libraries; retrieval-augmented generation; OCR; historical archives; semantic retrieval  \nIntroduction  \nDigitized historical archives represent some of the most important publicly accessible cultural heritage collections. National, regional, and university libraries have invested heavily in the digitization of newspapers, magazines, and other printed material spanning several centuries. In most cases, these collections are made available through a combination of scanned page images and OCR-derived text, together with search interfaces intended to facilitate access to their contents. The Cantonal and University Library of Lausanne (Bibliothèque Cantonaleet Universitaire de Lausanne, BCUL), for example, maintains a major digitized newspaper archive with issues dating back to the eighteenth century. Such repositories are invaluable for historians, journalists, students, and the wider public. However, despite their scale and documentary value, these repositories remain underused because their contents are not always easily accessible in practice (Kumpulainen & Late, 2022) .  \nA major reason for this limitation is that archival access still depends on fragile textual representations and rigid retrieval paradigms. First, OCR applied to historical documents often introduces substantial transcription errors. These errors arise from multiple sources, including document degradation, low-quality scans, typographic variability, unusual page layouts, and the graphical complexity of older print traditions. As a result, the OCR text stored alongside digitized pages is frequently noisy and incomplete. This affects not only readabil","cbCaioCnH2QcptcX","https://ap.wps.com/l/cbCaioCnH2QcptcX","pdf",5182981,5,1,26,"English","en",105,"# Abstract\n# Introduction\n## Motivation: Limitations of OCR and Keyword Search\n## Opportunity: LLMs and RAG for Natural-Language Access\n## Challenges: Noisy Text and Latency in Modern RAG Pipelines","[{\"question\":\"What problem does the document address in accessing historical archives?\",\"answer\":\"Access is limited by OCR transcription noise and rigid keyword-based retrieval, which together reduce retrieval quality and make relevant content harder to find.\"},{\"question\":\"What are the two core components of the proposed framework?\",\"answer\":\"The system includes an LLM-based OCR refinement module and a semantic retrieval pipeline with cross-encoder reranking for natural-language question answering via RAG.\"},{\"question\":\"How do the experiments evaluate effectiveness and what improvements are reported?\",\"answer\":\"Evaluations are run on 500,000 Swiss newspaper segments with 384 natural-language queries, showing reduced OCR errors (up to 44.52% CER and 60.95% WER) and improved retrieval metrics such as NDCG@10 increasing by 31.9%.\"}]",1784189971,66,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"improving-access-to-historical-archives-with-real-time-rag-based-systems","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/improving-access-to-historical-archives-with-real-time-rag-based-systems/83721/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-25","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What problem does the document address in accessing historical archives?","Question",{"text":76,"@type":77},"Access is limited by OCR transcription noise and rigid keyword-based retrieval, which together reduce retrieval quality and make relevant content harder to find.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"What are the two core components of the proposed framework?",{"text":81,"@type":77},"The system includes an LLM-based OCR refinement module and a semantic retrieval pipeline with cross-encoder reranking for natural-language question answering via RAG.",{"name":83,"@type":74,"acceptedAnswer":84},"How do the experiments evaluate effectiveness and what improvements are reported?",{"text":85,"@type":77},"Evaluations are run on 500,000 Swiss newspaper segments with 384 natural-language queries, showing reduced OCR errors (up to 44.52% CER and 60.95% WER) and improved retrieval metrics such as NDCG@10 increasing by 31.9%.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":20,"slug":138},19,"General","general"]