[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85095-en":3,"doc-seo-85095-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},85095,1099514067438,"River Wang","https://ap-avatar.wpscdn.com/avatar/100002539ee87300030?x-image-process=image/resize,m_fixed,w_180,h_180&k=1780474512215547542",8,"Research & Report","Conversational Retrieval and On-the-Fly Knowledge Modeling of Historical Penitentiary Repression Records","Recent developments in digital libraries increasingly favor conversational and natural language access to information through Retrieval-Augmented Generation (RAG). However, extractive approaches tied to individual records struggle to interpret document collections holistically and incorporate expert knowledge dynamically. This article presents a document analysis system for historical digital libraries with on-the-fly knowledge modeling. Facts from experts or retrieval are stored in a graph-based structure, enabling long-term queries, link discovery, and richer synthesized information.","arXiv :2607 .08459v 1 [ cs .IR] 9 Jul 2026  \nConversational Retrieval and On-the-Fly Knowledge Modeling of Historical Penitentiary Repression Records  \nPaula Font Solà, Adrià Molina RodriguezB, Josep Lladós Canet  \nComputer Vision Center, Universitat Autònoma de Barcelona [paula.fonts@autonoma.cat](paula.fonts@autonoma.cat) , {amolina, [josep}@cvc.uab.cat](josep}@cvc.uab.cat)  \nAbstract. Recent developments in digital libraries increasingly favor conversational and natural language access to information through RetrievalAugmented Generation (RAG) . Although these approaches are effective for extractive tasks grounded in individual records, they remain limited in their ability to interpret document collections holistically and to incorporate expert knowledge dynamically. In this article, we present a document analysis system designed for the management of historical digital libraries that supports on-the-fly knowledge modeling. The system is equipped with the capability to store facts produced either by expert archivists or derived from document retrieval processes within a graph-based structure. Through continuous professional interaction, the system can retrieve information not only from primary sources such as documents, but also from previously modeled knowledge, with the graphbased index acting as a memory for the language model to access. This enables increasingly complex queries involving long-term dependencies across documents, link discovery, and the integration of expert knowledge that may not be explicitly present in the original sources. As a result, the proposed approach facilitates the generation of richer and more comprehensive information.  \nKeywords: Document Analysis Systems · Historical Documents  \n1 Introduction  \nIn recent years, historical document analysis systems have shifted from indices and enumerations of content to increasingly conversational and adaptive interfaces. This evolution is not merely a superficial redesign of how users access archives and portals; it represents a deeper transformation in how archival data is organized, interpreted, and communicated. The advent of Generative AI has accelerated this transition, transforming both how users interact with systems, using natural language queries and conversational exploration, and how results are synthesized and delivered. Instead of static document lists, modern systems generate summaries and reveal relationships in archival collections. Although it may be tempting to attribute this transformation solely to the rapid development of Large Language Models (LLMs) and Retrieval-Augmented Generation  \n2 P. Font [et.al](et.al)  \n(RAG), it is important to recognize that archival science had already begun moving beyond document enumeration toward more holistic paradigms. The International Council on Archives’s Records in Contexts (RiC) framework [17] exemplifies this shift, emphasizing relationships, provenance, and contextual interdependencies across collections rather than isolated descriptions of individual records.  \nFrom the perspective of Document Analysis Systems, however, sustaining such a rich and interconnected framework raises practical challenges. The construction and maintenance of contextual knowledge representations (e.g. knowledge graphs) have traditionally relied on the manual labor of experienced historians and archivists. These professionals are expected not only to identify the source and context of a given document, but also to interpret its contents, resolve ambiguities, trace links across collections, and determine the role each record plays within a broader historical narrative. While this work is essential, it is timeconsuming, difficult to scale, and inevitably constrained by available resources. This challenge becomes particularly critical in large historical archives containing heterogeneous, partially digitized, and often degraded materials, where access to personalized, well-focused, and detailed information is increasing","cbCaisWRsbBPj6QU","https://ap.wps.com/l/cbCaisWRsbBPj6QU","pdf",25417493,4,1,17,"English","en",105,"# Introduction\n## Shift toward conversational archival access\n## Records in Contexts and contextual knowledge modeling\n## Challenges in historical archives\n## Contributions: graph-based knowledge and transcription enhancement","[{\"question\":\"What limitation do conventional RAG-based conversational systems have for historical archives?\",\"answer\":\"They work well for extracting answers grounded in individual records, but they struggle to interpret document collections holistically and to incorporate expert knowledge dynamically.\"},{\"question\":\"How does the proposed system model knowledge during conversation?\",\"answer\":\"It creates and continuously refines a persistent graph that stores facts produced by expert archivists or derived from document retrieval, serving as structured memory for the language model.\"},{\"question\":\"Why is transcription quality important for downstream reasoning in this setting?\",\"answer\":\"Historical materials are often handwritten or degraded, and transcriptions are noisy or incomplete; the reliability of reasoning therefore depends on uncertainty management in transcription and extraction.\"}]",1784201081,43,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"conversational-retrieval-and-on-the-fly-knowledge-modeling-of-historical-penitentiary-repression-records","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/conversational-retrieval-and-on-the-fly-knowledge-modeling-of-historical-penitentiary-repression-records/85095/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What limitation do conventional RAG-based conversational systems have for historical archives?","Question",{"text":75,"@type":76},"They work well for extracting answers grounded in individual records, but they struggle to interpret document collections holistically and to incorporate expert knowledge dynamically.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the proposed system model knowledge during conversation?",{"text":80,"@type":76},"It creates and continuously refines a persistent graph that stores facts produced by expert archivists or derived from document retrieval, serving as structured memory for the language model.",{"name":82,"@type":73,"acceptedAnswer":83},"Why is transcription quality important for downstream reasoning in this setting?",{"text":84,"@type":76},"Historical materials are often handwritten or degraded, and transcriptions are noisy or incomplete; the reliability of reasoning therefore depends on uncertainty management in transcription and extraction.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]