[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82300-en":3,"doc-seo-82300-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},82300,13056703019662,"Evangeline","https://ap-avatar.wpscdn.com/avatar/be000253a8e92610077?_k=1778726343310543188",8,"Research & Report","Automatic Thematic Indexing of Large Literary Corpora: A Machine Learning Approach to Voltaire’s Complete Works","Thematic indexing—assigning structured conceptual labels to text sections—is crucial for scholarly access to large literary and historical editions, but is often manual and labour-intensive. This paper applies machine learning to automatic thematic indexing using two sub-corpora from Voltaire’s Complete Works: Essai sur les mœurset l’esprit des nations and Questions sur l’Encyclopédie. The problem is posed as multi-label classification, comparing encoder-based models and fine-tuned generative LLMs via LoRA across ~3–120B parameters. Best results reach F1 up to 0.67, and the study evaluates cross-corpus generalisation and qualitative failures.","arXiv :2607 .093 16v 1 [ cs .CL] 10 Jul 2026  \nResearch Article  \nMiguel Arana-Catania, Gillian Pink*, and Glenn Roe  \nAutomatic Thematic Indexing of Large Literary Corpora: A Machine Learning Approach to Voltaire’s Complete Works  \nAbstract: Thematic indexing—the practice of assigning structured conceptual labels to sections of text—is essential to scholarly access in large-scale literary and historical editions, yet it remains a largely manual, labour-intensive process. This paper explores the application of machine learning to automatic thematic indexing, using two substantial sub-corpora of the Complete Works of Voltaire as a test case: the Essai sur les mœurset l’esprit des nations and the Questions sur l’Encyclopédie. The task is framed as a multi-label classification problem, in which a model must assign the set of index entries that a professional indexer would apply to a given page of text. We compare a range of approaches—from encoder-based models with classification heads to generative large language models (LLMs) fine-tuned via Low-Rank Adaptation (LoRA)—spanning model sizes from approximately 3 to 120 billion parameters. Our best-performing model, from the Mistral family in a 4-bit quantised configuration, achieves F1 scores of up to 0.67; we argue that these figures represent lower bounds, given the inherent subjectivity of professional indexing and the frequency with which model predictions prove semantically valid despite diverging from the print index. We further evaluate cross-corpus generalisation and conduct a detailed qualitative analysis of model behaviour on literary and rhetorical features of the source texts that prove particularly resistant to automated treatment. Our findings have implications for the broader challenge of providing structured thematic access to large-scale literary and historical corpora.  \nKeywords: automatic indexing, multi-label classification, large language models, finetuning, Voltaire, digital humanities, historical text classification  \n1 Introduction  \nOne of the founding assumptions of digital editions, and of digitised text more generally, is that they are inherently accessible —not only because digital channels have dramatically expanded their availability, but because they are, in a word, searchable. Yet the equation  \nMiguel Arana-Catania, Glenn Roe, University of Oxford, UK *Corresponding author: Gillian Pink, University of Oxford, UK  \n2 ~~ ~~  \nof searchability with access deserves scrutiny. Full text search, for all its power, offers a surprisingly narrow mode of engagement with a textual corpus: it presupposes that the reader already knows what terms to look for. The user arrives with a query in hand; the text responds or it does not. Discovery, in any richer sense, is left largely to chance (Whitelaw 2012) .  \nTraditional scholarly publishing, by contrast, has long relied on a rather different instrument. A well-constructed analytical index is not a concordance, nor a finding aid in any merely mechanical sense: it is an interpretative map of a text’s contents, compiled by a professional indexer, often with specialist knowledge, who reads the work in its entirety, identifies key concepts and headwords, and progressively refines a structure of categories and sub-categories as the work takes shape. The result is qualitatively different from what can be achieved using a search function—a structured invitation to lateral reading, browsing, conceptual navigation, and the kind of serendipitous discovery that no keyword query can reliably produce (Stephen 2009; Hjørland 2018) .  \nFor large corpora, however, manual indexing quickly becomes either prohibitively expensive or simply impractical. The Œuvres complètes de Voltaire / Complete Works of Voltaire (Voltaire 1968–2022) presents an instructive case in point. Published by the Voltaire Foundation over fifty-five years, the edition runs to 205 print volumes and approximately fifteen million words, encompassing the full ran","cbCaitPyncaoAGi1","https://ap.wps.com/l/cbCaitPyncaoAGi1","pdf",879610,1,22,"English","en",105,"# Introduction\n## Searchability vs. access\n## Analytical indexing as interpretative mapping\n## Scale challenge in large corpora\n## Automatic indexing as text classification","[{\"question\":\"Why is thematic indexing important for large literary and historical editions?\",\"answer\":\"Thematic indexing provides structured conceptual labels that support scholarly access and navigation. It goes beyond simple keyword search by enabling interpretative browsing and discovery.\"},{\"question\":\"How is the automatic thematic indexing task defined in the paper?\",\"answer\":\"The task is framed as multi-label classification where a model assigns the set of index entries a professional indexer would apply to a given page of text.\"},{\"question\":\"What model approaches does the paper compare, and what is the best reported performance?\",\"answer\":\"The study compares encoder-based models with classification heads and generative LLMs fine-tuned with LoRA, evaluated across model sizes from about 3 to 120 billion parameters. The best-performing model (Mistral, 4-bit quantised) achieves F1 scores up to 0.67.\"}]",1784179475,55,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"automatic-thematic-indexing-of-large-literary-corpora-a-machine-learning-approach-to-voltaires-complete-works","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/automatic-thematic-indexing-of-large-literary-corpora-a-machine-learning-approach-to-voltaires-complete-works/82300/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-21","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why is thematic indexing important for large literary and historical editions?","Question",{"text":75,"@type":76},"Thematic indexing provides structured conceptual labels that support scholarly access and navigation. It goes beyond simple keyword search by enabling interpretative browsing and discovery.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How is the automatic thematic indexing task defined in the paper?",{"text":80,"@type":76},"The task is framed as multi-label classification where a model assigns the set of index entries a professional indexer would apply to a given page of text.",{"name":82,"@type":73,"acceptedAnswer":83},"What model approaches does the paper compare, and what is the best reported performance?",{"text":84,"@type":76},"The study compares encoder-based models with classification heads and generative LLMs fine-tuned with LoRA, evaluated across model sizes from about 3 to 120 billion parameters. The best-performing model (Mistral, 4-bit quantised) achieves F1 scores up to 0.67.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":45,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":45,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":45,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]