[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86373-en":3,"doc-seo-86373-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86373,962075006959,"Anda","https://ap-avatar.wpscdn.com/avatar/e0002397efbe92a78e?_k=1776741047341049297",8,"Research & Report","Loci Similes: A Benchmark for Extracting Intertextualities in Latin Literature","Loci Similes introduces a standardized benchmark for detecting intertextual connections in Latin literature by tracing how later historical texts reuse earlier passages through quotation, allusion, paraphrase, and morphological variation. The benchmark provides a curated dataset of about 176k text segments alongside 1,490 expert-verified parallels, including 945 labeled references from an existing dataset. Using this data, the work establishes baselines for retrieval and classification of intertextualities with pretrained encoder language models, addressing prior limitations caused by scarce benchmark resources and hard-to-use datasets.","Loci Similes: A Benchmark for Extracting Intertextualities  \nin Latin Literature∗  \nJulian Schelb†, Michael Wittweiler⋄ , Marie Revellio‡, Barbara Feichtinger‡, and Andreas Spitz†  \n†Department of Computer and Information Science, University of Konstanz ‡Department of Latin Philology, University of Konstanz ⋄Institute of Archaology, Classical Philology and Ancient Studies, University of Zurich  \n{ [firstname.lastname}@uzh.ch](firstname.lastname}@uzh.ch)[ ](firstname.lastname}@uzh.ch){ [firstname.lastname}@uni-konstanz.de](firstname.lastname}@uni-konstanz.de)  \narXiv :2601 .07533v2 [ cs .IR] 11 Jul 2026  \nAbstract  \nTracing connections between historical texts isan important part of intertextual research, enabling scholars to reconstruct the virtual library of a writer and identify the sources influencing their creative process. These intertextual links manifest in diverse forms, ranging from direct verbatim quotations to subtle allusionsand paraphrases disguised by morphological variation. Language models offer a promising path forward due to their capability of capturing semantic similarity beyond lexical overlap. However, the development of new methods for this task is held back by the scarcity of standardized benchmarks and easy-to-use datasets. We address this gap by introducing Loci Similes, a benchmark for Latin intertextuality detection comprising a curated dataset of ∼ 176k text segments and 1,490 expert-verified parallels, including 945 labeled references from an existing dataset. Using this data, we establish baselines for retrieval and classification of intertextualities with pretrained encoder language models.  \n1 Introduction  \nIdentifying intertextual connections between documents is an important task in classical philology, as it reveals how later works engage with earlier textsand traditions. For centuries, scholars detected intertextual references by relying on memory and the manual collation of Loci Similes, i.e., parallel passages that exhibit lexical, semantic, or thematic resemblance. Although digitization has augmented this process through lexical search tools, most approaches still depend on exact n-gram matching or heuristic filtering (Schropp et al., 2024a) . This limits discovery rates in ancient texts, where intertextuality typically manifests itself not as verbal quotation, but as subtle allusion, paraphrase, or  \node and data are available at [https://anonymous](https://anonymous) . [4open.science/r/locisimiles-2338](4open.science/r/locisimiles-2338)  \n(anonymized for review) .  \nExample of Historical Text Reuse  \nSource: Virgil, Aeneid 2.774  \nContext: Aeneas is terrified when the ghost of his wife Creusa appears to him during the burning of Troy.  \n“... obstipui, steteruntque comae et uox faucibus haesit.”  \n(... I was stupefied, my hair stood on end, and my voice stuck in my throat.)  \nReuse: Jerome, Epistula 130.5.5  \nContext: A family’s shocked reaction to Demetrias’s vow of Christian virginity.  \n“Haesit uox faucibus et inter ruborem atque pallorem metumque ac laetitiam cogitationes uariae mutabantur.”  \n( The voice stuck in their throat, and between blushing and pallor, fear and joy manifold thoughts kept shifting.)  \nFigure 1: Example of intertextual reference. Reuse of a classic Virgilian phrase for speechlessness by Jerome. While retaining the semantic core, the author alters the word order to adapt the expression to a different context.  \nthematic variation (Manjavacas et al., 2019 ; Gonget al., 2025), often complicated by orthographic volatility (Miller et al., 2025) .  \nRecovering such textual reuses is not merely a matter of identifying sources. It facilitates research on broader cultural-historical phenomena (Tangherlini and Chen, 2024) . In particular, it supports work on reception and cultural hybridization in Late Antiquity, where pagan texts persist as the rhetorical substrate of elite writing while being recontextualized within emerging Christian discourse. Classical forms often","cbCaibZ1Rc8wPV01","https://ap.wps.com/l/cbCaibZ1Rc8wPV01","pdf",2728292,3,1,25,"English","en",105,"# Abstract\n# Introduction\n## Example of Historical Text Reuse","[{\"question\":\"What problem does Loci Similes target in Latin intertextual research?\",\"answer\":\"It targets the difficulty of reliably extracting intertextual links between historical Latin texts, especially when reuse appears as subtle allusion or paraphrase rather than exact quotation.\"},{\"question\":\"What does the Loci Similes benchmark include?\",\"answer\":\"It provides a curated dataset of roughly 176k text segments and 1,490 expert-verified parallels, including 945 labeled references from an existing dataset.\"},{\"question\":\"How is the benchmark used to support model development?\",\"answer\":\"The work uses the dataset to establish baselines for retrieval and classification of intertextualities using pretrained encoder language models.\"}]",1784211209,63,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"loci-similes-a-benchmark-for-extracting-intertextualities-in-latin-literature","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/loci-similes-a-benchmark-for-extracting-intertextualities-in-latin-literature/86373/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-27","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does Loci Similes target in Latin intertextual research?","Question",{"text":75,"@type":76},"It targets the difficulty of reliably extracting intertextual links between historical Latin texts, especially when reuse appears as subtle allusion or paraphrase rather than exact quotation.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What does the Loci Similes benchmark include?",{"text":80,"@type":76},"It provides a curated dataset of roughly 176k text segments and 1,490 expert-verified parallels, including 945 labeled references from an existing dataset.",{"name":82,"@type":73,"acceptedAnswer":83},"How is the benchmark used to support model development?",{"text":84,"@type":76},"The work uses the dataset to establish baselines for retrieval and classification of intertextualities using pretrained encoder language models.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]