[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82291-en":3,"doc-seo-82291-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},82291,13056703019662,"Evangeline","https://ap-avatar.wpscdn.com/avatar/be000253a8e92610077?_k=1778726343310543188",8,"Research & Report","Letter Lemmatization: One-to-one and Banded RNNs for Reversing Character-set Simplification and Abbreviation in Medieval Text","Medieval document transcription faces highly variable practices and heterogeneous digitization policies that make the character set behave as a fluid resource across corpora. The work addresses flexible character-set conversion, training character-level one-to-one RNNs to undo one-to-one mappings with self-supervision and recover substantial CER even from limited text. These networks improve HTR post-correction, then expand medieval abbreviations using banded RNNs with alignment ground truth. A heuristic letter-lemmatization metric defines semantic similarity between arbitrary character sets, supported by an efficient Python library.","arXiv :2607 .0929 1v 1 [ cs .CL] 10 Jul 2026  \nLetter Lemmatization: One-to-one and Banded RNNs for Reversing Character-Set Simplification and Abbreviation in Medieval  \nText  \nAnguelos Nicolaou 1 , Maria Pia Tiseo2 ,3 , Tam´as Kov´acs 1 , Nicolas Renet 1 ,  \nand Georg Vogeler 1   \n1 University of Graz, Graz, Austria  \n{[firstname.lastname](firstname.lastname}@uni-graz.at)[}](firstname.lastname}@uni-graz.at)[@uni-graz.at](firstname.lastname}@uni-graz.at)  \n2 University of Basel, Basel, Switzerland  \n3 Universit`a degli Studi di Napoli Federico II, Naples, Italy  \nAbstract. Medieval document transcribers have very different practices; on top of that, heterogeneous digitization policies have resulted in corpora where the character-set must be viewed as fluid. In this paper we address the problem of changing between character-sets in a flexible manner. We focus on one-to-one character mappings and train characterlevel one-to-one RNNs to undo them with self-supervision; recovering half the CER even with 20 text lines. We analyse the use of these one-to-one networks for HTR post-correction and we see that they obtain significant improvements while totally ignoring ins-dels. We then use the exact same networks with character-level alignment groundtruth compiled from parallel corpora in a training and inference mode we call Banded RNNs. We use such networks to successfully expand abbreviations in medieval charter transcriptions. Finally we introduce an elaborate heuristic which takes the characters of two arbitrary character-sets and defines a metric encapsulating what we consider to be semantic similarity of characters. We call the construction of such mappings letter lemmatization and present a rich Python library that efficiently performs all presented methods.  \nKeywords: Letter Lemmatization · Charset Simplification · One-to-one RNN · Banded RNN · Abbreviation Expansion · HTR Post-correction · CER  \n1 Introduction  \nIn this paper, we address the question of which character set to use when digitizing medieval documents and how to effectively change between character sets.  \nWhile acknowledging normalization issues that predate the digital era [2], as well as the perils of making transcriptions conform to an elusive, ideal state of a language in the late Middle Ages [15], we do not intend to take a linguist’s stance on these topics but rather to serve the pragmatic needs of the curators  \n2 A. Nicolaou et al.  \nand remote readers of digital corpora. Are there ways to align character sets that make it easier to train digitization tools and to evaluate their performance consistently, without erasing the variations over time and space that make these corpora relevant to the digital humanists in the first place?  \nConflicting needs and tradeoffs make it impossible to have a clear answer to which approach is optimal, yet we try to provide quantifiable insights on the question. Optical Character Recognition (OCR) has a long history, with open-source OCR tools driving research on the topic [21,3,18,14] . With the introduction of Neural Networks (NN) for OCR, OCR performance increased [4] to the extent that they could also perform usable Handwritten Text Recognition (HTR) [20,19], a much harder task. Unlike OCR, HTR cannot be trained on synthetic data effectively and is more sensitive to the handwriting style. Many digitization projects nowadays employ HTR to digitize the vast majority of their corpora, but they usually need manual transcription of representative parts of their corpus to adapt HTR engines to their data and have a comprehensive quality control of the process. Our approach assumes that HTR training and quality evaluation are a principal use-case for transcribing text. At the same time, we recognize that linguistic analysis, distant reading [17], and automatic indexing of corpora are also important use cases and they are not all best served by the same character set when it comes to text encoding. Character sets are also a k","cbCaiaXIcwCcgmWZ","https://ap.wps.com/l/cbCaiaXIcwCcgmWZ","pdf",402693,2,1,15,"English","en",105,"# Introduction\n## Character-set choice and alignment challenges\n## Charset simplification in the DH pipeline\n## Abbreviation handling and Banded RNNs","[{\"question\":\"What problem does “letter lemmatization” address in medieval text digitization?\",\"answer\":\"It provides a method to map between arbitrary character sets by defining character correspondences and a metric for semantic similarity, enabling flexible conversion across digitization policies.\"},{\"question\":\"How do one-to-one character-level RNNs help reverse character-set simplification?\",\"answer\":\"The approach trains character-level one-to-one RNNs with self-supervision to undo one-to-one mappings and recovers a significant portion of CER loss even with a small number of text lines.\"},{\"question\":\"How are abbreviations expanded using these networks?\",\"answer\":\"Because abbreviation expansion is not strictly one-to-one, the paper uses the same idea with character-level alignment ground truth in a training/inference mode called Banded RNNs to model insertions and deletions and successfully expand abbreviations in medieval charter transcriptions.\"}]",1784179425,38,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"letter-lemmatization-one-to-one-and-banded-rnns-for-reversing-character-set-simplification-and-abbreviation-in-medieval-text","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/letter-lemmatization-one-to-one-and-banded-rnns-for-reversing-character-set-simplification-and-abbreviation-in-medieval-text/82291/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-22","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does “letter lemmatization” address in medieval text digitization?","Question",{"text":75,"@type":76},"It provides a method to map between arbitrary character sets by defining character correspondences and a metric for semantic similarity, enabling flexible conversion across digitization policies.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How do one-to-one character-level RNNs help reverse character-set simplification?",{"text":80,"@type":76},"The approach trains character-level one-to-one RNNs with self-supervision to undo one-to-one mappings and recovers a significant portion of CER loss even with a small number of text lines.",{"name":82,"@type":73,"acceptedAnswer":83},"How are abbreviations expanded using these networks?",{"text":84,"@type":76},"Because abbreviation expansion is not strictly one-to-one, the paper uses the same idea with character-level alignment ground truth in a training/inference mode called Banded RNNs to model insertions and deletions and successfully expand abbreviations in medieval charter transcriptions.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]