[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-122594-en":3,"doc-seo-122594-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},122594,1099513958607,"Jiven","https://ap-avatar.wpscdn.com/avatar/100002390cf8733938c?x-image-process=image/resize,m_fixed,w_180,h_180&k=1778829742770036399",8,"Research & Report","Transcribing Medieval Manuscripts for Machine Learning","The article examines how new automation in digital archives reshapes medieval studies, focusing on transcription practices for machine processing rather than solely for human reading. It argues that future computational text analysis requires intentionally theorized transcriptions, informed by historic methods and by the implications of handwritten text recognition systems such as Transkribus and eScriptorium. Through a brief historiography of transcription norms for thirteenth-century Latin Bibles, it highlights normalization, research-use planning, and how general HTR models can alter humanities interpretation.","Transcribing Medieval Manuscripts for Machine Learning  \nKeywords  \nParis Bible; Latin Bible; handwritten text recognition (HTR) ; Thirteenth Century Europe; bias; transcription norms; computational text analysis  \nAuthors:  \n● Estelle Guéville, Yale University, USA: [estelle.gueville@yale.edu](estelle.gueville@yale.edu)  \n● David Joseph Wrisley, New York University Abu Dhabi, [UAE: djw12@nyu.edu](UAE: djw12@nyu.edu)  \nIntroduction  \nIn the early twentieth century, many scholars focused on the preparation of editions and translations of texts previously available only to the few specialists able to read archaic hands and privileged enough to travel to work in person with them in manuscript. Valuable scholarship in its own right, the preparation of these editions and translations for particular texts deemed important enough to justify the effort and time, laid the foundation for generations of scholarship in medieval studies. On the other hand, for many materials in historical archival collections–including already digitised collections–medievalists have only had the time to create partial transcriptions, if any at all. Access to textual material from the medieval period has increased greatly in recent years with digitisation, and we are able to imagine many new research projects in decades to come. What challenges do new frontiers of automation in the archives raise with respect to medieval studies and in particular to the ways we transcribe? In this article, we argue that if medievalists hope to pursue the kinds of analysis that goes on in advanced computational research, we will need new kinds of transcriptions, intentionally theorized not only for human reading, but also for machine processing. We already have mature methods for remediating generations of editions of medieval works such as Optical Character Recognition (OCR), but we can ask ourselves if these are the kinds of text we want to use for future computational analysis. We suggest instead that one way forward is by going back to the scriptorium.  \nPractices of Transcribing Medieval Manuscripts: a Very Short Historiography  \nIn this section, we give a brief overview of different ways that editors and publishers of medieval texts have treated the question of the difference between the writing systems that we typically use today and those that are found in manuscripts. It is not meant to be a full assessment of historical trends, but rather away of situating our discussion of transcription. We frame that discussion by referring to work done specifically with thirteenth-century Latin Bibles, but we trust that our contextualized discussion of transcription will benefit other use cases for communities who may be considering automatic forms of text creation from manuscript.  \na. Historicizing Normalisation  \nA transcription of an old text is both a theoretical and a practical endeavour. We all have inherited multiple methodologies for transcribing, but machine learning systems for handwritten text recognition (HTR) such as Transkribus, eScriptorium and others that will no doubt emerge in coming years foreground three main issues related to transcription. First, their emergence emphasises the question of normalisation as a historically contingent and changing category. Second, their rise in popularity foregrounds the necessity for anticipating how target transcriptions will be used in research before the transcription begins. Third, we are confronted by the question of how \"general\" HTR models emerging in the digital GLAM sector, that is, general-purpose models for large amounts of text, will change modes of analysis and interpretation in the humanities (Hodel et al., 2021) .  \nSo, what are some of the ways that we implicitly or explicitly normalise texts when we transcribe them? Scholars working on a particular source base might have a given set of transcription norms inherited from a publisher or a philological education that encourage the normalisation of letter forms ","cbCaitOpkFVJkPyn","https://ap.wps.com/l/cbCaitOpkFVJkPyn","pdf",2785482,1,25,"English","en",105,"# Introduction\n# Practices of Transcribing Medieval Manuscripts: a Very Short Historiography\n## Historicizing Normalisation","[{\"question\":\"Why does the article argue for new kinds of medieval transcriptions for machine processing?\",\"answer\":\"It argues that advanced computational research requires transcriptions that are theorized not only for human readability but also for machine processing and analysis in future projects.\"},{\"question\":\"What three main issues do HTR systems highlight for transcription?\",\"answer\":\"They foreground normalization as a historically contingent category, the need to anticipate how target transcriptions will be used in research before transcription begins, and how general-purpose HTR models may change analysis and interpretation in the humanities.\"},{\"question\":\"How do scholars distinguish between different transcription styles for medieval texts?\",\"answer\":\"They distinguish normalized, semi-diplomatic, and diplomatic transcriptions, though each transcriber may define different rules, producing significant variation in norms.\"}]","Transcribing Medieval Manuscripts for Machine Learning | PDF",1785811640,63,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"transcribing-medieval-manuscripts-for-machine-learning","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/transcribing-medieval-manuscripts-for-machine-learning/122594/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-04",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why does the article argue for new kinds of medieval transcriptions for machine processing?","Question",{"text":75,"@type":76},"It argues that advanced computational research requires transcriptions that are theorized not only for human readability but also for machine processing and analysis in future projects.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What three main issues do HTR systems highlight for transcription?",{"text":80,"@type":76},"They foreground normalization as a historically contingent category, the need to anticipate how target transcriptions will be used in research before transcription begins, and how general-purpose HTR models may change analysis and interpretation in the humanities.",{"name":82,"@type":73,"acceptedAnswer":83},"How do scholars distinguish between different transcription styles for medieval texts?",{"text":84,"@type":76},"They distinguish normalized, semi-diplomatic, and diplomatic transcriptions, though each transcriber may define different rules, producing significant variation in norms.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]