[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-119880-en":3,"doc-seo-119880-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},119880,1649267921044,"Ava Thompson","https://us-avatar.wpscdn.com/avatar/1800007509477c92dfb?_k=1782875107921204101",8,"Research & Report","CuneiML - A Cuneiform Dataset for Machine Learning","CuneiML introduces a curated dataset to advance machine learning for processing ancient cuneiform materials spanning over 3,000 years, from the mid-fourth millennium BCE to the late first millennium BCE. The dataset contains 38,947 high-resolution 2D photographs of Sumerian and Akkadian tablets, paired with cuneiform Unicode transcriptions, transliterations, lineart, and rich metadata. It is designed for consistent formatting through careful preprocessing, segmentation, filtering, and re-transliteration of data derived from the CDLI collection, enabling tasks such as genre, provenance, and period prediction from unannotated tablet images.","CuneiML: A Cuneiform Dataset for Machine Learning  \nDANLU CHEN  ADITI AGARWAL   \nTAYLOR BERG-KIRKPATRICK  JACOBO MYERSTON   \n*Author affiliations can be found in the back matter of this article  \nABSTRACT  \nThe cuneiform writing system holds a vast reservoir of ancient literature, encompassing over 3000 years of history. Originating around the mid-fourth millennium BCE and enduring until the late first millennium BCE, cuneiform writing spans various genres such as administrative, legal, medical, and scientific documents, among others. This article introduces a curated dataset, CuneiML, featuring 38,947 high-resolution 2D photos of Sumerian and Akkadian cuneiform tablets, accompanied by their cuneiform Unicode transcriptions, transliterations, lineart, and metadata. This dataset aims to support the development of machine learning tools for processing and analyzing Sumerian and Akkadian cuneiform artifacts – e.g. for automatically classifying genre, provenance, or period from unannotated tablet images. Thus, CuneiML is designed with consistency of format as a primary concern. Specifically, CuneiML is a result of meticulously preprocessing, segmenting, filtering, and re-transliterating data that is available online in the Cuneiform Digital Library Initiative (CDLI) collection.  \nCOLLECTION: REPRESENTING THE ANCIENT WORLD THROUGH DATA  \nDATA PAPER  \nCORRESPONDING AUTHOR:  \nDanlu Chen  \nComputer Science and Engineering, UC San Diego, La Jolla, US  \n[dac013@ucsd.edu](dac013@ucsd.edu)  \nKEYWORDS:  \ncuneiform; machine learning; computational paleography; image processing  \nTO CITE THIS ARTICLE:  \nChen, D., Agarwal, A., BergKirkpatrick, T., & Myerston, J.(2023) . CuneiML: A Cuneiform Dataset for Machine Learning. Journal of Open Humanities Data, 9: 30, pp. 1–9. DOI:  \n[https://doi.org/10.5334/](https://doi.org/10.5334/)[ ](https://doi.org/10.5334/)[johd.151](johd.151)  \n1 INTRODUCTION  \nIn this article we present a curated dataset of 38,947 2D photographs of Sumerian and Akkadian cuneiform tablets with their accompanying transcriptions in cuneiform Unicode – as well as lineart, transliterations, and metadata specifying attributes like period and genre. In contrast to the data provided by digital libraries which offer general access to cuneiform texts, our dataset was envisioned from the very beginning for machine learning with an emphasis on consistency of format. Therefore, we developed our dataset with strict preprocessing and filtering criteria and present preliminary baseline experiments for three classification tasks supported by our data: period, provenance, and genre prediction, conditioned on major face cutouts from tablet photographs.  \nThe CuneiML dataset was produced by processing photographs and transliterations available online in the Cuneiform Digital Library Initiative (CDLI) (Englund et al., 2023). This library gives access to 56,694 photographs of inscribed objects classified by time period, genre, provenance, and museum collection. Current digitized cuneiform archives like CDLI were designed as portals where experts can consult photographs of inscribed objects (tablets, seals, inscriptions, etc.), transliterations, dictionaries, and other working tools. Although the CDLI is an extraordinary resource which has proven to be of invaluable use for Assyriologists, it offers its data in a format not suitable for machine learning experiments. CDLI photographs are of varied quality: some are high-resolution while others do not meet the minimum requirements for machine learning tasks. In addition, CDLI images are composite; this means they contain multiple perspectives of the same object: front, back and sides of tablets (see Figure 1) . Another issue is that the transliterations of Sumerian and Akkadian that accompany CDLI images are missing their rendering into cuneiform Unicode. These aspects make the CDLI data unsuitable for machine learning. Thus, with machine learning in mind, we have meticulously filtered and processed ","cbCaivEIJrRN0Uao","https://ap.wps.com/l/cbCaivEIJrRN0Uao","pdf",2074421,1,9,"English","en",105,"# Introduction\n## Dataset Overview and Motivation\n## Source Data and Processing Pipeline\n## Supported Machine Learning Tasks","[{\"question\":\"What does CuneiML include besides tablet photographs?\",\"answer\":\"CuneiML provides 38,947 high-resolution 2D photos along with cuneiform Unicode transcriptions, transliterations, lineart, and metadata such as period and genre.\"},{\"question\":\"Why is CuneiML formatted specifically for machine learning?\",\"answer\":\"The dataset is produced with strict preprocessing and filtering criteria to ensure consistent format, addressing issues in the original CDLI data such as varied image quality, composite views, and missing Unicode rendering.\"},{\"question\":\"What kinds of classification tasks can be supported using CuneiML?\",\"answer\":\"The document describes baseline experiments for classification tasks including period, provenance, and genre prediction, conditioned on major face cutouts from tablet photographs.\"}]","CuneiML - A Cuneiform Dataset for Machine Learning | PDF",1785726810,23,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"cuneiml-a-cuneiform-dataset-for-machine-learning","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/cuneiml-a-cuneiform-dataset-for-machine-learning/119880/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-03",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What does CuneiML include besides tablet photographs?","Question",{"text":75,"@type":76},"CuneiML provides 38,947 high-resolution 2D photos along with cuneiform Unicode transcriptions, transliterations, lineart, and metadata such as period and genre.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"Why is CuneiML formatted specifically for machine learning?",{"text":80,"@type":76},"The dataset is produced with strict preprocessing and filtering criteria to ensure consistent format, addressing issues in the original CDLI data such as varied image quality, composite views, and missing Unicode rendering.",{"name":82,"@type":73,"acceptedAnswer":83},"What kinds of classification tasks can be supported using CuneiML?",{"text":84,"@type":76},"The document describes baseline experiments for classification tasks including period, provenance, and genre prediction, conditioned on major face cutouts from tablet photographs.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,127,130,134],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":21,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]