[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-120647-en":3,"doc-seo-120647-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},120647,1099513958607,"Jiven","https://ap-avatar.wpscdn.com/avatar/100002390cf8733938c?x-image-process=image/resize,m_fixed,w_180,h_180&k=1778829742770036399",8,"Research & Report","Automated metadata annotation - What is and is not possible with machine learning","Automated metadata annotation depends on the training dataset and on the domain rules that are available for modeling. Understanding what a pre-trained machine learning system was trained on is essential to evaluate both its limitations and likely biases. While large-scale, readily accessible content can support training, scholarly and historical resources often lack consumable, homogenized, and interoperable formats at the volumes required for machine learning. The paper surveys current practice in cultural heritage and research data, outlines challenges from real use cases, and proposes solutions.","REVIEW  \nAutomated metadata annotation: What is and is not possible with machine learning  \nMingfang Wu1, Hans Brandhorst2, Maria-Cristina Marinescu3, Joaquim More Lopez3,  \nMargorie Hlava4 & Joseph Busch5†  \n1Australian Research Data Commons, Australian Research Data Commons, Melbourne, Australia, Australia  \n2Iconclass, Voorscoten, The Netherlands  \n3Barcelona Supercomputing Center, Barcelona, Spain  \n4Access Innovations, Albuquerque, New Mexico, USA  \n5Taxonomy Strategies, Washington, DC, USA  \nKeywords: Metadata annotation; Metadata, Machine learning; Culture heritage; Research data  \nCitation: Wu, M.F., Brandhorst, H., Marinescu, [M.-C. et](M.-C. et) al.: Automated metadata annotation: What is and is not possible with  \nmachine learning. Data Intelligence 5 (2023) . doi: 10.1162/dint_a_00162  \nReceived: March 15, 2022; Revised: April 18, 2022; Accepted: June 7, 2022  \nABSTRACT  \nAutomated metadata annotation is only as good as training dataset, or rules that are available for the domain. It’s important to learn what type of data content a pre-trained machine learning algorithm has been trained on to understand its limitations and potential biases. Consider what type of content is readily available to train an algorithm—what’s popular and what’s available. However, scholarly and historical content is often not available in consumable, homogenized, and interoperable formats at the large volume that is required for machine learning. There are exceptions such as science and medicine, where large, well documented collections are available. This paper presents the current state of automated metadata annotation in cultural heritage and research data, discusses challenges identified from use cases, and proposes solutions.  \n1. INTRODUCTION  \nArtificial Intelligence (AI) is often defined as “The simulation of human intelligence process by machines, especially computer systems” [1, 2] . This means it can be many things: statistical, rules engines, or other forms of intelligence and it encompasses several other domains including natural language processing,  \n† Corresponding author: Joseph Busch ([Email: jbusch@taxonomystrategies.com](Email: jbusch@taxonomystrategies.com); OICID: 0000-0003-4775-8225).  \ncomputer vision, speech recognition, object recognition, and cognitive linguistics to name a few. It also intersects with cross language retrieval and automatic translations.  \nMachine Learning (ML), a subset of AI, trains mathematical models by learning from provided data and making predictions. The learning can be  \n• supervised by learning from labeled data and predicting labels for unlabeled data,  \n• unsupervised by discovering structure or groups in data without predefined labels, or  \n• semi-supervised by taking advantage of both supervised and unsupervised and learning from a small number of labeled data.  \nWe have seen successful applications of ML, for example, image recognition for identifying objects from digital images or videos, speech recognition for translating spoken words into the text, personalized recommendation systems, and classification systems in email filtering, identification of fake news from social media, and many more. The ability of commercial companies to collect large volumes of data, labeled and unlabeled, through user interactions, is a key factor in the success of these applications.  \nIn this paper, we examine the ML applications supporting curation and discovery of digital objects from cultural, archival, and research data catalogs. These digital objects could be images, videos, digitized historical handwritten records, scholarly papers, news articles, dataset files, etc. What these catalogs have in common is that metadata needs to be generated for describing digital objects in curation, especially descriptive metadata that provides information about the intellectual contents of a digital object [3] . Descriptive metadata includes title, description, subject, author etc.  \nTraditionally, curators ","cbCainDKsIpHHfp0","https://ap.wps.com/l/cbCainDKsIpHHfp0","pdf",3148966,1,17,"English","en",105,"# Abstract\n# Introduction\n## What AI and machine learning are\n## ML methods for labeling and discovery\n## Metadata generation in cultural and research catalogs\n# Data Intelligence\n## Barriers and dataset limitations\n## Implications for curator substitution","[{\"question\":\"What determines how well automated metadata annotation works?\",\"answer\":\"It depends on the training dataset and on the domain-specific rules available for the task. The quality and suitability of what the model learned from strongly affect results.\"},{\"question\":\"Why is it difficult to apply machine learning to scholarly and historical metadata?\",\"answer\":\"Scholarly and historical content often is not available in consumable, homogenized, interoperable formats at the large volumes required for machine learning.\"},{\"question\":\"What kinds of challenges does the paper highlight from use cases?\",\"answer\":\"It identifies barriers such as a lack of commonly available high-quality annotated datasets and the cognitive effort involved in manual metadata annotation.\"}]","Automated metadata annotation - What is and is not possible with machine learning | PDF",1785731155,43,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"automated-metadata-annotation-what-is-and-is-not-possible-with-machine-learning","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/automated-metadata-annotation-what-is-and-is-not-possible-with-machine-learning/120647/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-04","2026-08-03",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What determines how well automated metadata annotation works?","Question",{"text":76,"@type":77},"It depends on the training dataset and on the domain-specific rules available for the task. The quality and suitability of what the model learned from strongly affect results.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"Why is it difficult to apply machine learning to scholarly and historical metadata?",{"text":81,"@type":77},"Scholarly and historical content often is not available in consumable, homogenized, interoperable formats at the large volumes required for machine learning.",{"name":83,"@type":74,"acceptedAnswer":84},"What kinds of challenges does the paper highlight from use cases?",{"text":85,"@type":77},"It identifies barriers such as a lack of commonly available high-quality annotated datasets and the cognitive effort involved in manual metadata annotation.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,121,124,129,132,136],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":46,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":46,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":46,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":133,"doc_module":4,"doc_module_name":46,"category_name":134,"show_sort_weight":133,"slug":135},10,"Lifestyle","lifestyle",{"id":137,"doc_module":4,"doc_module_name":46,"category_name":138,"show_sort_weight":107,"slug":139},19,"General","general"]