[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82305-en":3,"doc-seo-82305-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},82305,1374391974564,"Clementine","https://ap-avatar.wpscdn.com/avatar/14000253aa45c000a9e?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779874745381141002",8,"Research & Report","Letting the Data Speak: Extracting Keywords from Crowdsourced Collections with AI","Identifying and assigning keywords at scale presents technical, practical, and ethical challenges for crowdsourced collections. This article summarizes findings from the “Extracting Keywords from Crowdsourced Collections” project, using the Their Finest Hour Online Archive as a case study. Three NLP approaches—Named Entity Recognition, keyword extraction, and topic modelling—are tested across methods from traditional statistics to GenAI neural networks. Results show strong potential for scalable keyword extraction, yet no single method fully solves the task and model choice strongly affects outcomes, while automated tagging adds stewardship responsibilities.","Letting the Data Speak: Extracting Keywords from Crowdsourced Collections with AI  \nMiguel Arana-Catania* , Catherine Conisbee*°, Matthew Kidd*  \nAbstract  \nIdentifying and assigning keywords at scale is a technical, practical, and ethical challenge for crowdsourced collections. This article reports the findings of the“Extracting Keywords from Crowdsourced Collections” project, which used the Their Finest Hour Online Archive, a crowdsourced Second WorldWar digital collection hosted by the University of Oxford, as a case study. The project evaluated three Natural Language Processing approaches to automate keyword extraction: Named Entity Recognition, Keyword Extraction, and Topic Modelling. It tested these approaches across a range of artificial intelligence techniques, from traditional statistical methods to modern GenAI neural networks. Our quantitative and qualitative findings indicate that Natural Language Processing approaches offer real potential for keyword extraction at scale in crowdsourced collections, but that no single method offers a complete solution and that model choice significantly shapes results. We argue that in crowdsourced collections, where metadata is the direct product of engagement with living contributors, automated keyword extraction raises distinct stewardship responsibilities that must be addressed alongside technical performance. Open-weight, extractive models emerge from our evaluation as best placed to support responsible deployment, while generative AI, despite its abstractive potential, introduces accountability risks that anyone managing crowdsourced collections should weigh carefully.  \nKeywords  \nKeywords; Crowdsourced; NLP; NER; Keyword Extraction; Topic Modelling  \n1. Introduction  \nThe problem of keywords  \nKeywords are central to search and discovery across digital collections.1 They provide entry points into records, enable filtering and browsing, align with user search behaviour (Ehrmann et al. 2023), and support FAIR principles (Wilkinson et al. 2016) by helping to make data more findable, accessible, interoperable, and reusable. Commonly adopted metadata standards such as DataCite and Dublin Core therefore recommend their  \n1 This article uses the term ‘keywords’ broadly to refer to searchable descriptive terms attached to digital records in order to improve retrieval and discovery. These may include manually assigned tags, automatically extracted terms, named entities, or other searchable subject descriptors. Other literature refer to key terms or keyphrases.  \ninclusion (Salse et al. 2024) . Yet identifying, extracting, and assigning keywords at scale presents two interconnected challenges. The first is technical and practical: the process is labour-intensive and time-consuming, often necessitating additional resourcing or automation. This problem is compounded by wider GLAM (Galleries, Libraries, Archives, Museums)-sector challenges: digitised collections have grown faster than institutions ’capacity to connect users with specific content (Newman et al. 2010), while insufficient investment, constrained infrastructure, digital skills gaps, and pressure to digitise at scale all influence what metadata can be created and how discovery can be improved (Gosling et al. 2022; Bailey et al. 2024; Gooch et al. 2025; Drabczyk et al. 2025; Korn et al. 2024; Cebr 2025) . The second challenge is ethical: keywords necessarily simplify represented content, lack contextual depth and nuance, and, like all collections metadata, are not neutral (Long et al. 2017; Schwartz and Cook, 2002) . Those responsible for this process must therefore both make, and be answerable for, decisions about which terms are, and are not, included (Pacheco et al. 2023) .  \nOur case study  \nThese challenges are particularly pertinent in the context of crowdsourced collections, where metadata is the direct result of engagement with living contributors and imposed terms have the potential to misrepresent or mischaracterise","cbCaihBBoAwojXSb","https://ap.wps.com/l/cbCaihBBoAwojXSb","pdf",512933,1,45,"English","en",105,"# Introduction\n## The problem of keywords\n## Our case study\n## Purpose of the article","[{\"question\":\"为什么在数字化/众包文献中需要“关键词”？\",\"answer\":\"关键词在检索与发现中起核心作用：为记录提供入口、支持筛选与浏览，并帮助对齐用户搜索行为，同时有助于满足 FAIR 原则对可发现、可访问、可互操作与可再利用的要求。\"},{\"question\":\"本文指出在大规模生成关键词面临哪些主要挑战？\",\"answer\":\"挑战包括两个相互关联方面：技术与实践上，关键词识别与分配往往劳动密集、耗时并需要额外资源或自动化；伦理上，关键词会简化内容、缺乏情境与细节，并且像所有元数据一样并非中立。\"},{\"question\":\"EKCC 项目评估了哪些 NLP 方法，并得出了什么总体结论？\",\"answer\":\"项目评估三类 NLP 技术：命名实体识别（NER）、关键词提取与主题建模，并覆盖从传统统计到现代 GenAI 神经网络的多种实现方式。结论表明 NLP 在众包语料中具备规模化关键词提取的真实潜力，但不存在单一方法能完全解决问题，且模型选择会显著影响结果。\"}]",1784179502,113,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"letting-the-data-speak-extracting-keywords-from-crowdsourced-collections-with-ai","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/letting-the-data-speak-extracting-keywords-from-crowdsourced-collections-with-ai/82305/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-19","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"为什么在数字化/众包文献中需要“关键词”？","Question",{"text":75,"@type":76},"关键词在检索与发现中起核心作用：为记录提供入口、支持筛选与浏览，并帮助对齐用户搜索行为，同时有助于满足 FAIR 原则对可发现、可访问、可互操作与可再利用的要求。","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"本文指出在大规模生成关键词面临哪些主要挑战？",{"text":80,"@type":76},"挑战包括两个相互关联方面：技术与实践上，关键词识别与分配往往劳动密集、耗时并需要额外资源或自动化；伦理上，关键词会简化内容、缺乏情境与细节，并且像所有元数据一样并非中立。",{"name":82,"@type":73,"acceptedAnswer":83},"EKCC 项目评估了哪些 NLP 方法，并得出了什么总体结论？",{"text":84,"@type":76},"项目评估三类 NLP 技术：命名实体识别（NER）、关键词提取与主题建模，并覆盖从传统统计到现代 GenAI 神经网络的多种实现方式。结论表明 NLP 在众包语料中具备规模化关键词提取的真实潜力，但不存在单一方法能完全解决问题，且模型选择会显著影响结果。","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":45,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":45,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":45,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]