[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86211-en":3,"doc-seo-86211-105":29,"detail-sidebar-cat-0-en-105":90},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},86211,1374391974564,"Clementine","https://ap-avatar.wpscdn.com/avatar/14000253aa45c000a9e?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779874745381141002",8,"Research & Report","Characterising AI Models for Cataloguing","The creation of digital collections requires not only digitisation but also catalogue records, a slow and costly expert task. This project evaluates applying AI models to cataloguing by comparing different implementations and models. It includes qualitative and quantitative assessment of experiments and offers recommendations for using AI models beyond the immediate use case. Experiments focus on automated metadata creation from digitised content images.","Characterising AI Models for Cataloguing  \nMiguel Arana-Catania1 and Neil Jefferies1  \nJournal Title XX(X):1–7  \n©The Author(s) 2026  \nReprints and permission: [sagepub.co.uk/journalsPermissions.nav](sagepub.co.uk/journalsPermissions.nav)[ ](sagepub.co.uk/journalsPermissions.nav)DOI: 10.1177/ToBeAssigned [www.sagepub.com/](www.sagepub.com/)  \nSAGE  \nAbstract  \nThe creation of digital collections involves not only the digitisation of content, but also the creation of catalogue records for it. This often-overlooked task requires slow and costly expert manual work. In this project, we have evaluated the application of AI models to this task, comparing different implementations and models. This work includes a qualitative and quantitative evaluation of the experiments carried out, as well as recommendations on the use of AI models that go beyond the specific use case.  \narXiv :2607 . 1 1353v 1 [ cs .CL] 13 Jul 2026  \nKeywords  \ncataloguing, llms, multimodal, digital collections  \nIntroduction  \nThe broad objective of this work at the Bodleian Libraries was to build in-house understanding and expertise in AI tools based on Large Language Models (LLMs), with a particular focus on cataloguing-related tasks. In this project, this has been approached from two main angles:  \n1. Understand the performance of various AI and other automated tools in creating or extracting metadata from a variety of sources, in comparison with the benchmark records from the previous activity.  \n2. Understand the broader behaviour of AI tools that may influence future deployment, such as their reliability, consistency, and dependency on prompt construction.  \nAI technologies have been applied in diverse ways to the task of library cataloguing and classification. [1, 2] provide comprehensive reviews of these explorations. In this regard, prior to our project, [3] evaluated the generation of subject headings using human-generated metadata from the works as input; [4] used LLMs for topic classification using a controlled vocabulary; [5] worked with university theses, using AI to identify the ‘college’ field in the metadata record; [6] applied traditional machine learning techniques to the generation of Dewey Decimal Classification classes; [7] used machine learning to classify documents based on their images; [8] fine-tuned LLMs to generate specific metadata fields such as title, language or date; and [9] used LLMs to generate Library of Congress Classification classes.  \nWhilst these previous works offer interesting insights into the use of AI for specific tasks within the cataloguing process, our project aims for a much more ambitious goal. In this work, we employ an end-to-end approach using AI models for the entire cataloguing process. Our experiments employ scanned images of the content as input and produce the complete metadata record as output, in the appropriate encoding format (MARC, BIBTEX, JSON), which includes not only the content of the relevant fields but also the correct formatting code.  \nMethodology  \nIn this section, we describe the methodology used in this article. In brief, this has involved applying various LLMs to a collection of digitised document images with the aim of automating the creation of catalogue records for these documents.  \nBelow, we present details of the data used, as well as the experimental and evaluation methodology.  \nDataset  \nThe dataset used in this project was drawn from the Bodleian Libraries’ Global Dissertations collection that had been newly scanned. This, mainly 16th & 17th Century European, dataset had the advantage that it was very unlikely that the texts had previously been made available online, so the AI algorithms under test would not have been trained on them. The results obtained would thus be an accurate reflection of cataloguing performance rather than the retrieval of information already embedded in the AI model. A card catalogue for the collection was also scanned and used to provide a human-generat","cbCaiulT2wwUUlkO","https://ap.wps.com/l/cbCaiulT2wwUUlkO","pdf",146012,1,7,"English","en",105,"# Abstract\n## Introduction\n## Methodology\n### Dataset\n### Preprocessing\n### AI inference","[{\"question\":\"What task does the project target in digital collections?\",\"answer\":\"The project targets the creation of catalogue records for digitised content, aiming to automate metadata production that is usually done manually by experts.\"},{\"question\":\"How did the researchers evaluate AI models for cataloguing?\",\"answer\":\"They conducted both qualitative and quantitative evaluations of experiments, comparing AI outputs against benchmark records and human-generated ground truth from a scanned card catalogue.\"},{\"question\":\"What data and input format were used for AI inference?\",\"answer\":\"The experiments used scanned images from the Bodleian Libraries’ Global Dissertations collection, preprocessing them from archival TIFF to JPEG and resizing them before running the AI models.\"}]",1784209494,18,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":85,"head_meta":87,"extra_data":89,"updated_unix":27},"characterising-ai-models-for-cataloguing","",{"@graph":35,"@context":84},[36,53,67],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/characterising-ai-models-for-cataloguing/86211/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":61,"encodingFormat":60,"isAccessibleForFree":62,"interactionStatistic":63},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-16",true,{"@type":64,"interactionType":65,"userInteractionCount":4},"InteractionCounter",{"@type":66},"ViewAction",{"@type":68,"mainEntity":69},"FAQPage",[70,76,80],{"name":71,"@type":72,"acceptedAnswer":73},"What task does the project target in digital collections?","Question",{"text":74,"@type":75},"The project targets the creation of catalogue records for digitised content, aiming to automate metadata production that is usually done manually by experts.","Answer",{"name":77,"@type":72,"acceptedAnswer":78},"How did the researchers evaluate AI models for cataloguing?",{"text":79,"@type":75},"They conducted both qualitative and quantitative evaluations of experiments, comparing AI outputs against benchmark records and human-generated ground truth from a scanned card catalogue.",{"name":81,"@type":72,"acceptedAnswer":82},"What data and input format were used for AI inference?",{"text":83,"@type":75},"The experiments used scanned images from the Bodleian Libraries’ Global Dissertations collection, preprocessing them from archival TIFF to JPEG and resizing them before running the AI models.","https://schema.org",{"og:url":51,"og:type":86,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":88,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":91},[92,96,100,104,109,114,118,121,126,129,133],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":93,"show_sort_weight":94,"slug":95},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":97,"show_sort_weight":98,"slug":99},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":101,"show_sort_weight":102,"slug":103},"Exam",70,"exam",{"id":105,"doc_module":4,"doc_module_name":45,"category_name":106,"show_sort_weight":107,"slug":108},5,"Comic",60,"comic",{"id":110,"doc_module":4,"doc_module_name":45,"category_name":111,"show_sort_weight":112,"slug":113},6,"Technology",50,"technology",{"id":21,"doc_module":4,"doc_module_name":45,"category_name":115,"show_sort_weight":116,"slug":117},"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":119,"slug":120},30,"research-report",{"id":122,"doc_module":4,"doc_module_name":45,"category_name":123,"show_sort_weight":124,"slug":125},9,"Religion & Spirituality",20,"religion-spirituality",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":127,"show_sort_weight":124,"slug":128},"World Cup","world-cup",{"id":130,"doc_module":4,"doc_module_name":45,"category_name":131,"show_sort_weight":130,"slug":132},10,"Lifestyle","lifestyle",{"id":134,"doc_module":4,"doc_module_name":45,"category_name":135,"show_sort_weight":105,"slug":136},19,"General","general"]