[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-121659-en":3,"doc-seo-121659-105":29,"detail-sidebar-cat-0-en-105":89},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":20,"language":21,"language_code":22,"site_id":23,"html_lang":22,"table_of_contents":24,"faqs":25,"seo_title":26,"seo_description":14,"update_tm":27,"read_time":28},121659,687197100911,"Himbo","https://ap-avatar.wpscdn.com/avatar/a000239b6f1da00475?x-image-process=image/resize,m_fixed,w_180,h_180&k=1785132997149421697",8,"Research & Report","Investigations Into Using Machine Learning Models to Automate the Sorting of Digitized Texas State Publications","For the past decade, UNT Libraries digitized Texas State Publications through the Texas State Depository Program, receiving thousands of items annually that must be manually categorized by content type before creating descriptive metadata. This work tests machine learning methods to automate cover-page-based sorting, using OCR-extracted text and experiments including Bag of Words, TF-IDF, BERT, Word2vec, Doc2vec, and CNN. Cover image preprocessing and CNN-based image classification are compared with text classification approaches.","Investigations Into Using Machine Learning Models to Automate the Sorting of Digitized Texas State Publications.  \nPraneeth Rikka (Graduate Research Assistant)  \nMark E. Phillips (Associate Dean of Digital Libraries)  \nAbstract Data Experiments  \nFor the past ten years, the UNT Libraries has digitized Texas State Publications as part of their role in the Texas State Depository Program . Each year, UNT receives a diverse collection of published materials from the program to digitize . We categorize these items according to content type before creating descriptive metadata. Every year, over 2000 items must be manually sorted and grouped to facilitate metadata creation, a task that is time consuming for the content specialist.  \nHere, we seek to test a machine learning model to automate this sorting. Cover page images are used as input for our model training as content experts typically sort published materials depending on the cover page . In this poster, we look into the popular text classification methods Bag of Words, TF-IDF, BERT, Word2vec, Doc2vec, and CNN.  \nBackground  \nThe Portal to Texas history is a repository operated by the UNT Libraries which works with partners across the state of Texas to host and provide access to their unique materials .  \nOne of these collections, The Texas State Publications Collection is focused on collecting and digitizing documents published by agencies and commissions in Texas . Currently there are over 19 , 000 records available in this collection. To provide access to these publications, we need to organize and create descriptive metadata for each item.  \nDuring this project, thousands of publications are sent to a vendor for digitization. As part of that workflow these digitized items are returned in the order they were scanned, which is not always optimal for processing, grouping like items, and creating metadata Sorting these digitized publications is a timeconsuming task carried out by metadata librarians .  \nThis project explored the use of machine models to assist librarians in the sorting and grouping of documents as they were returned from the vendor after digitization.  \nThe data used for this project were cover images from documents returned from the digitization workflow. These documents contain different kinds of records which are reports, books, budgets, serials, strategic plans and other miscellaneous types . The sample of them shown below fig:1 .  \nFig 1: Sample of the publication cover pages  \n\n| Model | Number of Features | Classification Algorithms |\n| --- | --- | --- |\n| Bag of Words | 400 | Decision Tree, Multinomial Naive Bayes |\n| TF-IDF | 400 | Decision Tree, Multinomial Naive Bayes |\n| BERT | 10, 20, 30, 40, 50 epochs |  |\n| Word2Vec | 100[vector size] | CBOW,SG |\n| Doc2Vec | 100[vector size] | DM, DOW,\u003Cbr>DM+DOW |\n\nTable 1: Models used in the text classification  \nImage Processing :  \nAt first, we conducted experiments using image-based models . Cover pages are in different rectangular shapes that required conversion into 500x500 pixel images . We used a Contrast Limited Adaptive Histogram equalization to improve the clarity of the different images . Next, we used a pretrained CNN model with 2 convolution layers and 3 dense layer . Here max pooling is applied after every convolution layer with size 3x3, the size of the dense layers used were 250 , 100 , and 8 . SoftMax is used because of categorical data with Adam as optimizer .  \nText Classification :  \nWe leveraged available textual information contained on the cover pages to help differentiate between categories . We applied Optical Character Recognition to extract text from the cover pages . After the extraction we applied text cleaning to remove stop words and non-standard characters . Then we tested different algorithms which are shown in Table 1 .  \nResults  \n\n| Model | Classification Algorithm | Accuracy(%) |\n| --- | --- | --- |\n| Bag of Words | Decision Tree | 66 |\n|  | Multinominal Naïve Bayes | 72 |\n| ","cbCaifvnB0JGl2EE","https://ap.wps.com/l/cbCaifvnB0JGl2EE","pdf",336533,1,"English","en",105,"# Abstract\n# Background\n# Image Processing\n# Text Classification\n# Results\n# Conclusion\n# Discussion","[{\"question\":\"What problem does this poster address?\",\"answer\":\"It addresses the manual time-consuming process of sorting and grouping digitized Texas State Publications before metadata creation.\"},{\"question\":\"What input data is used to train and evaluate the models?\",\"answer\":\"Cover page images returned from the digitization workflow are used, with OCR employed to extract text for text classification experiments.\"},{\"question\":\"Which model performed best in the results?\",\"answer\":\"BERT achieved the highest accuracy, reaching 94% at 40–50 epochs.\"}]","Investigations Into Using Machine Learning Models to Automate the Sorting of Digitized Texas State Publications | PDF",1785806022,3,{"code":4,"msg":30,"data":31},"ok",{"site_id":23,"language":22,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":84,"head_meta":86,"extra_data":88,"updated_unix":27},"investigations-into-using-machine-learning-models-to-automate-the-sorting-of-digitized-texas-state-publications","",{"@graph":35,"@context":83},[36,52,66],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,49],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":28},"https://docshare.wps.com/document/research-report/",{"item":50,"name":13,"@type":42,"position":51},"https://docshare.wps.com/document/investigations-into-using-machine-learning-models-to-automate-the-sorting-of-digitized-texas-state-publications/121659/",4,{"url":50,"name":13,"@type":53,"author":54,"headline":13,"publisher":56,"fileFormat":59,"inLanguage":22,"description":14,"dateModified":60,"datePublished":60,"encodingFormat":59,"isAccessibleForFree":61,"interactionStatistic":62},"DigitalDocument",{"name":9,"@type":55},"Person",{"url":40,"name":57,"@type":58},"DocShare","Organization","application/pdf","2026-08-04",true,{"@type":63,"interactionType":64,"userInteractionCount":4},"InteractionCounter",{"@type":65},"ViewAction",{"@type":67,"mainEntity":68},"FAQPage",[69,75,79],{"name":70,"@type":71,"acceptedAnswer":72},"What problem does this poster address?","Question",{"text":73,"@type":74},"It addresses the manual time-consuming process of sorting and grouping digitized Texas State Publications before metadata creation.","Answer",{"name":76,"@type":71,"acceptedAnswer":77},"What input data is used to train and evaluate the models?",{"text":78,"@type":74},"Cover page images returned from the digitization workflow are used, with OCR employed to extract text for text classification experiments.",{"name":80,"@type":71,"acceptedAnswer":81},"Which model performed best in the results?",{"text":82,"@type":74},"BERT achieved the highest accuracy, reaching 94% at 40–50 epochs.","https://schema.org",{"og:url":50,"og:type":85,"og:title":13,"og:site_name":57,"og:description":14},"article",{"robots":87,"canonical":50},"index,follow",{"doc_id":7,"site_id":23},{"code":4,"msg":5,"data":90},[91,95,99,103,108,113,118,121,126,129,133],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":92,"show_sort_weight":93,"slug":94},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":96,"show_sort_weight":97,"slug":98},"Literature",80,"literature",{"id":51,"doc_module":4,"doc_module_name":45,"category_name":100,"show_sort_weight":101,"slug":102},"Exam",70,"exam",{"id":104,"doc_module":4,"doc_module_name":45,"category_name":105,"show_sort_weight":106,"slug":107},5,"Comic",60,"comic",{"id":109,"doc_module":4,"doc_module_name":45,"category_name":110,"show_sort_weight":111,"slug":112},6,"Technology",50,"technology",{"id":114,"doc_module":4,"doc_module_name":45,"category_name":115,"show_sort_weight":116,"slug":117},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":119,"slug":120},30,"research-report",{"id":122,"doc_module":4,"doc_module_name":45,"category_name":123,"show_sort_weight":124,"slug":125},9,"Religion & Spirituality",20,"religion-spirituality",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":127,"show_sort_weight":124,"slug":128},"World Cup","world-cup",{"id":130,"doc_module":4,"doc_module_name":45,"category_name":131,"show_sort_weight":130,"slug":132},10,"Lifestyle","lifestyle",{"id":134,"doc_module":4,"doc_module_name":45,"category_name":135,"show_sort_weight":104,"slug":136},19,"General","general"]