[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-121287-en":3,"doc-seo-121287-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},121287,1099514067415,"Rowan","https://ap-avatar.wpscdn.com/avatar/100002539d78ffe74a7?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779092875211072502",8,"Research & Report","Information Extraction and Machine Learning for Archaeological Texts - Chapter 11 - practical approaches","Archaeology increasingly generates large volumes of textual data, making manual reading and inspection impractical. This chapter explains how computational methods can mine that information to support more efficient research and new analyses at scale. It introduces and discusses machine learning methods for extracting information from archaeological texts, supported by practical examples. Readers gain insight into the capabilities of text mining in archaeology, the current state of research, and the information needed to begin their own text analyses.","Information extraction and machine learning for archaeological texts  \nBrandsen, A.; Gonzalez-Perez, C.; Martin-Rodilla, P.; Pereira-Fariña, M.  \nCitation  \nBrandsen, A. (2023) . Information extraction and machine learning for archaeological texts. In C. Gonzalez-Perez, P. Martin-Rodilla, &  \nM. Pereira-Fariña (Eds.), Quantitative Archaeology and Archaeological Modelling (pp. 229-261) . Cham: Springer.  \ndoi:10.1007/978-3-031-37156-1_ 11  \nVersion: Publisher's Version  \nLicensed under Article 25fa Copyright  \nLicense:  \nAct/Law (Amendment Taverne)  \nDownloaded from:  [https://hdl.handle.net/1887/3714101](https://hdl.handle.net/1887/3714101)  \n[Note:](Note: To cite this publication please use the final published version)[ To cite this publication please use the final published version](Note: To cite this publication please use the final published version)[ ](Note: To cite this publication please use the final published version)(if applicable) .  \nChapter 11  \nInformation Extraction and Machine Learning for Archaeological Texts  \nAlex Brandsen  \nAbstract Archaeologists are creating ever-increasing amounts of textual data. So much in fact, that manual reading and inspection has become practically impossible. By leveraging computational approaches, it is possible to extract relevant information from this big data, allowing for more efﬁcient research and new analyses. In this chapter, methods and techniques to extract information from archaeological texts through Machine Learning are introduced and discussed, with a focus on practical examples. After reading the chapter, you should have a clear grasp on the possibilities of text mining in archaeology, the current state of research, and enough information to start your own text analyses.  \nKeywords Information extraction · Text mining · Machine learning · Data science  \n11.1 Introduction  \nIn the last ten years or so, archaeologists have started generating ‘big data’: information assets characterised by the four V’s: Volume, Velocity, Veracity, and Variety. Volume simply means the size of the data, generally meaning many gigabytes or terabytes of data. The Velocity is the speed at which data updates, and Veracity is a measure of how trustworthy data is, these V’s are generally less relevant to archaeology. Variety speaks to the level of heterogeneity in the data, and how fuzzy or unclear data is, something we do encounter regularly in archaeology. In short, big data is so unwieldy that it is not feasible to analyse it with conventional methods. This problem of having too much data has been described by multiple authors, with Bevan calling it a “data deluge”(Bevan, 2015, p. 1) and Vince noting “we are drowning in our own data”(Vince, 1996, p. 1) . Dealing with  \nA. Brandsen (􀀂)  \nFaculty of Archaeology, Leiden University, Leiden, Netherlands e-mail: [a.brandsen@arch.leidenuniv.nl](a.brandsen@arch.leidenuniv.nl)  \n© The Author(s), under exclusive license to Springer Nature Switzerland AG 2023 C. Gonzalez-Perez et al. (eds.), Discourse and Argumentation in Archaeology: Conceptual and Computational Approaches, Quantitative Archaeology and Archaeological Modelling, [https://doi.org/10.1007/978-3-031-37156-1_11](https://doi.org/10.1007/978-3-031-37156-1_11)  \n229  \n230 A. Brandsen  \nstructured data—such as databases and geospatial data—has received a fair share of our attention, but much less research is being done on processing and analysing unstructured information: the documents that archaeologists write (Bevan, 2015) .  \nThese texts do contain a wealth of information, and by using computational tools to access, extract, and combine information in the documents, we can perform new synthesising research on large scales. Due to the amount of text data, computational methods almost become a necessity: in the Netherlands alone more than 4000 excavation reports are produced each year, not to mention thousands of books, papers, and preprints as well. When we extrapolate that to the situation","cbCaiuXHlxYLLlxU","https://ap.wps.com/l/cbCaiuXHlxYLLlxU","pdf",691011,1,34,"English","en",105,"# 11.1 Introduction\n## Big data and why text is difficult to analyze\n# 11.2 Information Extraction Techniques\n## Natural Language Processing (NLP) and corpora\n## Text mining and Information Extraction (IE)","[{\"question\":\"Why is manual analysis of archaeological texts becoming impractical?\",\"answer\":\"Archaeologists face “big data” conditions where text volume is too large for conventional manual inspection. Computational processing becomes necessary to extract and synthesize information efficiently.\"},{\"question\":\"What are the main ideas behind information extraction (IE) in this chapter?\",\"answer\":\"IE turns unstructured text into a structured view of the information it contains. It includes tasks such as Named Entity Recognition and document classification.\"},{\"question\":\"Which machine learning-based approach does the chapter emphasize for archaeological texts?\",\"answer\":\"The chapter focuses on using machine learning to extract information from archaeological texts, discussing methods and techniques with practical examples to support real analyses.\"}]","Information Extraction and Machine Learning for Archaeological Texts - Chapter 11 - practical approaches | PDF",1785734920,86,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"information-extraction-and-machine-learning-for-archaeological-texts-chapter-11-practical-approaches","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/information-extraction-and-machine-learning-for-archaeological-texts-chapter-11-practical-approaches/121287/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-03",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why is manual analysis of archaeological texts becoming impractical?","Question",{"text":75,"@type":76},"Archaeologists face “big data” conditions where text volume is too large for conventional manual inspection. Computational processing becomes necessary to extract and synthesize information efficiently.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What are the main ideas behind information extraction (IE) in this chapter?",{"text":80,"@type":76},"IE turns unstructured text into a structured view of the information it contains. It includes tasks such as Named Entity Recognition and document classification.",{"name":82,"@type":73,"acceptedAnswer":83},"Which machine learning-based approach does the chapter emphasize for archaeological texts?",{"text":84,"@type":76},"The chapter focuses on using machine learning to extract information from archaeological texts, discussing methods and techniques with practical examples to support real analyses.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]