[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-121620-en":3,"doc-seo-121620-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},121620,1099513958607,"Jiven","https://ap-avatar.wpscdn.com/avatar/100002390cf8733938c?x-image-process=image/resize,m_fixed,w_180,h_180&k=1778829742770036399",8,"Research & Report","Predictive keywords - Using machine learning to explain document characteristics","Keyword analysis is widely used in corpus linguistics to study discourse-domain characteristics, yet reliable evaluation of extracted keywords remains challenging. This work reframes keyword analysis as a prediction task by distinguishing target-corpus texts from reference-corpus texts. Using linear support vector machines, the method both quantifies discrimination and extracts keywords. Evaluation is performed with systematic metrics grounded in machine-learning notions of usefulness and relevance, and results are compared with the text dispersion keyness measure.","TYPE Original Research PUBLISHED 05 January 2023 DOI 10. 3389/frai.2022.975729  \nOPEN ACCESS  \nEDITED BY  \nJonathan Dunn,  \nUniversity of Canterbury, New Zealand  \nREVIEWED BY  \nDaniel Keller,  \nNorthern Arizona University, United States  \nTove Larsson,  \nNorthern Arizona University, United States  \nSeda Acikara,  \nNorthern Arizona University, United States in collaboration with reviewer TL  \n*CORRESPONDENCE  \nAki-Juhani Kyröläinen  \n [akkyro@gmail.com](akkyro@gmail.com)  \nSPECIALTY SECTION  \nThis article was submitted to Language and Computation, a section of the journal Frontiers in Artiﬁcial Intelligence  \nRECEIVED 22 June 2022  \nACCEPTED 07 December 2022  \nPUBLISHED 05 January 2023  \nCITATION  \nKyröläinen A-J and Laippala V (2023) Predictive keywords: Using machine learning to explain document characteristics.  \nFront. Artif. Intell. 5:975729 .  \ndoi: 10.3389/frai.2022.975729  \nCOPYRIGHT  \n© 2023 Kyröläinen and Laippala. This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY) . The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.  \nPredictive keywords: Using machine learning to explain document characteristics  \nAki-Juhani Kyröläinen* and Veronika Laippala  \nSchool of Languages and Translation Studies, University of Turku, Turku, Finland  \nWhen exploring the characteristics of a discourse domain associated with texts, keyword analysis is widely used in corpus linguistics. However, one of the challenges facing this method is the evaluation of the quality of the keywords. Here, we propose casting keyword analysis as a prediction problem with the goal of discriminating the texts associated with the target corpus from the reference corpus. We demonstrate that, when using linear support vector machines, this approach can be used not only to quantify the discrimination between the two corpora, but also extract keywords. To evaluate the keywords, we develop a systematic and rigorous approach anchored to the concepts of usefulness and relevance used in machine learning. The extracted keywords are compared with the recently proposed text dispersion keyness measure. We demonstrate that that our approach extracts keywords that are highly useful and linguistically relevant, capturing the characteristics of their discourse domain.  \nKEYWORDS  \nkeyness, keyword, corpus linguistics, support vector machines, machine learning  \n1. Introduction  \nIntuitively, some elements of a text are more important than others in informing readers about the text’s characteristics. In corpus linguistics, this intuitive concept has been developed into a method that is referred to as keyword analysis (for recent overviews see Gabrielatos and Marchi, 2011; Egbert and Biber, 2019; Gries, 2021) . Over the years, keyword analysis has become an instrumental part of quantitative text analysis in corpus linguistics as a way to examine the characteristics of various text varieties ranging from news articles to erotic narratives, through the contribution of words or other linguistic elements (see Gabrielatos and Marchi, 2011; Egbert and Biber, 2019, fora comprehensive overview of studies) .  \nRecently, there has been an interest in methodological development of keyword analysis, as exempli􀀂ed by such studies as Egbert and Biber (2019) and Gries (2021) . The present study is situated against this backdrop. We present a new approach for a keyword analysis that is based on prediction rather than statistical calculation. We exemplify this approach by examining the characteristics of a corpus featuring two text varieties: news and blogs. By using linear support vector machines as classi􀀂ers, this approach allows us not only to predict the","cbCaioT8rfi3rZ2X","https://ap.wps.com/l/cbCaioT8rfi3rZ2X","pdf",1342230,1,23,"English","en",105,"# Introduction\n## Keywords and keyness in corpus linguistics","[{\"question\":\"How does the paper treat keyword analysis instead of traditional statistical keyword extraction?\",\"answer\":\"It recasts keyword analysis as a prediction problem, training a classifier to discriminate texts from a target corpus versus a reference corpus.\"},{\"question\":\"What model is used to predict document categories and extract keywords?\",\"answer\":\"The study uses linear support vector machines as classifiers, enabling both prediction of text variety and extraction of informative keywords.\"},{\"question\":\"How are the extracted keywords evaluated in the proposed framework?\",\"answer\":\"The paper develops rigorous evaluation metrics anchored in machine-learning concepts of usefulness and relevance, and compares the extracted keywords with the text dispersion keyness measure.\"}]","Predictive keywords - Using machine learning to explain document characteristics | PDF",1785805698,58,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"predictive-keywords-using-machine-learning-to-explain-document-characteristics","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/predictive-keywords-using-machine-learning-to-explain-document-characteristics/121620/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-04",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"How does the paper treat keyword analysis instead of traditional statistical keyword extraction?","Question",{"text":75,"@type":76},"It recasts keyword analysis as a prediction problem, training a classifier to discriminate texts from a target corpus versus a reference corpus.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What model is used to predict document categories and extract keywords?",{"text":80,"@type":76},"The study uses linear support vector machines as classifiers, enabling both prediction of text variety and extraction of informative keywords.",{"name":82,"@type":73,"acceptedAnswer":83},"How are the extracted keywords evaluated in the proposed framework?",{"text":84,"@type":76},"The paper develops rigorous evaluation metrics anchored in machine-learning concepts of usefulness and relevance, and compares the extracted keywords with the text dispersion keyness measure.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]