[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-121422-en":3,"doc-seo-121422-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},121422,962075114101,"Seraphina","https://ap-avatar.wpscdn.com/avatar/e000253a75eb197efd?x-image-process=image/resize,m_fixed,w_180,h_180&k=1780044092746381165",8,"Research & Report","Machine learning and rule-based embedding techniques for classifying text documents","Rapid growth of electronic document archives and online information makes text-document categorization increasingly difficult, despite its value for information retrieval grounded in conceptual organization. This study tackles efficient document classification by combining machine learning models with a new method, W2vRule, a word-to-vector rule-based framework. Performance is evaluated against traditional baselines using the Reuters Newswire dataset, stressing the need for hyperparameter tuning. Results indicate W2vRule and machine learning distinguish key categories effectively, with rule-based methods outperforming Naive Bayes and BayesNet on standard metrics.","Int J Syst Assur Eng Manag  \n[https://doi.org/10.1007/s13198-024-02555-w](https://doi.org/10.1007/s13198-024-02555-w)  \nORIGINAL ARTICLE  \nMachine learning and rule‑based embedding techniques for classifying text documents  \nAsmaa M. Aubaid1 · Alok Mishra2 · Atul Mishra3  \nReceived: 3 May 2024 / Revised: 12 September 2024 / Accepted: 5 October 2024 © The Author(s) 2024  \nAbstract Rapid expansion of electronic document archives and the proliferation of online information have made it incredibly difficult to categorize text documents. Classification helps in information retrieval from a conceptual framework. This study addresses the challenge of efficiently categorizing text documents amidst the vast electronic document landscape. Employing machine learning models and a novel document categorization method, W2vRule, we compare its performance with traditional methods. Emphasizing the importance of tuning hyperparameters for optimal performance, the research recommends the W2vRule, a word-to-vector rule-based framework, for improved association-based text classification. The study used the Reuters Newswire dataset. Findings show that W2vRule and machine learning can effectively tell apart important categories. Rule-based approaches perform better than Naive Bayes, BayesNet, Decision Tables, and others in terms of performance metrics.  \nSupplementary Information The online version contains supplementary material available at [https://doi.org/10.1007/](https://doi.org/10.1007/)[ ](https://doi.org/10.1007/)s13198-024-02555-w.  \n* Alok Mishra [alok.mishra@ntnu.no](alok.mishra@ntnu.no)  \n[Asmaa M. Aubaid](Asmaa M. Aubaid)  \n[asalhmuh@gmail.com](asalhmuh@gmail.com)  \nAtul Mishra  \n[atul.mishra@bmu.ed.in](atul.mishra@bmu.ed.in)  \n1 Ministry of Higher Education and Scientific  \nResearch/Science and Technology, Baghdad/Al-Jadriya, Iraq  \n2 Faculty of Engineering, Norwegian University of Science and Technology, Trondheim, Norway  \n3 BML Munjal University, Kapriwas, India  \nKeywords Bayesian · BayesNet · Lazy · IBK · Naïve Bayes · IBL · Reuters · Text analytics  \n1 Introduction  \nText classification is a type of machine learning that puts open-ended text into a set of predefined categories. (Agrawal & Batra 2013) . Text classifiers can organize, structure, and classify almost any text, including documents, medical studies, files, and text from the web. Classifying a large amount of text allows for platform standardization, more relevant and efficient searches, and an improved user experience.(Batrinca & Treleaven 2015) . Thus, technologies like artificial intelligence (AI) and machine learning (ML) are becoming useful in many fields. Interpretable machine learning models facilitate making rational, data-driven conclusions.(Stiglic et al. 2020) . Interpretability is how well people understand a model’s decisions or how distinct features (inputs) generate an inevitable conclusion (output) . The text is complex, with a wide range of topics and terms. Text classification can categorize emails, and text genres. (Onan 2018), topic sarcasm (Onan 2019), Sentiment analysis (Balliet al. 2022) and Web queries. New research utilizes cuttingedge machine learning techniques rather than ontology and rule-based approaches.  \nText classification is essential in applications like information retrieval (Dwivedi & Arya 2016), e-government (Ku & Leroy 2014), data filtration (Melville et al. 2009), text archives (Tao et al. 2020), and digital libraries (Deng et al. 2019) . It involves the categorization of text data into predefined classes based on its content. Machine learning algorithms, such as Support Vector Machines (SVMs), Naive Bayes, and Deep Neural Networks (DNNs), can be employed to perform text classification. Recent text classification  \nresearch has focused on categorizing news and related stories. Recent research on text categorization has concentrated mainly on classifying the information and other relevant items. This study compares the novel doc","cbCail5xs6UMbSkR","https://ap.wps.com/l/cbCail5xs6UMbSkR","pdf",1559993,1,16,"English","en",105,"# Introduction\n## Text classification background and applications\n## Motivation and study objective\n# Proposed approach: W2vRule\n## Word2vec-based feature selection via rules\n## Hyperparameter tuning for performance\n# Dataset and evaluation\n## Reuters Newswire (Reuters-21578)\n## Feature selection metrics and rules-based RBML strategies\n# Results and discussion\n## Comparison with traditional methods\n## Performance metrics and findings","[{\"question\":\"What problem does the study address in text document classification?\",\"answer\":\"It addresses the challenge of efficiently categorizing text documents in large electronic archives where online information makes classification difficult.\"},{\"question\":\"What is W2vRule and how does it work?\",\"answer\":\"W2vRule is a word-to-vector rule-based framework that leverages Word2vec for efficient feature selection, then applies rule-based strategies for document categorization.\"},{\"question\":\"Which dataset and baseline methods are used for evaluation?\",\"answer\":\"The study uses the Reuters Newswire dataset (Reuters-21578) and compares W2vRule and machine learning against traditional methods such as Naive Bayes, BayesNet, and Decision Tables.\"}]","Machine learning and rule-based embedding techniques for classifying text documents | PDF",1785735594,40,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"machine-learning-and-rule-based-embedding-techniques-for-classifying-text-documents","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/machine-learning-and-rule-based-embedding-techniques-for-classifying-text-documents/121422/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-03",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the study address in text document classification?","Question",{"text":75,"@type":76},"It addresses the challenge of efficiently categorizing text documents in large electronic archives where online information makes classification difficult.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What is W2vRule and how does it work?",{"text":80,"@type":76},"W2vRule is a word-to-vector rule-based framework that leverages Word2vec for efficient feature selection, then applies rule-based strategies for document categorization.",{"name":82,"@type":73,"acceptedAnswer":83},"Which dataset and baseline methods are used for evaluation?",{"text":84,"@type":76},"The study uses the Reuters Newswire dataset (Reuters-21578) and compares W2vRule and machine learning against traditional methods such as Naive Bayes, BayesNet, and Decision Tables.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,119,122,127,130,134],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":29,"slug":118},7,"Healthcare","healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]