[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-120675-en":3,"doc-seo-120675-105":30,"detail-sidebar-cat-0-en-105":83},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},120675,1099513958762,"Logic","https://ap-avatar.wpscdn.com/avatar/1000023916a998db790?x-image-process=image/resize,m_fixed,w_180,h_180&k=1784791008015729253",8,"Research & Report","Clinical Text Classification with Word Representation Features and Machine Learning Algorithms","Clinical text classification of electronic medical records is a difficult task due to irrelevant content, misspellings, semantic ambiguity, and abbreviation noise. The paper presents an intelligent machine-learning framework for classifying a medical transcription dataset through four phases: text preprocessing, word representation, feature reduction, and classification. Four algorithms—support vector machines, naïve bayes, logistic regression, and k-nearest neighbors—are tested with bag of words, TF-IDF, and word2vec. Evaluation uses precision, recall, accuracy, and F1.","Paper—Clinical Text Classification with Word Representation Features and Machine Learning Algorithms  \nClinical Text Classification with Word Representation Features and Machine Learning Algorithms  \n[https://doi.org/10.3991/ijoe.v19i04.36099](https://doi.org/10.3991/ijoe.v19i04.36099)  \nLaiali Almazaydeh1(􀀍), Mohammed Abuhelaleh2, Arar Al Tawil3, Khaled Elleithy4 1Department of Software Engineering, Al-Hussein Bin Talal University, Ma’an, Jordan 2Department of Computer Science, Al-Hussein Bin Talal University, Ma’an, Jordan 3AbdulAziz Al Ghurair School of Advanced Computing, Luminus Technical University  \nCollege, Amman, Jordan  \n4Department of Computer Science and Engineering, University of Bridgeport, Bridgeport, USA [laiali.almazaydeh@ahu.edu.jo](laiali.almazaydeh@ahu.edu.jo)  \nAbstract—Clinical text classification of electronic medical records is a challenging task. Existing electronic records suffer from irrelevant text, misspellings, semantic ambiguity, and abbreviations. The approach reported in this paper elaborates on machine learning techniques to develop an intelligent framework for classification of the medical transcription dataset. The proposed approach is based on four main phases: the text preprocessing phase, word representation phase, features reduction phase and classification phase. We have used four machine learning algorithms, support vector machines, naïve bayes, logistic regression and k-nearest neighbors in combination with different word representation models. We have applied the four algorithms to the bag of words, to TF-IDF, to word2vec. Experimental results were evaluated based on precision, recall, accuracy and F1 score. The best results were obtained with the combination of the k-NN classifier, and the word represented by Word2vec achieving an accuracy of 92% to correctly classify the medical specialties based on the transcription text.  \nKeywords—clinical text, classification, logistic regression, support vector machines, k-nearest neighbors, naïve bayes, word of bag, TF-IDF, word2vec  \n1 Introduction  \nNatural Language Processing (NLP) is a process utilized for analyzing programming languages to help develop tools and techniques that can allow computers to perform various tasks based on the knowledge of humans. It has gained widespread attention due to its numerous applications in various fields like machine translation, email spamming detection, medical inquiries, and question answering. Currently, NLP field consists variety of cognitive models, linguistic theories, and engineering issues; and its applications can be classified into different classifications. One of them is to classify it into: speech recognition, natural language interface, story understanding, text generations, discourse Management, and text classification [1–2] .  \nText classification is an important task in NLP which is a process of extracting structured and unstructured information from a text and then classify it according to  \nPaper—Clinical Text Classification with Word Representation Features and Machine Learning Algorithms  \nsome rules. One of the most important fields of text classification is the clinical text classification, where a text is extracted from the Electronic Medical Records (EMR) and classified into different classifications. The clinical text classification process involves different tasks that may differ from system to another. For example, clinical text classification may provide smoking-status detection, patient classifications for different purposes, and classification of obesity [3–4]. EMR are widely adapted in medical providers’ systems, which led to creating large containers of structured and unstructured clinical data. These data can be classified to be leveraged in biomedical research and in health care delivering. On the other side, Text classification of EMR involves specific challenges compared to other related systems. These challenges include: themisbalancing of the dataset, misspelling","cbCaijIyWmmokGdp","https://ap.wps.com/l/cbCaijIyWmmokGdp","pdf",964067,1,12,"English","en",105,"# 1 Introduction\n# 2 Related works\n# 3 Proposed methodology\n## 3.1 Text preprocessing\n## 3.2 Word representation and feature reduction\n## 3.3 Classification with ML algorithms\n# 4 Experimental results and evaluation\n# 5 Conclusion and future works","[{\"question\":\"Which machine learning algorithms and input representations performed best?\",\"answer\":\"The best reported results come from combining the k-NN classifier with word representations produced by Word2vec, achieving about 92% accuracy for correctly classifying medical specialties based on the transcription text.\"}]","Clinical Text Classification with Word Representation Features and Machine Learning Algorithms | PDF",1785731300,30,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":78,"head_meta":80,"extra_data":82,"updated_unix":28},"clinical-text-classification-with-word-representation-features-and-machine-learning-algorithms","",{"@graph":36,"@context":77},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/clinical-text-classification-with-word-representation-features-and-machine-learning-algorithms/120675/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-03",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71],{"name":72,"@type":73,"acceptedAnswer":74},"Which machine learning algorithms and input representations performed best?","Question",{"text":75,"@type":76},"The best reported results come from combining the k-NN classifier with word representations produced by Word2vec, achieving about 92% accuracy for correctly classifying medical specialties based on the transcription text.","Answer","https://schema.org",{"og:url":52,"og:type":79,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":81,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":84},[85,89,93,97,102,107,112,114,119,122,126],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":86,"show_sort_weight":87,"slug":88},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":90,"show_sort_weight":91,"slug":92},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Exam",70,"exam",{"id":98,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},5,"Comic",60,"comic",{"id":103,"doc_module":4,"doc_module_name":46,"category_name":104,"show_sort_weight":105,"slug":106},6,"Technology",50,"technology",{"id":108,"doc_module":4,"doc_module_name":46,"category_name":109,"show_sort_weight":110,"slug":111},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":29,"slug":113},"research-report",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},9,"Religion & Spirituality",20,"religion-spirituality",{"id":117,"doc_module":4,"doc_module_name":46,"category_name":120,"show_sort_weight":117,"slug":121},"World Cup","world-cup",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":123,"slug":125},10,"Lifestyle","lifestyle",{"id":127,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":98,"slug":129},19,"General","general"]