[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-120784-en":3,"doc-seo-120784-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},120784,34359740700684,"Finn","https://ap-avatar.wpscdn.com/avatar/1f400023980c374ae676?_k=1777273430885731487",8,"Research & Report","Machine Learning Model for Language Classification - Bag-of-words and Multilayer Perceptron","The study addresses the difficulty of classifying unstructured text by proposing a supervised machine learning model for language identification. Four language categories are considered: English, Indonesian, German, and French. It combines Bag-of-words for straightforward text representation with Multilayer Perceptron for learning more complex patterns. Data are collected through text mining by crawling Twitter, using 4,000 records. The resulting model reaches 98% accuracy with 0.14% loss, indicating strong performance for language classification from text data.","JITE, 7 (1) July 2023 ISSN 2549-6247 (Print) ISSN2549-6255 (Online)  \nJITE (Journal of Informatics and Telecommunication Engineering)  \nAvailable online [http://ojs.uma.ac.id/index.php/jite](http://ojs.uma.ac.id/index.php/jite) DOI : 10.31289/jite.v7i1.10114  \n| Received: 20 July 2023 | Accepted: 28 July 2023 | Published: 28 July 2023 |\n| --- | --- | --- |\n\nMachine Learning Model for Language Classification: Bag-of-words  \nand Multilayer Perceptron  \nDevi Hawana Lubis 1), Sawaluddin 2) & Ade Candra 3)  \n1,3) Master of Informatics Program, Universitas Sumatera Utara  \n2) Department of Mathematics, Universitas Sumatera Utara  \n*Coresponding Email: [devihawana@gmail.com](devihawana@gmail.com)  \nAbstrak  \nKetersediaan data saat ini telah menjadi aset besar bagi penelitianyang digunakan untuk berbagai keperluan sepertiuntuk pembelajaran mesin. Salah satu metode pembelajaran mesin dasar untuk pemrosesan bahasa alami adalah bag-of-words. Masalah dalam penelitian ini adalah sulitnya mengklasifikasikan teks karena teks masih memiliki karakteristik yang tidak terstruktur, sehingga penelitian ini akan menerapkan model untuk mengklasifikasikan bahasa teks. Teks akan ditempatkan dalam empat kategori, Inggris, Indonesia, Jerman dan Perancis. Penelitian dilakukan dengan menggunakan Bag-of-words dan Multilayer Perceptron untuk mengatasi masalah pembelajaranmesin terawasi ini. Penggunaan Bag-of-words untuk melakukan representasi teks untuk pola yang sederhana, pemrosesan yang mudah, dan kinerja yang baik. Di sisi lain, perceptron multilayer memiliki kemampuan untuk mempelajari pola data yang kompleks dalam bentuk gambar, teks, atau video. Penelitian ini akan mengumpulkan data dengan menggunakan teknik text mining yaitu crawling media sosial Twitter sebanyak 4000 record data. Penelitian ini menghasilkan model dengan akurasi 98 persen dengan loss 0,14 persen yang menunjukkan kinerja model yang baik dalam mengklasifikasikan bahasa berdasarkan data teks.  \nKata kunci: klasifikasi teks, multilayer perceptron, bag-of-words, model pembelajaran mesin  \nAbstract  \nThe availability of data today has become a great assetfor research that is used for various purposes such as for machine learning. One of the basic machine learning methods for natural language processing is bag-of-words. The problem in this study is the difficulty in classifying texts because texts still have unstructured characteristics, so this study will apply a model to classify the language of texts. Texts will be placed in four categories, English, Indonesian, German and French. Research was conducted using Bag-of-words and Multilayer Perceptron to solve this supervised machine learning problem. The use of Bag-of-words to perform text representation for simple patterns, easy processing and good performance. On the other hand, a multilayer perceptron has the ability to study complex data patterns in the form of images, text or videos. This study will collect data using text mining techniques, namely crawling Twitter social media as many as 4000 data records. This study produces a model with an accuracy of 98 percent with a loss of 0.14 percent which shows good model performance in classifying languages based on text data.  \nKeywords: text classification, multilayer perceptron, bag-of-words, machine learning model  \nHow to Cite: Lubis, D. H., Sawaluddin, S., & Candra, A. (2023) . Machine Learning Model for Language Classification: Bag-of-words and Multilayer Perceptron. JITE (Journal of Informatics and Telecommunication Engineering), 7(1), 356-365.  \nI. INTRODUCTION  \nMoment interference is happening in the brain when multilingual communication takes place (Grundy et al., 2017); (Liu et al., 2019); (García et al., 2017). Switching language from one to another is resulting in a dynamic interaction that affects both the speaker and listener. Furthermore, studies said that there is an occurrence of general adaptation when switching language on cross-language communication(Wang et al., 2022","cbCairYxc6pIXZE6","https://ap.wps.com/l/cbCairYxc6pIXZE6","pdf",758250,1,10,"English","en",105,"# Introduction\n# Literature Review\n## Related Research","[{\"question\":\"Which languages are used for the language classification task in this study?\",\"answer\":\"The model classifies texts into four languages: English, Indonesian, German, and French.\"},{\"question\":\"How is the text data collected for training the model?\",\"answer\":\"The study collects data by crawling Twitter using text mining techniques, resulting in 4,000 data records.\"},{\"question\":\"What model components are used and why?\",\"answer\":\"Bag-of-words is used to represent text for simple patterns with easy processing, while Multilayer Perceptron is used to learn complex data patterns from text.\"}]","Machine Learning Model for Language Classification - Bag-of-words and Multilayer Perceptron | PDF",1785732025,25,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"machine-learning-model-for-language-classification-bag-of-words-and-multilayer-perceptron","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/machine-learning-model-for-language-classification-bag-of-words-and-multilayer-perceptron/120784/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-03",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Which languages are used for the language classification task in this study?","Question",{"text":75,"@type":76},"The model classifies texts into four languages: English, Indonesian, German, and French.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How is the text data collected for training the model?",{"text":80,"@type":76},"The study collects data by crawling Twitter using text mining techniques, resulting in 4,000 data records.",{"name":82,"@type":73,"acceptedAnswer":83},"What model components are used and why?",{"text":84,"@type":76},"Bag-of-words is used to represent text for simple patterns with easy processing, while Multilayer Perceptron is used to learn complex data patterns from text.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,134],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":21,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":21,"slug":133},"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]