[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-125257-en":3,"doc-seo-125257-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},125257,687197207057,"Sage","https://ap-avatar.wpscdn.com/davatar_29158cc5080c5b710cf443261637dec0",8,"Research & Report","Sentiment Analysis of Multilingual Roman Text for E-Commerce Reviews using Machine Learning Approaches","Sentiment analysis, a natural language processing technique, extracts subjective information such as attitudes, opinions, and emotions from text. This paper builds a machine learning-based sentiment analysis model for multilingual e-commerce product reviews written in Roman Urdu and Roman Sindhi to assess whether public reactions to posts and products are negative, positive, or neutral. Reviews are collected from diverse online platforms including YouTube, Facebook, TikTok, Daraz, and Instagram. TF-IDF features and several classifiers (LR, NB, SVM, RF, KNN) are evaluated using precision, recall, and F1-score, with SVM and LR performing best; SMOTE improves accuracy under class imbalance.","VFAST Transactions on Software Engineering Volume 13, Issue 1, 2025  \nVFAST Transactions on Software Engineering  \n[https://vfast.org/journals/index.php/VTSE@ 2025](https://vfast.org/journals/index.php/VTSE@ 2025), ISSN(e): 2309-3978, ISSN(p): 2411-6246  \nVolume 13, Number 1, January-March 2025 pp: 131-140  \nKeywords: Sentiment Analysis, Multilingual Roman Text Reviews, Product Reviews Sentiment Analysis,  \nSentiment Analysis  \nusing Machine Learning.  \nJournal Info:  \nSubmitted:  \nFebruary 2, 2025 Accepted:  \nMarch 16, 2025  \nPublished:  \nMarch 28, 2025  \nSentiment Analysis of Multilingual Roman Text for E-Commerce Reviews using Machine Learning Approaches  \nSana Riaz 1 , Sarfraz Natha 1, 2* , Asghar Ali Chandio3 , Mehwish Leghari4 , Abeer Javed Syed5  \n1 Department of Information Technology, Quaid e Awam University of Engineering, Science & Technology, Nawabshah, Pakistan; 2* Department of Software Engineering, Sir Syed University of Engineering & Technology, Karachi, Pakistan; 3 School of Engineering and Information Technology, University of New South Wales, Canberra, Australia; 3 Department of Artiﬁcial Intelligence, Quaid-e-Awam University of Engineering, Science & Technology, Pakistan; 4 Department of Data Science, Quaid-e-Awam University of Engineering, Science & Technology, Pakistan; 5 Department of Computer Science, IQRA University, Pakistan.  \nAbstract Sentiment analysis, a type of natural language processing (NLP) analyzes the text data to extract and identify subjective information including attitudes, opinions, and feelings. Sentiment analysis can be used to examine audience feedback and reviews in the context of multilingual product reviews. In this paper, a sentiment analysis model using machine learning approaches has been developed for multilingual product reviews in Roman Urdu or Sindhi to determine how the public feels about certain posts, products, etc. The importance of sentiment analysis for product context reviews in many languages in Roman is multifaceted. It can offer insightful information on the preferences of the likes and dislikes of the audience. To accomplish multilingual sentiment analysis, adataset of reviews in Roman Urdu and Sindhi languages from diverse online platforms and social media sources like YouTube, Facebook, TikTok, Daraz, and Instagram was collected. To identify pertinent features essential for categorizing reviews into negative, positive, or neutral sentiments based on polarity, the Term Frequency Inverse Document Frequency (TF-IDF) method was used. For classiﬁcation, ﬁve different machine learning classiﬁers including Linear Regression (LR), Naive Bayes (NB), Support Vector Machine (SVM), Random Forest (RF), and K-nearest neighbors (KNN) were used. The classiﬁcation results were measured in terms of precision score, recall score, and F1-score. With TF-IDF, the SVM, and LR outperformed than other classiﬁers and obtained an F1-score of 0.77%, and 0.78% . To further improve the classiﬁcation accuracy, the Synthetic Minority Over-sampling TEchnique (SMOTE) was used to manage the class imbalance problem. With SMOTE, the classiﬁcation accuracy of LR and SVM was improved to 0.79% and 0.80% .  \n*Correspondence author email address: [sasattar@ssuet.edu.pk](sasattar@ssuet.edu.pk)[ ](sasattar@ssuet.edu.pk)DOI: 10.21015/vtse.v13i1 .2067  \nThis work is licensed under a Creative Commons Attribution 3.0 License.  \nVFAST Transactions on Software Engineering Volume 13, Issue 1, 2025  \n1 Introduction  \nOpinion prospecting, another name for Sentiment analysis, has gained more attraction and importance in the last few years [1, 2] . This ﬁeld stands out for the potential beneﬁts it offers as well as its increasing popularity. This expansion has been made possible largely by technological advancements and the growth of the internet. As a result, there is now much more data that is easily accessible for analysis, which presents both new opportunities and challenges [3] . As social media, online revie","cbCaimkTXnbk0edB","https://ap.wps.com/l/cbCaimkTXnbk0edB","pdf",235547,1,10,"English","en",105,"# Abstract\n# Introduction\n## Sentiment analysis overview and motivation\n## Applications and challenges","[{\"question\":\"What languages and data sources are used for the sentiment analysis model?\",\"answer\":\"The model targets multilingual Roman Urdu and Roman Sindhi reviews. The dataset is collected from multiple online platforms such as YouTube, Facebook, TikTok, Daraz, and Instagram.\"},{\"question\":\"How are review texts converted into features for classification?\",\"answer\":\"TF-IDF is used to identify pertinent features for categorizing reviews into negative, positive, or neutral sentiments based on polarity.\"},{\"question\":\"Which machine learning methods perform best, and how is class imbalance handled?\",\"answer\":\"Using TF-IDF, SVM and LR outperform other classifiers, achieving strong F1-scores. SMOTE is applied to manage class imbalance, improving classification accuracy for LR and SVM further.\"}]","Sentiment Analysis of Multilingual Roman Text for E-Commerce Reviews using Machine Learning Approaches | PDF",1785897761,25,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"sentiment-analysis-of-multilingual-roman-text-for-e-commerce-reviews-using-machine-learning-approaches","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/sentiment-analysis-of-multilingual-roman-text-for-e-commerce-reviews-using-machine-learning-approaches/125257/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-05",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What languages and data sources are used for the sentiment analysis model?","Question",{"text":75,"@type":76},"The model targets multilingual Roman Urdu and Roman Sindhi reviews. The dataset is collected from multiple online platforms such as YouTube, Facebook, TikTok, Daraz, and Instagram.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How are review texts converted into features for classification?",{"text":80,"@type":76},"TF-IDF is used to identify pertinent features for categorizing reviews into negative, positive, or neutral sentiments based on polarity.",{"name":82,"@type":73,"acceptedAnswer":83},"Which machine learning methods perform best, and how is class imbalance handled?",{"text":84,"@type":76},"Using TF-IDF, SVM and LR outperform other classifiers, achieving strong F1-scores. SMOTE is applied to manage class imbalance, improving classification accuracy for LR and SVM further.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,134],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":21,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":21,"slug":133},"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]