[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-127573-en":3,"doc-seo-127573-105":31,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":28,"seo_description":14,"update_tm":29,"read_time":30},127573,687207020761,"Patrick","https://ap-avatar.wpscdn.com/davatar_155a257f0dc6eb9ab79c44ca47cae57d",8,"Research & Report","Machine Learning NLP-based recommendation system on production issues - Master of Science thesis","Natural Language Processing (NLP) for information extraction has advanced rapidly in media, e-commerce, and online games, yet its use in manufacturing production quality control remains underdeveloped. This thesis builds a recommendation system that uses textual production issue descriptions to retrieve the most relevant cases. Manufacturing control data in Finnish is embedded into numerical vectors using TF-IDF, Word2Vec, spaCy, Sentence Transformers, and SBERT, then ranked by cosine distance. Turku NLP Sentence Transformer yields MAP@10 = 0.67, showing effective retrieval with comparatively limited data, and supporting optimization via additional data and online testing.","Xiaotian Bi  \nMachine Learning NLP-based recommendation system on production issues  \nSchool of Technology and Innovations Master of Science thesis  \nIndustrial Systems Analytics  \nUNIVERSITY OF VAASA  \nSchool of Technology and Innovations  \nAuthor: Xiaotian Bi  \nTitle of the thesis: Machine Learning NLP-based recommendation system on produc  \ntion issues  \nDegree: Master of Science in Technology  \nDiscipline: Industrial Systems Analytics  \nSupervisor: Mohammed Elmusrati  \nPetri Välisuo  \nYear: 2023 Pages: 77  \nABSTRACT :  \nThe techniques related to Natural Language Processing (NLP) as information extraction are increasingly popular in media, E-commerce, and online games. However, the application with such techniques is yet to be established for production quality control in the manufacturing industry.  \nThe goal of this research is to build a recommendation system based on production issue descriptions in a textual format. The data was extracted from a manufacturing control system where it has been collected in Finnish on a relatively good scale for years. Five different NLP methods (TF-IDF, Word2Vec, spaCy, Sentence Transformers and SBERT) are used for modelling, converting human digital written texts into numerical feature vectors. The most relevant issue cases could be retrieved by calculating the cosine distance between the query sentence vector and corpus embed matrix which represents the whole dataset. Turku NLP-based Sentence Transformer achieves the best result with Mean Average Precision @10 equal to 0.67, inferring that the initial dataset is large enough using deep learning algorithms competing with machine learning methods. Even though a categorical variable were chosen as a target variable to compute evaluation metrics, this research is not a classification problem with single variable for model training. Additionally, the metric selected for performance evaluation measures for every issue case. Therefore, it is not necessary to balance and split the dataset.  \nThis research work achieves a relatively good result with less data available compared to the size of data used for other businesses. The recommendation system can be optimized by feeding more data and implementing online testing. It also has the possibility to transform into collaborative filtering to find patterns of users instead of simply focusing on items, in the condition of comprehensive user information included.  \nKEYWORDS: NLP, Recommendation System, Mean Average Precision @K, Sentence Transformers, SBERT.  \nAcknowledgements  \nI would like to express my gratitude to my academic advisor, Prof. Mohammed Elmusrati and Petri Välisuo, providing incredible support and useful insights to make my thesis research proceeding smoothly.  \nI also appreciate that Mr. Christian Sundman offered this interesting topic to me, and other employees from company side gave valuable advice from MES and data science point of view.  \nContents  \n1 Introduction 7  \n2 Natural Language Processing (NLP) and Recommendation System 10  \n2.1 Natural Language Processing (NLP) 10  \n2.2 Recommendation System 12  \n2.3 NLP-based Recommendation system 14  \n3 Methodology 16  \n3.1 Content-based Recommendation System 16  \n3.2 Python Libaries for NLP 20  \n3.3 NLP Methods 21  \n3.3.1 Term Frequency-Inverse Document Frequency 21  \n3.3.2 Word2Vec 23  \n3.3.3 SpaCy 25  \n3.3.4 Sentence Transformers 28  \n3.3.5 SBERT 29  \n3.4 Evaluation Metrics 31  \n3.4.1 Decision Support Metrics 32  \n3.4.2 Ranking-based Metrics 34  \n3.4.3 Other Metircs 36  \n4 Case Study 38  \n4.1 Data Collection and Introduction 38  \n4.2 Data Pre-processing 39  \n4.3 Data Analysis and Visualization 41  \n4.4 Model Training and Prediction 44  \n4.4.1 Term Frequency-Inverse Document Frequency 45  \n4.4.2 Word2Vec 46  \n4.4.3 SpaCy 47  \n4.4.4 Sentence Transformers 48  \n4.4.5 SBERT 50  \n4.5 Results 51  \n5 Conclusions, Discussions and Future Works 55  \nReferences 58  \nAppendices 61  \nAppendix 1. Wordcloud for Each Class in “REASON_CLASS” 61  ","cbCainMTZ6EKa7qN","https://ap.wps.com/l/cbCainMTZ6EKa7qN","pdf",2520448,2,1,77,"English","en",105,"# Introduction\n# Natural Language Processing (NLP) and Recommendation System\n## Natural Language Processing (NLP)\n## Recommendation System\n## NLP-based Recommendation system\n# Methodology\n## Content-based Recommendation System\n## Python Libraries for NLP\n## NLP Methods\n## Evaluation Metrics\n# Case Study\n## Data Collection and Introduction\n## Data Pre-processing\n## Data Analysis and Visualization\n## Model Training and Prediction\n## Results\n# Conclusions, Discussions and Future Works\n# References\n# Appendices","[{\"question\":\"What is the main goal of this research?\",\"answer\":\"To build a recommendation system that retrieves relevant production issue cases from textual issue descriptions for manufacturing quality control.\"},{\"question\":\"How are the text descriptions converted for modelling?\",\"answer\":\"The thesis encodes written descriptions into numerical feature vectors using multiple NLP methods including TF-IDF, Word2Vec, spaCy, Sentence Transformers, and SBERT.\"},{\"question\":\"Which NLP approach performs best, and how is it evaluated?\",\"answer\":\"Turku NLP-based Sentence Transformer achieves the best results with Mean Average Precision at 10 (MAP@10) equal to 0.67, evaluated using ranking-oriented performance metrics per issue case.\"}]","Machine Learning NLP-based recommendation system on production issues - Master of Science thesis | PDF",1785940055,194,{"code":4,"msg":32,"data":33},"ok",{"site_id":25,"language":24,"slug":34,"title":13,"keywords":35,"description":14,"schema_data":36,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":29},"machine-learning-nlp-based-recommendation-system-on-production-issues-master-of-science-thesis","",{"@graph":37,"@context":86},[38,54,69],{"@type":39,"itemListElement":40},"BreadcrumbList",[41,45,48,51],{"item":42,"name":43,"@type":44,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":46,"name":47,"@type":44,"position":20},"https://docshare.wps.com/document/","Document",{"item":49,"name":12,"@type":44,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":44,"position":53},"https://docshare.wps.com/document/machine-learning-nlp-based-recommendation-system-on-production-issues-master-of-science-thesis/127573/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":42,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-23","2026-08-05",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What is the main goal of this research?","Question",{"text":76,"@type":77},"To build a recommendation system that retrieves relevant production issue cases from textual issue descriptions for manufacturing quality control.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How are the text descriptions converted for modelling?",{"text":81,"@type":77},"The thesis encodes written descriptions into numerical feature vectors using multiple NLP methods including TF-IDF, Word2Vec, spaCy, Sentence Transformers, and SBERT.",{"name":83,"@type":74,"acceptedAnswer":84},"Which NLP approach performs best, and how is it evaluated?",{"text":85,"@type":77},"Turku NLP-based Sentence Transformer achieves the best results with Mean Average Precision at 10 (MAP@10) equal to 0.67, evaluated using ranking-oriented performance metrics per issue case.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,121,124,129,132,136],{"id":21,"doc_module":4,"doc_module_name":47,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":47,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":47,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":47,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":47,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":47,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":47,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":125,"doc_module":4,"doc_module_name":47,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":47,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":133,"doc_module":4,"doc_module_name":47,"category_name":134,"show_sort_weight":133,"slug":135},10,"Lifestyle","lifestyle",{"id":137,"doc_module":4,"doc_module_name":47,"category_name":138,"show_sort_weight":107,"slug":139},19,"General","general"]