[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-120699-en":3,"doc-seo-120699-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},120699,687197100911,"Himbo","https://ap-avatar.wpscdn.com/avatar/a000239b6f1da00475?x-image-process=image/resize,m_fixed,w_180,h_180&k=1785132997149421697",8,"Research & Report","A Cloud-based Machine Learning Pipeline for the Efficient Extraction of Insights from Customer Reviews - Paper","Cloud-based machine learning enables scalable natural language processing, yet extracting domain-specific insights from customer reviews remains challenging due to text noise and varying content characteristics. This work proposes an end-to-end pipeline integrating transformer-based topic modeling, vector-embedding keyword extraction, and clustering to improve efficient information extraction and topic modeling quality. The system is evaluated against state-of-the-art methods on publicly available datasets, demonstrating improved results for both topic modeling and keyword extraction while supporting practical service use.","A Cloud-based Machine Learning Pipeline for the Efficient Extraction of Insights from Customer Reviews  \nRóbert Lakatos, Gerg Bogacsovics*, Balázs Harangi*, István Lakatos*, Attila Tiba*, János Tóth*, Marianna Szabó*, András Hajdu*  \narXiv :2306 .07786v1 [ cs .CL] 13 Jun 2023  \nAbstract—The efficiency of natural language processing has improved dramatically with the advent of machine learning models, particularly neural networkbased solutions. However, some tasks are still challenging, especially when considering specific domains. In this paper, we present a cloud-based system that can extract insights from customer reviews using machine learning methods integrated into a pipeline. For topic modeling, our composite model uses transformer-based neural networks designed for natural language processing, vector embedding-based keyword extraction, and clustering. The elements of our model have been integrated and further developed to meet better the requirements of efficient information extraction, topic modeling of the extracted information, and user needs. Furthermore, our system can achieve better results than this task’s existing topic modeling and keyword extraction solutions. Our approach is validated and compared with other stateof-the-art methods using publicly available datasets for benchmarking.  \nIndex Terms—natural language processing; machine learning; neural networks; unsupervised learning; clustering; keyphrase extraction; topic modeling  \nI. INTRODUCTION  \nUsers of social platforms, forums, and online stores generate a significant amount of textual data. One of the most useful applications of machine learning-based text processing is to find words and phrases that describe the content of these texts. In e-commerce, the knowledge contained in data such as customer reviews can be of great value and provide a tangible and measurable financial return. However, it is impossible to efficiently extract information from such large amounts of data using human labor alone.  \nThe difficulty in solving this problem effectively with automated methods is that human-generated texts often contain a lot of noise in addition to substantive details. Filtering the relevant information is further complicated by the fact that different texts can have different characteristics. For example, the document to be analyzed may contain words too common to be distinctive or, for example, information irrelevant to the analysis objective. In fact, different parts of  \n* These authors contributed equally.  \nthe text may be considered noise, depending on how we view the data and what we think is relevant. This in turn makes it difficult to solve this task: it is not enough to find some specific information in texts, but we also have to decide what information we need based on the texts.  \nOur aim is to extract information from textual data in the field of e-commerce. Our application is an end-to-end system that runs on a cloud-based infrastructure and can be used as a service by small and medium-sized businesses. Our system uses machine learning tools developed for natural language processing and can identify those sets of words and phrases in customer reviews that characterize their opinion. We have built a system based on machine learning solutions that effectively handles such text-processing tasks and, in some aspects, outperforms currently available approaches.  \nTo build an application that can be used in an e-commerce environment, we needed a model that could identify topics in texts and provide a way to determine which topics are relevant, given our analysis goals. Therefore, before developing our system, we investigated the N-gram model [1], dependency parsing [2] and embedded vector space-based keyword extraction solutions, and various distance or density-based and hierarchical clustering [3] [4] techniques. In addition, we tested the LDA [5], [6], Top2Vec [7], and BERTopic [8] complex topic modeling methods. We focused on these tools beca","cbCaidSUu1gzWyq4","https://ap.wps.com/l/cbCaidSUu1gzWyq4","pdf",999877,1,12,"English","en",105,"# Abstract\n# Introduction\n## Problem background: noisy customer text\n## End-to-end cloud application for e-commerce\n## Prior methods and modeling choices\n## Goals: extract opinion phrases and group topics\n## Handling noise and sentence-level difficulties","[{\"question\":\"What is the main goal of the proposed system?\",\"answer\":\"Extract actionable insights from customer reviews by identifying word/phrase sets that characterize customer opinions and organizing them into relevant topics.\"},{\"question\":\"How does the pipeline perform topic modeling and keyword extraction?\",\"answer\":\"It uses a composite approach: transformer-based neural networks for topic modeling, vector embedding-based keyword extraction, and clustering to group extracted information.\"},{\"question\":\"What was used to validate and compare the system?\",\"answer\":\"Publicly available datasets were used for benchmarking, and the results were compared with other state-of-the-art topic modeling and keyword extraction methods.\"}]","A Cloud-based Machine Learning Pipeline for the Efficient Extraction of Insights from Customer Reviews - Paper | PDF",1785731610,30,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"a-cloud-based-machine-learning-pipeline-for-the-efficient-extraction-of-insights-from-customer-reviews-paper","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/a-cloud-based-machine-learning-pipeline-for-the-efficient-extraction-of-insights-from-customer-reviews-paper/120699/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-03",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What is the main goal of the proposed system?","Question",{"text":75,"@type":76},"Extract actionable insights from customer reviews by identifying word/phrase sets that characterize customer opinions and organizing them into relevant topics.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the pipeline perform topic modeling and keyword extraction?",{"text":80,"@type":76},"It uses a composite approach: transformer-based neural networks for topic modeling, vector embedding-based keyword extraction, and clustering to group extracted information.",{"name":82,"@type":73,"acceptedAnswer":83},"What was used to validate and compare the system?",{"text":84,"@type":76},"Publicly available datasets were used for benchmarking, and the results were compared with other state-of-the-art topic modeling and keyword extraction methods.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,122,127,130,134],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":29,"slug":121},"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]