[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-123920-en":3,"doc-seo-123920-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},123920,4398048950312,"Violet","https://ap-avatar.wpscdn.com/avatar/400002538284de19e3c?_k=1778320343897328908",8,"Research & Report","Machine Learning Approaches to Enhance Revenue Capture in E-commerce Platforms - Thesis Abstract","This thesis investigates machine learning methods for increasing revenue capture in e-commerce platforms amid rapidly changing online retail conditions. A Markov Chains-based algorithm assigns importance scores to individual webpages in terms of potential revenue generation. In parallel, clustering models for mixed data types identify detailed user groups and estimate potentially missing revenue linked to website malfunction. Variable selection is supported by Random Forest feature analysis to reveal latent opportunities. The findings characterize user behavior and platform efficiency, enabling data-driven personalization and growth strategies.","University of Padova  \nDepartment of Physics and Astronomy ”Galileo Galilei”  \nMaster Thesis in Physics of Data  \nMachine Learning Approaches to Enhance Revenue Capture in E-commerce  \nPlatforms  \nSupervisor Master Candidate  \nProf. Marco Baiesi Lorenzo Ausilio  \nUniversity of Padova  \nCo-supervisor Student ID  \nMattiaZoccarato 2046831  \nAcademic Year  \n2023-2024  \nii  \niv  \nAbstract  \nThis thesis explores techniques on how to increase revenue in e-commerce platforms through machine learning approaches, a critical endeavor in the face of rapidly evolving online retail landscapes.  \nAn algorithm utilizing Markov Chains was implemented to assign scores to individual webpages, reflecting their importance in terms of potential revenue generation.  \nIn a parallel analysis, the thesis delves into clustering techniques for mixed data types, utilizing k-prototype clustering and HDBSCAN coupled with UMAP for dimensionality reduction, to identify nuanced user clusters and compute the potential missing revenue from themalfunction of the e-commerce website. The selection of variables for this clustering was informed by a Random Forest analysis, ensuring an empirical approach to uncovering latent revenue opportunities. The results reveal significant insights into user behavior and platform efficiency, guiding e-commerce platforms in tailoring user experiences and maximizing revenue capture.  \nThe discussion of these findings highlights the transformative potential of integrating machine learning methodologies into e-commerce strategies. By computing critical engagement metrics using Random Forest and unveiling hidden user clusters, this research offers a blueprint for e-commerce platforms to sustain growth and competitiveness in a digital-first economy.  \nvi  \nContents  \nAbstract v  \nList of figures ix  \nList of tables xi  \nListing of acronyms xiii  \n1 Introduction 1  \n1.1 Techniques & Research Topics ........................ 2  \n1.2 Main Challenges & Trends .......................... 4  \n1.3 Thesis Structure ............................... 5  \n2 Markov ChainsAnd Page Values 7  \n2.1 Multi-touch attribution models ........................ 8  \n2.1.1 Heuristic Attribution Models .................... 9  \n2.1.2 Probabilistic Attribution Models ................... 9  \n2.1.3 Choosing The Right Model ..................... 10  \n2. 2 Markov Chain Model . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 10  \n2.2.1 Deciphering The Algorithm ..................... 11  \n2.2.2 Implementation & Results ...................... 13  \n3 Clustering Users 17  \n3.1 Exploratory Data Analysis ........................... 18  \n3.2 Random Forest Feature Extraction ...................... 21  \n3.2.1 Gini Importance ........................... 21  \n3.2.2 Results ................................ 22  \n3.3 K-prototypes Clustering ........................... 23  \n3.3.1 Deciphering The Algorithm ..................... 24  \n3.3.2 Implementation & Results ...................... 25  \n3.4 Density-based Clustering ........................... 28  \n3.4.1 UMAP & HDBSCAN ........................ 29  \n3.4.2 Normality Assumption & Sample Size ................ 33  \n3.4.3 Implementation & Results ...................... 36  \n3.5 Comparison Of The Methods ......................... 42  \n4 Conclusion 43  \nReferences 45  \nListing of figures  \n2.1 Sankey Diagram of User paths : Red transitions are self-loops and thicker lines indicate more transitions. Height of the rectangles is arbitrary ........ 14  \n2.2 Page value by channel over time frames Ti ................... 15  \n3.1 Bar plots of categorical variables: referrer, country code, referrer type, user agent device, user agent browser info and is campaign ............. 19  \n3.2 Histogram of numerical variables in log scale: total interactions, total session click, total page viewed abd total session time ................. 20  \n3.3 Bar plot of Feature Importance Scores ..................... 23  \n3.4 Elbow method plot; value of cost funct","cbCaibJAsvLJ7WFm","https://ap.wps.com/l/cbCaibJAsvLJ7WFm","pdf",1828701,1,66,"English","en",105,"# Introduction\n## Techniques & Research Topics\n## Main Challenges & Trends\n## Thesis Structure\n# Markov Chains and Page Values\n## Multi-touch attribution models\n## Markov Chain Model\n## Deciphering the Algorithm\n## Implementation & Results\n# Clustering Users\n## Exploratory Data Analysis\n## Random Forest Feature Extraction\n## K-prototypes Clustering\n## Density-based Clustering\n## Comparison of the Methods\n# Conclusion\n## References","[{\"question\":\"How does the thesis use Markov Chains to support revenue generation on e-commerce platforms?\",\"answer\":\"It implements a Markov Chains-based algorithm that assigns scores to individual webpages, reflecting their importance for potential revenue generation.\"},{\"question\":\"What clustering approaches are used to find user groups and estimate missing revenue?\",\"answer\":\"The thesis applies k-prototypes for mixed data types and uses UMAP combined with HDBSCAN for density-based clustering to identify nuanced user clusters and compute potentially missing revenue.\"},{\"question\":\"How are features selected for clustering in the thesis?\",\"answer\":\"A Random Forest analysis guides variable selection, using the empirical importance of features to uncover latent revenue opportunities.\"}]","Machine Learning Approaches to Enhance Revenue Capture in E-commerce Platforms - Thesis Abstract | PDF",1785819241,166,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"machine-learning-approaches-to-enhance-revenue-capture-in-e-commerce-platforms-thesis-abstract","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/machine-learning-approaches-to-enhance-revenue-capture-in-e-commerce-platforms-thesis-abstract/123920/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-04",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"How does the thesis use Markov Chains to support revenue generation on e-commerce platforms?","Question",{"text":75,"@type":76},"It implements a Markov Chains-based algorithm that assigns scores to individual webpages, reflecting their importance for potential revenue generation.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What clustering approaches are used to find user groups and estimate missing revenue?",{"text":80,"@type":76},"The thesis applies k-prototypes for mixed data types and uses UMAP combined with HDBSCAN for density-based clustering to identify nuanced user clusters and compute potentially missing revenue.",{"name":82,"@type":73,"acceptedAnswer":83},"How are features selected for clustering in the thesis?",{"text":84,"@type":76},"A Random Forest analysis guides variable selection, using the empirical importance of features to uncover latent revenue opportunities.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]