[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-116931-en":3,"doc-seo-116931-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},116931,962075006959,"Anda","https://ap-avatar.wpscdn.com/avatar/e0002397efbe92a78e?_k=1776741047341049297",8,"Research & Report","Lead Scoring with Machine Learning - Master Thesis 2023","This master thesis evaluates the effectiveness of four machine learning algorithms—linear regression, decision tree, random forest, and neural network—for the lead scoring task. Model performance is tested on datasets built without sampling and with random under-sampling or random over-sampling, including SMOTE. Evaluation uses accuracy, AUC-ROC, specificity, sensitivity, precision, recall, F1 score, and G-mean. Results show higher accuracy for models trained without sampling, while the neural network delivers strong results across datasets. Findings support lead management decisions on algorithm and sampling choice and inform future research with broader models and data.","MASTER THESIS  \nMiss  \nSafa Binte Ayaz  \nLead Scoring with Machine Learning  \n2023  \nFaculty of Applied Computer Sciences and Biosciences  \nMASTER THESIS  \nLead Scoring with Machine Learning  \nAuthor:  \nSafa Binte Ayaz  \nStudy Programme:  \nApplied Mathematics in Networks and Data Science  \nSeminar Group: MA18w1-M  \nFirst Referee:  \nProf. Dr. Thomas Villmann  \nSecond Referee: Michael Olschimke  \nMittweida, Febuary 2023  \nBibliographic Information  \nBinte Ayaz, Safa: Lead Scoring with Machine Learning, 61 pages, 24 figures, Hochschule Mittweida, University of Applied Sciences, Faculty of Applied Computer Sciences and Biosciences  \nMaster Thesis, 2023  \nAbstract  \nThis thesis investigates the efficacy of four machine learning algorithms, namely linear regression, decision tree, random forest and neural network in the task of lead scoring. Specifically, the study evaluates the performance of these algorithms using datasets without sampling and with random under-sampling and over-sampling using SMOTE. The performance of each algorithm is measure using various performance metrics, including accuracy, AUC-ROC, specificity, sensitivity, precision, recall, F1 score, and G-mean . The results indicate that models trained on the dataset without sampling achieved higher accuracy than those trained on the dataset with either random under-sampling or random over-sampling using SMOTE. However, the neural network demonstrated remarkable results on each dataset compared to the other algorithms. These findings provide valuable insights into the effectiveness of machine learning algorithms for lead scoring tasks, particularly when using different sampling techniques.  \nThe findings of this study can aid lead management practices in selecting the most suitable algorithm and sampling technique for their needs. Furthermore, the study contributes to the literature by providing a comprehensive evaluation of the performance of machine learning algorithms for lead scoring tasks. This thesis has practical implications for businesses looking to improve their lead management practices, and future research could extend the analysis to other machine learning algorithms or more extensive datasets.  \nI. Contents  \nContents . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . I  \nList of Figures ............................................................................... II  \nList of Tables ................................................................................ III  \nNomenclature ................................................................................ IV  \nPreface ....................................................................................... V  \n1 Introduction ........................................................................... 1  \n1.1 Related Work ......................................................................... 2  \n1.2 Problem Statement ................................................................... 4  \n1.3 Thesis Structure ...................................................................... 5  \n2 Theory ................................................................................ 7  \n2.1 Lead Scoring .......................................................................... 7  \n2.2 Predictive Lead Scoring ............................................................... 9  \n2.3 The significance of customer attributes in the process of lead scoring .............. 10  \n2.4 Classification Algorithms used in Predictive Analysis ................................ 11  \n3 Machine learning ...................................................................... 13  \n3.1 Classifiers ............................................................................. 13  \n3.1.1 Logistic Regression (LR) .............................................................. 13  \n3.1.2 Decision Trees (DT): ..................","cbCaidp490Yh6loC","https://ap.wps.com/l/cbCaidp490Yh6loC","pdf",2320857,1,77,"English","en",105,"# Contents\n## Introduction\n## Theory\n## Machine learning\n## Data Preprocessing and Analysis","[{\"question\":\"Which machine learning algorithms are compared for lead scoring?\",\"answer\":\"The thesis compares linear regression, decision tree, random forest, and neural network for the lead scoring task.\"},{\"question\":\"How are sampling techniques used in the experiments?\",\"answer\":\"Experiments use datasets without sampling and datasets with random under-sampling or random over-sampling, including SMOTE.\"},{\"question\":\"What performance metrics are used to evaluate the models?\",\"answer\":\"Models are evaluated using accuracy, AUC-ROC, specificity, sensitivity, precision, recall, F1 score, and G-mean.\"}]","Lead Scoring with Machine Learning - Master Thesis 2023 | PDF",1785672607,194,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"lead-scoring-with-machine-learning-master-thesis-2023","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/lead-scoring-with-machine-learning-master-thesis-2023/116931/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-02",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Which machine learning algorithms are compared for lead scoring?","Question",{"text":75,"@type":76},"The thesis compares linear regression, decision tree, random forest, and neural network for the lead scoring task.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How are sampling techniques used in the experiments?",{"text":80,"@type":76},"Experiments use datasets without sampling and datasets with random under-sampling or random over-sampling, including SMOTE.",{"name":82,"@type":73,"acceptedAnswer":83},"What performance metrics are used to evaluate the models?",{"text":84,"@type":76},"Models are evaluated using accuracy, AUC-ROC, specificity, sensitivity, precision, recall, F1 score, and G-mean.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]