[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-117815-en":3,"doc-seo-117815-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},117815,549758146520,"Patrick","https://ap-avatar.wpscdn.com/avatar/80002397d8c0411e94?_k=1775819394049821470",8,"Research & Report","Demographics Imputation in Marketing Sector by Means of Machine Learning - Internship Report","Develop a predictive machine-learning model to impute missing demographics values from survey data and assess model performance for business use. The project addresses incomplete or absent demographic information per user, using data cleaning, normalization, and feature selection before sampling and training. Random Forest and Gradient Boosting models are trained and evaluated with metrics aligned to the target variables and production constraints. Results for ethnicity-related targets and household income show limitations under default hyperparameters, motivating further work on feature selection and additional data sources.","Master’s degree Program in  \nData Science and Advanced Analytics  \nMDSAA  \nDEMOGRAPHICS IMPUTATION IN MARKETING SECTOR BY MEANS OF MACHINE LEARNING  \nMargarita Venediktova  \nInternship Report  \npresented as partial requirement for obtaining the Master Degree Program in Data Science and Advanced Analytics  \nNOVA Information Management School Instituto Superior de Estatística e Gestão de Informação  \nUniversidade Nova de Lisboa  \nNOVA Information Management School Instituto Superior de Estatística e Gestão de Informação  \nUniversidade Nova de Lisboa  \nDEMOGRAPHICS IMPUTATION IN MARKETING SECTOR BY MEANS  \nOF MACHINE LEARNING  \nby  \nMargarita Venediktova  \nInternship report presented as partial requirement for obtaining the Master’s degree in Advanced Analytics, with a Specialization in Data Science  \nSupervisor: Prof. Doutor Flávio Luis Portas Pinheiro  \nNovember, 2022  \nSTATEMENT OF INTEGRITY  \nI hereby declare having conducted this academic work with integrity. I confirm that I have not used plagiarism or any form of undue use of information or falsification of results along the process leading to its elaboration. I further declare that I have fully acknowledge the Rules of Conduct and Code of Honor from the NOVA Information Management School.  \nLisbon, 25th November 2022  \nABSTRACT  \nThe goal of this project is to develop a predictive model in order to impute missing values in data collected through surveys (demographics data) and evaluate its performance. Currently there are two existing issues: demographics data for each user is either incomplete or missing entirely. Current POCis an attempt to exploit the capabilities of machine learning in order to impute missing demographics data.  \nData cleaning, normalization, feature selection was performed prior to applying sampling techniques and training several machine learning models. The following machine learning models were trained and tested: Random Forest and Gradient Boosting. After, the metrics appropriate for the current business purposes were selected and models’ performance was evaluated.  \nThe results for the targets ‘Ethnicity’,‘Hispanic’and ‘Household income’ are not within the acceptable range and therefore could not be used in production at the moment. The metrics obtained with the default hyperparameters indicate that both models demonstrate similar results for ‘ Hispanic’ and‘Ethnicity’ response variables. ‘Household income’ variable seems to have the poorest results, not allowing to predict the variable with adequate accuracy. Current POC suggests that the accurate prediction of demographic variable is complex task and is accompanied by certain challenges: weak relationship between demographic variables and purchase behavior, purchase location and neighborhood and its demographic characteristics, unreliable data, sparse feature set. Further investigations on feature selection and incorporation of other data sources for the training data should be considered.  \nKEYWORDS  \nMachine Learning; Missing Values Imputation  \nINDEX  \n1. Introduction ............................................................................................................................. 1  \n1.1. Problem Statement .......................................................................................................... 1  \n1.2. Objective of the present project ...................................................................................... 1  \n1.3. Structure of the present project....................................................................................... 2  \n1.4. Contribution to the company ........................................................................................... 2  \n2. Literature review ...................................................................................................................... 3  \n2.1. Literature review of demographics imputation techniques............................................. 3  \n2.1.1. Missing data, missing data t","cbCaieX2UsExAN1R","https://ap.wps.com/l/cbCaieX2UsExAN1R","pdf",1009639,1,37,"English","en",105,"# Introduction\n## Problem Statement\n## Objective of the present project\n## Structure of the present project\n## Contribution to the company\n# Literature review\n## Literature review of demographics imputation techniques\n### Missing data, missing data types and missing patterns\n### Main methods of accounting for missing values\n## Machine learning for missing data imputation\n## Features pre-processing and selection\n## Performance evaluation metrics\n# Data and Methods\n## Data\n## Data pre-processing and normalization\n## Feature selection\n## Data sampling techniques for imbalanced data\n## Machine learning methods\n## Models’ performance evaluation criteria\n# Results and discussion","[{\"question\":\"What problem does the project address in marketing demographics data?\",\"answer\":\"Demographics data for each user is either incomplete or missing entirely. The project aims to impute these missing values using a predictive model.\"},{\"question\":\"Which machine learning models were trained and tested?\",\"answer\":\"Random Forest and Gradient Boosting were trained and tested after data cleaning, normalization, feature selection, and sampling.\"},{\"question\":\"Why were some predictions not suitable for production?\",\"answer\":\"The results for targets such as Ethnicity, Hispanic, and Household income are not within the acceptable range, and Household income shows the poorest performance under default hyperparameters.\"}]","Demographics Imputation in Marketing Sector by Means of Machine Learning - Internship Report | PDF",1785679720,93,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"demographics-imputation-in-marketing-sector-by-means-of-machine-learning-internship-report","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/demographics-imputation-in-marketing-sector-by-means-of-machine-learning-internship-report/117815/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-02",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the project address in marketing demographics data?","Question",{"text":75,"@type":76},"Demographics data for each user is either incomplete or missing entirely. The project aims to impute these missing values using a predictive model.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"Which machine learning models were trained and tested?",{"text":80,"@type":76},"Random Forest and Gradient Boosting were trained and tested after data cleaning, normalization, feature selection, and sampling.",{"name":82,"@type":73,"acceptedAnswer":83},"Why were some predictions not suitable for production?",{"text":84,"@type":76},"The results for targets such as Ethnicity, Hispanic, and Household income are not within the acceptable range, and Household income shows the poorest performance under default hyperparameters.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]