[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-128508-en":3,"doc-seo-128508-105":31,"detail-sidebar-cat-0-en-105":96},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":28,"seo_description":14,"update_tm":29,"read_time":30},128508,687207017582,"Himbo","https://ap-avatar.wpscdn.com/davatar_994ba38a5ba835b3df7d355c54d3ed8d",8,"Research & Report","Multilingual Fine-grained Sentiment Text Classification of Web Content for Advertising Technology Applications - Master Thesis","Programmatic advertising matches publishers’ web content with advertisers’ strategies, often filtering topics based on audience interests. This can reduce reach and ignore subtle differences in how content is expressed emotionally. The thesis explores fine-grained sentiment analysis to map web pages into richer emotion groups, moving beyond the limited positive/negative/neutral framing. Models are trained on the GoEmotions dataset and evaluated on annotated real-world web texts, including multilingual language-agnostic approaches.","University of Padova  \nDepartment of Mathematics Tullio Levi-Civita Master Thesis in Big Data Management and Analytics  \nMultilingual fine-grained sentiment text  \nclassification of web content for  \nadvertising technology applications  \nSupervisor Master Candidate  \nProfessor Nicolò Navarin Nicole Zafalon Kovacs  \nUniversity of Padova  \nCo-supervisor Giovanni Vedana Anonymised  \nSeptember 1st 2023  \nii  \nvi  \nAbstract  \nProgrammatic advertising involves matching publishers’ web page content with advertisers’strategies, usually by filtering out content topics based on audience interests. However, this approach can limit audience reach and overlook content nuances. Sentiment analysis comes into place to enhance content understanding by categorizing web pages into sentiment groups. While existing research often focuses on the “positive”,“negative”, and “neutral” labels, abroader range of emotion categories can provide more details, better aligning content with advertisers’ goals. This study investigates sentiment classification models that use fine-grained emotion categories to classify the content of publishers’ web pages. Considering the multilingual nature of the business, this research also explores and compares language-agnostic models. Several models were trained using the GoEmotions dataset, composed of 58k English Reddit comments annotated by humans using 28 emotion categories. These models were assessed across various sentiment groupings and evaluated on manually annotated real-world web page texts. The final proposed emotion groups encompass six categories: “repudiation”,“sadness”,“neutral”, “curiosity”,“appreciation”, and “positive experience”. Among the models compared, transformer models, particularly BERT-large, exhibited superior performance. The best English-only model achieved a weighted F1-score of 44% on the annotated web data. Languageagnostic models showed lower metrics on Italian texts but were comparable to English-only models for English text. The leading language-agnostic model, Multilingual BERT, achieved a weighted F1-score of 37% on the same data translated into Italian. The study achieved promising outcomes across the six emotion groups, surpassing the traditional “positive,” “negative,”and “neutral” categories. Additionally, the evaluation of multilingual models demonstrated their applicability to multiple languages despite being trained solely on English data.  \nviii  \nContents  \nAbstract v  \nList of figures xi  \nList of tables xiii  \nListing of acronyms xv  \n1 Introduction 1  \n1.1 Background and need ............................. 1  \n1.2 Statement of the problem ........................... 2  \n1.3 Purpose of the study ............................. 3  \n1.4 Methodology ................................. 4  \n1.5 Results and findings .............................. 5  \n2 Background 9  \n2.1 Preliminaries ................................. 9  \n2.1.1 Programmatic advertising ...................... 9  \n2.1.2 Web scraping ............................. 10  \n2.1.3 Emotion taxonomies . . . . . . . . . . . . . . . . . . . . . . . . . 11  \n2.1.4 Classifiers ............................... 13  \n2.1.5 Sentiment analysis .......................... 24  \n2.2 Related work ................................. 29  \n2.2.1 Machine learning models ....................... 29  \n2.2.2 Deep learning models ......................... 30  \n2.2.3 Transformers models ......................... 32  \n2.2.4 Fine-grained sentiment analysis .................... 33  \n3 Problem exposition 35  \n3.1 Business Understanding ............................ 35  \n3.2 Detailed Research Questions ......................... 37  \n3.3 Data Understanding .............................. 37  \n3.4 Resource constraints . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 38  \n4 Methods, Results and Findings 39  \n4. 1 Data sources . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 40  \n4. 1. 1 GoEmotions Dataset . . . . . . . . . . . . . . . . .","cbCaid1ZOXPhYRAg","https://ap.wps.com/l/cbCaid1ZOXPhYRAg","pdf",3033496,2,1,95,"English","en",105,"# Abstract\n# Introduction\n## Background and need\n## Statement of the problem\n## Purpose of the study\n## Methodology\n## Results and findings\n# Background\n## Preliminaries\n## Related work\n# Problem exposition\n## Business Understanding\n## Detailed Research Questions\n## Data Understanding\n## Resource constraints\n# Methods, Results and Findings\n## Data sources\n## Emotion groups\n## Preprocessing\n## Text classification\n## Evaluation on web data\n## Multilingual evaluation on web data\n# Conclusion\n## Future work","[{\"question\":\"Why does programmatic advertising benefit from sentiment analysis?\",\"answer\":\"Because filtering by audience interests can miss emotional nuances and reduce effective reach. Sentiment analysis improves understanding by grouping web pages according to emotional signals.\"},{\"question\":\"What data and emotion scheme are used to train the models?\",\"answer\":\"Several transformer models are trained on the GoEmotions dataset, which contains 58k English Reddit comments annotated with 28 emotion categories, then adapted to multiple sentiment groupings.\"},{\"question\":\"What are the final emotion groups proposed in the study?\",\"answer\":\"The final proposed emotion groups cover six categories: repudiation, sadness, neutral, curiosity, appreciation, and positive experience.\"},{\"question\":\"How do language-agnostic models perform for Italian versus English?\",\"answer\":\"Language-agnostic models achieve lower metrics on Italian texts, but remain comparable to English-only models for English. The best language-agnostic approach, Multilingual BERT, attains a weighted F1-score of 37% on Italian data translated from the same source.\"}]","Multilingual Fine-grained Sentiment Text Classification of Web Content for Advertising Technology Applications - Master Thesis | PDF",1786001456,239,{"code":4,"msg":32,"data":33},"ok",{"site_id":25,"language":24,"slug":34,"title":13,"keywords":35,"description":14,"schema_data":36,"social_meta":91,"head_meta":93,"extra_data":95,"updated_unix":29},"multilingual-fine-grained-sentiment-text-classification-of-web-content-for-advertising-technology-applications-master-thesis","",{"@graph":37,"@context":90},[38,54,69],{"@type":39,"itemListElement":40},"BreadcrumbList",[41,45,48,51],{"item":42,"name":43,"@type":44,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":46,"name":47,"@type":44,"position":20},"https://docshare.wps.com/document/","Document",{"item":49,"name":12,"@type":44,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":44,"position":53},"https://docshare.wps.com/document/multilingual-fine-grained-sentiment-text-classification-of-web-content-for-advertising-technology-applications-master-thesis/128508/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":42,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-22","2026-08-06",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82,86],{"name":73,"@type":74,"acceptedAnswer":75},"Why does programmatic advertising benefit from sentiment analysis?","Question",{"text":76,"@type":77},"Because filtering by audience interests can miss emotional nuances and reduce effective reach. Sentiment analysis improves understanding by grouping web pages according to emotional signals.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"What data and emotion scheme are used to train the models?",{"text":81,"@type":77},"Several transformer models are trained on the GoEmotions dataset, which contains 58k English Reddit comments annotated with 28 emotion categories, then adapted to multiple sentiment groupings.",{"name":83,"@type":74,"acceptedAnswer":84},"What are the final emotion groups proposed in the study?",{"text":85,"@type":77},"The final proposed emotion groups cover six categories: repudiation, sadness, neutral, curiosity, appreciation, and positive experience.",{"name":87,"@type":74,"acceptedAnswer":88},"How do language-agnostic models perform for Italian versus English?",{"text":89,"@type":77},"Language-agnostic models achieve lower metrics on Italian texts, but remain comparable to English-only models for English. The best language-agnostic approach, Multilingual BERT, attains a weighted F1-score of 37% on Italian data translated from the same source.","https://schema.org",{"og:url":52,"og:type":92,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":94,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":97},[98,102,106,110,115,120,125,128,133,136,140],{"id":21,"doc_module":4,"doc_module_name":47,"category_name":99,"show_sort_weight":100,"slug":101},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":47,"category_name":103,"show_sort_weight":104,"slug":105},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":47,"category_name":107,"show_sort_weight":108,"slug":109},"Exam",70,"exam",{"id":111,"doc_module":4,"doc_module_name":47,"category_name":112,"show_sort_weight":113,"slug":114},5,"Comic",60,"comic",{"id":116,"doc_module":4,"doc_module_name":47,"category_name":117,"show_sort_weight":118,"slug":119},6,"Technology",50,"technology",{"id":121,"doc_module":4,"doc_module_name":47,"category_name":122,"show_sort_weight":123,"slug":124},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":47,"category_name":12,"show_sort_weight":126,"slug":127},30,"research-report",{"id":129,"doc_module":4,"doc_module_name":47,"category_name":130,"show_sort_weight":131,"slug":132},9,"Religion & Spirituality",20,"religion-spirituality",{"id":131,"doc_module":4,"doc_module_name":47,"category_name":134,"show_sort_weight":131,"slug":135},"World Cup","world-cup",{"id":137,"doc_module":4,"doc_module_name":47,"category_name":138,"show_sort_weight":137,"slug":139},10,"Lifestyle","lifestyle",{"id":141,"doc_module":4,"doc_module_name":47,"category_name":142,"show_sort_weight":111,"slug":143},19,"General","general"]