[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83857-en":3,"doc-seo-83857-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},83857,8796095462418,"Noah","https://ap-avatar.wpscdn.com/avatar/80000253c1241d02b47?x-image-process=image/resize,m_fixed,w_180,h_180&k=1778826106357471780",8,"Research & Report","Semantic Homogenization in Italian Popular Music: A Diachronic Analysis","Recent research reports a decline in semantic variety in popular music lyrics, especially for English songs on streaming platforms. This study tests whether a comparable shift occurs in a different linguistic and cultural setting by analyzing the lyrics of all finalist songs across 75 editions of Italy’s Sanremo Music Festival. It introduces a flexible, efficient methodology for tracking semantic similarity over time, combining full-text, segment-, topic-, and word-level analyses using embedding techniques and large language models. Applied to the Sanremo corpus, the framework reveals a gradual increase in semantic uniformity, aligning with earlier global findings and highlighting long-term change in musical language.","arXiv :2607 .04832v 1 [ cs .CL] 6 Jul 2026  \nSemantic Homogenization in Italian Popular Music: A Diachronic Analysis  \nLorenzo Canale 1,2*†, Stefano Scotta 1,3*† and Alberto Messina 1,4*  \n1 Centro Ricerche, Innovazione Tecnologica e Sperimentazione, RAI, Via Giovanni Carlo Cavalli 6, Turin, 10138, Italy.  \n2 ORCID: 0000-0002-7556-595X; Google Scholar: hsgYvg0AAAAJ.  \n3 ORCID: 0000-0003-1078-2985; Google Scholar: mJRjh1sAAAAJ.  \n4 ORCID: 0000-0002-8262-2449; Google Scholar: hmd7668AAAAJ.  \n*Corresponding author(s). E-mail(s): [lorenzo.canale@rai.it](lorenzo.canale@rai.it) ; [stefano.scotta@rai.it](stefano.scotta@rai.it) ; [alberto.messina@rai.it](alberto.messina@rai.it) ;  \n†These authors contributed equally to this work.  \nAbstract  \nIn recent years, studies have revealed a decline in semantic variety across popular music lyrics, particularly in English-language songs on streaming platforms like Spotify. This research examines whether a similar trend can be observed ina different linguistic and cultural context: the lyrics of all finalist songs from the  \n75 editions of the Sanremo Music Festival, Italy’s most renowned music competition. What sets this work apart is the development of a flexible and efficient methodology for tracking changes in semantic similarity over time, which can be applied to different datasets to study similar phenomena. Drawing on a combination of full-text, segment-based, topic-based, and word-level analyses, the approach leverages both embedding techniques and large language models. When applied to the Sanremo corpus, this framework reveals a gradual move toward increasing semantic uniformity, echoing the global patterns identified in previous studies. These findings underscore the value of natural language processing tools in uncovering long-term shifts in musical language and cultural expression.  \nKeywords: Semantic similarity, Embedding models, Song lyrics analysis, Large  \nlanguage models (LLMs), Natural Language Processing, Sanremo Festival  \n1  \nThis preprint has not undergone peer review (when applicable) or any postsubmission improvements or corrections. The Version of Record of this article is published in Journal of Computational Social Science, and is available online at [https://doi.org/10.1007/s42001-026-00468-1](https://doi.org/10.1007/s42001-026-00468-1) .  \n1 Introduction  \nUnderstanding textual meaning has long been a central concern in both literary studies and computational linguistics. In literary theory, semantics is the study of how words and structures convey meaning, encompassing aspects such as denotation, connotation, and intertextuality [1, 2] . Computational approaches to semantics aim to model these nuances through mathematical representations, with word and sentence embeddingsemerging as a powerful technique to capture semantic similarities between texts [3, 4] .  \nSemantic embeddings map textual units (e.g., words, sentences, or entire documents) into high-dimensional vector spaces, where distances between vectors correspond to semantic similarity [5] . Models such as Word2Vec [6], GloVe [7], and transformer-based embeddings [8] have demonstrated remarkable capabilities in capturing linguistic patterns, including synonymy, topical coherence, and stylistic variation.  \nIn this study, we present two main contributions:  \n• A methodology that enables the analysis of semantic similarity within a musical collection by selecting specific variables for comparison, such as time. By combining different similarity measures, such as full-text, portion-based, topic-based, and wordbased similarity—using embedding models and large language models (LLMs), this approach provides a more comprehensive analysis.  \n• An in-depth analysis of the Sanremo Festival lyrics over time: we apply the methodology to analyze lyrical diversity at the Sanremo Festival, a major Italian music competition. Our findings indicate a trend of increasing semantic homogeneity in lyrics over time, suggesting a","cbCaiheTxqbOdA3S","https://ap.wps.com/l/cbCaiheTxqbOdA3S","pdf",3445563,1,18,"English","en",105,"# Introduction\n# Motivation","[{\"question\":\"What research question does the paper address about semantic variety in music?\",\"answer\":\"The paper investigates whether semantic variety declines in Italian popular music lyrics, as reported for English-language songs on streaming platforms. It focuses on Sanremo Festival finalist lyrics over multiple editions.\"},{\"question\":\"How does the proposed methodology measure semantic similarity over time?\",\"answer\":\"It tracks changes using multiple similarity views—full-text, segment-based, topic-based, and word-level—implemented with embedding techniques and large language models. This combination aims to provide a comprehensive perspective on meaning changes.\"},{\"question\":\"What key trend is found when the framework is applied to the Sanremo corpus?\",\"answer\":\"The analysis shows a gradual move toward increasing semantic uniformity in lyrics over time. The results suggest a shift toward more homogeneous semantic expression during the festival’s history.\"}]",1784191009,45,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"semantic-homogenization-in-italian-popular-music-a-diachronic-analysis","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/semantic-homogenization-in-italian-popular-music-a-diachronic-analysis/83857/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-20","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What research question does the paper address about semantic variety in music?","Question",{"text":75,"@type":76},"The paper investigates whether semantic variety declines in Italian popular music lyrics, as reported for English-language songs on streaming platforms. It focuses on Sanremo Festival finalist lyrics over multiple editions.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the proposed methodology measure semantic similarity over time?",{"text":80,"@type":76},"It tracks changes using multiple similarity views—full-text, segment-based, topic-based, and word-level—implemented with embedding techniques and large language models. This combination aims to provide a comprehensive perspective on meaning changes.",{"name":82,"@type":73,"acceptedAnswer":83},"What key trend is found when the framework is applied to the Sanremo corpus?",{"text":84,"@type":76},"The analysis shows a gradual move toward increasing semantic uniformity in lyrics over time. The results suggest a shift toward more homogeneous semantic expression during the festival’s history.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":45,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":45,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":45,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]