[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-118887-en":3,"doc-seo-118887-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},118887,16904993612988,"Olivia Brown","https://ap-avatar.wpscdn.com/davatar_a8503ba1806abce46bf441b54a3ca4cd",8,"Research & Report","The Song Describer Dataset - A Corpus of Audio Captions for Music-and-Language Evaluation","The Song Describer dataset (SDD) introduces a crowdsourced collection of high-quality audio–caption pairs for evaluating music-and-language models. It provides 1.1k human-written natural language descriptions spanning 706 music recordings, using openly accessible recordings released under Creative Commons licenses. The work benchmarks popular models across music captioning, text-to-music generation, and music-language retrieval, emphasizing cross-dataset evaluation and practical guidance for researchers seeking a broader, more reliable view of model performance.","The Song Describer Dataset: a Corpus of Audio Captions for Music-and-Language Evaluation  \nIlaria Manco∗1 ,2 , Benno Weck∗3, Seungheon Doh4 , Minz Won5 , Yixiao Zhang 1 , Dmitry Bodganov3 , Yusong Wu6 , Ke Chen7 , Philip Tovstogan3 Emmanouil Benetos 1 , Elio Quinton2 , George Fazekas 1 , Juhan Nam4  \n1 QMUL, 2UMG, 3UPF, 4 KAIST, 5ByteDance, 6Mila, 7UCSD  \nAbstract  \nWe introduce the Song Describer dataset (SDD), a new crowdsourced corpus of high-quality audio-caption pairs, designed for the evaluation of music-andlanguage models. The dataset consists of 1.1k human-written natural language descriptions of 706 music recordings, all publicly accessible and released under Creative Common licenses. To showcase the use of our dataset, we benchmark popular models on three key music-and-language tasks (music captioning, text-tomusic generation and music-language retrieval) . Our experiments highlight the importance of cross-dataset evaluation and offer insights into how researchers can use SDD to gain a broader understanding of model performance.  \n1 Introduction  \nMultimodal approaches that jointly process audio and language are becoming increasingly important within music understanding and generation, giving rise to a new area of research, which we refer to as music-and-language (M&L) . Several recent works have emerged in this domain, proposing methods to automatically generate music descriptions [21, 7, 9], synthesise music from a text prompt [2, 14, 30, 6], search for music based on language queries [8, 22, 13], and more [20, 17] . However, evaluating M&L models remains a challenge due to a lack of public and accessible datasets with paired audio and language, resulting in the widespread use of private data [21, 22, 23, 14, 2, 13] and inconsistent evaluation practices. To mitigate this, we release the Song Describer dataset (SDD), anew high-quality evaluation dataset of crowdsourced captions paired with openly licensed music recordings. Through our dataset, we allow for the first standardised comparison of models on realistic music-language data and expand the number and variety of datasets available to the research community. We propose SDD as an evaluation-only dataset to promote the development and use of out-of-domain data and counteract the tendency to overfit to a particular dataset [5, 29, 32] . Alongside SDD, there are currently two other human-curated datasets of music-caption pairs: MusicCaps [2] and YT8M-MusicTextClips [25], composed of 5.5k and 4k 10-second clips from AudioSet [11] and YouTube8M [1] respectively. More recently, Doh et al. [7] have released LP-MusicCaps, a dataset of 2.2M music descriptions generated by prompting a large language model with a set of instructions and ground-truth tags. While synthetic data presents an opportunity to artificially scale up training data, limitations such as LLM hallucinations still pose significant challenges to their use in evaluation, where reliable data plays a crucial role.  \n∗Equal contribution  \nMachine Learning for Audio Workshop (NeurIPS 2023) .  \nTable 1: An overview of the SDD compared to other music-caption datasets, MusicCaps (MC) [2] and YT8MMusicTextClips (MTC) [25] . * denotes the validated subset.  \n\n| Dataset Annotators |  | Text |  |  | Audio |  |  |  |\n| --- | --- | --- | --- | --- | --- | --- | --- | --- |\n|  |  | Caption \\# | Avg length | Vocab size | Audio \\# | Length | Source | Public |\n| MC | 10 | 5,521 | 54.8 ± 18.7 | 6,144 | 5,521 | 10 sec | [11] | ✗ |\n| MTC | - | 4,169 | 16.9 ± 4.4 | 2,599 | 4,169 | 10 sec | [1] | ✗ |\n| SDD* | 114 | 746 | 18.2 ± 7.6 | 1,942 | 547 | 2 min | [4] | ✓ |\n| SDD | 142 | 1106 | 21.7 ± 12.4 | 2,859 | 706 | 2 min | [4] | ✓ |\n\nTable 2: The main task instruction given to the annotators, together with an example of three captions produced by different annotators describing the same track in the dataset.  \n\n| Instructions | Example caption |\n| --- | --- |\n| Write one sentence that describes the track you just listened to: foc","cbCaipCnnoJVRQBd","https://ap.wps.com/l/cbCaipCnnoJVRQBd","pdf",379933,1,13,"English","en",105,"# Introduction\n## Evaluation challenges in music-and-language\n## Song Describer dataset overview\n# The Song Describer Dataset\n## Data source and licensing\n## Caption characteristics and metadata linkage\n## Key differences vs. MusicCaps and YT8M-MusicTextClips","[{\"question\":\"What is the Song Describer Dataset (SDD)?\",\"answer\":\"SDD is a crowdsourced dataset of high-quality audio-caption pairs designed to evaluate music-and-language models.\"},{\"question\":\"What does SDD contain in terms of scale and captions?\",\"answer\":\"SDD includes 1.1k human-written natural language descriptions for 706 music recordings, with captions crafted as rich single-sentence descriptions covering musical features.\"},{\"question\":\"How does SDD support evaluation compared with other music-caption datasets?\",\"answer\":\"SDD offers standardized, openly licensed audio and longer track segments, enabling more reliable cross-dataset model comparisons and evaluation with automatic captioning metrics.\"}]","The Song Describer Dataset - A Corpus of Audio Captions for Music-and-Language Evaluation | PDF",1785720789,33,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"the-song-describer-dataset-a-corpus-of-audio-captions-for-music-and-language-evaluation","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/the-song-describer-dataset-a-corpus-of-audio-captions-for-music-and-language-evaluation/118887/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-03",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What is the Song Describer Dataset (SDD)?","Question",{"text":75,"@type":76},"SDD is a crowdsourced dataset of high-quality audio-caption pairs designed to evaluate music-and-language models.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What does SDD contain in terms of scale and captions?",{"text":80,"@type":76},"SDD includes 1.1k human-written natural language descriptions for 706 music recordings, with captions crafted as rich single-sentence descriptions covering musical features.",{"name":82,"@type":73,"acceptedAnswer":83},"How does SDD support evaluation compared with other music-caption datasets?",{"text":84,"@type":76},"SDD offers standardized, openly licensed audio and longer track segments, enabling more reliable cross-dataset model comparisons and evaluation with automatic captioning metrics.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]