[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85113-en":3,"doc-seo-85113-105":29,"detail-sidebar-cat-0-en-105":90},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":11,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},85113,2336464648746,"Skyler","https://ap-avatar.wpscdn.com/davatar_276721f389ce27ea32af1340a28f341c",8,"Research & Report","MulTTiPop Multitrack Transcription Dataset for Pop Music","MulTTiPop is a multitrack music transcription dataset designed to evaluate automatic music transcription models on real-world pop music. It includes 572 pop music segments totaling 3.5 hours, spanning genres and decades from the 1930s to the 2000s. Dataset construction uses metadata-based matching between song segments in Lakh MIDI and TheoryTab, manual selection of an anchor beat, and beat tracking with MIDI time-warping. Experiments show substantial headroom, with the best Onset F1 reaching 38%.","MulTTiPop: A MULTITRACK TRANSCRIPTION DATASET FOR POP MUSIC  \nNathan Pruyne∗ Chien-yu Huang  \nBenjamin Stoler Shinji Watanabe  \nWilliam Chen Chris Donahue  \nCarnegie Mellon University  \narXiv :2607 .08756v 1 [ cs . SD] 9 Jul 2026  \nABSTRACT  \nWe present MulTTiPop, a dataset of pop music segments and their associated multitrack MIDI recordings for the evaluation of automatic music transcription models. MulTTiPop contains 572 segments of popular music totaling 3.5 hours of audio, and contains songs from diverse genres and decades from the 1930s to 2000s. To collect this dataset, we perform metadata-based matching on song segments from the Lakh MIDI and TheoryTab datasets, manually identify an anchor beat between the audio and MIDI, then use beat tracking on the audio and warp the MIDI to match its tempo and timing. We evaluate state-of-the-art automatic music transcription models on MulTTiPop and find substantial room for improvement, with the best model achieving 38% Onset F1 . More details and sound examples of MulTTiPop are available at [https://gclef-cmu.org/multtipop](https://gclef-cmu.org/multtipop).  \n1. INTRODUCTION  \nIn recent years, automatic music transcription (AMT) systems for converting audio to note-level symbolic representations of music have evolved from transcribing solo piano music [9] to targeting performance on a wide variety of instruments and musical styles [10] . However, current AMT models do not yet meet the task of multitrack transcription on real-world pop music. Systems such as YourMT3+ [11] are only trained on synthetic recordings of pop music, and perform poorly on commercially-produced songs. In contrast, Sheet Sage [6] performs well on commercial recordings, but only generates melodies and chord labels, not full multitrack transcriptions.  \nWhile current AMT models require substantial improvement for use on commercial pop music, their systematic evaluation remains difficult due to the lack of a dataset that matches pop audio with ground truth, time-aligned transcriptions. Datasets like MAESTRO [2] and MusicNet [1] provide transcription labels for the acoustically narrow domains of solo piano and classical music, respectively. Slakh2100 [3] contains multitrack MIDI, but only provides synthesized audio, which frequently does not closely match commercial recordings in timbre, especially for vocals. TheoryTab [6] contains original, fully-produced pop music audio, but only  \n∗ Corresponding author: [npruyne@cmu.edu](npruyne@cmu.edu)  \nPop Music (YouTube)  \nTime-aligned, multitrack MIDI  \nFig. 1. MulTTiPop contains segments of YouTube audio beataligned with multitrack MIDI transcriptions.  \nprovides weakly-aligned melody and chord annotations, not full multitrack transcriptions. The closest existing dataset is RWC-Pop [8], which contains pop recordings and multitrack MIDI. However, it has limited genre diversity and uses songs composed specifically for RWC-Pop, rather than real-world pop music for which AMT systems may be used for.  \nTo address this gap, we introduce MulTTiPop, a dataset of multitrack MIDI transcriptions aligned to segments of commercial pop music. We compile multitrack MIDI from the LMD-matched subset of the Lakh MIDI Dataset [12], and pair MIDI files with audio segments from YouTube by performing metadata matching with the TheoryTab [6] dataset. TheoryTab is sourced from user-provided audio and chord transcriptions, thus indicating segments of audio that users would be interested in transcribing. MulTTiPop has the following key properties:  \n• Diverse, Popular Music: By sourcing audio from TheoryTab, we provide labels for segments of commercial recordings of popular music from the past decades. This enables MulTTiPop to be a representative sample of use cases for multitrack AMT.  \n\n| Dataset | Music Genre | Annotation Type | Audio Type | Size (Hours) | Total Samples |\n| --- | --- | --- | --- | --- | --- |\n| MusicNet [1] | Classical | Multitrack MIDI | Commercial | 34 | 330 |\n| MAESTR","cbCaitcvbInL63jX","https://ap.wps.com/l/cbCaitcvbInL63jX","pdf",899466,2,1,"English","en",105,"# Introduction\n# Methodology","[{\"question\":\"What does MulTTiPop provide for automatic music transcription evaluation?\",\"answer\":\"MulTTiPop provides pop music audio segments paired with time-aligned multitrack MIDI recordings, serving as ground truth for evaluating automatic music transcription models.\"},{\"question\":\"How is the audio aligned with the multitrack MIDI in MulTTiPop?\",\"answer\":\"The process performs metadata matching between TheoryTab audio segments and Lakh MIDI files, then manually identifies an anchor beat, followed by beat tracking on the audio and warping the MIDI to match timing and tempo.\"},{\"question\":\"What performance results are reported when evaluating existing transcription models on MulTTiPop?\",\"answer\":\"The evaluation finds substantial room for improvement, with the best model achieving an Onset F1 of 38%.\"}]",1784201187,20,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":85,"head_meta":87,"extra_data":89,"updated_unix":27},"multtipop-multitrack-transcription-dataset-for-pop-music","",{"@graph":35,"@context":84},[36,52,67],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,46,49],{"item":40,"name":41,"@type":42,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":20},"https://docshare.wps.com/document/","Document",{"item":47,"name":12,"@type":42,"position":48},"https://docshare.wps.com/document/research-report/",3,{"item":50,"name":13,"@type":42,"position":51},"https://docshare.wps.com/document/multtipop-multitrack-transcription-dataset-for-pop-music/85113/",4,{"url":50,"name":13,"@type":53,"author":54,"headline":13,"publisher":56,"fileFormat":59,"inLanguage":23,"description":14,"dateModified":60,"datePublished":61,"encodingFormat":59,"isAccessibleForFree":62,"interactionStatistic":63},"DigitalDocument",{"name":9,"@type":55},"Person",{"url":40,"name":57,"@type":58},"DocShare","Organization","application/pdf","2026-07-19","2026-07-16",true,{"@type":64,"interactionType":65,"userInteractionCount":20},"InteractionCounter",{"@type":66},"ViewAction",{"@type":68,"mainEntity":69},"FAQPage",[70,76,80],{"name":71,"@type":72,"acceptedAnswer":73},"What does MulTTiPop provide for automatic music transcription evaluation?","Question",{"text":74,"@type":75},"MulTTiPop provides pop music audio segments paired with time-aligned multitrack MIDI recordings, serving as ground truth for evaluating automatic music transcription models.","Answer",{"name":77,"@type":72,"acceptedAnswer":78},"How is the audio aligned with the multitrack MIDI in MulTTiPop?",{"text":79,"@type":75},"The process performs metadata matching between TheoryTab audio segments and Lakh MIDI files, then manually identifies an anchor beat, followed by beat tracking on the audio and warping the MIDI to match timing and tempo.",{"name":81,"@type":72,"acceptedAnswer":82},"What performance results are reported when evaluating existing transcription models on MulTTiPop?",{"text":83,"@type":75},"The evaluation finds substantial room for improvement, with the best model achieving an Onset F1 of 38%.","https://schema.org",{"og:url":50,"og:type":86,"og:title":13,"og:site_name":57,"og:description":14},"article",{"robots":88,"canonical":50},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":91},[92,96,100,104,109,114,119,122,126,129,133],{"id":21,"doc_module":4,"doc_module_name":45,"category_name":93,"show_sort_weight":94,"slug":95},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":45,"category_name":97,"show_sort_weight":98,"slug":99},"Literature",80,"literature",{"id":51,"doc_module":4,"doc_module_name":45,"category_name":101,"show_sort_weight":102,"slug":103},"Exam",70,"exam",{"id":105,"doc_module":4,"doc_module_name":45,"category_name":106,"show_sort_weight":107,"slug":108},5,"Comic",60,"comic",{"id":110,"doc_module":4,"doc_module_name":45,"category_name":111,"show_sort_weight":112,"slug":113},6,"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":45,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":45,"category_name":124,"show_sort_weight":28,"slug":125},9,"Religion & Spirituality","religion-spirituality",{"id":28,"doc_module":4,"doc_module_name":45,"category_name":127,"show_sort_weight":28,"slug":128},"World Cup","world-cup",{"id":130,"doc_module":4,"doc_module_name":45,"category_name":131,"show_sort_weight":130,"slug":132},10,"Lifestyle","lifestyle",{"id":134,"doc_module":4,"doc_module_name":45,"category_name":135,"show_sort_weight":105,"slug":136},19,"General","general"]