[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83536-en":3,"doc-seo-83536-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83536,962075006959,"Anda","https://ap-avatar.wpscdn.com/avatar/e0002397efbe92a78e?_k=1776741047341049297",8,"Research & Report","Evaluating Pretrained Music Embeddings for Cross-Performance Jazz Standard Recognition","Jazz standard recognition from audio is a tune-level music retrieval task complicated by variability across performances in tempo, key, arrangement, instrumentation, and improvisation, including whether the head melody is present. This study uses a curated subset of the Jazz Trio Database for cross-performance standard recognition and compares a from-scratch Harmonic CNN baseline with frozen pretrained embeddings from music foundation models via supervised probing and nearest-neighbor retrieval. Results show spectrogram models overfit training performances, while pretrained embeddings improve top-k accuracy yet remain sensitive to performer identity, partially mitigated by a lightweight contrastive projection.","Evaluating Pretrained Music Embeddings for Cross-Performance Jazz Standard Recognition  \nCagri Eser 1  \narXiv :2607 .00777v 1 [ cs . SD] 1 Jul 2026  \nAbstract  \nRecognizing jazz standards from audio is a challenging form of tune-level music retrieval: different performances of the same standard may vary in tempo, key, arrangement, instrumentation, improvisational content, and even whether the head melody is present. We study this problem using a curated subset of the Jazz Trio Database designed for cross-performance standard recognition. We compare a from-scratch trained Harmonic CNN baseline against frozen pretrained music representations from recent music understanding foundation models, using both supervised probing and nearest-neighbor retrieval.  \nOur results suggest that from-scratch spectrogram models overfit strongly to training performances, while pretrained embeddings provide better top-k results but are sensitive to performer identity, which can be partially reduced with a lightweight contrastive projection. Our findings motivate jazz standard recognition as a useful stress test for music representation models and asa step toward retrieval-based standard identification. Project page: [https://github.com/](https://github.com/)[ ](https://github.com/)cagries/tipofmyear.  \n1. Introduction  \nAudio recognition systems such as Shazam (Wang, 2003) are highly effective for identifying exact recordings, but recognizing the underlying tune across different performances is a different problem. This distinction is especially important in jazz: a standard such as Autumn Leaves can appear in many keys, different tempo, different arrangements, and in many improvisational contexts. In this paper, we study whether modern audio representations support this form of  \n1Department of Computer Engineering, Middle East Technical University, Ankara, Turkey. Correspondence to: Cagri Eser \u003C[cagri.eser@ceng.metu.edu.tr](cagri.eser@ceng.metu.edu.tr) >.  \nAccepted to the Workshop on Machine Learning for Audio atthe 43rd International Conference on Machine Learning (ICML), Seoul, South Korea, 2026 .  \ncross-performance tune recognition. The task is difficult for several reasons. First, jazz performances often devote long stretches to improvisation, where cropped local windows may not contain the head melody. Second, different standards can share common harmonic progressions or similar melodic fragments. Third, recordings by the same group can be acoustically and stylistically similar across performances of different standards, which causes problems for nearest-neighbor retrieval methods.  \nOur work makes three contributions. First, we construct a curated filtered benchmark subset from the Jazz Trio Database (Cheston et al., 2024) for standard recognition across performances. Second, we compare spectrogrambased training, supervised probing of frozen embeddings, and embedding-based retrieval under the same evaluation protocol. Third, we analyze retrieval-based classification and demonstrate that nearest neighbors often retrieve performer identity rather than tune identity, and explore a lightweight supervised contrastive projection to reduce performer-biased retrieval.  \n2. Related Work  \nFingerprinting systems for audio such as Shazam (Wang, 2003) are very effective for exact recording identification from short and noisy excerpts, however they are designed to match a given signal to a recording already present in the database rather than recognizing a new performance of the same underlying composition. With this in mind, jazz standard recognition is therefore closer to cover or version identification (Serr et al., 2011 ; Xun et al., 2023), which is concerned with retrieval of different renditions of the same work using features that are invariant to changes in key, tempo, timbre, and arrangement. Recent versionidentification systems have increasingly adopted embeddingbased retrieval formulations to improve scalability while preserving work-level ","cbCaieC0mt4cqxRI","https://ap.wps.com/l/cbCaieC0mt4cqxRI","pdf",327662,2,1,6,"English","en",105,"# Abstract\n# Introduction\n# Related Work\n# Dataset, Evaluation and Experiment Setup","[{\"question\":\"What makes cross-performance jazz standard recognition challenging?\",\"answer\":\"Different performances vary in tempo, key, arrangement, instrumentation, and improvisation, and query windows may omit the head melody. Additionally, different standards can share harmonic or melodic fragments, and recordings by the same group can be acoustically similar across standards.\"},{\"question\":\"Which models and evaluation methods are compared in the study?\",\"answer\":\"The work compares a from-scratch trained Harmonic CNN baseline against frozen pretrained music representations from recent music understanding foundation models. It evaluates using supervised probing and nearest-neighbor retrieval under a unified protocol.\"},{\"question\":\"What do the results indicate about spectrogram models versus pretrained embeddings?\",\"answer\":\"From-scratch spectrogram models strongly overfit to training performances. Pretrained embeddings yield better top-k retrieval results but are sensitive to performer identity, which can be reduced using a lightweight contrastive projection.\"}]",1784188678,15,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"evaluating-pretrained-music-embeddings-for-cross-performance-jazz-standard-recognition","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/evaluating-pretrained-music-embeddings-for-cross-performance-jazz-standard-recognition/83536/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What makes cross-performance jazz standard recognition challenging?","Question",{"text":75,"@type":76},"Different performances vary in tempo, key, arrangement, instrumentation, and improvisation, and query windows may omit the head melody. Additionally, different standards can share harmonic or melodic fragments, and recordings by the same group can be acoustically similar across standards.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"Which models and evaluation methods are compared in the study?",{"text":80,"@type":76},"The work compares a from-scratch trained Harmonic CNN baseline against frozen pretrained music representations from recent music understanding foundation models. It evaluates using supervised probing and nearest-neighbor retrieval under a unified protocol.",{"name":82,"@type":73,"acceptedAnswer":83},"What do the results indicate about spectrogram models versus pretrained embeddings?",{"text":84,"@type":76},"From-scratch spectrogram models strongly overfit to training performances. Pretrained embeddings yield better top-k retrieval results but are sensitive to performer identity, which can be reduced using a lightweight contrastive projection.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,114,119,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":111,"show_sort_weight":112,"slug":113},"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]