[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85857-en":3,"doc-seo-85857-105":28,"detail-sidebar-cat-0-en-105":90},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":11,"language":21,"language_code":22,"site_id":23,"html_lang":22,"table_of_contents":24,"faqs":25,"seo_title":13,"seo_description":14,"update_tm":26,"read_time":27},85857,1099514068035,"Ezra","https://ap-avatar.wpscdn.com/davatar_276721f389ce27ea32af1340a28f341c",8,"Research & Report","MeloBottleneck: Self-Supervised Melody Skeleton Extraction with a Latent Subsequence Bottleneck","Melody skeleton extraction derives a shorter melody that keeps structural notes while removing ornaments. Existing approaches depend on hand-crafted reduction rules or note-wise salience classifiers trained on heuristic or procedural pseudo-labels, which can carry generator bias and optimize local keep/delete decisions rather than the reduced melody itself. MeloBottleneck learns skeletons as length-controlled, order-preserving latent subsequences using hard bottleneck extraction, rhythmic closure, and re-ornamentation reconstruction, with ornament-invariant self-supervision. Evaluations across three benchmarks show stronger transfer than pseudo-label imitation and improved BM25 fragment retrieval.","MELOBOTTLENECK: SELF-SUPERVISED MELODY SKELETON EXTRACTION WITH A LATENT SUBSEQUENCE BOTTLENECK  \nFan Bu 1 Rongfeng Li 1 ,∗ Linfeng Fan2  \n1 Beijing University of Posts and Telecommunications  \n2 Central Conservatory of Music  \n[m.july@qq.com](m.july@qq.com) , [lirongfeng@bupt.edu.cn](lirongfeng@bupt.edu.cn) , [linfeng@ccom.edu.cn](linfeng@ccom.edu.cn)  \n∗ Corresponding author: [lirongfeng@bupt.edu.cn](lirongfeng@bupt.edu.cn)  \narXiv :2607 . 10233v 1 [ cs . SD] 11 Jul 2026  \nABSTRACT  \nMelody skeleton extraction aims to derive a shorter melody that preserves structural notes while removing ornaments. Prior methods rely on hand-crafted reduction rules or notewise salience classifiers trained with heuristically or procedurally generated pseudo-labels. Such supervision can inherit generator bias and does not explicitly optimize a coherent reduced melody. We introduce MeloBottleneck, a self-supervised framework that represents a skeleton as a length-controlled, order-preserving latent subsequence. A hard-bottleneck extractor selects note events, a rhythmic-closure operator produces a self-consistent skeleton, and a re-ornamentation decoder reconstructs the input melody. Training combines reconstruction, a frozen autoregressive melody prior, ornament-invariant consistency across procedurally ornamented views, and ornament exclusion. We evaluate three regimes: synthetic out-ofdistribution ornament-to-skeleton, TAVERN variation-totheme, and Jiugong ornamented-to-gongche. A matched pseudo-label classifier excels on the synthetic benchmark, while MeloBottleneck transfers better, achieving competitive selection quality on TAVERN and Jiugong. Skeletonized melodies also improve BM25-based fragment retrieval, boosting Recall@K and MRR while reducing query time. Overall, the results suggest that learning skeletons as latent subsequences yields more robust transfer than pseudo-label imitation.  \n1. INTRODUCTION  \nIn monophonic symbolic music, melody skeleton extraction aims to derive a shorter melody that preserves structural notes while removing ornaments. Such skeletons can support symbolic comparison, retrieval, and theme tracing across variants, all tasks where surface elaboration can obscure shared melodic material. A practical extractor should therefore do more than identify locally important notes: it should retain information, produce a standalone reduced melody, remain stable under ornamentation, and allow explicit control over reduction length.  \nExisting approaches only partially meet these requirements. Music-theoretic and heuristic reducers are often interpretable, but they encode strong analytical priors that may be style-dependent [1–3] . More recent learning-  \nbased work moves closer to note-wise structural prediction [4] . More generally, when skeleton extraction is trained from heuristic or procedural note-wise keep/delete pseudolabels, the model can inherit generator bias. Moreover, note-wise prediction optimizes local decisions rather than the output melody itself—a selected subset is not yet a standalone melody.  \nWe model a melody skeleton as a length-controlled, order-preserving latent subsequence, followed by a deterministic rhythmic-closure step that turns selected note events into a self-consistent reduced melody. Based on this view, we introduce MeloBottleneck, a selfsupervised framework that learns the bottleneck through re-ornamentation reconstruction, a frozen autoregressive melody prior, and ornament-invariant learning.  \nExperiments on three benchmarks support this formulation, showing that latent-subsequence learning transfers more robustly than pseudo-label imitation and benefits downstream retrieval. A pseudo-label note classifier is strongest on synthetic out-of-distribution ornamentto-skeleton data, but MeloBottleneck transfers better to zero-shot and cross-domain settings. The extracted skeletons also improve BM25-based fragment retrieval under ornamentation and corruption, increasing retrieval quality ","cbCaik4ZVN2oTsYm","https://ap.wps.com/l/cbCaik4ZVN2oTsYm","pdf",833673,1,"English","en",105,"# Abstract\n# Introduction\n# Melody Skeleton Extraction: Task Formulation and Prior Work\n## Task Formulation\n## Relation to Prior Work","[{\"question\":\"What problem does melody skeleton extraction address?\",\"answer\":\"It produces a shorter monophonic melody that preserves structural note events while removing ornamental elaborations that can hide shared melodic material.\"},{\"question\":\"How does MeloBottleneck represent the melody skeleton?\",\"answer\":\"It models the skeleton as a length-controlled, order-preserving latent subsequence, followed by a deterministic rhythmic-closure step that converts selected note events into a coherent reduced melody.\"},{\"question\":\"How does MeloBottleneck training improve over pseudo-label-based methods?\",\"answer\":\"Training combines reconstruction, a frozen autoregressive melody prior, ornament-invariant consistency across ornamented views, and ornament exclusion, avoiding direct imitation of heuristic keep/delete pseudo-labels and improving cross-domain transfer.\"}]",1784206737,20,{"code":4,"msg":29,"data":30},"ok",{"site_id":23,"language":22,"slug":31,"title":13,"keywords":32,"description":14,"schema_data":33,"social_meta":85,"head_meta":87,"extra_data":89,"updated_unix":26},"melobottleneck-self-supervised-melody-skeleton-extraction-with-a-latent-subsequence-bottleneck","",{"@graph":34,"@context":84},[35,52,67],{"@type":36,"itemListElement":37},"BreadcrumbList",[38,42,46,49],{"item":39,"name":40,"@type":41,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":43,"name":44,"@type":41,"position":45},"https://docshare.wps.com/document/","Document",2,{"item":47,"name":12,"@type":41,"position":48},"https://docshare.wps.com/document/research-report/",3,{"item":50,"name":13,"@type":41,"position":51},"https://docshare.wps.com/document/melobottleneck-self-supervised-melody-skeleton-extraction-with-a-latent-subsequence-bottleneck/85857/",4,{"url":50,"name":13,"@type":53,"author":54,"headline":13,"publisher":56,"fileFormat":59,"inLanguage":22,"description":14,"dateModified":60,"datePublished":61,"encodingFormat":59,"isAccessibleForFree":62,"interactionStatistic":63},"DigitalDocument",{"name":9,"@type":55},"Person",{"url":39,"name":57,"@type":58},"DocShare","Organization","application/pdf","2026-07-22","2026-07-16",true,{"@type":64,"interactionType":65,"userInteractionCount":20},"InteractionCounter",{"@type":66},"ViewAction",{"@type":68,"mainEntity":69},"FAQPage",[70,76,80],{"name":71,"@type":72,"acceptedAnswer":73},"What problem does melody skeleton extraction address?","Question",{"text":74,"@type":75},"It produces a shorter monophonic melody that preserves structural note events while removing ornamental elaborations that can hide shared melodic material.","Answer",{"name":77,"@type":72,"acceptedAnswer":78},"How does MeloBottleneck represent the melody skeleton?",{"text":79,"@type":75},"It models the skeleton as a length-controlled, order-preserving latent subsequence, followed by a deterministic rhythmic-closure step that converts selected note events into a coherent reduced melody.",{"name":81,"@type":72,"acceptedAnswer":82},"How does MeloBottleneck training improve over pseudo-label-based methods?",{"text":83,"@type":75},"Training combines reconstruction, a frozen autoregressive melody prior, ornament-invariant consistency across ornamented views, and ornament exclusion, avoiding direct imitation of heuristic keep/delete pseudo-labels and improving cross-domain transfer.","https://schema.org",{"og:url":50,"og:type":86,"og:title":13,"og:site_name":57,"og:description":14},"article",{"robots":88,"canonical":50},"index,follow",{"doc_id":7,"site_id":23},{"code":4,"msg":5,"data":91},[92,96,100,104,109,114,119,122,126,129,133],{"id":20,"doc_module":4,"doc_module_name":44,"category_name":93,"show_sort_weight":94,"slug":95},"Story & Novel",90,"story-novel",{"id":45,"doc_module":4,"doc_module_name":44,"category_name":97,"show_sort_weight":98,"slug":99},"Literature",80,"literature",{"id":51,"doc_module":4,"doc_module_name":44,"category_name":101,"show_sort_weight":102,"slug":103},"Exam",70,"exam",{"id":105,"doc_module":4,"doc_module_name":44,"category_name":106,"show_sort_weight":107,"slug":108},5,"Comic",60,"comic",{"id":110,"doc_module":4,"doc_module_name":44,"category_name":111,"show_sort_weight":112,"slug":113},6,"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":44,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":44,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":44,"category_name":124,"show_sort_weight":27,"slug":125},9,"Religion & Spirituality","religion-spirituality",{"id":27,"doc_module":4,"doc_module_name":44,"category_name":127,"show_sort_weight":27,"slug":128},"World Cup","world-cup",{"id":130,"doc_module":4,"doc_module_name":44,"category_name":131,"show_sort_weight":130,"slug":132},10,"Lifestyle","lifestyle",{"id":134,"doc_module":4,"doc_module_name":44,"category_name":135,"show_sort_weight":105,"slug":136},19,"General","general"]