[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86204-en":3,"doc-seo-86204-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86204,1374391974564,"Clementine","https://ap-avatar.wpscdn.com/avatar/14000253aa45c000a9e?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779874745381141002",8,"Research & Report","Verifier-Guided Twelve-Tone Composition: Generate–Verify–Repair Harness for Symbolic Music Generation","Large language models can produce twelvetone scores that meet surface constraints yet collapse into degenerate musical textures. The work proposes a neuro-symbolic generate–verify–repair–trace harness that wraps an LLM proposer with symbolic verification. The pipeline improves event-local consistency without asserting whole-piece legality, explicitly abstaining when checks fail. Across 40 controlled tasks and four paired models, audited delivery increases from 13.3% to 48.1%, while degeneracy stays near 0.05. Expert-blinded evaluation shows stronger preference toward harness candidates on adherence, legality perception, coherence, and quality.","Verifier-Guided Twelve-Tone Composition: A Generate–Verify–Repair Harness for Symbolic Music Generation  \nCongren Dai1,2 , Danni Zhao1 , Enyang Liu1 , Michael Ching Yam1 ,  \nZhancheng Guo1,2 , Siyi Gu1 , Wentao Yang1 , Bo Dai1 , Xiaobing Li1 , and Maosong Sun1,2 B  \n1 Central Conservatory of Music [congren.dai@mail.ccom.edu.cn](congren.dai@mail.ccom.edu.cn)  \n2 Tsinghua University [sms@tsinghua.edu.cn](sms@tsinghua.edu.cn)  \narXiv :2607 . 1 1334v 1 [ cs .AI] 13 Jul 2026  \nAbstract  \nLarge language models can produce superficially legal twelvetone scores that collapse into degenerate textures. We introduce a neuro-symbolic harness that wraps a language-model proposer in a generate–verify–repair–trace loop with symbolic verification. The complete pipeline improves event-local consistency without claiming whole-piece legality. Across 40 controlled tasks and four paired models, audited delivery yield rises from 13 .3% under raw generation to 48 . 1% with the harness, which explicitly abstains otherwise. The pass rate of a narrower collision and serialisation-consistency check rises from 33.5% to 58.3%, while degeneracy remains near 0.05, including under exploratory adversarial prompting. A blinded evaluation by five experts also shows a descriptive aggregate preference for harness candidates over raw generation in adherence, perceived legality, coherence, and overall quality.  \n1 Introduction  \nTwelve-tone (serial) music, a compositional method in which an ordered tone row of pitch classes governs harmony and melody through transposition, inversion, and retrograde (Schoenberg 1950; Straus 2016), is an attractive testbed for controllable generation: its core rules are crisp and formal, yet musically interesting serial realisations often depend on global organisation (Perle 1991) . More generally, maintaining long-range coherence is a central challenge in symbolic music generation (Huang et al. 2019) . This combination exposes a failure mode that is easy to miss when “rule legality”is the only metric. A model instructed to prioritise legality can satisfy the letter of the rules with a degenerate score (a few sustained chords, collapsed voices, near-empty bars) that is nominally compliant but musically empty. This behaviour is analogous to specification gaming: it satisfies the stated rule-based proxy while defeating the intended musical objective. Reward-hacking work studies the related, but not identical, case in which optimisation exploits a misspecified objective or reward proxy (Amodei et al. 2016; Pan, Bhatia, and Steinhardt 2022; Skalse et al. 2022); here we elicit shortcuts through adversarial prompts rather than trainingtime reward optimisation.  \nWe argue that a structural interventionis more reliable than prompting alone. A deterministic verifier records hard-rule  \nBCorresponding author.  \nviolations, and a row-aware repair operator projects proposed material back onto the active row slice whenever possible. The large language model (LLM) is retained where it is genuinely useful (high-level planning and note-level proposals that drive musical quality), while symbolic components control pitch-class assignment and surface unresolved violations explicitly. Unlike prompt-only self-correction, the symbolic layer evaluates executable events rather than modelgenerated claims. Its trace records which row positions survive repair, making dropped or modified material inspectable in the delivered artefact.  \nThis setting is a clean microcosm of a broader challenge in neuro-symbolic generation: automatic proxies can be satisfied without producing useful outputs, and process-level checks need to be distinguished from independent checks on the delivered artefact. Twelve-tone music sharpens the tension because legality is a crisp Boolean property, whereas musical quality is global and is precisely what a legalitymaximiser sacrifices. Practically, the harness returns either an audit-verified score or an explicit failure with an inspectab","cbCaipSCITXL9Ukm","https://ap.wps.com/l/cbCaipSCITXL9Ukm","pdf",1823448,5,1,16,"English","en",105,"# Introduction\n## Challenges in Rule-Legality Maximization\n## Neuro-Symbolic Structural Intervention\n# Related Work\n## Symbolic Music Generation","[{\"question\":\"What problem does the paper address in twelve-tone symbolic music generation?\",\"answer\":\"Language models may generate scores that look rule-compliant but collapse into degenerate textures, effectively gaming legality while failing musically useful coherence. The paper targets this gap between local “rule legality” and meaningful musical output.\"},{\"question\":\"How does the generate–verify–repair–trace harness work?\",\"answer\":\"An LLM proposes candidate events, then a deterministic verifier checks hard rules and applies a row-aware repair operator that projects proposals back onto the active pitch-class row when possible. A trace records which row positions survive, enabling inspectable diagnostics and a final delivered-score audit gate.\"},{\"question\":\"What improvements are reported over raw generation?\",\"answer\":\"On 40 controlled tasks, audited delivery rises from 13.3% to 48.1% when using the harness, with an increased pass rate for collision and serialisation-consistency checks from 33.5% to 58.3%. Degeneracy remains low (near 0.05), including under adversarial prompting, and expert evaluation shows overall preference for harness outputs.\"}]",1784209414,40,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"verifier-guided-twelve-tone-composition-generateverifyrepair-harness-for-symbolic-music-generation","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/verifier-guided-twelve-tone-composition-generateverifyrepair-harness-for-symbolic-music-generation/86204/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-27","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What problem does the paper address in twelve-tone symbolic music generation?","Question",{"text":76,"@type":77},"Language models may generate scores that look rule-compliant but collapse into degenerate textures, effectively gaming legality while failing musically useful coherence. The paper targets this gap between local “rule legality” and meaningful musical output.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How does the generate–verify–repair–trace harness work?",{"text":81,"@type":77},"An LLM proposes candidate events, then a deterministic verifier checks hard rules and applies a row-aware repair operator that projects proposals back onto the active pitch-class row when possible. A trace records which row positions survive, enabling inspectable diagnostics and a final delivered-score audit gate.",{"name":83,"@type":74,"acceptedAnswer":84},"What improvements are reported over raw generation?",{"text":85,"@type":77},"On 40 controlled tasks, audited delivery rises from 13.3% to 48.1% when using the harness, with an increased pass rate for collision and serialisation-consistency checks from 33.5% to 58.3%. Degeneracy remains low (near 0.05), including under adversarial prompting, and expert evaluation shows overall preference for harness outputs.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,110,115,119,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":29,"slug":118},7,"Healthcare","healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":20,"slug":137},19,"General","general"]