[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82191-en":3,"doc-seo-82191-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},82191,2336464648746,"Skyler","https://ap-avatar.wpscdn.com/davatar_276721f389ce27ea32af1340a28f341c",8,"Research & Report","Event-Based Token Sequences for Audio-Conditioned Music-Game Level Modeling","Procedural generation of music game levels translates musical structure into interactive sequences of timed gameplay events. Frame-based methods discretize audio into uniform time grids, leaving event-level timing relations and long-range structure implicit and hard to model. The work proposes an event-level, token-sequence formulation inspired by event-based symbolic music modeling, casting level generation as multimodal sequence-to-sequence learning conditioned on audio excerpts and metadata. A Transformer alternates gameplay-event and beat-shift tokens, improving event-level evaluation and enabling analysis of rhythm-aligned event prediction support from audio.","Event-Based Token Sequences for Audio-Conditioned Music-Game Level Modeling  \nKe Zhang∗ Japan Advanced Institute of Science and Technology Nomi, Ishikawa, Japan[s2660002@jaist.ac.jp](s2660002@jaist.ac.jp)  \nChu-Hsuan Hsueh  \nJapan Advanced Institute of Science and Technology Nomi, Ishikawa, Japan [hsuehch@jaist.ac.jp](hsuehch@jaist.ac.jp)  \nKokolo Ikeda  \nJapan Advanced Institute of Science and Technology Nomi, Ishikawa, Japan [kokolo@jaist.ac.jp](kokolo@jaist.ac.jp)  \narXiv :2607 .09095v 1 [ cs . SD] 10 Jul 2026  \nAbstract  \nProcedural generation of music game levels is an exciting yet challenging problem, as levels must translate musical structure into interactive sequences of timed gameplay events. Most existing approaches formulate this task by frame-based representations, dividing audio into uniform time grids and predicting events at each frame. This makes gameplay events implicit across many frames. Asa result, it is hard to describe event-level timing relations and longerrange structure found in human-authored levels. We use procedural generation as a practical setting to study how musical cues map to interactive event sequences. Inspired by event-based symbolic music modeling, we propose a token-level sequence formulation that casts level generation as a multimodal sequence-to-sequence problem. Conditioned on an audio excerpt and level metadata, the model generates a token sequence alternating gameplay-event and beat-shift tokens. This explicitly represents actions and their relative timing in beat space. Based on this formulation, we build a Transformer model. It outperforms representative frame-level baselines under event-level evaluation. It also enables systematic analysis of how audio supports rhythm-aligned event prediction beyond metadata conditioning.  \nCCS Concepts  \n• Applied computing → Sound and music computing; • Information systems → Music retrieval; • Computing methodologies → Neural networks.  \nKeywords  \nmusic game, rhythm game, multimodal sequence modeling, tokenlevel modeling, audio-conditioned generation  \nACM Reference Format:  \nKe Zhang, Chu-Hsuan Hsueh, and Kokolo Ikeda. 2026. Event-Based Token Sequences for Audio-Conditioned Music-Game Level Modeling. In International Conference on Multimedia Retrieval (ICMR’26), June 16–19, 2026, Amsterdam, Netherlands. ACM, New York, NY, USA, 9 pages. [https:](https:)//[doi.org/10.1145/3805622.3810623](doi.org/10.1145/3805622.3810623)  \n∗ Corresponding author.  \nThis work is licensed under a Creative Commons Attribution 4 .0 International License. ICMR’26, Amsterdam, Netherlands  \n© 2026 Copyright held by the owner/author(s) .  \nACM ISBN 979-8-4007-2617-0/2026/06  \n[https://doi.org/10.1145/3805622.3810623](https://doi.org/10.1145/3805622.3810623)  \n1 Introduction  \nUnderstanding how musical cues map to structured event sequences is a core challenge in music-game level modeling, and relates to broader questions in music information retrieval on representing and interpreting audio [26]. Procedural content generation (PCG) provides a practical setting to test this audio–event correspondence, because the mapping must be realized as playable interactive content [29]. A music game level is not merely a sequence of actions, but a temporally structured artefact shaped by musical rhythm, phrasing, and difficulty design. Well-designed levels translate musical structure into expressive and playable event sequences, requiring both precise local timing and coherent long-range organization. Generating such content demands models that bridge continuous audio signals and discrete symbolic actions, while preserving temporal structure across multiple scales. At its core, level modeling requires identifying salient temporal cues from rich, time-varying audio and organizing them into structured event sequences that players can directly interact with. This perspective emphasizes event patterns, rather than dense uniform time grids. While the objective is to synthesize new interac","cbCaibn9a7BqhzZ3","https://ap.wps.com/l/cbCaibn9a7BqhzZ3","pdf",828836,1,9,"English","en",105,"# Introduction\n## Audio-to-event correspondence in music-game level modeling\n## Limitations of frame-based temporal representations\n## Motivation from event-based symbolic music modeling","[{\"question\":\"Why are frame-based representations less suitable for music-game level modeling?\",\"answer\":\"They discretize time into uniform grids, making event-level timing relations and long-range structure implicit across many frames, so rhythmic consistency and long-range organization are not modeled explicitly.\"},{\"question\":\"How does the proposed method represent level generation?\",\"answer\":\"It formulates generation as a multimodal sequence-to-sequence problem using token-level sequences that alternate gameplay-event tokens and beat-shift tokens in beat space.\"},{\"question\":\"What inputs condition the token generation model?\",\"answer\":\"The model conditions on an audio excerpt together with level metadata to guide the generated sequence of interactive events.\"}]",1784178715,23,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"event-based-token-sequences-for-audio-conditioned-music-game-level-modeling","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/event-based-token-sequences-for-audio-conditioned-music-game-level-modeling/82191/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why are frame-based representations less suitable for music-game level modeling?","Question",{"text":75,"@type":76},"They discretize time into uniform grids, making event-level timing relations and long-range structure implicit across many frames, so rhythmic consistency and long-range organization are not modeled explicitly.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the proposed method represent level generation?",{"text":80,"@type":76},"It formulates generation as a multimodal sequence-to-sequence problem using token-level sequences that alternate gameplay-event tokens and beat-shift tokens in beat space.",{"name":82,"@type":73,"acceptedAnswer":83},"What inputs condition the token generation model?",{"text":84,"@type":76},"The model conditions on an audio excerpt together with level metadata to guide the generated sequence of interactive events.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,127,130,134],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":21,"doc_module":4,"doc_module_name":45,"category_name":124,"show_sort_weight":125,"slug":126},"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":45,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":45,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":45,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]