[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82633-en":3,"doc-seo-82633-105":29,"detail-sidebar-cat-0-en-105":83},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},82633,16904993612988,"Olivia Brown","https://ap-avatar.wpscdn.com/davatar_a8503ba1806abce46bf441b54a3ca4cd",8,"Research & Report","ICME 2026 Academic Text-to-Music Generation Text-to-Audio Batch Sampling Submission","This work studies batch sampling strategies for text-to-audio music generation when training data are scarce and models are small, focusing on the ICME 2026 Grand Challenge submission. Training samples are clustered by text embeddings or audio embeddings, then grouped into mini-batches to reduce gradient interference caused by heterogeneous data. The study analyzes how modality choice and cluster granularity affect both objective evaluation and perceptual listening outcomes. Text-embedding clustering yields stronger objective metrics, while moderate cluster counts score best overall; more clusters improve coherent musical structure.","UT-AISTimprt submission for ICME 2026 Grand Challenge on Academic Text-to-Music Generation  \nShunsuke Yoshida∗ , Yu-Hua Chen†, Satoru Fukayama†  \n∗ University of Tokyo, Japan  \n†National Institute of Advanced Industrial Science and Technology (AIST), Japan  \narXiv :2607 .0 1669v 1 [ cs . SD] 2 Jul 2026  \nAbstract—This work investigates the effect of batch sampling strategies during training for text-to-audio music generation under low-data and small-scale model settings. This paper describes our approach and findings for the ICME 2026 Grand Challenge on Academic Text-to-Music Generation. Training data are clustered using either text embeddings or audio embeddings, and samples with similar characteristics are grouped within the same mini-batch to mitigate gradient interference. The effects of modality and cluster granularity on clustering are analyzed. Results show that clustering based on text embeddings achieves better performance on objective evaluation metrics than clustering based on audio embeddings. In addition, different cluster granularity leads to different behaviors across evaluation criteria: a moderate number of clusters performs best on objective metrics, while a larger number of clusters tends to exhibit music with more coherent structure in listening tests.  \nIndex Terms—music generation, text-to-audio, batch sampling, clustering  \nI. INTRODUCTION  \nMachine learning for music generation often suffers from a limited amount of publicly available training data compared to text or image domains. As a result, models for music generation are frequently trained on small-scale datasets, particularly in conditional and personalized generation settings. In practice, training small-scale models from scratch remains a realistic and commonly adopted setting.  \nWith a limited amount of training data, using small-scale models can help mitigate overfitting, but their representational capacity is limited. Consequently, the way training data are selected and presented during learning has a strong impact on generation performance. Restricting training data to a narrow genre or condition often reduces the diversity of output. On the other hand, naively mixing heterogeneous data can destabilize training of small-scale models due to gradient interference among diverse training data. The limited model capacity of the small-scale models have less flexibility to disentangle heterogeneous generative factors [1], [2] .  \nSimilar issues have been reported in the natural language processing (NLP) community, particularly in the context of multi-task learning and instruction tuning [1]–[3] . When multiple heterogeneous tasks or instructions are jointly learned, gradient conflicts can degrade both training efficiency and model performance. To address this problem, recent studies have proposed focusing on batch construction during training rather than on global data mixing ratios [3] . Commonality-aware Instruction Tuning (CommonIT) clusters training samples in  \nadvance and constructs each mini-batch from a single cluster, thereby increasing intra-batch homogeneity while preserving diversity across batches.  \nInspired by CommonIT, this work investigates clusteringbased batch sampling for text-to-audio music generation under low-resource and scratch-training conditions. The conditions are aligned with the ICME 2026 Grand Challenge on Academic Text-to-Music Generation [4] . Since text-toaudio models involve both text and audio modalities, the choice of modality used to define similarity becomes a key decision. In this submission, training samples are clustered using embedding representations derived from either a text encoder or an audio encoder, and mini-batches are constructed from samples within the same cluster.  \nThis work reports three main findings:  \n• clustering-based batch sampling improves the performance of small-scale music generation models under lowdata settings,  \n• text-based clustering achieves better performance in objective e","cbCaieEqppA1ZIYD","https://ap.wps.com/l/cbCaieEqppA1ZIYD","pdf",145435,1,5,"English","en",105,"# Abstract\n# Introduction\n# Clustering-based Batch Sampling for Text-to-Audio Music Generation\n## Overview\n## Baseline Text-to-Audio Model","[{\"question\":\"What do the results say about choosing text vs. audio embeddings for clustering and about cluster granularity?\",\"answer\":\"Clustering based on text embeddings performs better on objective evaluation metrics than clustering based on audio embeddings. A moderate number of clusters performs best on objective metrics, while a larger number of clusters tends to yield more coherent structure in listening tests.\"}]",1784181936,13,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":78,"head_meta":80,"extra_data":82,"updated_unix":27},"icme-2026-academic-text-to-music-generation-text-to-audio-batch-sampling-submission","",{"@graph":35,"@context":77},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/icme-2026-academic-text-to-music-generation-text-to-audio-batch-sampling-submission/82633/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71],{"name":72,"@type":73,"acceptedAnswer":74},"What do the results say about choosing text vs. audio embeddings for clustering and about cluster granularity?","Question",{"text":75,"@type":76},"Clustering based on text embeddings performs better on objective evaluation metrics than clustering based on audio embeddings. A moderate number of clusters performs best on objective metrics, while a larger number of clusters tends to yield more coherent structure in listening tests.","Answer","https://schema.org",{"og:url":51,"og:type":79,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":81,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":84},[85,89,93,97,101,106,111,114,119,122,126],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":86,"show_sort_weight":87,"slug":88},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":90,"show_sort_weight":91,"slug":92},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Exam",70,"exam",{"id":21,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Comic",60,"comic",{"id":102,"doc_module":4,"doc_module_name":45,"category_name":103,"show_sort_weight":104,"slug":105},6,"Technology",50,"technology",{"id":107,"doc_module":4,"doc_module_name":45,"category_name":108,"show_sort_weight":109,"slug":110},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":112,"slug":113},30,"research-report",{"id":115,"doc_module":4,"doc_module_name":45,"category_name":116,"show_sort_weight":117,"slug":118},9,"Religion & Spirituality",20,"religion-spirituality",{"id":117,"doc_module":4,"doc_module_name":45,"category_name":120,"show_sort_weight":117,"slug":121},"World Cup","world-cup",{"id":123,"doc_module":4,"doc_module_name":45,"category_name":124,"show_sort_weight":123,"slug":125},10,"Lifestyle","lifestyle",{"id":127,"doc_module":4,"doc_module_name":45,"category_name":128,"show_sort_weight":21,"slug":129},19,"General","general"]