[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85538-en":3,"doc-seo-85538-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},85538,8796095360427,"Lucas Martin","https://ap-avatar.wpscdn.com/davatar_994ba38a5ba835b3df7d355c54d3ed8d",8,"Research & Report","MUGEN Evaluating and Improving Multi audio Understanding of Large Audio Language Models","Multi-audio understanding is essential for large audio-language models (LALMs) but remains insufficiently studied. This work presents MUGEN, a comprehensive benchmark spanning speech, general audio, and music, covering seven auditory dimensions. Experiments show systematic weaknesses in multi-audio settings and rapid performance drops as concurrent audio inputs increase, making input scaling a core bottleneck. Training-free methods are explored, where Audio-Permutational Self-Consistency boosts accuracy up to 6.28%, and combining with Chain-of-Thought reaches 6.74%.","MUGEN: Evaluating and Improving Multi-audio Understanding of Large  \nAudio-Language Models  \nChih-Kai Yang 1 ,∗, Yun-Shao Tsai 1 ,∗, Yu-Kai Guo2 ,†, Ping-Le Tsai2 ,†, Yen-Ting Piao2 ,‡, Hung-Wei Chen2 ,‡, Ting-Lin Hsiao2 ,‡, Yun-Man Hsu2 ,‡, Ke-Han Lu 1, Hung-yi Lee3 ,∗∗  \n1 Graduate Institute of Communication Engineering, National Taiwan University, Taiwan  \n2 National Taiwan University, Taiwan  \n3 NTU Artificial Intelligence Center of Research Excellence (NTU AI-CoRE), Taiwan  \n[chihkaiyang1124@gmail.com](chihkaiyang1124@gmail.com) , [r14942093@ntu.edu.tw](r14942093@ntu.edu.tw) , [hungyilee@ntu.edu.tw](hungyilee@ntu.edu.tw)  \narXiv :2603 .097 14v2 [ cs . SD] 12 Jul 2026  \nAbstract  \nWhile multi-audio understanding is critical for large audiolanguage models (LALMs), it remains underexplored. We introduce MUGEN, a comprehensive benchmark evaluating this capability across speech, general audio, and music. Our experiments reveal consistent weaknesses in multi-audio settings, and performance degrades sharply as the number of concurrent audio inputs increases, identifying input scaling as a fundamental bottleneck. We further investigate training-free strategies and observe that Audio-Permutational Self-Consistency, which diversifies the order of audio candidates, helps models form more robust aggregated predictions, yielding up to 6.28% accuracy gains. Combining this permutation strategy with Chain-ofThought further improves performance to 6.74% . These results expose blind spots in current LALMs and provide a foundation for evaluating complex auditory comprehension.  \nIndex Terms: large audio-language model, multi-audio understanding, benchmark  \n1. Introduction  \nLarge language models (LLMs) [1–3] have achieved remarkable progress in language understanding and have been extended to multimodal domains such as vision [4] and speech [5, 6] . Building on this trend, large audio-language models (LALMs) [7–14] integrate auditory perception [15, 16] with strong reasoning capability, enabling flexible interfaces for audio-centric applications such as voice agents [17] . However, current research predominantly evaluates these models in isolated and single-audio environments.  \nIn real-world deployments, LALMs require reasoning over multiple audio segments simultaneously. This requirement appears in settings such as audio-based in-context learning [18], where multiple audio demonstrations are required to guide adaptation, as well as in applications including speech retrievalaugmented generation (RAG) [19], multi-speaker analytics, and cross-utterance event matching. These scenarios all involve jointly understanding multiple audio segments, requiring models to compare, aggregate, and reconcile information across clips. Consequently, multi-audio understanding is not merely an advanced feature but a strict prerequisite for practical LALMs.  \nDespite its importance, current evaluation remains largely focused on single-audio settings. Existing benchmarks [20] assess general understanding [21, 22], reasoning [23–25], dialogue [26], bias [27], and safety [28, 29], but they largely  \n*†‡Equal Contribution.  \n**indicates the corresponding author.  \noverlook multi-audio scenarios. Recent attempts at multi-audio evaluation still exhibit two limitations: (1) narrow coverage of auditory attributes [30–32], often emphasizing semantic content or sound events while underrepresenting non-semantic aspects such as emotion; and (2) limited input scale [24, 32–35], typically involving only two to three audio clips per sample. As a result, systematic evaluation across diverse auditory dimensions and larger input scales remains underexplored.  \nTo fill this gap, we introduce MUGEN (Multi-audio Grounding and Understanding Benchmark), comprising 35 audio-grounding tasks across seven dimensions spanning speech, general audio, and music (Figure 1) . Each task requires selecting the audio that best satisfies a constraint from five candidates, enforcing cross-audio c","cbCaiusGl6r4GysV","https://ap.wps.com/l/cbCaiusGl6r4GysV","pdf",1383886,2,1,6,"English","en",105,"# Introduction\n# MUGEN Benchmark\n## Overview","[{\"question\":\"What is MUGEN, and what does it evaluate?\",\"answer\":\"MUGEN is a multi-audio grounding and understanding benchmark for large audio-language models. It evaluates performance across speech, general audio, and music with 35 tasks covering seven auditory dimensions.\"},{\"question\":\"What main limitation do experiments reveal for LALMs in multi-audio settings?\",\"answer\":\"The experiments show consistent weaknesses in multi-audio scenarios. Performance degrades sharply as the number of concurrent audio inputs increases, indicating input scaling as a fundamental bottleneck.\"},{\"question\":\"How do training-free strategies improve multi-audio performance in MUGEN?\",\"answer\":\"Audio-Permutational Self-Consistency diversifies the order of audio candidates to produce more robust aggregated predictions, yielding up to 6.28% accuracy gains. When combined with Chain-of-Thought reasoning, performance improves further to 6.74%.\"}]",1784204294,15,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"mugen-evaluating-and-improving-multi-audio-understanding-of-large-audio-language-models","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/mugen-evaluating-and-improving-multi-audio-understanding-of-large-audio-language-models/85538/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-18","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What is MUGEN, and what does it evaluate?","Question",{"text":75,"@type":76},"MUGEN is a multi-audio grounding and understanding benchmark for large audio-language models. It evaluates performance across speech, general audio, and music with 35 tasks covering seven auditory dimensions.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What main limitation do experiments reveal for LALMs in multi-audio settings?",{"text":80,"@type":76},"The experiments show consistent weaknesses in multi-audio scenarios. Performance degrades sharply as the number of concurrent audio inputs increases, indicating input scaling as a fundamental bottleneck.",{"name":82,"@type":73,"acceptedAnswer":83},"How do training-free strategies improve multi-audio performance in MUGEN?",{"text":84,"@type":76},"Audio-Permutational Self-Consistency diversifies the order of audio candidates to produce more robust aggregated predictions, yielding up to 6.28% accuracy gains. When combined with Chain-of-Thought reasoning, performance improves further to 6.74%.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,114,119,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":111,"show_sort_weight":112,"slug":113},"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]