[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82808-en":3,"doc-seo-82808-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},82808,4398048950312,"Violet","https://ap-avatar.wpscdn.com/avatar/400002538284de19e3c?_k=1778320343897328908",8,"Research & Report","Auto-AEG Scalable Data Construction for Open-Vocabulary Audio Event Grounding","Large Audio-Language Models can reason about sound while failing to precisely localize when events occur, whereas classical Sound Event Detection achieves frame-level precision only within closed label vocabularies. Open-Vocabulary Audio Event Grounding predicts all onset/offset intervals of a target sound event specified by arbitrary natural-language queries, but progress is constrained by scarce training data and expensive manual boundary annotation. Auto-AEG introduces scalable automatic data construction plus model fine-tuning to generate exact supervision from synthesized clips and reward-driven pseudo-label learning from real audio, improving temporal localization benchmarks.","arXiv :2607 .04383v 1 [ cs . SD] 5 Jul 2026  \nAuto-AEG: Scalable Data Construction for Open-Vocabulary Audio Event Grounding  \nZihan Zhang 1, * , Xize Cheng 1, * , Wenhao Yan2, * , Tong Zhang 1 , Dongjie Fu 1 , Boyun Zhang 1 , Yongbo He 1 , Tao Jin 1,†  \n1 Zhejiang University, 2 Tsinghua University  \njint [zju@zju.edu.cn](zju@zju.edu.cn)  \n* Equal contribution. †Corresponding author.  \nAbstract  \nLarge Audio-Language Models (LALMs) reason fluently about sound yet struggle to localize precisely when events occur, while classical Sound Event Detection attains frame-level precision only over a closed label set. At the intersection of these paradigms lies the task of Open-Vocabulary Audio Event Grounding: predicting all time intervals of a target sound event described by an arbitrary natural language query. While this task is crucial for real-world audio understanding and LALM adaptation, it is bottlenecked by data scarcity. Few large-scale resources provide open-vocabulary onset/offset supervision, and manual temporal annotation is prohibitively expensive.  \nTo address this, we introduce Auto-AEG, a scalable pipeline that constructs such supervision by automatic data construction and model fine-tuning. It pairs programmatically synthesized clips, which carry exact ground-truth intervals for supervised cold-start, with multi-model pseudo-labels on realworld audio that supply the reward signal for reinforcement learning. Training with this pipeline yields promising performance gains on both the DESED SED benchmark and AEGBench, an independent difficulty-stratified benchmark we release. Our results show that automatically constructed data, coupled with interval-aware reward function design, is an effective data-side route to expanding the temporal localization capability of LALMs.  \nKeywords: audio event grounding, large audio-language models, temporal localization, data construction, reinforcement learning, sound event detection  \n1. Introduction  \nLarge Audio Language Models (LALMs) [7, 12, 28, 4, 10] have demonstrated remarkable capabilities in audio understanding and reasoning, bridging the acoustic and linguistic domains by aligning audio encoders with large language models. As these models grow increasingly proficient at describing, classifying, and reasoning about sound content, the ability to precisely localize sound events in time  \nemerges as a critical yet underexplored frontier.  \nClassical Sound Event Detection (SED) systems achieve frame-level temporal precision yet remain constrained by closed-label vocabularies that cannot generalize to the richness of natural language, while LALMs handle arbitrary queries with remarkable flexibility but struggle to produce finegrained temporal predictions. We study the task  \nthat sits between these paradigms, which we call Open-Vocabulary Audio Event Grounding: given an audio clip and a natural language query describing a target sound event, predict all onset/offset intervals where that event occurs, supporting multiple occurrences and an open event vocabulary.  \nProgress on this field is fundamentally limited by data scarcity. Existing temporal audio grounding datasets [37, 21] are small, narrow in label diversity, and the cost of manually annotating finegrained boundaries for arbitrary sound events is prohibitive. The field therefore lacks a viable training resource even though the task is within reach of current LALMs. Our central claim is that this gap can be closed on the data side, without architectural change, by constructing supervision automatically and exploiting it with reinforcement learning.  \nFigure 1: Auto-AEG pipeline overview. The pipeline annotates real-world audio with pseudo-labels for GRPO, while programmatically synthesizing clips with exact ground-truth intervals for SFT cold-start.  \nWe instantiate this with Auto-AEG (Figure 1), a scalable pipeline whose key insight is that the two kinds of data a grounding policy needs have complementary, annotation-free ac","cbCaiufDJkNgCBqI","https://ap.wps.com/l/cbCaiufDJkNgCBqI","pdf",1697878,2,1,15,"English","en",105,"# Introduction\n# Related Work","[{\"question\":\"What problem does Open-Vocabulary Audio Event Grounding address?\",\"answer\":\"It predicts all onset/offset time intervals where a queried sound event occurs in an audio clip, supporting multiple occurrences and an open event vocabulary described in natural language.\"},{\"question\":\"Why is progress in this task limited?\",\"answer\":\"Existing resources are small and narrow in label diversity, and manually annotating fine-grained temporal boundaries for arbitrary sound events is prohibitively expensive.\"},{\"question\":\"How does Auto-AEG construct training supervision?\",\"answer\":\"It uses two complementary, annotation-free sources: programmatically synthesized clips with exact ground-truth intervals for supervised cold-start, and multi-model pseudo-labels on real audio that provide rewards for reinforcement learning via GRPO.\"}]",1784183092,38,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"auto-aeg-scalable-data-construction-for-open-vocabulary-audio-event-grounding","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/auto-aeg-scalable-data-construction-for-open-vocabulary-audio-event-grounding/82808/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-19","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does Open-Vocabulary Audio Event Grounding address?","Question",{"text":75,"@type":76},"It predicts all onset/offset time intervals where a queried sound event occurs in an audio clip, supporting multiple occurrences and an open event vocabulary described in natural language.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"Why is progress in this task limited?",{"text":80,"@type":76},"Existing resources are small and narrow in label diversity, and manually annotating fine-grained temporal boundaries for arbitrary sound events is prohibitively expensive.",{"name":82,"@type":73,"acceptedAnswer":83},"How does Auto-AEG construct training supervision?",{"text":84,"@type":76},"It uses two complementary, annotation-free sources: programmatically synthesized clips with exact ground-truth intervals for supervised cold-start, and multi-model pseudo-labels on real audio that provide rewards for reinforcement learning via GRPO.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]