[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85614-en":3,"doc-seo-85614-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},85614,3848291630094,"Emma Wilson","https://eur-avatar.wpscdn.com/davatar_085a072bc5b1113ac321206ff7593b45",8,"Research & Report","Towards Event-Robust Acoustic Scene Classification","A study presents the Event-Shifted Acoustic Scene (ESAS) dataset, designed to benchmark Acoustic Scene Classification (ASC) robustness against unknown or unexpected sound events. Existing ASC benchmarks emphasize clean, consistent audio, while real environments contain diverse foreground events. ESAS simulates realistic acoustic variability by injecting foreground sound events into background scenes using large language models, and provides construction methodology, dataset statistics, and evaluation protocols. Experiments show substantial performance degradation for current ASC models under event-shift conditions, motivating future event-robust research.","Towards Event-Robust Acoustic Scene Classification  \nYiqiang Cai  1 ,∗, Bohan Hu 1 ,∗, Yu Yang2, Pengwei Lu3, Shengchen Li 1 ,∗∗, Xi Shao4  \n1 Xi’an Jiaotong-Liverpool University, Suzhou, China  \n2 Zhongdian Zhiheng Information Technology Service Co., Ltd, Nanjing, China  \n3 China Telecom Jiangsu Branch, Nanjing, China  \n4 Nanjing University of Posts and Telecommunications, Nanjing, China  \n[caiyiqaing0902@gmail.com](caiyiqaing0902@gmail.com) , [shengchen.li@xjtlu.edu.cn](shengchen.li@xjtlu.edu.cn)  \narXiv :2606 .0692 1v 3 [ cs . SD] 13 Jul 2026  \nAbstract  \nThis paper introduces the Event-Shifted Acoustic Scene (ESAS) dataset, a novel benchmark for evaluating the robustness of Acoustic Scene Classification (ASC) systems against unknown sound events. Existing ASC datasets typically contain recordings of clean and consistent audio, while real-world environments often include diverse and unexpected sound events. To bridge this gap, ESAS simulates real-world acoustic variability by injecting foreground sound events into background scenes with the assistance of large language models. In this work, we present the construction methodology, dataset statistics, and evaluation protocols. Furthermore, a comprehensive evaluation of state-of-the-art ASC systems is conducted using the ESAS benchmark. Experimental results reveal that existing ASC models suffer significant performance degradation when facing the event-shift challenge. The introduction of the ESAS dataset aims to drive future research toward event-robust ASC. Index Terms: Acoustic scene classification, dataset, event shift.  \n1. Introduction  \nAcoustic Scene Classification (ASC) is a fundamental task in the field of computational sound scene analysis [1, 2] . ASC aims to recognize the environment in which an audio recording was captured, such as a park, airport, or metro station. While ASC models have achieved remarkable progress with the help of large-scale datasets [3, 4, 5] and deep learning architectures [6, 7, 8], real-world acoustic scenes are often far more complex than those represented in current benchmarks.  \nAn acoustic scene typically consists of foreground sound events and background noise [9] . In practical environments, the foreground sound events within a scene can vary drastically depending on time, season, and location [10] . For example, a park during the day may be dominated by children playing and birds chirping, whereas at night it may contain footsteps or traffic noise from nearby roads. Similarly, the same residential area may sound very different across regions or countries, reflecting distinct cultural or environmental activities. These variations cause what we refer to as event shift—a phenomenon where the foreground events within an acoustic scene change significantly while the underlying scene category remains the same. Event shift is common and inevitable in the real world, and it poses a major challenge to ASC systems that often rely heavily on event-related information [11, 12, 13] . Studying event shift is therefore crucial not only for improving robustness but also for revealing how acoustic scene features encode and separate background ambience from foreground sound events.  \n*These authors contributed equally.  \n**indicates the corresponding author.  \nPrevious research has explored other dimensions of acoustic variations, such as cross-city setting [5, 14], cross-device setting [15, 16], and time-variant setting [10] . These studies have provided valuable insights into geographic, channel and temporal mismatches. However, such factors are relatively limited to recording conditions or hardware settings. In contrast, event shift is a more general and fundamental problem, as it naturally occurs in all real-world acoustic environments regardless of device, city, or recording setup. It reflects the intrinsic dynamic nature of soundscapes, where the acoustic composition of a scene continually evolves with human activity and environmental context.","cbCaijHZH6TQlHKH","https://ap.wps.com/l/cbCaijHZH6TQlHKH","pdf",1524713,3,1,5,"English","en",105,"# Introduction\n## Acoustic Scene Classification and Event Shift\n## Limitations of Existing Benchmarks\n## ESAS Dataset Design","[{\"question\":\"What problem does the ESAS benchmark target?\",\"answer\":\"ESAS targets the event-shift challenge in acoustic scene classification, where foreground sound events change significantly while the scene category remains the same. It evaluates how accurately ASC systems recognize scenes under unknown or unfamiliar events.\"},{\"question\":\"How is the ESAS dataset constructed?\",\"answer\":\"ESAS mixes background scenes from CochlScene with foreground sound events from FSD50K to form polyphonic sound scenes with overlapping events. Large language models guide event-scene grouping to preserve semantic consistency.\"},{\"question\":\"What do the experiments show about current ASC models under event shift?\",\"answer\":\"The experiments indicate existing ASC models suffer significant performance degradation when faced with the event-shift challenge. This supports the need for event-robust ASC research and evaluation.\"}]",1784204933,13,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"towards-event-robust-acoustic-scene-classification","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/towards-event-robust-acoustic-scene-classification/85614/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the ESAS benchmark target?","Question",{"text":75,"@type":76},"ESAS targets the event-shift challenge in acoustic scene classification, where foreground sound events change significantly while the scene category remains the same. It evaluates how accurately ASC systems recognize scenes under unknown or unfamiliar events.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How is the ESAS dataset constructed?",{"text":80,"@type":76},"ESAS mixes background scenes from CochlScene with foreground sound events from FSD50K to form polyphonic sound scenes with overlapping events. Large language models guide event-scene grouping to preserve semantic consistency.",{"name":82,"@type":73,"acceptedAnswer":83},"What do the experiments show about current ASC models under event shift?",{"text":84,"@type":76},"The experiments indicate existing ASC models suffer significant performance degradation when faced with the event-shift challenge. This supports the need for event-robust ASC research and evaluation.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,109,114,119,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":106,"show_sort_weight":107,"slug":108},"Comic",60,"comic",{"id":110,"doc_module":4,"doc_module_name":46,"category_name":111,"show_sort_weight":112,"slug":113},6,"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":22,"slug":137},19,"General","general"]