[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82649-en":3,"doc-seo-82649-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},82649,1649267921044,"Ava Thompson","https://us-avatar.wpscdn.com/avatar/1800007509477c92dfb?_k=1782875107921204101",8,"Research & Report","A Multi-Branch Hierarchy-Aware Framework for Heterogeneous Audio Classification","Technical report for DCASE 2026 Challenge Task 1 on heterogeneous audio classification under the Broad Sound Taxonomy (BST), requiring both accurate second-level predictions and top-level taxonomy consistency. The system leverages CLAP-based audio-text representations and improves performance through training set expansion with a filtered BSD35k subset, feature-specific acoustic branches (MFCC, log-Mel, log-STFT), and hierarchy-aware classifiers with KNN post-processing. Single-model results reach Hier. F1 80.84% on BSD10k-v1.2, while ensembles achieve 81.25% and 81.18% Hier. F1 using complementary features and heads.","A MULTI-BRANCH HIERARCHY-AWARE FRAMEWORK FOR HETEROGENEOUS AUDIO  \nCLASSIFICATION  \nTechnical Report  \nBeile Ning 1 ,†, Jiayi Yu 1 ,†, Zitong Wang 1 ,†, Yufei Hu 1 ,†, Wenjun Xu 1 ,†  \nYuanhang Qian 1, Zhongxin Bai2, Gongping Huang 1 ,∗  \n1 Wuhan University, Wuhan, China  \n2 Harbin Engineering University, Harbin, China  \n{ningbeile, [gongpinghuang](gongpinghuang}@whu.edu.cn)[}](gongpinghuang}@whu.edu.cn)[@whu.edu.cn](gongpinghuang}@whu.edu.cn)  \narXiv :2607 .0 1974v 1 [ cs . SD] 2 Jul 2026  \nABSTRACT  \nThis technical report describes our system for Task 1 of the DCASE 2026 Challenge, which aims to classify heterogeneous audio recordings according to the Broad Sound Taxonomy (BST) . The task requires both accurate second-level prediction and consistency with the top-level taxonomy. Our system is built on CLAP-based audio-text representations and is improved along three strategies: expanding the training set with a filtered subset of BSD35k, enhancing acoustic modeling with feature-specific branches, and refining predictions using hierarchy-aware classifiers and KNN-based postprocessing. Among the acoustic features considered, the log-STFT branch provides the strongest single-model performance. With KNN-based post-processing, our best single system achieves a hierarchical F1 score (Hier. F1) of 80 . 84% on the BSD10k-v1 .2 set under the same evaluation protocol as the baseline. We further construct ensemble systems by combining models with complementary acoustic features and classification heads, achieving Hier. F1 of 81.25% and 81.18%, respectively.  \nIndex Terms— DCASE2026, CLAP, Heterogeneous Audio Classification, Broad Sound Taxonomy (BST)  \n1. INTRODUCTION  \nHeterogeneous audio classification aims to recognize sound events and acoustic scenes from real-world recordings with diverse content, recording conditions, and metadata quality. In DCASE 2026 Task 1, audio samples are annotated according to the Broad Sound Taxonomy (BST), where 23 second-level categories are grouped into 5 top-level classes [6] . Since the official evaluation metric is based on the Hier. F1 [10], a system should not only predict the correct second-level class, but also preserve consistency with the corresponding top-level category.  \nOur system uses CLAP audio-text representations [11] as the main semantic representation. To improve data diversity and reduce the impact of label noise, we construct an expanded training set, denoted as BSD-Grand, by incorporating a filtered subset of BSD35k [5] into the BSD10k-v1.2 training data [7, 8] . The additional samples are selected through category-aware metadata cleaning, teacher-model filtering, and uploader-level constraints, which help reduce noisy annotations and uploader-specific bias. To complement the high-level semantic information captured by CLAP,  \n† These authors contributed equally to this work.  \n∗ Corresponding author.  \nwe further introduce feature-specific acoustic branches based on MFCC, log-Mel spectrogram, and log-STFT features. Each acoustic feature is encoded by a corresponding branch and fused with the CLAP audio-text embeddings for audio classification. In addition, we investigate several prediction heads, including flat, globalclassifier-based, and local-classifier-per-level-based heads, to better exploit the hierarchical relationship between the 5 top-level groups and the 23 second-level categories. Finally, we refine the model predictions using a KNN-based post-processing derived from the training embedding bank, and further incorporate this post-processing as soft supervision in a knowledge distillation framework.  \nBased on these components, we construct the following four submitted systems:  \n• A single log-STFT-based model with KNN-based postprocessing. This is our best single model and achieves a Hier. F1 of 80 . 84% .  \n• An ensemble of KD-log-STFT, log-Mel, Flat, and LCL. This system is trained on the full training set.  \n• An ensemble of KD-log-STFT, log-Mel, Flat, and LCL. Thi","cbCaiftkFPdi8RiU","https://ap.wps.com/l/cbCaiftkFPdi8RiU","pdf",638114,2,1,5,"English","en",105,"# Abstract\n# Introduction\n# Model Architecture and Training Strategy\n## Feature-Specific Acoustic Branch Framework","[{\"question\":\"What is the main goal of the DCASE 2026 Task 1 system described in the report?\",\"answer\":\"The system classifies heterogeneous audio recordings according to the Broad Sound Taxonomy (BST), achieving accurate second-level prediction while maintaining consistency with the corresponding top-level taxonomy.\"},{\"question\":\"How does the system expand and clean the training data to reduce label noise?\",\"answer\":\"It constructs an expanded training set (BSD-Grand) by adding a filtered subset of BSD35k into BSD10k-v1.2, using category-aware metadata cleaning, teacher-model filtering, and uploader-level constraints to mitigate noisy annotations and uploader-specific bias.\"},{\"question\":\"Which modeling and inference strategies are used to improve hierarchical performance?\",\"answer\":\"The report introduces feature-specific acoustic branches (MFCC, log-Mel, log-STFT), explores multiple hierarchy-exploiting classification heads, and refines predictions with KNN-based post-processing from a training embedding bank, further used as soft supervision in a knowledge distillation framework.\"}]",1784182064,13,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"a-multi-branch-hierarchy-aware-framework-for-heterogeneous-audio-classification","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/a-multi-branch-hierarchy-aware-framework-for-heterogeneous-audio-classification/82649/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-23","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What is the main goal of the DCASE 2026 Task 1 system described in the report?","Question",{"text":75,"@type":76},"The system classifies heterogeneous audio recordings according to the Broad Sound Taxonomy (BST), achieving accurate second-level prediction while maintaining consistency with the corresponding top-level taxonomy.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the system expand and clean the training data to reduce label noise?",{"text":80,"@type":76},"It constructs an expanded training set (BSD-Grand) by adding a filtered subset of BSD35k into BSD10k-v1.2, using category-aware metadata cleaning, teacher-model filtering, and uploader-level constraints to mitigate noisy annotations and uploader-specific bias.",{"name":82,"@type":73,"acceptedAnswer":83},"Which modeling and inference strategies are used to improve hierarchical performance?",{"text":84,"@type":76},"The report introduces feature-specific acoustic branches (MFCC, log-Mel, log-STFT), explores multiple hierarchy-exploiting classification heads, and refines predictions with KNN-based post-processing from a training embedding bank, further used as soft supervision in a knowledge distillation framework.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,109,114,119,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":106,"show_sort_weight":107,"slug":108},"Comic",60,"comic",{"id":110,"doc_module":4,"doc_module_name":46,"category_name":111,"show_sort_weight":112,"slug":113},6,"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":22,"slug":137},19,"General","general"]