[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86352-en":3,"doc-seo-86352-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86352,1099513958762,"Logic","https://ap-avatar.wpscdn.com/avatar/1000023916a998db790?x-image-process=image/resize,m_fixed,w_180,h_180&k=1784791008015729253",8,"Research & Report","OpenBEATs A Fully Open-Source General-Purpose Audio Encoder","Masked token prediction has become an effective pretraining objective across language, vision, and speech, yet general audio understanding is still underexplored beyond BEATs. BEATs lacks open-source pretraining code and is trained only on AudioSet, limiting downstream transfer. OpenBEATs provides an open-source framework extending BEATs through multidomain audio pretraining, evaluated on six task types, twenty-five datasets, and three domains. It reaches strong results on bioacoustics, environmental sounds, and reasoning tasks, enabling reproducible research via released code and checkpoints.","OpenBEATs: A Fully Open-Source General-Purpose Audio Encoder  \nShikhar Bharadwaj 1, Samuele Cornell 1, Kwanghee Choi 1, Satoru Fukayama2, Hye-jin Shim 1 , Soham Deshmukh 1, Shinji Watanabe 1  \n1 Carnegie Mellon University, USA  \n2National Institute of Advanced Industrial Science and Technology (AIST), Japan  \n[sbharad2@andrew.cmu.edu](sbharad2@andrew.cmu.edu)  \narXiv :2507 . 14 129v2 [ cs . SD] 13 Jul 2026  \nAbstract—Masked token prediction has emerged as a powerful pretraining objective across language, vision, and speech, offering the potential to unify these diverse modalities through a single pre-training task. However, its application for general audio understanding remains underexplored, with BEATs being the only notable example. BEATs has seen limited modifications due to the absence of open-source pre-training code. Furthermore, BEATs was trained only on AudioSet, restricting its broader downstream applicability. To address these gaps, we present OpenBEATs, an open-source framework that extends BEATs via multidomain audio pre-training. We conduct comprehensive evaluations across six types of tasks, twenty five datasets, and three audio domains, including audio reasoning tasks such as audio question answering, entailment, and captioning. OpenBEATs achieves state-of-the-art performance on six bioacoustics datasets, two environmental sound datasets and five reasoning datasets, performing better than models exceeding a billion parameters at one-fourth their parameter size. These results demonstrate the effectiveness of multi-domain datasets and masked token prediction task to learn general-purpose audio representations. To promote further research and reproducibility, we release all pre-training and evaluation code, pretrained and fine-tuned checkpoints, and training logs 1.  \n1. INTRODUCTION  \nSelf-supervised learning (SSL) has shown significant promise across a wide range of audio processing tasks. It allows models to learn generalpurpose representations that transfer effectively to various downstream applications. Notable examples of SSL-based audio encoders (AEs) include BEATs [1], SS-AST [2], MAE-AST [3], and Audio-MAE [4] . Among these, BEATs stands out due to the wide adoption of its pretrained checkpoints, strong performance across DCASE challenges [5]–[7] and the potential to unify vision, language and audio encoders through masked token prediction based learning. However, the pretraining pipeline for BEATs remains closed-source, limiting such broader research impact. In addition, BEATs uses a multistage pretraining setup, combining teacher-student distillation and masked audio modeling, making the pre-training more complex than other AEs. In the speech community, open-sourcing models and their training pipelines has been a crucial practice [8]–[11] that leads to enhanced reproducibility and accessibility. Motivated by these gaps, this work aims to completely open-source the BEATs pre-training using the ESPnet toolkit [12], [13] to foster further advancements.  \nDespite advances in SSL-based audio modeling, the training and evaluation of audio encoders remain fragmented across distinct domains – environmental sound, bioacoustics and music. As a result, state-of-the-art (SOTA) performance is typically achieved by domainspecific models. For instance, BEATs excels in diverse environmental sound benchmarks, GPM-BT [14] leads in bioacoustics tasks, and MERT [15] sets the bar for music-related tasks. In contrast, the speech community has demonstrated that multitask and multilingual training [16]–[18] produces more generalizable and transferable representations. Similar trends are evident in Audio-Language Models  \n1 [https://github.com/Shikhar-S/OpenBEATs](https://github.com/Shikhar-S/OpenBEATs)  \n(ALMs), which unify diverse audio tasks under a shared sequence-tosequence framework. The performance of ALMs [19]–[21] is often strongly correlated with the quality of the underlying audio encoder [22], underscoring the need ","cbCaipWnrAtYbYyo","https://ap.wps.com/l/cbCaipWnrAtYbYyo","pdf",361044,5,1,6,"English","en",105,"# Introduction\n## Self-supervised learning for audio encoders\n## Motivation for open-sourcing BEATs and cross-domain evaluation\n## Goals and contributions overview","[{\"question\":\"What problem does OpenBEATs address compared with BEATs?\",\"answer\":\"BEATs has closed-source pretraining code and is limited to AudioSet training, restricting broader research and downstream applicability. OpenBEATs provides an open framework with multidomain audio pretraining to improve general-purpose audio understanding.\"},{\"question\":\"How is OpenBEATs evaluated in the study?\",\"answer\":\"The work evaluates OpenBEATs across six task types, twenty-five datasets, and three audio domains. The benchmark includes reasoning-oriented tasks such as audio question answering, entailment, and captioning.\"},{\"question\":\"What resources does the paper release to support reproducibility?\",\"answer\":\"The release includes all pre-training and evaluation code, pretrained and fine-tuned checkpoints, and training logs. It also uses the ESPnet toolkit to support an accessible training and testing workflow.\"}]",1784210759,15,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"openbeats-a-fully-open-source-general-purpose-audio-encoder","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/openbeats-a-fully-open-source-general-purpose-audio-encoder/86352/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-26","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What problem does OpenBEATs address compared with BEATs?","Question",{"text":76,"@type":77},"BEATs has closed-source pretraining code and is limited to AudioSet training, restricting broader research and downstream applicability. OpenBEATs provides an open framework with multidomain audio pretraining to improve general-purpose audio understanding.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How is OpenBEATs evaluated in the study?",{"text":81,"@type":77},"The work evaluates OpenBEATs across six task types, twenty-five datasets, and three audio domains. The benchmark includes reasoning-oriented tasks such as audio question answering, entailment, and captioning.",{"name":83,"@type":74,"acceptedAnswer":84},"What resources does the paper release to support reproducibility?",{"text":85,"@type":77},"The release includes all pre-training and evaluation code, pretrained and fine-tuned checkpoints, and training logs. It also uses the ESPnet toolkit to support an accessible training and testing workflow.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,110,114,119,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},"Comic",60,"comic",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":111,"show_sort_weight":112,"slug":113},"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":20,"slug":137},19,"General","general"]