[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84756-en":3,"doc-seo-84756-105":29,"detail-sidebar-cat-0-en-105":82},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":11,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},84756,4398048950312,"Violet","https://ap-avatar.wpscdn.com/avatar/400002538284de19e3c?_k=1778320343897328908",8,"Research & Report","CARD Cross-component Audio Representation Distillation for Encoder-free Audio Captioning","Modern audio captioning systems rely on a frozen audio encoder plus a trainable projector to bridge audio features into a large language model (LLM), but this keeps encoder inference in the loop and constrains acoustic understanding to fixed encoder representations. CARD introduces encoder-free audio captioning by removing the audio encoder at inference. A 13.2M projector feeds a frozen LLM with merged LoRA adapters. Training distills knowledge from CLAP-HTSAT, distributing teacher representations across projector perceptual stages and LLM semantic stages. This placement yields +12.18 CIDEr-D on AudioCaps and +5.21 on Clotho while eliminating encoder cost.","CARD: Cross-component Audio Representation Distillation for Encoder-Free Audio Captioning  \nGanesh Pavan Kartikeya Bharadwaj Kolluri 1 , Yuchen Zhang 1,2 , Michael Kampouridis 1 , Ravi Shekhar 1,2  \n1 School of Computer Science and Electronic Engineering, University of Essex  \n2 Institute for Analytics and Data Science, University of Essex  \n{karthik.kolluri, yuchen.zhang, mkampo, [r.shekhar](r.shekhar}@essex.ac.uk)[}](r.shekhar}@essex.ac.uk)[@essex.ac.uk](r.shekhar}@essex.ac.uk)  \narXiv :2607 .046 19v 1 [ cs . SD] 6 Jul 2026  \nAbstract—Modern automated audio captioning systems pair a frozen audio encoder with a large language model (LLM) via a trainable projector, incurring the encoder’s inference cost and bottlenecking the model through its fixed acoustic features. We present CARD, an encoder-free audio captioning model that removes the encoder at inference: a 13.2M projector feeds a frozen LLM with merged LoRA adapters, while the teacher used to train it is discarded. CARD distills a pretrained audio teacher (CLAP-HTSAT) into the model, but rather than injecting it into the LLM alone, it routes the teacher’s representations across components: perceptual stages to the projector and semantic stages to the LLM. This placement improves CIDEr-D by +12.18 over an LLM-only distilled model on AudioCaps and by +5.21 on Clotho, reaching 55.4 against a 66.4 encoder-kept upper bound with no encoder at inference, showing that where a teacher’s knowledge is placed matters as much as its presence.  \nIndex Terms—audio captioning, encoder-free models, crosscomponent distillation, audio-language models.  \nI. INTRODUCTION  \nAutomated audio captioning (AAC) is the task of generating a natural-language description of the sound events in an audio clip, such as “rain falls as thunder rumbles in the distance.”Unlike automatic speech recognition (ASR), which transcribes spoken words from the speech audio, AAC aims to describe the broader acoustic scene, capturing non-speech sounds and their relationships in free-form natural language. This ability to summarize what an audio clip contains has made AAC valuable for applications such as accessibility, audio retrieval, and content understanding [1], [2] .  \nMost recent AAC systems follow a common architectural design: A frozen pre-trained audio encoder, such as CLAP [3], first converts the raw audio clip into a sequence of acoustic feature vectors. A trainable projector then maps these features into the input space of the Large Language Model (LLM), which decodes them into a caption. Since both the audio encoder and the LLM are pre-trained separately and usually kept frozen, only the projector is trained to bridge the two components together. This design has been widely adopted because it inherits the strengths of a powerful audio encoderand a capable language model while learning little more than the alignment between them [4]–[6] . Its main limitation, however, is that inference still depends on the dedicated audio encoder. Every clip must pass through a full encoder forward pass before the language model can act, which adds computation and memory, and confines the language model to perceiving audio only through the encoder’s fixed features.  \nA natural way to remove this dependency is to discard the audio encoder from the pipeline and let the model caption audio on its own. However, without the encoder, the projector and the language model must learn to perceive audio directly from the raw acoustic signal, rather than building on representations from a model already trained on large amounts of audio. Learning to extract useful acoustic features from scratch, while at the same time learning to generate captions, places a much heavier burden on training and tends to be unstable. For this reason, encoder-free models have generally struggled to reach the performance of their encoder-based counterparts [7], [8] .  \nA promising way to ease this difficulty is knowledge distillation. Instead of learning to ","cbCaim5iYUHsqkLu","https://ap.wps.com/l/cbCaim5iYUHsqkLu","pdf",3098325,3,1,"English","en",105,"# Introduction\n## Automated Audio Captioning and Existing Encoder-based Designs\n## Encoder-free Challenges\n## Knowledge Distillation and the Need for Cross-component Placement\n## CARD Approach and Contributions","[{\"question\":\"What changes occur at inference time in CARD?\",\"answer\":\"The teacher and distillation heads are removed, and the LoRA adapters are merged into the base LLM. The deployed model contains only the audio projector and the language model, with no audio encoder used at inference.\"}]",1784198065,20,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":77,"head_meta":79,"extra_data":81,"updated_unix":27},"card-cross-component-audio-representation-distillation-for-encoder-free-audio-captioning","",{"@graph":35,"@context":76},[36,52,67],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,49],{"item":40,"name":41,"@type":42,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":20},"https://docshare.wps.com/document/research-report/",{"item":50,"name":13,"@type":42,"position":51},"https://docshare.wps.com/document/card-cross-component-audio-representation-distillation-for-encoder-free-audio-captioning/84756/",4,{"url":50,"name":13,"@type":53,"author":54,"headline":13,"publisher":56,"fileFormat":59,"inLanguage":23,"description":14,"dateModified":60,"datePublished":61,"encodingFormat":59,"isAccessibleForFree":62,"interactionStatistic":63},"DigitalDocument",{"name":9,"@type":55},"Person",{"url":40,"name":57,"@type":58},"DocShare","Organization","application/pdf","2026-07-23","2026-07-16",true,{"@type":64,"interactionType":65,"userInteractionCount":20},"InteractionCounter",{"@type":66},"ViewAction",{"@type":68,"mainEntity":69},"FAQPage",[70],{"name":71,"@type":72,"acceptedAnswer":73},"What changes occur at inference time in CARD?","Question",{"text":74,"@type":75},"The teacher and distillation heads are removed, and the LoRA adapters are merged into the base LLM. The deployed model contains only the audio projector and the language model, with no audio encoder used at inference.","Answer","https://schema.org",{"og:url":50,"og:type":78,"og:title":13,"og:site_name":57,"og:description":14},"article",{"robots":80,"canonical":50},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":83},[84,88,92,96,101,106,111,114,118,121,125],{"id":21,"doc_module":4,"doc_module_name":45,"category_name":85,"show_sort_weight":86,"slug":87},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":89,"show_sort_weight":90,"slug":91},"Literature",80,"literature",{"id":51,"doc_module":4,"doc_module_name":45,"category_name":93,"show_sort_weight":94,"slug":95},"Exam",70,"exam",{"id":97,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},5,"Comic",60,"comic",{"id":102,"doc_module":4,"doc_module_name":45,"category_name":103,"show_sort_weight":104,"slug":105},6,"Technology",50,"technology",{"id":107,"doc_module":4,"doc_module_name":45,"category_name":108,"show_sort_weight":109,"slug":110},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":112,"slug":113},30,"research-report",{"id":115,"doc_module":4,"doc_module_name":45,"category_name":116,"show_sort_weight":28,"slug":117},9,"Religion & Spirituality","religion-spirituality",{"id":28,"doc_module":4,"doc_module_name":45,"category_name":119,"show_sort_weight":28,"slug":120},"World Cup","world-cup",{"id":122,"doc_module":4,"doc_module_name":45,"category_name":123,"show_sort_weight":122,"slug":124},10,"Lifestyle","lifestyle",{"id":126,"doc_module":4,"doc_module_name":45,"category_name":127,"show_sort_weight":97,"slug":128},19,"General","general"]