[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82486-en":3,"doc-seo-82486-105":29,"detail-sidebar-cat-0-en-105":90},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},82486,1099513958762,"Logic","https://ap-avatar.wpscdn.com/avatar/1000023916a998db790?x-image-process=image/resize,m_fixed,w_180,h_180&k=1784791008015729253",8,"Research & Report","Adaptive Perturbation Selection for Contrastive Audio Decoding","Large audio-language models (LALMs) often hallucinate by letting linguistic priors overpower acoustic evidence during decoding. Contrastive decoding mitigates this without training, but existing audio negative branches use coarse perturbations such as masking or additive noise, leaving structured transformations underexplored. This work evaluates a targeted perturbation library and adaptively selects the best negative branch per example. Binary yes/no prompt constraints reduce false confirmations, while results show strong task dependence across temporal, spectral, frequency, and amplitude domains; an adaptive selector yields further gains on existence accuracy.","Adaptive Perturbation Selection for Contrastive  \nAudio Decoding  \nAaron Isidore Grace (Wang)  \nDepartment of Computer Science University of Iowa Iowa City, IA, USA [aaron-wang@uiowa.edu](aaron-wang@uiowa.edu)  \nZhouyuan Huo  \nGoogle Mountain View, CA, USA [huozhouyuan@gmail.com](huozhouyuan@gmail.com)  \nWeiran Wang  \nDepartment of Computer Science University of Iowa Iowa City, IA, USA [werian-wang@uiowa.edu](werian-wang@uiowa.edu)  \narXiv :2607 .00247v1 [ cs . SD] 30 Jun 2026  \nAbstract—Large audio-language models (LALMs) frequently hallucinate by overriding acoustic evidence with language priors. While contrastive decoding (CD) offers training-free mitigation, existing methods rely on blunt perturbations like masking or noise, leaving structured audio transformations unexplored. We explore this design space by evaluating a diverse library of targeted audio perturbations and adaptively selecting the optimal negative branch for each task and example. First, we improve upon earlier prompt engineering by showing that a simple binary yes/no constraint reduces the model’s tendency to falsely confirm absent audio features. Second, evaluating our library across temporal, spectral, frequency, and amplitude domains reveals that optimal transformations are highly task-dependent; for instance, reversing the audio array disrupts temporal coherence, raising accuracy on the temporal order task from 74.7% to 81.4%. Finally, we trained a light-weight perturbation selector on model hidden states to dynamically route negative branches, yielding an additional +4 .3% gain on the existence task.  \nIndex Terms—large audio-language models, audio hallucination, contrastive decoding, adaptive perturbation  \nI. INTRODUCTION  \nLarge audio-language models (LALMs) [1]–[6] score well on benchmarks but routinely hallucinate [7]–[11] . These errors typically emerge during decoding when strong textual priors dominate the output space, causing the model to generate linguistically probable text that overrides the actual acoustic reality. Contrastive decoding (CD) offers a training-free remedy [12] . By amplifying the log-probability difference between an expert model (normal input branch) and a weaker, perturbed negative branch, CD penalizes tokens driven purely by language priors. However, current applications rely on a limited set of negative branches—such as total audio masking or basic additive noise [13]—leaving the broader design space of audio perturbations largely unexplored.  \nIn this work, we perform a systematic exploration of this uncharted design space by evaluating a diverse library of targeted audio perturbations for CD. Instead of using generic noise, we design these perturbations to target specific model failure modes. Since diverse audio tasks depend on varying acoustic characteristics, the design of a negative branch must be task-specific. It is crucial to generate meaningful contrastive examples while avoiding severe acoustic distortions that could inadvertently induce further hallucinations. Matching  \nthe perturbation to the task reveals a systematic way to suppress language priors across varied audio settings.  \nWe evaluate our method across two LALMs: Qwen2-Audio- 7B-Instruct [14] and Audio Flamingo 3 [15](henceforth AF3) . Testing spans four tasks: Clotho-AQA [16], along with the existence, temporal order, and object attribute variants of the Audio Hallucination dataset [7](henceforth AH Existence, AH Order, and AH Attribute) . Our contributions are:  \n• Prompt calibration. Constraining model outputs to a single yes/no token calibrates the model’s innate affirmative bias, raising AH Existence accuracy by +11% before any CD—more than four times the prompt engineering gain of prior work on the same model.  \n• Perturbation library. We introduce an extensive perturbation library for audio CD, comprising 105 perturbations across 38 types covering temporal, spectral, frequency, and amplitude transformations. We show that optimal perturbation","cbCaiiqJgBxOuCDv","https://ap.wps.com/l/cbCaiiqJgBxOuCDv","pdf",768299,1,7,"English","en",105,"# Introduction\n# Related Work\n# Method and System Overview\n# Experiments and Results","[{\"question\":\"Why do large audio-language models hallucinate during decoding?\",\"answer\":\"Hallucinations occur when strong textual priors dominate token generation, overriding acoustic evidence and producing linguistically probable outputs unrelated to the input audio.\"},{\"question\":\"What does contrastive decoding do to reduce audio hallucinations?\",\"answer\":\"Contrastive decoding amplifies the log-probability gap between an expert prediction on the normal branch and a perturbed negative branch, penalizing tokens driven mainly by language priors.\"},{\"question\":\"How does adaptive perturbation selection improve results compared with fixed perturbations?\",\"answer\":\"Because the best negative transformation varies by task and by individual examples, a lightweight selector routes each input to its optimal negative branch using model hidden states, providing additional gains over the best fixed branch.\"}]",1784180858,18,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":85,"head_meta":87,"extra_data":89,"updated_unix":27},"adaptive-perturbation-selection-for-contrastive-audio-decoding","",{"@graph":35,"@context":84},[36,53,67],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/adaptive-perturbation-selection-for-contrastive-audio-decoding/82486/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":61,"encodingFormat":60,"isAccessibleForFree":62,"interactionStatistic":63},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-16",true,{"@type":64,"interactionType":65,"userInteractionCount":4},"InteractionCounter",{"@type":66},"ViewAction",{"@type":68,"mainEntity":69},"FAQPage",[70,76,80],{"name":71,"@type":72,"acceptedAnswer":73},"Why do large audio-language models hallucinate during decoding?","Question",{"text":74,"@type":75},"Hallucinations occur when strong textual priors dominate token generation, overriding acoustic evidence and producing linguistically probable outputs unrelated to the input audio.","Answer",{"name":77,"@type":72,"acceptedAnswer":78},"What does contrastive decoding do to reduce audio hallucinations?",{"text":79,"@type":75},"Contrastive decoding amplifies the log-probability gap between an expert prediction on the normal branch and a perturbed negative branch, penalizing tokens driven mainly by language priors.",{"name":81,"@type":72,"acceptedAnswer":82},"How does adaptive perturbation selection improve results compared with fixed perturbations?",{"text":83,"@type":75},"Because the best negative transformation varies by task and by individual examples, a lightweight selector routes each input to its optimal negative branch using model hidden states, providing additional gains over the best fixed branch.","https://schema.org",{"og:url":51,"og:type":86,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":88,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":91},[92,96,100,104,109,114,118,121,126,129,133],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":93,"show_sort_weight":94,"slug":95},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":97,"show_sort_weight":98,"slug":99},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":101,"show_sort_weight":102,"slug":103},"Exam",70,"exam",{"id":105,"doc_module":4,"doc_module_name":45,"category_name":106,"show_sort_weight":107,"slug":108},5,"Comic",60,"comic",{"id":110,"doc_module":4,"doc_module_name":45,"category_name":111,"show_sort_weight":112,"slug":113},6,"Technology",50,"technology",{"id":21,"doc_module":4,"doc_module_name":45,"category_name":115,"show_sort_weight":116,"slug":117},"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":119,"slug":120},30,"research-report",{"id":122,"doc_module":4,"doc_module_name":45,"category_name":123,"show_sort_weight":124,"slug":125},9,"Religion & Spirituality",20,"religion-spirituality",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":127,"show_sort_weight":124,"slug":128},"World Cup","world-cup",{"id":130,"doc_module":4,"doc_module_name":45,"category_name":131,"show_sort_weight":130,"slug":132},10,"Lifestyle","lifestyle",{"id":134,"doc_module":4,"doc_module_name":45,"category_name":135,"show_sort_weight":105,"slug":136},19,"General","general"]