[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86538-en":3,"doc-seo-86538-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86538,962075114765,"Quinn","https://ap-avatar.wpscdn.com/davatar_a8503ba1806abce46bf441b54a3ca4cd",8,"Research & Report","CHARM Charge Calibration and Acoustic Rescue for LLM-based Multimodal Sarcasm Detection","Sarcasm detection resolves conflicts between literal wording and intended pragmatic meaning, yet zero-shot instruction-tuned LLMs often over-predict the sarcastic (positive) class and underuse prosodic cues, with inconsistent transfer across languages. CHARM introduces a training-free multimodal framework combining BiCAL bidirectional charge calibration and ALFR acoustic late-fusion rescue. Bidirectional prompt charge symmetry cancels bias while recovering an unbiased pragmatic signal. ALFR integrates calibrated votes with prosodic descriptors and auditory-perception probes, suppressing saturated text evidence in favor of acoustic cues. Experiments report improved zero-shot Macro-F1 on MUStARD and stronger gains on CMMA, supported by statistical significance and cross-cultural prosodic decoupling findings.","CHARM: Charge Calibration and Acoustic Rescue for LLM-based Multimodal Sarcasm Detection  \nQiyang Sun Student Member, IEEE, Yi Chang*, Yupei Li, Xi Shao Member, IEEE, Zixing Zhang* Senior  \nMember, IEEE, and Bjrn W. Schuller Fellow, IEEE  \narXiv :2607 . 1 1 102v 1 [ cs . SD] 13 Jul 2026  \nAbstract—Sarcasm detection, the identification of discrepancies between literal and intended meaning, is a fundamental task in affective computing. However, zero-shot instruction-tuned Large Language Models (LLMs) systematically over-predict the positive (sarcastic) class across the entire capability spectrum, while the prosodic cues humans rely on remain underexploited and transfer unevenly across languages. We introduce CHARM (Charge Calibration and Acoustic Rescue for Multimodal Sarcasm Detection), a training-free framework that couples two modules. Bidirectional Charge Calibration (BiCAL) steers the LLM toward opposing sarcastic and literal verdicts along a symmetric axis of charged prompts; the induced directional biases cancel by construction, and a simple aggregation recovers an unbiased pragmatic signal. Acoustic Late-Fusion Rescue (ALFR) then fuses the calibrated votes with prosodic descriptors and LLM-generated auditory-perception probes through a shallow classifier, actively down-weighting saturated text votes in favour of acoustic evidence. Without fine-tuning any backbone, BiCAL attains the highest reported zero-shot text-only Macro-F1 of 0.787 on MUStARD, while ALFR lifts weak backbones by up to +0.382 Macro-F1 on CMMA. A Stouffer meta-analysis confirms statistical significance on MUStARD and CMMA (Z = 13 .89 and Z = 34 .64, respectively; p \u003C 10 −43). Our analysis further uncovers a cross-cultural prosodic decoupling: low-level acoustics fail to transfer across languages, whereas high-level perceptual abstractions remain robust. Together, these components yield an explainable, cross-lingual multimodal detector.  \nIndex Terms—Multimodal Sarcasm Detection, Large Language Models (LLMs), Zero-Shot Learning, Charge Calibration, Acoustic Late Fusion, Cross-Cultural Analysis, Prosodic Decoupling, Explainability  \nI. INTRODUCTION  \nHUMANS recognise sarcasm almost instantly from a  \ndry tone or a knowing pause. However, contemporary instruction-tuned Large Language Models (LLMs) systematically err in the opposite direction. Sarcasm detection involves  \nQiyang Sun, Yi Chang, and Yupei Li are with GLAM – the Group on Language, Audio, & Music, Imperial College London, UK. e-mail:  \n[q.sun23@imperial.ac.uk](q.sun23@imperial.ac.uk); [yichang312@gmail.com](yichang312@gmail.com); [yupei.li22@imperial.ac.uk](yupei.li22@imperial.ac.uk)  \nXi Shao is with the College of Telecommunications and Information Engineering, Nanjing University of Posts and Telecommunications, China,  \ne-mail: [shaoxi@njupt.edu.cn](shaoxi@njupt.edu.cn)  \nZixing Zhang is with the College of Computer Science and Electronic Engineering, Hunan University, China; Zixing Zhang is also with the Shenzhen  \nResearch Institute, Hunan University, China. e-mail: [zixingzhang@hnu.edu.cn](zixingzhang@hnu.edu.cn)  \nBjrn W. Schuller is with GLAM – the Group on Language, Audio, & Music, Imperial College London, UK; CHI – Chair of Health Informatics, TUM University Hospital, Germany; relAI – the Konrad Zuse School of Excellence in Reliable AI, Germany; MDSI – Munich Data Science Institute, Germany; and MCML – Munich Center for Machine Learning, Germany.  \n[e-mail: bjoern.schuller@imperial.ac.uk](e-mail: bjoern.schuller@imperial.ac.uk)[ ](e-mail: bjoern.schuller@imperial.ac.uk)Corresponding author: Yi Chang, Zixing Zhang  \nManuscript received April 19, 2021; revised August 16, 2021 .  \nidentifying the deliberate mismatch between the literal and intended meaning of a speaker [1] . This operation serves asa cornerstone task within affective computing. This field aims to enable machines to perceive, understand, interpret, and express human emotions [2], [3] . Unlike standard sentiment analysis [","cbCaif1gDksLJ0g7","https://ap.wps.com/l/cbCaif1gDksLJ0g7","pdf",703814,4,1,18,"English","en",105,"# Introduction\n## Motivation and problem setup\n## Sycophancy bias and charge calibration intuition","[{\"question\":\"What problem does CHARM address in LLM-based sarcasm detection?\",\"answer\":\"It addresses LLMs’ tendency to over-predict the sarcastic class in zero-shot settings and the limited, uneven transfer of prosodic cues across languages.\"},{\"question\":\"How does Bidirectional Charge Calibration (BiCAL) reduce bias without fine-tuning?\",\"answer\":\"BiCAL steers the LLM using charged prompts along a symmetric axis for opposing sarcastic and literal verdicts, then aggregates the induced directional biases so they cancel while an unbiased pragmatic signal remains.\"},{\"question\":\"What role does Acoustic Late-Fusion Rescue (ALFR) play?\",\"answer\":\"ALFR fuses BiCAL-calibrated text votes with prosodic descriptors and LLM-generated auditory-perception probes using a shallow classifier, actively down-weighting saturated text evidence and emphasizing acoustic cues.\"}]",1784212484,45,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"charm-charge-calibration-and-acoustic-rescue-for-llm-based-multimodal-sarcasm-detection","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/charm-charge-calibration-and-acoustic-rescue-for-llm-based-multimodal-sarcasm-detection/86538/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-27","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does CHARM address in LLM-based sarcasm detection?","Question",{"text":75,"@type":76},"It addresses LLMs’ tendency to over-predict the sarcastic class in zero-shot settings and the limited, uneven transfer of prosodic cues across languages.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does Bidirectional Charge Calibration (BiCAL) reduce bias without fine-tuning?",{"text":80,"@type":76},"BiCAL steers the LLM using charged prompts along a symmetric axis for opposing sarcastic and literal verdicts, then aggregates the induced directional biases so they cancel while an unbiased pragmatic signal remains.",{"name":82,"@type":73,"acceptedAnswer":83},"What role does Acoustic Late-Fusion Rescue (ALFR) play?",{"text":84,"@type":76},"ALFR fuses BiCAL-calibrated text votes with prosodic descriptors and LLM-generated auditory-perception probes using a shallow classifier, actively down-weighting saturated text evidence and emphasizing acoustic cues.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]