[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86145-en":3,"doc-seo-86145-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86145,962075114765,"Quinn","https://ap-avatar.wpscdn.com/davatar_a8503ba1806abce46bf441b54a3ca4cd",8,"Research & Report","Simple Features and Honest Calibration for Ambivalence and Hesitancy Recognition in Video","Ambivalence and hesitancy (A/H) recognition on short interview videos is targeted for the ABAW 2026 BAH Challenge, predicting whether a person shows A/H signs. The approach fuses affect-specialised text, audio, and visual representations with interpretable hesitation cues via a reliability gate called Affective Marker Fusion (AMF), then applies an AP-weighted ensemble at a fixed decision threshold. ASR-erased time exploits ASR chunk timestamps to recover a strong 16-dimensional non-verbal signal. Experiments show cross-modal conflict design gives limited benefit, language dominates, and calibration outweighs architecture, with thresholded AP-weighting achieving stronger test macro-F1.","Simple Features and Honest Calibration for Ambivalence and Hesitancy Recognition in Video  \nVikas Kumar  \nIndian Institute of Science Education and Research Bhopal, India [vikas25@iiserb.ac.in](vikas25@iiserb.ac.in)  \nAditya Mishra  \nIndian Institute of Science Education and Research Bhopal, India [aditya21@iiserb.ac.in](aditya21@iiserb.ac.in)  \nHaroon R. Lone  \nIndian Institute of Science Education and Research Bhopal, India [haroon@iiserb.ac.in](haroon@iiserb.ac.in)  \narXiv :2607 . 11120v1 [ cs .CV] 13 Jul 2026  \nAbstract  \nWe address ambivalence and hesitancy (A/H) recognition in the ABAW 2026 BAH Challenge: given a short interview video, predict whether the person shows signs of A/H. Our system combines affect-specialised text, audio, and visual representations with a small set of readable linguistic hesitation cues, fused by a reliability gate we call Affective Marker Fusion (AMF), and finished with a simple AP-weighted ensemble at a fixed decision threshold. We also introduce ASR-erased time: speech recognisers delete fillers and hesitation pauses from the transcript, but the chunk timestamps keep the time those events took, and sixteen features built from these gaps form the strongest and most independent non-verbal channel we measured (AP 0.718, correlation 0.11–0.36 with all other members) . Across controlled experiments we find three things: cross-modal conflict design does not reliably help on BAH; language is by far the strongest channel while affect-specialised audio is a useful second; and calibration matters more than architecture. Fitting ensemble weights and a threshold on the small validation split overfits: it scores 0.741 macro-F1 on validation but only 0.690 on the untouched test set. AP-weighting at a fixed threshold instead reaches 0.731 on test. The full pipeline is deterministic and runs from one script.1  \n1 Introduction  \nAmbivalence and hesitancy (A/H) are states of internal conflict: a person is unsure, or pulled two ways at once. Detecting these moments in short answer videos is useful for behaviourchange tools, where a moment of hesitation is a good time to intervene. The ABAW 2026 challenge frames this as a binary, video-level task on the BAH dataset [1], scored by macro-F1 (the unweighted mean of the present-and absent-class F1) with average precision (AP) as a secondary measure.  \nA/H is widely treated as a cross-modal conflict problem, and the strongest prior systems build hand-designed conflict features on top of three pre-trained encoders [2, 3] . We reproduce that recipe, test its central claim, and then ask a simpler question: where is the signal actually? Three results organise the paper.  \n1 Code: [https://github.com/wmivikas/](https://github.com/wmivikas/)[ ](https://github.com/wmivikas/)ABAW2026-Task2-ECCV.  \n1. Conflict design is not the lever. With encoders and training fixed, swapping the conflict operator (absolute difference, orthogonal split, or none) moves public-test AP by at most 0.05 and does not order the same way on validation and test (Sec. 4.5) .  \n2. Features help, but narrowly. With general encoders the video and audio channels are near chance. An emotionspecialised audio encoder is a real gain; no facial representation we tried rescues the video channel (Sec. 4.4) . Calibration is the biggest gain. On 778 training videos, fitting ensemble weights and the decision threshold on the 124-video validation split scores 0.741 there but only  \n0.690 on test—the classic overfitting signature. Removing that search and using AP weights at a fixed threshold lifts public-test macro-F1 to 0.731 (Sec. 4.6) .  \nContributions.  \n• AMF, a compact reliability-gated fusion of affectspecialised streams and eleven interpretable hesitation markers.  \n• ASR-erased time, a new 16-dimensional signal recovered from ASR chunk timestamps; it is the strongest nonverbal channel on BAH and nearly uncorrelated with every model member.  \n• A controlled study showing that conflict design does not help,","cbCaipghvCh6KWOu","https://ap.wps.com/l/cbCaipghvCh6KWOu","pdf",1385906,3,1,6,"English","en",105,"# Abstract\n# Introduction\n# Related Work","[{\"question\":\"What task does the paper address in the ABAW 2026 BAH Challenge?\",\"answer\":\"It addresses ambivalence and hesitancy recognition as a binary, video-level prediction task from short interview videos, producing whether the person exhibits A/H signs.\"},{\"question\":\"How does the proposed system fuse information from different modalities?\",\"answer\":\"It combines affect-specialised text, audio, and visual features with interpretable linguistic hesitation cues using a reliability-gated fusion module named Affective Marker Fusion (AMF), followed by an AP-weighted ensemble at a fixed decision threshold.\"},{\"question\":\"What is the main finding about improvements: model architecture or calibration?\",\"answer\":\"Calibration matters more than architecture: searching ensemble weights and thresholds on the small validation split overfits, while using AP-weighting at a fixed threshold yields better test macro-F1.\"}]",1784208896,15,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"simple-features-and-honest-calibration-for-ambivalence-and-hesitancy-recognition-in-video","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/simple-features-and-honest-calibration-for-ambivalence-and-hesitancy-recognition-in-video/86145/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What task does the paper address in the ABAW 2026 BAH Challenge?","Question",{"text":75,"@type":76},"It addresses ambivalence and hesitancy recognition as a binary, video-level prediction task from short interview videos, producing whether the person exhibits A/H signs.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the proposed system fuse information from different modalities?",{"text":80,"@type":76},"It combines affect-specialised text, audio, and visual features with interpretable linguistic hesitation cues using a reliability-gated fusion module named Affective Marker Fusion (AMF), followed by an AP-weighted ensemble at a fixed decision threshold.",{"name":82,"@type":73,"acceptedAnswer":83},"What is the main finding about improvements: model architecture or calibration?",{"text":84,"@type":76},"Calibration matters more than architecture: searching ensemble weights and thresholds on the small validation split overfits, while using AP-weighting at a fixed threshold yields better test macro-F1.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,114,119,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":111,"show_sort_weight":112,"slug":113},"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]