[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84431-en":3,"doc-seo-84431-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},84431,1099513958607,"Jiven","https://ap-avatar.wpscdn.com/avatar/100002390cf8733938c?x-image-process=image/resize,m_fixed,w_180,h_180&k=1778829742770036399",8,"Research & Report","VIBE Annotation-Free Video-to-Text Information Bottleneck Evaluation for TL;DR","Many decision-making workflows still require human supervision when accuracy and efficiency both matter, such as reviewing long dashcam footage or screening research videos. Existing vision-language model (VLM) evaluation relies on costly human-annotated references and often ignores downstream task utility, while generating verbose summaries can slow users. VIBE introduces an annotation-free selection method that ranks VLM outputs using grounding and utility scores derived from an information bottleneck principle. User studies show up to 61.23% accuracy gains and 75.77% faster responses.","arXiv :2505 . 17423v4 [ cs .CV] 13 Jul 2026  \nVIBE: Annotation-Free Video-to-Text Information Bottleneck Evaluation for TL;DR  \nShenghui Chen Po-han Li Sandeep Chinchali, Ufuk Topcu The University of Texas at Austin  \n{shenghui.chen, pohanli, sandeepc, [utopcu}@utexas.edu](utopcu}@utexas.edu)  \nAbstract  \nMany decision-making tasks, where both accuracy and efficiency matter, still require human supervision. For example, tasks like traffic officers reviewing hour-long dashcam footage or researchers screening conference videos can benefit from concise summaries that reduce cognitive load and save time. Yet current vision-language models (VLMs) often produce verbose, redundant outputs that hinder task performance. Existing video caption evaluation depends on costly human annotations and overlooks the summaries’ utility in downstream tasks. We address these gaps with Video-to-text Information Bottleneck Evaluation (VIBE), an annotation-free method that scores VLM outputs using two metrics: grounding (how well the summary aligns with visual content) and utility (how informative it is for the task) . VIBE selects from randomly sampled VLM outputs by ranking them according to the two scores to support effective human decision-making. Human studies on LearningPaper24, SUTD-TrafficQA, and LongVideoBench show that summaries selected by VIBE consistently improve performance—boosting task accuracy by up to 61.23% and reducing response time by 75.77% compared to naive VLM summaries or raw video. 2  \n1 Introduction  \nEfficiently extracting relevant information from extensive video is a major bottleneck for human decision-making, where both accuracy and efficiency matter [1–4] . Tasks demanding human supervision, such as a traffic officer analyzing hours of dashcam footage to determine fault or a researcher distilling key insights from a lengthy oral presentation, are often limited by the time and cognitive load required to process raw video streams. In this work, we aim to improve the quality and brevity of video summaries to boost human task performance compared to existing vision-language model (VLM) outputs and raw video, especially for longer clips where summarization offers greater utility.  \nExisting video caption evaluation metrics, however, rely heavily on reference-based comparisons to human-annotated summaries [5–8] . These metrics face two main issues. First, these works require human annotators to watch video clips and write gold-standard captions, which contradicts the goal of reducing human response time and limits generalization to unseen video clips. Second, they are oblivious to downstream tasks and fail to measure how well captions support the tasks.  \nWe propose Video-to-text Information Bottleneck Evaluation (VIBE), an annotation-free method for selecting task-relevant video summaries without model retraining. As shown in Figure 1, VIBE defines two metrics—grounding and utility scores—based on the information bottleneck principle [9] . It uses pointwise mutual information to quantify how well a summary reflects video evidence and supports the downstream task. We leverage next-token prediction in VLMs to access the probability  \n∗Equal contribution (Order determined by coin toss) .  \n2Project Website, Code, and LearningPaper24 Dataset.  \n39th Conference on Neural Information Processing Systems (NeurIPS 2025) .  \nFigure 1: VIBE for Video-to-Text Summary Selection. Given a video, a task, and VLM-generated summaries, VIBE ranks the summaries using the proposed grounding and utility scores, which assess video alignment and task relevance. It selects the summary most conducive to helping human users achieve higher task accuracy and lower response time compared to watching the full video.  \nof generating summaries or task answers. By comparing these probabilities with and without key information, we measure how well one modality (text or video) compensates for missing information in the other to assess grounding and task releva","cbCaigOrW3Ug3pKF","https://ap.wps.com/l/cbCaigOrW3Ug3pKF","pdf",1689769,1,20,"English","en",105,"# Abstract\n# Introduction\n## Motivation and limitations of existing evaluation\n## Proposed method: VIBE\n## Evaluation setup and results\n## Contributions\n## Critique and open problems","[{\"question\":\"What problem does VIBE address in video-to-text summarization evaluation?\",\"answer\":\"VIBE addresses the need for fast, accurate decision support when existing evaluation methods rely on costly human annotations and fail to measure usefulness for downstream tasks.\"},{\"question\":\"How does VIBE score and select VLM-generated summaries?\",\"answer\":\"VIBE computes a grounding score and a utility score, derived from an information bottleneck perspective using next-token prediction probabilities, then ranks randomly sampled VLM outputs accordingly.\"},{\"question\":\"What do the user studies on multiple datasets show about VIBE-selected summaries?\",\"answer\":\"VIBE-selected summaries improve human task accuracy by up to 61.23% and reduce response time by up to 75.77% compared with naive VLM summaries or raw video.\"}]",1784195593,50,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"vibe-annotation-free-video-to-text-information-bottleneck-evaluation-for-tldr","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/vibe-annotation-free-video-to-text-information-bottleneck-evaluation-for-tldr/84431/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does VIBE address in video-to-text summarization evaluation?","Question",{"text":75,"@type":76},"VIBE addresses the need for fast, accurate decision support when existing evaluation methods rely on costly human annotations and fail to measure usefulness for downstream tasks.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does VIBE score and select VLM-generated summaries?",{"text":80,"@type":76},"VIBE computes a grounding score and a utility score, derived from an information bottleneck perspective using next-token prediction probabilities, then ranks randomly sampled VLM outputs accordingly.",{"name":82,"@type":73,"acceptedAnswer":83},"What do the user studies on multiple datasets show about VIBE-selected summaries?",{"text":84,"@type":76},"VIBE-selected summaries improve human task accuracy by up to 61.23% and reduce response time by up to 75.77% compared with naive VLM summaries or raw video.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,114,119,122,126,129,133],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":28,"slug":113},6,"Technology","technology",{"id":115,"doc_module":4,"doc_module_name":45,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":45,"category_name":124,"show_sort_weight":21,"slug":125},9,"Religion & Spirituality","religion-spirituality",{"id":21,"doc_module":4,"doc_module_name":45,"category_name":127,"show_sort_weight":21,"slug":128},"World Cup","world-cup",{"id":130,"doc_module":4,"doc_module_name":45,"category_name":131,"show_sort_weight":130,"slug":132},10,"Lifestyle","lifestyle",{"id":134,"doc_module":4,"doc_module_name":45,"category_name":135,"show_sort_weight":106,"slug":136},19,"General","general"]