[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83387-en":3,"doc-seo-83387-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83387,13056703020460,"Valentina","https://ap-avatar.wpscdn.com/avatar/be000253dac470eee5d?_k=1778207105932848923",8,"Research & Report","VEGAS: Human-Aligned Video Caption Evaluation via Gaze","Vision-language models excel at video captioning but often fail to represent a viewer’s individual attention, producing broadly visible descriptions that may not match what a person wants to retrieve or act on. VEGAS (Video caption Evaluation via GAze Score) introduces a training-free, cross-modal, information-theoretic metric that evaluates candidate captions by test-time gaze alignment. It uses synchronized gaze and reference annotations from curated egocentric activities and instructional slides, selecting captions via rejection sampling. Experiments show improved alignment with human focus and better caption-to-video retrieval, demonstrating practical value of viewer-attention-aware inference.","arXiv :2607 .08489v 1 [ cs .CV] 9 Jul 2026  \nVEGAS: Human-Aligned Video Caption Evaluation via Gaze  \nShenghui Chen  \nThe University of Texas  \nPo-han Li  \nThe University of Texas  \nXimeng Sun  \nAMD  \nShijia Yang  \nAMD  \nEmad Barsoum  \nAMD  \nZicheng Liu  \nAMD  \nSandeep Chinchali  \nThe University of Texas  \nUfuk Topcu  \nThe University of Texas  \n[shenghui. chen@utexas. edu](shenghui. chen@utexas. edu)  \nat Austin  \n[pohanli@utexas. edu](pohanli@utexas. edu)  \nat Austin  \n[Ximeng.Sun@amd. com](Ximeng.Sun@amd. com)  \n[Shijia. Yang@amd. com](Shijia. Yang@amd. com)  \n[Emad.Barsoum@amd. com](Emad.Barsoum@amd. com)  \n[Zicheng.Liu@amd. com](Zicheng.Liu@amd. com)  \n[sandeepc@utexas. edu](sandeepc@utexas. edu)  \nat Austin  \n[utopcu@utexas. edu](utopcu@utexas. edu)  \nat Austin  \nAbstract  \nVision-language models excel at video captioning, yet typically generate descriptions that fail to capture individual viewers’ attention. We propose VEGAS (Video caption Evaluation via GAze Score), a training-free metric that leverages test-time gaze to sample personalized, attention-aligned text. It is a cross-modal, information-theoretic metric that quantifies how well a candidate caption matches a viewer’s focus. To evaluate VEGAS, we curate a dataset of egocentric activities and instructional slides paired with synchronized gaze and reference annotations. We then select captions based on VEGAS via rejection sampling without model retraining. Experiments show that VEGAS-selected captions align significantly better with human focus and improve downstream caption-to-video retrieval, demonstrating the practical utility of incorporating viewer attention during inference.  \n1 Introduction  \nIn many applications, video captions serve not merely as descriptions but as interfaces to user intent. Vision-language models (VLMs) generate semantically accurate captions, yet they are typically trained on crowd-sourced annotations that aggregate across diverse human interpretations. As a result, they often describe what is broadly visible while ignoring viewer-specific attention and subjective perception. Consider an egocentric setting: two users observing the same kitchen scene may attend to different objects or actions, but a conventional VLM may produce a single globally correct caption that is pragmatically misaligned with what either user intends to retrieve, revisit, or act upon. In caption-indexed retrieval, users often search for videos by referring to attended objects or actions rather than the full scene. When an index caption emphasizes unattended background content, the relevant video can become harder to retrieve.  \nGaze from native wearer (Project Aria glasses)  \nGaze from crowdsourced viewer (webcam eye tracking)  \nFigure 1: Overview of our framework. We curate a multimodal dataset pairing synchronized human gaze and annotated captions across egocentric video clips and instructional slides. VEGAS then leverages VLM’s conditional probabilities to evaluate the alignment between a caption and a viewer’s gaze.  \nGaze provides a measurable, though imperfect, proxy for visual attention Just & Carpenter (1976) . We use gaze at test time to select among plausible captions, favoring descriptions that better reflect the viewer’s attended referents without retraining or directly modifying the captioning model. Unlike prior work that uses human gaze only as a training-time supervision signal, we leverage gaze at test-time as a conditioning signal to generate personalized captions that reflect an individual’s specific attention patterns. We introduce the Video caption Evaluation via GAze Score (VEGAS), a cross-modal, information-theoretic metric that quantifies how well a caption reflects a viewer’s gaze patterns (Figure 1) . We then use VEGAS to select gaze-aligned captions via rejection sampling, without requiring model retraining.  \nTo enable controlled evaluation across diverse video stimuli, we curate a multimodal dataset that aligns human gaze signals with","cbCaiaKjdVpUE4NJ","https://ap.wps.com/l/cbCaiaKjdVpUE4NJ","pdf",4745199,3,1,21,"English","en",105,"# Abstract\n# Introduction\n# Related Works\n# Method and Dataset\n# Experiments and Results","[{\"question\":\"What problem does VEGAS address in video captioning?\",\"answer\":\"VEGAS targets the gap between captions generated from crowd-trained vision-language models and the way individual viewers attend to different objects or actions. Standard captions can therefore misalign with a viewer’s intent for retrieval or action.\"},{\"question\":\"How does VEGAS work without retraining the captioning model?\",\"answer\":\"VEGAS is a training-free metric that uses test-time gaze to score how well a candidate caption matches a viewer’s attention. Captions are selected using rejection sampling guided by VEGAS rather than modifying or retraining the model.\"},{\"question\":\"What data is used to evaluate VEGAS?\",\"answer\":\"The evaluation relies on a curated multimodal dataset that pairs synchronized human gaze with annotated captions. It covers egocentric activities (Aria Everyday Activities, AEA) and instructional slide decks (SlideVQA), enabling gaze-conditioned analysis.\"}]",1784187141,53,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"vegas-human-aligned-video-caption-evaluation-via-gaze","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/vegas-human-aligned-video-caption-evaluation-via-gaze/83387/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-25","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does VEGAS address in video captioning?","Question",{"text":75,"@type":76},"VEGAS targets the gap between captions generated from crowd-trained vision-language models and the way individual viewers attend to different objects or actions. Standard captions can therefore misalign with a viewer’s intent for retrieval or action.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does VEGAS work without retraining the captioning model?",{"text":80,"@type":76},"VEGAS is a training-free metric that uses test-time gaze to score how well a candidate caption matches a viewer’s attention. Captions are selected using rejection sampling guided by VEGAS rather than modifying or retraining the model.",{"name":82,"@type":73,"acceptedAnswer":83},"What data is used to evaluate VEGAS?",{"text":84,"@type":76},"The evaluation relies on a curated multimodal dataset that pairs synchronized human gaze with annotated captions. It covers egocentric activities (Aria Everyday Activities, AEA) and instructional slide decks (SlideVQA), enabling gaze-conditioned analysis.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]