[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86533-en":3,"doc-seo-86533-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86533,1374391974468,"Eden","https://ap-avatar.wpscdn.com/davatar_29158cc5080c5b710cf443261637dec0",8,"Research & Report","Do Video-LLMs Actually Watch? Diagnosing Character-Tracking Failures in Long-Form Video","Benchmarked Video-LLMs are evaluated on their ability to follow a named person through a full TV-length video and report, in order, how that person’s outfit changes. This paper tests whether high InfiniBench global-appearance accuracy truly requires character identity tracking. Using a nine-condition diagnostic protocol on three distinct open-source Video-LLMs and Gemini 2.5 Flash, it shows answers change only 4–31% when the queried character name is swapped, indicating models largely ignore identity. Gender-based substitution reveals coarse gender cues rather than same-person discrimination; open-ended questioning further reduces correct outputs without restoring full tracking.","Do Video-LLMs Actually Watch? Diagnosing Character-Tracking  \nFailures in Long-Form Video  \nMohammad Al-Ratrout Shayla Sharmin Aditya Raikwar Roghayeh Leila Barmaki  \narXiv :2607 . 11078v1 [ cs .CV] 13 Jul 2026  \nSame-Gender Swap Cross-Gender Swap  \nOriginal question (Sheldon)  \nQ: In what order does SHELDON change outfits in this episode?  \nCorrect answer: D = (purple t-shirt, lab coat, beige Bookman Old Style jacket Model picks: E = (purple t-shirt, beige jacket, lab coat) -> Wrong Answer  \n↓ Name-swapped question (same video, same options)  \nQ: In what order does LEONARD change outfits in this episode?  \nModel Sheldon Leonard Result  \nInterVL2-8BQwen2.5-VL-7B  \nLLAVA-NeXT-7B  \nE ~~ ~~E ~~ ~~A ~~ ~~  \nE  \nE  \nA  \nUnchanged  \nUnchanged  \nUnchanged  \nOriginal question (Penny)  \nQ: In what order does PENNY change outfits in this episode?  \nCorrect answer: C = (yellow sleeveless top, red patterned top, blue top)  \nModel picks: D = (blue top, yellow sleeveless top, red patterned top) -> Wrong Answer ↓ Swap to a different gender name (Howard) (same video, same options)  \nQ: In what order does HOWARD change outfits in this episode?  \nModel Penny Howard Result  \nInterVL2-8BQwen2.5-VL-7B  \nLLAVA-NeXT-7B  \nD ~~ ~~D ~~ ~~D~~ ~~  \nE  \nD  \nE  \nChanged Unchanged Changed  \nLetter changes, but still wrong (not C)  \nFigure 1: Name substitution diagnostic. Same-gender (left): on Sheldon→Leonard, all three models keep the same letter, they do not condition on the name. Cross-gender (right): on Penny→Howard, two of three change letter but land on a wrong answer, a coarse gender cue, not identification. Bottom: sensitivity over n=124 BBT swap pairs by gender. Takeaway: models respond to gender, not identity, same-gender swaps move the letter in only 7–17% of cases vs. 20–43% cross-gender.  \nAbstract  \nCan a Video Large Language Model (Video-LLM) follow one person through a long video, keeping track of who they are well enough to report, in order, how their outfit changes across a full TV episode? Benchmarks increasingly score this kind of task, and the strongest open-source 7–8B models now reach 37–38% on InfiniBench’s global appearance task, which asks exactly that. But does that score come from tracking the named character, or from something easier? We test this with a nine-condition diagnostic protocol applied to three architecturally distinct open-source Video-LLMs, with Gemini 2. 5 Flash as a frontier reference, and find the accuracy does not come from character tracking. When we change the character named in the question to a different cast member, leaving the video and answer options untouched, the models change their answer only 4–31% of the time, so they are largely ignoring who the question asks about. Break-  \ning that test down by the gender of the swapped name shows why: the models react more when the name is changed toa different-gender character than to a same-gender one (a 13– 28 point gap), picking up coarse gender cues but unable to tell same-gender individuals apart. This shallow processing surfaces again when we drop the multiple-choice options and ask the same questions open-endedly: open-source accuracy drops 18–25 points, with none of 151 answers fully correct, versus a 12-point drop for Gemini. Further checks rule out the obvious innocent explanations, adding subtitles, using the most informative frames, or doubling the number of frames all leave character tracking unimproved, so the bottleneck is not how much video the model sees but how it ties that video to the person the question names. We release a diagnostic toolkit for auditing what such benchmark scores actually measure: link to our toolkit.  \n1. Introduction  \nOpen-source 7–8B Video Large Language Models (VideoLLMs) score 37–38% accuracy in our evaluation on InfiniBench’s global appearance task [1], exceeding the previous generation of open-source long-form video specialists. On its face this is rapid progress. But a headline number is only as meaningful as the c","cbCaiattBfTCQpfS","https://ap.wps.com/l/cbCaiattBfTCQpfS","pdf",380938,5,1,9,"English","en",105,"# Introduction\n# Related Work","[{\"question\":\"What task do the evaluations on InfiniBench’s global appearance benchmark require?\",\"answer\":\"The benchmark asks models to follow a single named character through a 20–40 minute episode and output the chronological order of outfit changes.\"},{\"question\":\"How do the authors test whether Video-LLMs actually track the named character’s identity?\",\"answer\":\"They use a nine-condition diagnostic protocol that swaps the character name in each question while keeping the video and answer options unchanged, then measure how often the model’s answers change.\"},{\"question\":\"What do the gender-based name-swap results reveal about model behavior?\",\"answer\":\"When swapping to a different-gender character, models change answers more than with same-gender swaps, suggesting they use coarse gender cues instead of distinguishing same-gender individuals by identity.\"}]",1784212452,23,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"do-video-llms-actually-watch-diagnosing-character-tracking-failures-in-long-form-video","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/do-video-llms-actually-watch-diagnosing-character-tracking-failures-in-long-form-video/86533/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-27","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What task do the evaluations on InfiniBench’s global appearance benchmark require?","Question",{"text":76,"@type":77},"The benchmark asks models to follow a single named character through a 20–40 minute episode and output the chronological order of outfit changes.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How do the authors test whether Video-LLMs actually track the named character’s identity?",{"text":81,"@type":77},"They use a nine-condition diagnostic protocol that swaps the character name in each question while keeping the video and answer options unchanged, then measure how often the model’s answers change.",{"name":83,"@type":74,"acceptedAnswer":84},"What do the gender-based name-swap results reveal about model behavior?",{"text":85,"@type":77},"When swapping to a different-gender character, models change answers more than with same-gender swaps, suggesting they use coarse gender cues instead of distinguishing same-gender individuals by identity.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,110,115,120,123,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":20,"slug":137},19,"General","general"]