[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86327-en":3,"doc-seo-86327-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86327,7971461741311,"Ophelia","https://ap-avatar.wpscdn.com/avatar/74000253aff267980c6?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779345379180704826",8,"Research & Report","StoryTeller: Training-Free Narrative Grounding for Long-Form Audio Description","Long-form audio description (AD) for blind and low-vision audiences must preserve narrative meaning beyond frame-by-frame action, including who the characters are, how events relate, and how each scene connects to earlier context. Existing video–language models often describe moments independently and therefore miss story coherence. StoryTeller proposes a training-free framework that carries forward a verified narrative memory across scenes and can retrieve public movie metadata optionally. It avoids subtitles, transcripts, character banks, or fine-tuning, and improves narrative coherence and factual grounding on benchmarks via automatic, QA-based, and human evaluations.","arXiv :2607 . 11798v1 [ cs .CV] 13 Jul 2026  \nStoryTeller: Training-Free Narrative Grounding for Long-Form Audio Description  \nSeung Hyun Hahm, Minh T. Dinh, and SouYoung Jin  \nDartmouth College, USA  \n{Seung.Hyun.Hahm.GR,Minh.T.Dinh.GR,[SouYoung.Jin}@dartmouth.edu](SouYoung.Jin}@dartmouth.edu)  \nAbstract. Long-form audio description (AD) requires more than describing visible actions: it must preserve characters, events, relationships, and story context across scenes so that blind and low-vision (BLV) audiences can follow a film. Modern video–language models (VLMs) are effective on short clips, but they often treat each moment independently, producing descriptions that miss who characters are, why events matter, and how the current scene connects to earlier narrative context. We propose StoryTeller, a training-free framework for story-aware long-form AD. Instead of relying only on local visual cues, StoryTeller maintains a verified narrative memory that carries forward story-relevant information across scenes, enabling later descriptions to remain coherent, grounded, and contextually informative. Given only raw video and a movie title, StoryTeller can optionally retrieve public movie metadata to resolve names and story context, while accepting only facts that are supported by the video through semantic filtering and VLM verification. The method requires no subtitles, scripts, AD transcripts, aligned captions, character banks, precomputed face identities, or task-specific fine-tuning. To evaluate whether generated AD preserves narrative information, we introduce StoryAD-QA 1 , a question-answering benchmark that tests whether a language model can answer story-context questions using only the generated descriptions. Experiments on standard AD benchmarks and diverse long-form videos show that StoryTeller consistently improves narrative coherence, factual grounding, and story comprehension over strong baselines in automatic, QA-based, and human evaluations.  \nKeywords: audio description · long-form video understanding · retrieval-augmented generation · narrative grounding · accessibility  \n1 Introduction  \nAudio description (AD) enables blind and low-vision (BLV) audiences to access visual media by narrating important visual information between dialogue and other sounds. Effective AD goes beyond describing objects or actions in the current frame: it identifies who is present, explains what has just happened,  \n1 StoryAD-QA benchmark dataset and evaluation code: [https://github.com/SEE-AI-Lab/ECCV2026_StoryTeller_StoryAD_QA](https://github.com/SEE-AI-Lab/ECCV2026_StoryTeller_StoryAD_QA)  \n2 S. H. Hahm et al.  \nBaseline AD  \nAlbus Dumbledore makes his wand move with a flick of his wrist while conversing with someone. Harry is hit by a teacher, while Ron and Hermione look on.  \nFig. 1: Story-level audio description requires narrative memory. Existing audio description (AD) systems typically generate descriptions using only local visual context or static character banks, which often fails to preserve long-range narrative relationships. Our StoryTeller instead maintains a persistent narrative state consisting of an identity graph that tracks characters across scenes and a salience-weighted memory that accumulates narrative facts over time. This enables consistent character references and long-range story reasoning across an entire film. The baseline example is the output of AutoAD-Zero.  \nconveys why a moment matters, and connects the scene to the broader narrative. Professional AD writers are able to provide this narrative grounding because they watch the film and write with the whole story in mind. However, producing high-quality AD remains costly and time-intensive, and copyright restrictions make large publicly distributable AD datasets difficult to obtain.  \nRecent advances in video–language models (VLMs) [2,19,30,36] and large language models have motivated research on automatic AD generation. State-of-theart identity-aware AD sy","cbCaigdeIUYwEWlx","https://ap.wps.com/l/cbCaigdeIUYwEWlx","pdf",25235044,3,1,41,"English","en",105,"# Introduction\n## Problem: story-aware long-form audio description\n## Related work and limitations\n## Training-free approach\n## Proposed StoryTeller framework\n## Evaluation and StoryAD-QA benchmark","[{\"question\":\"Why is long-form audio description more difficult than describing short video clips?\",\"answer\":\"Long-form AD must maintain narrative continuity across scenes, preserving character identity, relationships, and event context. Many models treat each moment independently, so they can miss who characters are and how current events connect to earlier story information.\"},{\"question\":\"What is StoryTeller and how does it preserve narrative context?\",\"answer\":\"StoryTeller is a training-free framework that maintains an explicit narrative state while processing clips in chronological order. It uses a verified narrative memory that carries forward story-relevant facts to keep later descriptions coherent and contextually grounded.\"},{\"question\":\"What resources does StoryTeller avoid using?\",\"answer\":\"StoryTeller does not rely on subtitles, scripts, AD transcripts, aligned captions, character banks, precomputed face identities, or task-specific fine-tuning. It can optionally retrieve public movie metadata, but only accepts facts supported by the video through semantic filtering and VLM verification.\"}]",1784210501,103,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"storyteller-training-free-narrative-grounding-for-long-form-audio-description","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/storyteller-training-free-narrative-grounding-for-long-form-audio-description/86327/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-27","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why is long-form audio description more difficult than describing short video clips?","Question",{"text":75,"@type":76},"Long-form AD must maintain narrative continuity across scenes, preserving character identity, relationships, and event context. Many models treat each moment independently, so they can miss who characters are and how current events connect to earlier story information.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What is StoryTeller and how does it preserve narrative context?",{"text":80,"@type":76},"StoryTeller is a training-free framework that maintains an explicit narrative state while processing clips in chronological order. It uses a verified narrative memory that carries forward story-relevant facts to keep later descriptions coherent and contextually grounded.",{"name":82,"@type":73,"acceptedAnswer":83},"What resources does StoryTeller avoid using?",{"text":84,"@type":76},"StoryTeller does not rely on subtitles, scripts, AD transcripts, aligned captions, character banks, precomputed face identities, or task-specific fine-tuning. It can optionally retrieve public movie metadata, but only accepts facts supported by the video through semantic filtering and VLM verification.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]