[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85116-en":3,"doc-seo-85116-105":30,"detail-sidebar-cat-0-en-105":83},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},85116,2336464648746,"Skyler","https://ap-avatar.wpscdn.com/davatar_276721f389ce27ea32af1340a28f341c",8,"Research & Report","OpenCoF Learning to Reason Through Video Generation","Reasoning is a core capability for large models when reliable decisions depend on logical consequences. Video generation introduces Chain-of-Frame (CoF) reasoning, where logical steps unfold through temporally connected frames, yet existing generators lack dedicated supervision and designs for CoF. OpenCoF provides the OpenCoF-17K dataset across 11 task families and a fine-tuned model Wan-CoF to study temporal supervision effects. Experiments across four benchmarks show strong gains over the Wan2.2-I2V-A14B baseline, and further token-based designs improve intermediate reasoning organization. ","arXiv :2607 .08763v 1 [ cs .CV] 9 Jul 2026  \nOpenCoF: Learning to Reason Through  \nVideo Generation  \nXinyan Chen 1 ,2 ,∗ , Ziyu Guo3 ,∗ , Renrui Zhang 1 ,†, Dongzhi Jiang 1 ,2 , Hongsheng Li2  \n1 ByteDance Seed, 2 CUHK MMLab, 3 CUHK IMIXR  \n∗ Equal contribution, †Corresponding author  \nAbstract  \nReasoning has become a core capability for large models, especially when reliable decisions require understanding logical consequences. Recent video generation models offer a reasoning path distinct from previous Chain-of-Thought (CoT): reasoning can unfold through temporally connected frames, known as Chain-of-Frame (CoF) reasoning. However, existing video generators are primarily trained on general video corpora, still lacking diverse supervision and dedicated designs for CoF reasoning. To address this gap, we introduce OpenCoF, a framework comprising the OpenCoF-17K dataset, a reasoning video dataset spanning 11 task families, and Wan-CoF, a fine-tuned video model for studying whether diverse temporal supervision improves CoF behavior. Across four video reasoning benchmarks, Wan-CoF achieves considerable gains over the Wan2.2-I2V-A14B baseline. Building on this, we empirically explore more advanced designs for CoF capabilities, i.e., equipping the model with visual and textual reasoning tokens. This mechanism respectively captures low-level visual cues and high-level semantic priors for spatial and temporal reasoning. Through performance comparisons and attention analysis, we examine how these tokens contribute across model depth, denoising steps, space, and time. Our results suggest that stronger video reasoning requires both broad temporal supervision and explicit mechanisms for organizing intermediate reasoning state. We open-source the dataset, model, and code to facilitate future research on reasoning-oriented video generation.  \nDate: July 10, 2026  \nCorrespondence: Renrui Zhang at [renruizhang@bytedance.com](renruizhang@bytedance.com)  \nProject Page: [https://opencof.github.io/](https://opencof.github.io/)  \n1 Introduction  \nEnhancing reasoning ability [16, 24 , 48 , 49] has emerged as a central objective in the development of large language (LLM) and multimodal (LMM) models. Prior works demonstrate that robust reasoning capabilities are indispensable for reliable decision-making in vision-language contexts. However, mainstream visual Chain-of-Thought (CoT) [17, 27 , 43] pipelines remain largely anchored to static visual observations. These approaches typically rely on extracting localized evidence [4, 34], invoking external tools [11, 13 , 15 , 50], or employing auxiliary image-generation steps [20, 46], which limits their ability to intrinsically model dynamic transitions and multi-step visual consequences.  \nThe rapid advancement of video generation models offers a promising alternative: reasoning via video [10 , 23 , 30 , 44] . Rather than relying solely on textual steps or static visual content, models can reason through  \nSudoku  \nInstance-based Rendering  \n2D Geometry  \nExpert-guided Rendering  \nExternal Video  \nRepurposing  \nEmbodied Manipulation  \nPhysics Motion  \nFigure 1 Overview of OpenCoF. We construct the OpenCoF-17K dataset, comprising 17,312 videos across 11 tasks, via four complementary pipelines, to provide diverse temporal supervision for Chain-of-Frame (CoF) reasoning.  \nthe temporal evolution of frames, a process formalized as Chain-of-Frame (CoF) reasoning. Recent studies suggest that video models already exhibit nontrivial world knowledge and early signs of physical and causal understanding, indicating significant potential for CoF [44] . Nevertheless, this potential remains far from fully realized. In reasoning-intensive tasks, current video models still struggle with long-term temporal coherence, physical and spatial consistency, and logical continuity [10] . Ultimately, high visual realism does not guarantee reliable reasoning.  \nStarting from MME-CoF [10], recent works have predominantly focused o","cbCaitmNF4tZqVDJ","https://ap.wps.com/l/cbCaitmNF4tZqVDJ","pdf",3467134,2,1,19,"English","en",105,"# Abstract\n# Introduction\n## Reasoning for large models and limitations of CoT in vision\n## Chain-of-Frame (CoF) reasoning and challenges in video generators\n## Related work and the gap in enhancing video reasoning\n## OpenCoF framework and dataset/model contributions","[{\"question\":\"How does Wan-CoF evaluate the impact of temporal supervision?\",\"answer\":\"Wan-CoF is a fine-tuned video model trained on OpenCoF-17K without reasoning-specific techniques, enabling comparison across benchmarks to test whether diverse temporal supervision improves CoF reasoning behavior.\"}]",1784201200,48,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":78,"head_meta":80,"extra_data":82,"updated_unix":28},"opencof-learning-to-reason-through-video-generation","",{"@graph":36,"@context":77},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/opencof-learning-to-reason-through-video-generation/85116/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-19","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71],{"name":72,"@type":73,"acceptedAnswer":74},"How does Wan-CoF evaluate the impact of temporal supervision?","Question",{"text":75,"@type":76},"Wan-CoF is a fine-tuned video model trained on OpenCoF-17K without reasoning-specific techniques, enabling comparison across benchmarks to test whether diverse temporal supervision improves CoF reasoning behavior.","Answer","https://schema.org",{"og:url":51,"og:type":79,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":81,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":84},[85,89,93,97,102,107,112,115,120,123,127],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":86,"show_sort_weight":87,"slug":88},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":90,"show_sort_weight":91,"slug":92},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Exam",70,"exam",{"id":98,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},5,"Comic",60,"comic",{"id":103,"doc_module":4,"doc_module_name":46,"category_name":104,"show_sort_weight":105,"slug":106},6,"Technology",50,"technology",{"id":108,"doc_module":4,"doc_module_name":46,"category_name":109,"show_sort_weight":110,"slug":111},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":113,"slug":114},30,"research-report",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},9,"Religion & Spirituality",20,"religion-spirituality",{"id":118,"doc_module":4,"doc_module_name":46,"category_name":121,"show_sort_weight":118,"slug":122},"World Cup","world-cup",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":124,"slug":126},10,"Lifestyle","lifestyle",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":98,"slug":129},"General","general"]