[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82155-en":3,"doc-seo-82155-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},82155,1099514067438,"River Wang","https://ap-avatar.wpscdn.com/avatar/100002539ee87300030?x-image-process=image/resize,m_fixed,w_180,h_180&k=1780474512215547542",8,"Research & Report","Video Generation Models are General-Purpose Vision Learners","Large-scale text-to-video generation is proposed as a pretraining paradigm that enables general-purpose computer vision. The paper introduces GenCeption, which uses a pretrained video generative diffusion backbone as feed-forward perception to perform diverse vision tasks guided by text. Experiments show state-of-the-art results across depth, surface normals, camera pose, expression-refer ring segmentation, and 3D keypoint prediction. The approach also matches or exceeds specialized models with 7×–500× less data and demonstrates sim-to-real transfer from synthetic to real videos.","arXiv :2607 .09024v1 [ cs .CV] 10 Jul 2026  \nVideo Generation Models are General-Purpose Vision Learners  \nLetian Wang1,2, Chuhan Zhang1, Rishabh Kabra1,3, Jasper Uijlings1, Steven Waslander2, Andrew Zisserman1,4, Joao Carreira1 , Kaiming He1,5, Misha Andriluka1, Eduard Gabriel Bazavan1, Andrei Zanfir1 and Cristian Sminchisescu1,6,∗  \n1 Google DeepMind, 2University of Toronto, 3University College London, 4University of Oxford, 5MIT, 6Lund University  \nDriven by next-token prediction, NLP shifted from task-specific models into powerful generalist foundation models. What, then, is the equivalent catalyst needed to achieve a general-purpose model in computer vision? In this paper, we contend that large-scale text-to-video generation serves as a strong pre-training paradigm for computer vision, providing the necessary spatiotemporal priors, vision-language alignment, and scalability required for general visual intelligence. We introduce GenCeption, which leverages a pre-trained video generative diffusion backbone to define a feed-forward perception model, capable of performing various vision tasks steered by text instructions. Empirical results demonstrate that GenCeption achieves state-of-the-art performance across a diverse suite of tasks, including depth, surface normal, and camera pose estimation, expression-referring segmentation, and 3D keypoint prediction, often matching or surpassing specialized models (e.g. DepthAnything3, SAM3, D4RT, VGGT-Ω, Sapiens, David, Genmo, and Lotus-2). Furthermore, the video generative pretrained backbone outperforms alternative pretraining paradigms (e.g., V-JEPA, and Video MAE) under comparable settings. Importantly, GenCeption exhibits preliminary data and model scaling properties along with exceptional data efficiency where it achieves comparable performance with leading models like D4RT and VGGT-Ω with 7× to 500× less training data. Finally, GenCeption also exhibits intriguing emergent behaviors: a model trained exclusively on synthetic human videos generalizes to real-world footage and out-of-distribution object categories (e.g., animals and robots). These findings suggest that video generation is not merely a synthesis tool, but a foundational path toward generalist vision intelligence for the physical world.  \nProject page: [https://genception.github.io](https://genception.github.io)  \n[1](1). Introduction  \nNatural language processing (NLP) evolved from an era of specialized models—where separate models were developed for each language task (e.g. translation, summarization)—to a singular, unified foundation model paradigm. This paradigm, supported by large-scale next-token prediction pre-training followed by task-aligned post-training, has effectively collapsed thousands of disparate linguistic challenges into a single generalist intelligence, ultimately unlocking emerging behaviors such as chain-of-thought and in-context learning.  \nIn stark contrast, computer vision is still lingering in its \"specialized model\" stage. While recent years have yielded powerful foundation models such as Segment Anything series [10, 36, 53] for localization or Depth Anything series [41, 76, 77] for geometry, these remain fundamentally taskspecific models that need customized architectures for each task. We have mastered specialized perception, but we have yet to realize a unified vision foundation model—a unified task-agnostic architecture that mirrors LLMs from general pre-training to versatile, emerging intelligence. To this end, we posit that the quest for a generalist vision model is essentially a search for a universal pre-training objective—a visual analog to next-token prediction. Such a pre-training paradigm should satisfy three core imperatives:  \n1) Spatio-Temporal Evolution: The world is a 4D continuum. The pre-training objective should force the model to internalize the 4D temporal causality and physics of a world in motion.  \n2) Vision-Language Alignment: To inherit the instruction-following ","cbCailm8tvBuYLS3","https://ap.wps.com/l/cbCailm8tvBuYLS3","pdf",17867194,3,1,19,"English","en",105,"# Introduction\n## Video Generative Pre-training\n## GenCeption Framework\n# Empirical Results and Scaling\n## State-of-the-art Task Performance\n## Data Efficiency and Emergent Behaviors","[{\"question\":\"What pretraining objective does the paper identify as a visual analog to next-token prediction?\",\"answer\":\"The paper argues that large-scale text-to-video generation provides a universal pretraining objective for computer vision by learning spatiotemporal priors, aligning vision with language, and supporting scalable training.\"},{\"question\":\"How does GenCeption work at a high level?\",\"answer\":\"GenCeption leverages a pretrained video generative diffusion backbone and defines a feed-forward perception model that can perform multiple vision tasks steered by text instructions.\"},{\"question\":\"What performance and efficiency benefits does GenCeption report?\",\"answer\":\"It achieves state-of-the-art results on tasks such as depth, surface normals, camera pose, expression-referring segmentation, and 3D keypoint prediction, and often matches leading models using 7× to 500× less training data.\"}]",1784178484,48,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"video-generation-models-are-general-purpose-vision-learners","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/video-generation-models-are-general-purpose-vision-learners/82155/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-22","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What pretraining objective does the paper identify as a visual analog to next-token prediction?","Question",{"text":75,"@type":76},"The paper argues that large-scale text-to-video generation provides a universal pretraining objective for computer vision by learning spatiotemporal priors, aligning vision with language, and supporting scalable training.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does GenCeption work at a high level?",{"text":80,"@type":76},"GenCeption leverages a pretrained video generative diffusion backbone and defines a feed-forward perception model that can perform multiple vision tasks steered by text instructions.",{"name":82,"@type":73,"acceptedAnswer":83},"What performance and efficiency benefits does GenCeption report?",{"text":84,"@type":76},"It achieves state-of-the-art results on tasks such as depth, surface normals, camera pose, expression-referring segmentation, and 3D keypoint prediction, and often matches leading models using 7× to 500× less training data.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},"General","general"]