[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82449-en":3,"doc-seo-82449-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},82449,7971461741311,"Ophelia","https://ap-avatar.wpscdn.com/avatar/74000253aff267980c6?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779345379180704826",8,"Research & Report","Scalable Visual Pretraining for Language Intelligence","Large foundation models have advanced primarily through pretraining on text corpora, yet many knowledge forms are carried by figures, equations, and page layouts that text-only conversion cannot preserve. This work introduces Visual Pretraining (VP), a scalable framework that trains on raw visual documents without text extraction or image-text pairing supervision. A systematic study across multiple LLM backbones and scientific reasoning benchmarks shows VP improves over text-only pretraining on matched corpora, achieves efficiency with reduced token budgets, and strengthens cross-modal alignment.","arXiv :2607 .09657v1 [ cs .CV] 10 Jul 2026  \n2026-07-13  \nScalable Visual Pretraining for Language Intelligence  \nYiming Zhang1 ,2 ,* , Zhonghan Zhao 1 ,3 ,* , Wenwei Zhang1 ,* , Haiteng Zhao 1 , Tianyang Lin 1 , Yunhua Zhou 1 , Demin Song1 , Kuikun Liu 1 , Haochen Ye 1 , Haian Huang1 , Yuzhe Gu 1 ,4 , Haijun Lv1 , Qipeng Guo 1 , Bin Liu2 , Gaoang Wang3 ,†, Kai Chen 1 ,†  \n1 Shanghai Artificial Intelligence Laboratory 2 University of Science and Technology of China  \n3 Zhejiang University 4 Shanghai Jiao Tong University  \n*These authors contributed equally to this work. † Corresponding authors.  \n[gaoangwang@zju.edu.cn](gaoangwang@zju.edu.cn) , [chenkai@pjlab.org.cn](chenkai@pjlab.org.cn)  \nThe rapid progress of large foundation models has been driven predominantly by pretraining on large-scale text corpora. However, many forms of knowledge are conveyed through visual representations, where figures, typeset equations, and page layouts carry rich information that cannot be faithfully or completely captured by text alone. Yet current pretraining approaches discard these visual cues by converting visually rich sources, such as documents and web pages, into plain text for learning language intelligence. This paper challenges the default assumption that language models must be trained on text-only representations and shows that Visual Pretraining is a scalable learner for foundation model intelligence. To this end, we conduct a systematic study of unsupervised visual pretraining paradigms that directly leverage visual documents without text extraction. Across multiple backbones and benchmarks, visual pretraining on the same underlying corpora consistently outperforms text-only pretraining, offering an efficient pathway to scalable language intelligence.  \n1. Introduction  \nThe striking advances of large foundation models have been driven by pretraining on text corpora at unprecedented scale [4, 9, 12] . While effective, this paradigm rests on a strong implicit assumption: that all knowledge worth learning can be losslessly encoded as a linear sequence of text tokens. However, cognitive science has long shown that humans routinely reason with diagrams, spatial layouts, mathematical notation, and other representational forms whose visual cues make certain relations directly available for inference [1, 2, 15, 38] . Converting these visual cues into plain text before training is therefore inherently lossy. Here we challenge this assumption and show that foundation models can learn directly from visual corpora without text extraction or image-text pairing supervision, yielding stronger language intelligence than text pretraining on the same underlying corpus.  \nScientific documents are a particularly acute case of this loss. Papers, textbooks, and technical reports communicate complex content through figures, tables, formula layouts, and page-level spatial organization, all of which encode geometric constraints, symbolic topologies, and structural correspondences essential to scientific reasoning. The recently proposed Platonic Representation Hypothesis [10] formalizes a related intuition for machine learning, arguing that representations learned by different models and modalities converge toward shared abstractions of reality. These observations imply that the linguistic information extracted from a scientific document is only a projection of a richer underlying structure, and that this structure could in principle be learned directly from raw visual documents.  \nYet existing approaches do not exploit this possibility along either of two dimensions. At the data level, pretraining corpora reduce visual documents to plain text, either by parsing HTML or LaTeX source [7, 27] or by applying neural document-parsing models to PDFs as a preprocessing step [3, 13, 16, 32], after which the language model is trained exclusively on the resulting text. At the training level, recent multimodal foundation models do incorporate visual modalities duri","cbCairm266hshbfU","https://ap.wps.com/l/cbCairm266hshbfU","pdf",2590508,4,1,18,"English","en",105,"# Introduction\n## Visual information loss in text-only pretraining\n## Visual Pretraining (VP) framework\n## Contributions and experimental findings","[{\"question\":\"What problem does the paper address with current text-only pretraining?\",\"answer\":\"It argues that converting visual documents into plain text is inherently lossy, since diagrams, spatial layouts, equations, and tables contain structural cues needed for reasoning.\"},{\"question\":\"How does Visual Pretraining (VP) differ from text pretraining?\",\"answer\":\"VP trains on raw visual documents rendered as images and learns visual information directly, without text extraction and without requiring image-text pairing supervision.\"},{\"question\":\"What results does the paper report when comparing VP with text-only pretraining?\",\"answer\":\"On matched underlying corpora, visual pretraining consistently outperforms text-only pretraining across multiple backbones and benchmarks, is more efficient with a smaller token budget, and improves cross-modal alignment.\"}]",1784180445,45,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"scalable-visual-pretraining-for-language-intelligence","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/scalable-visual-pretraining-for-language-intelligence/82449/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-21","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the paper address with current text-only pretraining?","Question",{"text":75,"@type":76},"It argues that converting visual documents into plain text is inherently lossy, since diagrams, spatial layouts, equations, and tables contain structural cues needed for reasoning.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does Visual Pretraining (VP) differ from text pretraining?",{"text":80,"@type":76},"VP trains on raw visual documents rendered as images and learns visual information directly, without text extraction and without requiring image-text pairing supervision.",{"name":82,"@type":73,"acceptedAnswer":83},"What results does the paper report when comparing VP with text-only pretraining?",{"text":84,"@type":76},"On matched underlying corpora, visual pretraining consistently outperforms text-only pretraining across multiple backbones and benchmarks, is more efficient with a smaller token budget, and improves cross-modal alignment.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]