[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84465-en":3,"doc-seo-84465-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},84465,687197100911,"Himbo","https://ap-avatar.wpscdn.com/avatar/a000239b6f1da00475?x-image-process=image/resize,m_fixed,w_180,h_180&k=1782698725881665579",8,"Research & Report","Youtu-Parsing: Perception, Structuring and Recognition via High-Parallelism Decoding","Youtu-Parsing presents an efficient, versatile document parsing model for high-performance content extraction and layout understanding. It combines a native Vision Transformer with dynamic-resolution visual encoding and a prompt-guided Youtu-LLM-2B for region analysis, enabling feature reuse through a decoupled architecture. A high-parallelism decoding scheme—token parallelism with verification and query parallelism across multiple bounding boxes—achieves 5–11× speedup and an additional 2× acceleration for structured tasks like table recognition, while preserving output quality. The model supports text, formulas, tables, charts, seals, and hierarchical structures, and performs robustly on rare characters, multilingual, and handwritten content.","arXiv :2601 .20430v2 [ cs .CV] 13 Jul 2026  \n Youtu-Parsing Technical Report   \nYoutu-Parsing: Perception, Structuring and Recognition via  \nHigh-Parallelism Decoding  \nYoutu-Parsing Team∗  \nThis paper presents Youtu-Parsing, an efficient and versatile document parsing model designed for high-performance content extraction. The architecture employs a native Vision Transformer (ViT) featuring a dynamic-resolution visual encoder to extract shared document features, coupled with a prompt-guided Youtu-LLM-2B language model for layout analysis and region-prompted decoding. Leveraging this decoupled and feature-reusable framework, we introduce a high-parallelism decoding strategy comprising two core components: token parallelism and query parallelism. The token parallelism strategy concurrently generates up to 64 candidate tokens per inference step, which are subsequently validated through a verification mechanism. This approach yields a 5–11× speedup over traditional autoregressive decoding and is particularly well-suited for highly structured scenarios, such as table recognition. To further exploit the advantages of region-prompted decoding, the query parallelism strategy enables simultaneous content prediction for multiple bounding boxes (up to five), providing an additional 2 × acceleration while maintaining output quality equivalent to standard decoding. Youtu-Parsing encompasses a diverse range of document elements, including text, formulas, tables, charts, seals, and hierarchical structures. Furthermore, the model exhibits strong robustness when handling rare characters, multilingual text, and handwritten content. Extensive evaluations demonstrate that Youtu-Parsing achieves state-of-the-art (SOTA) performance on both the OmniDocBench and olmOCR-bench benchmarks. Overall, Youtu-Parsing demonstrates significant experimental value and practical utility for large-scale document intelligence applications.  \n Code: [https://github.com/TencentCloudADP/youtu-parsing](https://github.com/TencentCloudADP/youtu-parsing)  \n model: [https://huggingface.co/collections/tencent/youtu](https://huggingface.co/collections/tencent/youtu)  \n Demo: [https://huggingface.co/spaces/Tencent/Youtu-Parsing](https://huggingface.co/spaces/Tencent/Youtu-Parsing)  \n1 Introduction  \nDriven by the exponential expansion of digital information, documents have emerged as indispensable repositories for knowledge storage and transmission across diverse domains. However, the escalating complexity and volume of modern documents pose substantial challenges for effective information extraction, necessitating the advancement of sophisticated document parsing [Zhang et al., 2024, Feng et al., 2025, Liet al., 2025, Wang et al., 2024] techniques. The primary objectives of document parsing are manifold: precisely identifying and segmenting structural components—such as text blocks, columns, mathematical formulas, tables, and figures; establishing a logical reading order to preserve semantic coherence; and detecting auxiliary elements including footnotes and captions. Successfully achieving these goals is pivotal for facilitating efficient information retrieval and empowering downstream applications, such as content summarization, knowledge graph construction, and question answering.  \nThe rapid evolution of Large Language Models (LLMs) [Brown et al., 2020, Touvron et al., 2023, Achiam et al., 2023, Team et al., 2023, Jiang et al., 2023] and Vision-Language Models (VLMs) [Radford et al., 2021, Liu et al., 2023, Zhu et al., 2023, Dai et al., 2023, Tschannen et al., 2025] has catalyzed significant progress in document parsing. Nevertheless, the field continues to grapple with the challenge of balancing high-fidelity recognition with the stringent efficiency requirements of real-world applications. Modern documents encompass a  \n*  \nFull author list in contributions.  \nFigure 1. Performance of Youtu-Parsing on OmniDocBench v1.5. Youtu-Parsing surpasses both general-purpose vis","cbCaiq8XlEStaRzW","https://ap.wps.com/l/cbCaiq8XlEStaRzW","pdf",16050243,1,36,"English","en",105,"# Introduction\n## Problem and objectives in document parsing\n## LLM/VLM-driven progress and remaining efficiency challenges\n## Existing paradigms: pipeline vs end-to-end","[{\"question\":\"What core components does Youtu-Parsing use to parse documents?\",\"answer\":\"It uses a Vision Transformer with a dynamic-resolution visual encoder for shared document features and a prompt-guided Youtu-LLM-2B for layout analysis and region-prompted decoding.\"},{\"question\":\"How does high-parallelism decoding accelerate inference in Youtu-Parsing?\",\"answer\":\"Token parallelism generates up to 64 candidate tokens per step and verifies them, producing a 5–11× speedup over traditional autoregressive decoding. Query parallelism predicts multiple bounding boxes simultaneously (up to five), adding about 2× acceleration while keeping output quality.\"},{\"question\":\"What kinds of document elements and input difficulties can Youtu-Parsing handle?\",\"answer\":\"It supports text, formulas, tables, charts, seals, and hierarchical structures, and shows robustness for rare characters, multilingual text, and handwritten content.\"}]",1784195810,91,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"youtu-parsing-perception-structuring-and-recognition-via-high-parallelism-decoding","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/youtu-parsing-perception-structuring-and-recognition-via-high-parallelism-decoding/84465/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What core components does Youtu-Parsing use to parse documents?","Question",{"text":75,"@type":76},"It uses a Vision Transformer with a dynamic-resolution visual encoder for shared document features and a prompt-guided Youtu-LLM-2B for layout analysis and region-prompted decoding.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does high-parallelism decoding accelerate inference in Youtu-Parsing?",{"text":80,"@type":76},"Token parallelism generates up to 64 candidate tokens per step and verifies them, producing a 5–11× speedup over traditional autoregressive decoding. Query parallelism predicts multiple bounding boxes simultaneously (up to five), adding about 2× acceleration while keeping output quality.",{"name":82,"@type":73,"acceptedAnswer":83},"What kinds of document elements and input difficulties can Youtu-Parsing handle?",{"text":84,"@type":76},"It supports text, formulas, tables, charts, seals, and hierarchical structures, and shows robustness for rare characters, multilingual text, and handwritten content.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":45,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":45,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":45,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]