[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-227893-en":3,"doc-seo-227893-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},227893,2336477405376,"Stanley","https://ap-avatar.wpscdn.com/davatar_29158cc5080c5b710cf443261637dec0",8,"Research & Report","TIU-Bench: A Benchmark for Evaluating Large Multimodal Models on Text-rich Image Understanding - Research Summary","Text-rich images are a key medium for communicating information and improving accessibility, yet current multimodal benchmarks remain limited in scale, scenario coverage, and evaluation design for holistic understanding. TIU-Bench is introduced as a large-scale, multilingual benchmark with 100,000+ full-image annotations and 22,000 validated question-answer pairs spanning 18 subtasks. It adds a structured full-image output format and a two-stage framework (T2TIU) that first represents entire images then performs reasoning, supported by extensive experiments.","TIU-Bench: A Benchmark for Evaluating Large Multimodal Models on  \nText-rich Image Understanding  \nKun Zhang, Liqiang Niu , Zhen Cao, Fandong Meng* , Jie Zhou  \n1Pattern Recognition Center, WeChat AI, Tencent Inc  \n{peterkzhang, [zhenzcao}@tencent.com](zhenzcao}@tencent.com)  \nAbstract  \nText-rich images are ubiquitous in real-world applications, serving as a critical medium for conveying complex information and facilitating accessibility. Despite recent advances driven by Multimodal Large Language Models (MLLMs), existing benchmarks suffer from limited scale, fragmented scenarios, and evaluation protocols that fail to fully capture holistic image understanding. To address these gaps, we present TIU-Bench, a large-scale, multilingual benchmark comprising over 100,000 full-image annotations and 22,000 rigorously validated question-answer (QA) pairs that span  \n18 subtasks across diverse real-world scenarios. TIU-Bench introduces a novel full-image structured output format that jointly models geometric, textual, and relational information, enabling fine-grained evaluation of perception and reasoning capabilities. Furthermore, we propose a two-stage understanding framework named T2TIU, which first generates a structured representation of the entire image and subsequently conducts reasoning on this representation in order to address complex visualtextual queries. Extensive experiments on 10 state-of-the-art generative models highlight the challenges and opportunities in advancing textrich image understanding. Our benchmark and framework provide a comprehensive platform for developing and evaluating next-generation multimodal AI systems.  \n1 Introduction  \nText-rich images play a pivotal role in real-world scenarios by efficiently conveying complex information and improving accessibility (Biten et al., 2019) . Accurate interpretation of such images is essential for automating information extraction, advancing AI systems, and optimizing user interactions. To formalize this research domain, we define Text-rich Image Understanding (TIU) as consist-  \n* Corresponding author.  \n0 20 40 60 80 100 120 Instruction (k)  \nFigure 1: Comparison of the number of images and instructions between our dataset and existing datasets.  \ning of two core capabilities: perception and understanding. The perception dimension encompasses visual recognition tasks such as text detection (Liao et al., 2022), text recognition (Guan et al., 2025), formula recognition (Truong et al., 2024 ; Guan et al., 2024), and document layout analysis (Yupan et al., 2022) . In contrast, the understanding dimension involves semantic reasoning for downstream applications such as key information extraction and document-based visual question answering (e.g., DocVQA (Mathew et al., 2021), ChartQA (Masryet al., 2022), and TextVQA (Singh et al., 2019)) .  \nRecently, MLLMs have been proposed, which integrate large language models (LLMs) with visual encoders to jointly process visual tokens and linguistic elements through unified attention mechanisms, enabling end-to-end sequence modeling. Within the TIU domain, MLLMs have demonstrated impressive results in both perception and understanding. Nevertheless, despite recent advances, there are still two key challenges in current TIU research paradigms.  \nDataset Limitations. TIU tasks currently face challenges related to data diversity, scale, and quality. Existing datasets such as DocVQA and ChartQA focus on isolated scenarios (e.g., documents, tables, or charts), and their fragmented objectives and scenario-specific designs hinder the  \n24286  \nFindings of the Association for Computational Linguistics: EMNLP 2025 , pages 24286–24295 November 4-9, 2025 ©2025 Association for Computational Linguistics  \nFigure 2: Overview of sub-tasks in TIU-Bench datasets.  \ncomprehensive evaluation of perceptual and interpretative capabilities. Recent benchmarks such as OCRbench and its v2 version assess line-level text recognition in MLLMs but overlook ","cbCaibzMyC5SO4ei","https://ap.wps.com/l/cbCaibzMyC5SO4ei","pdf",3007014,1,10,"English","en",105,"# Introduction\n## Text-rich Image Understanding (TIU)\n## Dataset Limitations\n## Task Setting and Evaluation Constraints\n## TIU-Bench Design and Capabilities","[{\"question\":\"What problem does TIU-Bench target in text-rich image understanding benchmarks?\",\"answer\":\"It targets limited benchmark scale, fragmented scenarios, and evaluation protocols that do not capture holistic full-image perception and reasoning.\"},{\"question\":\"How does TIU-Bench structure evaluation for complex text-rich images?\",\"answer\":\"It uses a full-image structured output format and evaluates multiple capabilities including text recognition, relation extraction, and visual-textual reasoning across 18 subtasks.\"},{\"question\":\"What is the role of the T2TIU framework in TIU-Bench?\",\"answer\":\"T2TIU performs reasoning in two stages: generating a structured representation of the entire image first, then conducting reasoning on that representation for complex visual-textual queries.\"}]","TIU-Bench: A Benchmark for Evaluating Large Multimodal Models on Text-rich Image Understanding - Research Summary | PDF",1789016677,25,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"tiu-bench-a-benchmark-for-evaluating-large-multimodal-models-on-text-rich-image-understanding-research-summary","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/tiu-bench-a-benchmark-for-evaluating-large-multimodal-models-on-text-rich-image-understanding-research-summary/227893/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-09-10",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does TIU-Bench target in text-rich image understanding benchmarks?","Question",{"text":75,"@type":76},"It targets limited benchmark scale, fragmented scenarios, and evaluation protocols that do not capture holistic full-image perception and reasoning.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does TIU-Bench structure evaluation for complex text-rich images?",{"text":80,"@type":76},"It uses a full-image structured output format and evaluates multiple capabilities including text recognition, relation extraction, and visual-textual reasoning across 18 subtasks.",{"name":82,"@type":73,"acceptedAnswer":83},"What is the role of the T2TIU framework in TIU-Bench?",{"text":84,"@type":76},"T2TIU performs reasoning in two stages: generating a structured representation of the entire image first, then conducting reasoning on that representation for complex visual-textual queries.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,134],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":21,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":21,"slug":133},"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]