[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-133947-en":3,"doc-seo-133947-105":31,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":28,"seo_description":14,"update_tm":29,"read_time":30},133947,2336464648746,"Skyler","https://ap-avatar.wpscdn.com/davatar_276721f389ce27ea32af1340a28f341c",8,"Research & Report","MangaVQA and MangaLMM - A Benchmark and Specialized Model for Multimodal Manga Understanding","Manga is a richly multimodal narrative form that combines images and embedded text, challenging large multimodal models (LMMs) to understand panels at a human-like level. The work introduces two benchmarks: MangaOCR for in-page text recognition and MangaVQA for contextual visual question answering, with 526 manually built question–answer pairs. Based on these benchmarks, a manga-specialized model, MangaLMM, is fine-tuned from Qwen2.5-VL to jointly perform OCR and VQA, then evaluated against GPT-4o and Gemini 2.5.","MangaVQA and MangaLMM: A Benchmark and Specialized Model for  \nMultimodal Manga Understanding  \nJeonghun Baek* Kazuki Egashira* Shota Onohara* Atsuyuki Miyai*  \nYuki Imajuku Hikaru Ikuta Kiyoharu Aizawa  \nThe University of Tokyo  \n[baek@hal.t.u-tokyo.ac.jp](baek@hal.t.u-tokyo.ac.jp)  \n[https://manga109.github.io/MangaVQA_LMM/](https://manga109.github.io/MangaVQA_LMM/)  \narXiv :2505 .20298v3 [ cs .CL] 26 Jan 2026  \nAbstract  \nManga, or Japanese comics, is a richly multimodal narrative form that blends images and text in complex ways. Teaching large multimodal models (LMMs) to understand such narratives at a human-like level could help manga creators reflect on and refine their stories. To this end, we introduce two benchmarks for multimodal manga understanding: MangaOCR, which targets in-page text recognition, and MangaVQA, a novel benchmark designed to evaluate contextual understanding through visual question answering. MangaVQA consists of 526 high-quality, manually constructed question-answer pairs, enabling reliable evaluation across diverse narrative and visual scenarios. Building on these benchmarks, we develop MangaLMM, a manga-specialized model finetuned from the open-source LMM Qwen2.5-VL to jointly handle both tasks. Through extensive experiments, including comparisons with proprietary models such as GPT-4o and Gemini 2.5, we assess how well LMMs understand manga. Our benchmark and model provide a comprehensive foundation for evaluating and advancing LMMs in the richly narrative domain of manga.  \n1 Introduction  \nManga is a rich and distinctive form of multimodal narrative, combining complex panel layouts, expressive visual elements, and text embedded directly within images. As large multimodal models (LMMs) continue to advance in vision-language understanding, enabling them to understand manga presents an exciting opportunity, not only as a technical milestone, but also as a way to support human creativity. Such models could assist manga creatorsin reflecting on and refining their stories. To provide meaningful assistance, an LMM would need to function like a skilled editor or assistant, capable of reading and understanding manga in a way  \n*Equal contribution.  \nhuman does. This calls for evaluating models’ abilities to process visual-textual content and follow the context in a coherent and human-like manner.  \nAlthough recent efforts such as Magi (Sachdeva and Zisserman, 2024 ; Sachdeva et al., 2024 ; Sachdeva and Zisserman, 2025) and CoMix (Vivoliet al., 2024) have tackled comic understanding, they primarily focus on generating transcriptions from comic pages – they do not evaluate to what extent models can accurately read in-page text using optical character recognition (OCR), or understand the content based on that text through visual question answering (VQA) . As a result, it remains unclear to what extent models truly comprehend manga content in a human-like manner based on the embedded textual information.  \nTo pave a reliable path toward comprehensive manga understanding in LMMs, we believe it is essential to evaluate two core capabilities: OCRand VQA. To address these needs, we propose two benchmarks: MangaOCR and MangaVQA. MangaOCR focuses on detecting and recognizing textual content such as dialogue and sound effects. We consolidate existing annotations from the well-known Manga109 dataset (Matsui et al., 2017 ; Aizawa et al., 2020) and the manga onomatopoeia dataset (Baek et al., 2022) to construct this benchmark. Further, as our primary contribution, we propose MangaVQA, a novel benchmark designed to evaluate an LMM’s ability to accurately answer targeted, factual questions grounded in both visual and textual context. It consists of 526 high-quality, manually constructed question–answer pairs covering a diverse range of scenarios, enabling assessment of a model’s narrative understanding. Together, these benchmarks provide a comprehensive framework for evaluating a model’s ability to understand manga as","cbCaiipiBIFWjuHA","https://ap.wps.com/l/cbCaiipiBIFWjuHA","pdf",17938495,4,1,22,"English","en",105,"# Abstract\n# 1 Introduction\n# 2 Related Work","[{\"question\":\"What problem does MangaVQA and MangaLMM address?\",\"answer\":\"They target reliable multimodal understanding of manga by evaluating both in-page text reading (OCR) and context-aware visual question answering (VQA).\"},{\"question\":\"What are the key components of the proposed benchmarks?\",\"answer\":\"MangaOCR focuses on detecting and recognizing in-page text such as dialogue and sound effects, while MangaVQA provides 526 manually constructed visual question–answer pairs for contextual understanding.\"},{\"question\":\"How is MangaLMM trained and how is it evaluated?\",\"answer\":\"MangaLMM is fine-tuned from the open-source LMM Qwen2.5-VL to jointly handle OCR and VQA, and it is compared through extensive experiments against proprietary models like GPT-4o and Gemini 2.5.\"}]","MangaVQA and MangaLMM - A Benchmark and Specialized Model for Multimodal Manga Understanding | PDF",1787230346,55,{"code":4,"msg":32,"data":33},"ok",{"site_id":25,"language":24,"slug":34,"title":13,"keywords":35,"description":14,"schema_data":36,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":29},"mangavqa-and-mangalmm-a-benchmark-and-specialized-model-for-multimodal-manga-understanding","",{"@graph":37,"@context":86},[38,54,69],{"@type":39,"itemListElement":40},"BreadcrumbList",[41,45,49,52],{"item":42,"name":43,"@type":44,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":46,"name":47,"@type":44,"position":48},"https://docshare.wps.com/document/","Document",2,{"item":50,"name":12,"@type":44,"position":51},"https://docshare.wps.com/document/research-report/",3,{"item":53,"name":13,"@type":44,"position":20},"https://docshare.wps.com/document/mangavqa-and-mangalmm-a-benchmark-and-specialized-model-for-multimodal-manga-understanding/133947/",{"url":53,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":42,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-31","2026-08-20",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What problem does MangaVQA and MangaLMM address?","Question",{"text":76,"@type":77},"They target reliable multimodal understanding of manga by evaluating both in-page text reading (OCR) and context-aware visual question answering (VQA).","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"What are the key components of the proposed benchmarks?",{"text":81,"@type":77},"MangaOCR focuses on detecting and recognizing in-page text such as dialogue and sound effects, while MangaVQA provides 526 manually constructed visual question–answer pairs for contextual understanding.",{"name":83,"@type":74,"acceptedAnswer":84},"How is MangaLMM trained and how is it evaluated?",{"text":85,"@type":77},"MangaLMM is fine-tuned from the open-source LMM Qwen2.5-VL to jointly handle OCR and VQA, and it is compared through extensive experiments against proprietary models like GPT-4o and Gemini 2.5.","https://schema.org",{"og:url":53,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":53},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,121,124,129,132,136],{"id":21,"doc_module":4,"doc_module_name":47,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":48,"doc_module":4,"doc_module_name":47,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":47,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":47,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":47,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":47,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":47,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":125,"doc_module":4,"doc_module_name":47,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":47,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":133,"doc_module":4,"doc_module_name":47,"category_name":134,"show_sort_weight":133,"slug":135},10,"Lifestyle","lifestyle",{"id":137,"doc_module":4,"doc_module_name":47,"category_name":138,"show_sort_weight":107,"slug":139},19,"General","general"]