[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85560-en":3,"doc-seo-85560-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},85560,16904993612988,"Olivia Brown","https://ap-avatar.wpscdn.com/davatar_a8503ba1806abce46bf441b54a3ca4cd",8,"Research & Report","MG2-RAG Multi-Granularity Graph for Multimodal Retrieval-Augmented Generation","Retrieval-Augmented Generation (RAG) helps Multimodal Large Language Models (MLLMs) reduce hallucinations, yet current systems struggle with complex cross-modal reasoning. Flat vector retrieval misses structural dependencies, while graph-based solutions often depend on costly text-centric pipelines that discard fine-grained visual evidence. MG2-RAG introduces a lightweight multi-granularity multimodal knowledge graph, fusing textual entities and grounded visual regions into unified nodes for atomic evidence preservation. Multi-hop graph retrieval aggregates similarities and propagates relevance, delivering state-of-the-art results with substantial speed and cost improvements.","arXiv :2604 .04969v2 [ cs .IR] 12 Jul 2026  \nMG2-RAG: Multi-Granularity Graph for Multimodal Retrieval-Augmented Generation  \nSijun Dai 1 ,†, Qiang Huang 1 ,†, Xiaoxing You2, and Jun Yu 1 , *  \n1 School of Intelligence Science and Engineering,  \nHarbin Institute of Technology (Shenzhen)  \n[daisijun@stu.hit.edu.cn](daisijun@stu.hit.edu.cn) , {huangqiang, [yujun}@hit.edu.cn](yujun}@hit.edu.cn)  \n2 School of Computer Science, Hangzhou Dianzi University  \n[youxiaoxing@hdu.edu.cn](youxiaoxing@hdu.edu.cn)  \n†Equal contribution * Corresponding author  \nAbstract. Retrieval-Augmented Generation (RAG) mitigates hallucinations in Multimodal Large Language Models (MLLMs), yet existing systems struggle with complex cross-modal reasoning. Flat vector retrieval often ignores structural dependencies, while current graph-based methods rely on costly “translation-to-text” pipelines that discard fine-grained visual information. To address these limitations, we propose MG2-RAG, a lightweight Multi-Granularity Graph RAG framework that jointly improves graph construction, modality fusion, and cross-modal retrieval.  \nMG2-RAG constructs a hierarchical multimodal knowledge graph by combining lightweight textual parsing with entity-driven visual grounding, enabling textual entities and visual regions to be fused into unified multimodal nodes that preserve atomic evidence. Building on this representation, we introduce a multi-granularity graph retrieval mechanism that aggregates dense similarities and propagates relevance across the graph to support structured multi-hop reasoning. Extensive experiments across four representative multimodal tasks (i.e., retrieval, knowledgebased VQA, reasoning, and classification) demonstrate that MG2-RAG consistently achieves state-of-the-art performance while reducing graph construction overhead with an average 43.3 × speedup and 23.9 × cost reduction compared with advanced graph-based frameworks. The source code is publicly available at [https://github.com/Daboolu/MG2-RAG](https://github.com/Daboolu/MG2-RAG).  \nKeywords: Multimodal Knowledge Graph · Retrieval-Augmented Generation · Multimodal Large Language Models  \n1 Introduction  \nMultimodal Large Language Models (MLLMs) have achieved remarkable success in tasks requiring complex cross-modal understanding and reasoning, leading to their rapid adoption across numerous applications [20, 27, 62, 64] . Despite these advances, deploying MLLMs in knowledge-intensive scenarios still raises important reliability concerns. In particular, MLLMs may produce multimodal hallucinations and often lack access to private or domain-specific knowledge, since  \n2 Sijun Dai et al.  \ntheir parameters are primarily trained on large-scale public corpora [34,72,87] . To address these limitations, Multimodal Retrieval-Augmented Generation (MMRAG) augments MLLMs with external knowledge bases by retrieving relevant multimodal evidence to ground the generation process [5, 33, 78, 82, 85] . By incorporating factual context at inference time, MM-RAG significantly improves both the reliability and domain adaptability of MLLMs.  \nExisting MM-RAG methods generally fall into two paradigms: vector-based and graph-based approaches. Vector-based MM-RAG encodes multimodal inputs into a shared embedding space [61,65,68] and retrieves relevant evidence via similarity search [12, 13] . While effective in many scenarios, this paradigm typically retrieves isolated multimodal elements and largely ignores the structural relationships among different pieces of evidence [9, 32, 48, 74] . Consequently, two fundamental limitations arise: First, the semantic gap between modalities can exacerbate modality fragmentation, weakening cross-modal alignment. Second, similarity-based retrieval cannot model logical dependencies  \nFig. 1: Comparison between existing graph-based MM-RAG frameworks and MG2-RAG. Existing methods rely on costly text-centric graph construction that discards fine-grained visual information, whereas","cbCaip6zoqKHZvQp","https://ap.wps.com/l/cbCaip6zoqKHZvQp","pdf",4046007,4,1,34,"English","en",105,"# Introduction\n## Background: RAG and MLLMs\n## Limitations of existing vector-based and graph-based MM-RAG\n## Proposed approach: MG2-RAG\n## Contributions and experimental evaluation","[{\"question\":\"What problem does MG2-RAG address in multimodal RAG systems?\",\"answer\":\"MG2-RAG targets the gap between modalities and the inability of existing approaches to model structural and logical dependencies for reliable cross-modal multi-hop reasoning.\"},{\"question\":\"How does MG2-RAG construct its multimodal knowledge graph?\",\"answer\":\"It builds a hierarchical multimodal knowledge graph by fusing lightweight textual parsing with entity-driven visual grounding, creating unified multimodal nodes that preserve atomic evidence.\"},{\"question\":\"What retrieval mechanism does MG2-RAG use for multi-hop reasoning?\",\"answer\":\"It employs a multi-granularity graph retrieval strategy that aggregates dense similarity and propagates relevance across the graph to support structured multi-hop reasoning.\"}]",1784204569,86,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"mg2-rag-multi-granularity-graph-for-multimodal-retrieval-augmented-generation","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/mg2-rag-multi-granularity-graph-for-multimodal-retrieval-augmented-generation/85560/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-25","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does MG2-RAG address in multimodal RAG systems?","Question",{"text":75,"@type":76},"MG2-RAG targets the gap between modalities and the inability of existing approaches to model structural and logical dependencies for reliable cross-modal multi-hop reasoning.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does MG2-RAG construct its multimodal knowledge graph?",{"text":80,"@type":76},"It builds a hierarchical multimodal knowledge graph by fusing lightweight textual parsing with entity-driven visual grounding, creating unified multimodal nodes that preserve atomic evidence.",{"name":82,"@type":73,"acceptedAnswer":83},"What retrieval mechanism does MG2-RAG use for multi-hop reasoning?",{"text":84,"@type":76},"It employs a multi-granularity graph retrieval strategy that aggregates dense similarity and propagates relevance across the graph to support structured multi-hop reasoning.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]