[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82187-en":3,"doc-seo-82187-105":28,"detail-sidebar-cat-0-en-105":90},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":11,"language":21,"language_code":22,"site_id":23,"html_lang":22,"table_of_contents":24,"faqs":25,"seo_title":13,"seo_description":14,"update_tm":26,"read_time":27},82187,2336464648746,"Skyler","https://ap-avatar.wpscdn.com/davatar_276721f389ce27ea32af1340a28f341c",8,"Research & Report","AgentKGV Agentic LLM-RAG Framework with Two-Stage Training for the Fact Verification of Knowledge Graphs","Knowledge graphs built from large-scale corpora often include factual inaccuracies caused by noisy sources and extraction failures, creating a major bottleneck for reliable industrial deployment. AgentKGV introduces an agentic LLM-RAG framework that verifies KG triples using dynamic routing and iterative query rewriting, bridging the mismatch between compressed triples and natural-language evidence. A two-stage training strategy enhances accuracy and cost efficiency: turn-level distillation-based SFT stabilizes query rewriting and reasoning, while trajectory-level GRPO optimizes search behavior to reduce unnecessary retrieval.","AgentKGV: Agentic LLM-RAG Framework with Two-Stage Training for the Fact Verification of Knowledge Graphs  \nYumin Heo1 , Hyeon-gu Lee2 , Sumin Seo2 , Youngjoong Ko1 *  \n1 SungKyunKwan University, 2NAVER,  \nymheo1 [123@gmail.com](123@gmail.com) , [yjko@skku.edu](yjko@skku.edu)  \n{hyeongu.lee, [sumin.seo}@navercorp.com](sumin.seo}@navercorp.com)  \narXiv :2607 .09092v 1 [ cs .CL] 10 Jul 2026  \nAbstract  \nKnowledge graphs (KGs) are often automatically constructed from large-scale corpora, but they inevitably contain factual errors due to noisy sources and extraction failures, and verifying them reliably at industrial scale remains a critical challenge. To address this, we propose AgentKGV, the Agentic LLM-RAG framework for KG fact Verification, that integrates dynamic routing and iterative query rewriting, which handles surface-form mismatch in document-level retrieval. To make this framework more accurate and cost-efficient for industrial deployment, we further introduce a two-stage training strategy: turn-level distillation-based SFT that transfers reasoning ability from a large teacher model into a small model for stable query rewriting and reasoning, and trajectory-level GRPO that optimizes the search policy to reduce unnecessary retrieval at scale. On the long-tail-predicate split of the open-domain T-REx benchmark, our framework improves macro-F1 over single-turn RAG by 5.5 %p, and two-stage training does it further by 9.4 %p. GRPO also cuts the average number of search calls from 3 .24 to 1 .63 without lowering accuracy.  \n1 Introduction  \nKnowledge graphs encode entities and their relationships as triples in the form (subject, predicate, object), and they serve as core components across a wide range of knowledge-intensive applications such as search engines, recommendation systems, question answering, and decision support. In industrial settings, the demand for automatic construction of KGs from large volumes of documents has considerably grown. However, automatically constructed KGs inevitably contain incorrect information due to ambiguous sentence structures, low-reliability sources, and errors in NLP models. Therefore, the factual validity of extracted  \n* Corresponding author  \ntriples has become a critical bottleneck for the downstream reliability of industrial KGs.  \nResearchers have traditionally relied on isolated methodologies to evaluate the validity of automatically constructed triples. Graph-based methods (Bordes et al., 2013) assess validity through structural consistency, but they often fail to detect realworld factual errors because they rely solely on internal graph topology. LLM-based methods (Pan et al., 2024) offer broader semantic reasoning, but they remain vulnerable to domain-specific hallucinations. RAG-based methods (Lewis et al., 2020 ; Trivedi et al., 2023) utilize external document retrieval, but they fail when the system cannot retrieve the relevant information. A further difficulty is the structural modality gap between compressed triples and natural language documents, because factual information in documents rarely appears in the same standardized form as KG triples and this makes single-round retrieval unreliable.  \nTo overcome these limitations, recent research has adopted Agentic LLM frameworks (Yao et al., 2023 ; Schick et al., 2023), in which autonomous agents dynamically orchestrate reasoning and retrieval. On this paradigm, we propose AgentKGVin which the agent first decides whether it verifies a triple through internal parametric knowledge or through external retrieval through dynamic routing. When external retrieval is necessary, the agent then iteratively rewrites an input triple into natural language queries that are compatible with document-level retrieval. Because this iterative rewriting transforms the compressed triple into diverse natural language expressions, it grounds the verification in retrieved evidence.  \nHowever, deploying this framework reliably and efficiently in ind","cbCaikMMDB0aWPYV","https://ap.wps.com/l/cbCaikMMDB0aWPYV","pdf",575402,1,"English","en",105,"# Introduction\n## Challenges in KG fact verification\n## Existing approaches and limitations\n## AgentKGV framework overview\n## Two-stage training strategy\n## Contributions","[{\"question\":\"What problem does AgentKGV address in knowledge graph construction?\",\"answer\":\"AgentKGV targets factual errors in automatically constructed knowledge graphs caused by noisy sources, extraction failures, and ambiguous sentence structures, which undermine downstream reliability.\"},{\"question\":\"How does AgentKGV perform fact verification of a knowledge graph triple?\",\"answer\":\"The agent decides whether to verify using internal parametric knowledge or via external retrieval using dynamic routing, then iteratively rewrites the triple into natural-language queries aligned with document-level retrieval evidence.\"},{\"question\":\"What is the purpose of the two-stage training strategy?\",\"answer\":\"Turn-level distillation-based SFT transfers stable query rewriting and reasoning ability from a large teacher to a smaller model, while trajectory-level GRPO optimizes the search policy (including a per-turn search penalty) to reduce unnecessary retrieval and learn when to stop.\"}]",1784178688,20,{"code":4,"msg":29,"data":30},"ok",{"site_id":23,"language":22,"slug":31,"title":13,"keywords":32,"description":14,"schema_data":33,"social_meta":85,"head_meta":87,"extra_data":89,"updated_unix":26},"agentkgv-agentic-llm-rag-framework-with-two-stage-training-for-the-fact-verification-of-knowledge-graphs","",{"@graph":34,"@context":84},[35,52,67],{"@type":36,"itemListElement":37},"BreadcrumbList",[38,42,46,49],{"item":39,"name":40,"@type":41,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":43,"name":44,"@type":41,"position":45},"https://docshare.wps.com/document/","Document",2,{"item":47,"name":12,"@type":41,"position":48},"https://docshare.wps.com/document/research-report/",3,{"item":50,"name":13,"@type":41,"position":51},"https://docshare.wps.com/document/agentkgv-agentic-llm-rag-framework-with-two-stage-training-for-the-fact-verification-of-knowledge-graphs/82187/",4,{"url":50,"name":13,"@type":53,"author":54,"headline":13,"publisher":56,"fileFormat":59,"inLanguage":22,"description":14,"dateModified":60,"datePublished":61,"encodingFormat":59,"isAccessibleForFree":62,"interactionStatistic":63},"DigitalDocument",{"name":9,"@type":55},"Person",{"url":39,"name":57,"@type":58},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":64,"interactionType":65,"userInteractionCount":20},"InteractionCounter",{"@type":66},"ViewAction",{"@type":68,"mainEntity":69},"FAQPage",[70,76,80],{"name":71,"@type":72,"acceptedAnswer":73},"What problem does AgentKGV address in knowledge graph construction?","Question",{"text":74,"@type":75},"AgentKGV targets factual errors in automatically constructed knowledge graphs caused by noisy sources, extraction failures, and ambiguous sentence structures, which undermine downstream reliability.","Answer",{"name":77,"@type":72,"acceptedAnswer":78},"How does AgentKGV perform fact verification of a knowledge graph triple?",{"text":79,"@type":75},"The agent decides whether to verify using internal parametric knowledge or via external retrieval using dynamic routing, then iteratively rewrites the triple into natural-language queries aligned with document-level retrieval evidence.",{"name":81,"@type":72,"acceptedAnswer":82},"What is the purpose of the two-stage training strategy?",{"text":83,"@type":75},"Turn-level distillation-based SFT transfers stable query rewriting and reasoning ability from a large teacher to a smaller model, while trajectory-level GRPO optimizes the search policy (including a per-turn search penalty) to reduce unnecessary retrieval and learn when to stop.","https://schema.org",{"og:url":50,"og:type":86,"og:title":13,"og:site_name":57,"og:description":14},"article",{"robots":88,"canonical":50},"index,follow",{"doc_id":7,"site_id":23},{"code":4,"msg":5,"data":91},[92,96,100,104,109,114,119,122,126,129,133],{"id":20,"doc_module":4,"doc_module_name":44,"category_name":93,"show_sort_weight":94,"slug":95},"Story & Novel",90,"story-novel",{"id":45,"doc_module":4,"doc_module_name":44,"category_name":97,"show_sort_weight":98,"slug":99},"Literature",80,"literature",{"id":51,"doc_module":4,"doc_module_name":44,"category_name":101,"show_sort_weight":102,"slug":103},"Exam",70,"exam",{"id":105,"doc_module":4,"doc_module_name":44,"category_name":106,"show_sort_weight":107,"slug":108},5,"Comic",60,"comic",{"id":110,"doc_module":4,"doc_module_name":44,"category_name":111,"show_sort_weight":112,"slug":113},6,"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":44,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":44,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":44,"category_name":124,"show_sort_weight":27,"slug":125},9,"Religion & Spirituality","religion-spirituality",{"id":27,"doc_module":4,"doc_module_name":44,"category_name":127,"show_sort_weight":27,"slug":128},"World Cup","world-cup",{"id":130,"doc_module":4,"doc_module_name":44,"category_name":131,"show_sort_weight":130,"slug":132},10,"Lifestyle","lifestyle",{"id":134,"doc_module":4,"doc_module_name":44,"category_name":135,"show_sort_weight":105,"slug":136},19,"General","general"]