[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86497-en":3,"doc-seo-86497-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86497,13056703019404,"Miles","https://ap-avatar.wpscdn.com/davatar_29158cc5080c5b710cf443261637dec0",8,"Research & Report","CRiT-QA：使用反事实推理链与干扰陷阱评估多跳推理","Evaluating large language models’ multi-hop reasoning remains challenging: strong results on existing QA benchmarks can hide vulnerabilities to two key issues—(1) dependence on internal parametric knowledge rather than the given context, and (2) exploitation of dataset shortcuts such as single-document cues or type matching. CRiT-QA (Counterfactual Reasoning with Traps) addresses both by converting factual reasoning chains with counterfactual entities and inserting multi-anchor distractor chains. Experiments show significant performance drops versus standard datasets, making CRiT-QA a rigorous diagnostic for evidence-grounded reasoning.","arXiv :2607 . 10562v 1 [ cs .AI] 12 Jul 2026  \nCRiT-QA: Evaluating Multi-hop Reasoning with Counterfactual Chains and Distractor Traps  \nJungMin Yun 1* , JuneHyoung Kwon 1* , YoungBin Kim 1, 2  \n1 Department of Artificial Intelligence, Chung-Ang University  \n2 Graduate School of Advanced Imaging Sciences, Multimedia and Film, Chung-Ang University  \n{cocoro357, dirchdmltnv, [ybkim85}@cau.ac.kr](ybkim85}@cau.ac.kr)  \nAbstract  \nEvaluating the multi-hop reasoning capabilities of large language models remains a significant challenge. Although current models achieve strong results on existing multi-hop question answering datasets, such performance often masks two critical vulnerabilities: (1) reliance on internal parametric knowledge rather than adherence to the provided context, and (2) exploitation of dataset shortcuts, such as single-document cues or type-matching, that diminish the need for genuine evidence aggregation across multiple documents. We introduce CRiT-QA (Counterfactual Reasoning with Traps), a dataset explicitly designed to address both limitations. To neutralize reliance on memorized knowledge and enforce strict context dependency, CRiT-QA transforms factual reasoning chains with counterfactual entities. Furthermore, it injects multi-anchor distractor chains, plausible but incorrect reasoning paths that diverge at different hops. These traps require models to follow the entire reasoning process rather than exploiting shallow heuristics. Our experiments show that LLMs exhibit substantial performance degradation on CRiT-QA compared to standard datasets, exposing their vulnerability to counterfactual conditions and distractor traps. CRiT-QA thus serves as a rigorous diagnostic tool for evaluating genuine multi-hop reasoning and provides a foundation for developing more reliable, evidence-grounded LLMs.  \nKeywords: Question Answering, Multi-Hop Reasoning, Large Language Models  \n1. Introduction  \nRetrieval-Augmented Generation (RAG) has become a dominant approach for enhancing large language models (LLMs) by grounding outputsin external knowledge sources rather than relying solely on internal parametric memory (Lewis et al. , 2020 ; Fan et al. , 2024) . This paradigm has demonstrated promising improvements in factuality and adaptability across diverse domains (Tang and Yang, 2024 ; Gao et al. , 2024) . The effectiveness of RAG systems, however, critically depends on the model’s ability to identify, link, and aggregate multiple pieces of evidence distributed across different documents (Fang et al. , 2024 ; Suryawanshi et al. , 2025) .  \nMulti-hop Question Answering (QA) has thus become a standard task for evaluating the higherorder reasoning capabilities of LLMs. In principle, such datasets are designed to assess whether models can identify and synthesize intermediate evidence to derive logically coherent final answers (Kim et al. , 2024 ; Li et al. , 2024 ; Liu et al. , 2025) . However, recent studies have revealed that the strong performance of LLMs on existing multi-hop QA datasets does not necessarily reflect genuine reasoning ability (Wu et al. , 2024a; Jiang et al. , 2024 ; Parmar et al. , 2024) . Instead, models often exploit artifacts in the dataset or rely on shallow surface-level patterns, thereby inflat-  \n*Equal contribution.  \n\n| (A) | Question: Who founded the company that distributed the film UHF?\u003Cbr>——without any context——\u003Cbr>LLM-generated Answer: Mike Medavoy ✓ |\n| --- | --- |\n| (B) | Question: In which county is Mark Dismore’s birthplace located?\u003Cbr>Paragraph 1: Greenfield is a city in and the county seat of Hancock County, Indiana, United States, and a part of the Indianapolis metropolitan area.(...)\u003Cbr>LLM-generated Answer: Hancock County ✓\u003Cbr>Sub-question: What is the place of birth of Mark Dismore?\u003Cbr>LLM-generated Answer: Unanswerable |\n\nTable 1: Examples illustrating evaluation vulnerabilities in multi-hop QA: (A) an LLM answering correctly without context, and (B) a model exploiting single-para","cbCainMvH2W2xkh7","https://ap.wps.com/l/cbCainMvH2W2xkh7","pdf",1054341,3,1,10,"English","en",105,"# Abstract\n# Introduction\n## Retrieval-Augmented Generation and Multi-hop QA\n## Vulnerabilities: internal memory and dataset shortcuts\n## CRiT-QA design: counterfactual chains and distractor traps","[{\"question\":\"CRiT-QA主要用来评估什么能力？\",\"answer\":\"CRiT-QA用于严格评估大语言模型的多跳推理能力，重点检验模型是否真正完成证据聚合与推理过程，而不是依赖捷径。\"},{\"question\":\"为什么现有多跳QA数据集的评估结果可能不可信？\",\"answer\":\"文中指出，模型可能依赖内部记忆而不遵循给定上下文；同时也可能利用数据集中的捷径，如单文档线索或类型匹配，从而跳过跨文档的真实推理。\"},{\"question\":\"CRiT-QA通过哪些机制提升评测的严谨性？\",\"answer\":\"CRiT-QA用反事实实体转换推理链以抑制记忆依赖，并注入多锚点干扰推理链，要求模型在不同hop上都遵循完整推理路径。\"}]",1784212182,25,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"crit-qa-evaluating-multi-hop-reasoning-with-counterfactual-chains-and-distractor-traps","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/crit-qa-evaluating-multi-hop-reasoning-with-counterfactual-chains-and-distractor-traps/86497/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-26","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"CRiT-QA主要用来评估什么能力？","Question",{"text":75,"@type":76},"CRiT-QA用于严格评估大语言模型的多跳推理能力，重点检验模型是否真正完成证据聚合与推理过程，而不是依赖捷径。","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"为什么现有多跳QA数据集的评估结果可能不可信？",{"text":80,"@type":76},"文中指出，模型可能依赖内部记忆而不遵循给定上下文；同时也可能利用数据集中的捷径，如单文档线索或类型匹配，从而跳过跨文档的真实推理。",{"name":82,"@type":73,"acceptedAnswer":83},"CRiT-QA通过哪些机制提升评测的严谨性？",{"text":84,"@type":76},"CRiT-QA用反事实实体转换推理链以抑制记忆依赖，并注入多锚点干扰推理链，要求模型在不同hop上都遵循完整推理路径。","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":22,"slug":133},"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]