[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84019-en":3,"doc-seo-84019-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},84019,13056703019404,"Miles","https://ap-avatar.wpscdn.com/davatar_29158cc5080c5b710cf443261637dec0",8,"Research & Report","Mitigating Factual Hallucination in Large Reasoning Models via Mixed-Mode Advantage Regularization","Large reasoning models (LRMs) generate explicit thinking traces before producing final answers, which can enhance factuality-oriented question answering (QA). Yet instance-level analysis shows thinking may overturn correct direct answers, causing factual drift. The work frames explicit thinking as a residual over a model’s direct-answer tendency, which can either recover missing knowledge or add unsupported associations. It introduces MARGO, a reinforcement learning framework using non-thinking rollouts as same-model references in advantage estimation. Mixed-mode rollout groups suppress hallucination-prone thinking while retaining beneficial reasoning, improving factual reliability on multiple benchmarks while preserving general mathematical reasoning.","arXiv :2607 .0586 1v 1 [ cs .CL] 7 Jul 2026  \nMitigating Factual Hallucination in Large Reasoning Models via Mixed-Mode Advantage Regularization  \nKaishen Wang, Tong Zheng, Xuehao Cui, Ruibo Chen, Tianyi Xiong, Heng Huang  \nDepartment of Computer Science, University of Maryland, College Park  \n[kaishen@umd.edu](kaishen@umd.edu)  \nAbstract  \nLarge reasoning models (LRMs) improve language model capabilities by generating explicit thinking traces before final answers. In factuality-oriented question answering (QA), such thinking often improves overall performance by helping the model recover relevant knowledge and refine its answers. However, we find that this benefit is not uniform at the instance level: explicit thinking can also overturn correct non-thinking answers and lead to factual drift. We refer to this failure mode as thinking-induced hallucination. To explain this phenomenon, we formulate explicit thinking in factuality QA as a thinking residual over the model’s direct-answer tendency, which can either recover missing knowledge or introduce unsupported associations. Based on this formulation, we propose MARGO, Mixed-Mode Advantage Regularization for Grounded Optimization, a reinforcement learning framework that uses non-thinking rollouts as same-model references in advantage estimation.  \nBy constructing mixed-mode rollout groups with both thinking and non-thinking trajectories, MARGO evaluates whether explicit thinking adds factual value beyond direct answering, thereby suppressing hallucination-prone thinking while preserving beneficial thinking behaviors. Experiments across multiple factuality-oriented QA benchmarks demonstrate that MARGO improves factual reliability over strong baselines, while evaluations on mathematical benchmarks show that it preserves general reasoning ability.  \n1 Introduction  \nLarge reasoning models (LRMs) have recently achieved strong performance by generating explicit thinking traces before final answers [Jaech et al., 2024, Guo et al., 2025, Yang et al., 2025, Hou et al., 2025] . This paradigm is especially effective for mathematics, coding, and symbolic reasoning, where intermediate steps help decompose problems, verify partial results, and derive final answers [Wei et al., 2022, Wang et al., 2022, Zheng et al., 2023, Gao et al., 2023] . Explicit thinking can also benefit factuality-oriented question answering (QA), as reasoning traces may help models recall relevant knowledge, compare candidate facts, and refine final responses. As a result, thinking has become a widely used strategy for improving modern large language models (LLMs) .  \nHowever, in factuality QA, the benefit of thinking is not uniformly positive at the instance level. To examine how thinking changes model predictions, we compare non-thinking and thinking modeson TriviaQA [Joshi et al., 2017] . As shown in Table 1, on Qwen3-8B [Yang et al., 2025], thinking corrects 11.74% of examples that are originally incorrect under non-thinking, confirming its factual gains. At the same time, it overturns 7.50% of originally correct non-thinking answers into incorrect predictions. We call this failure mode thinking-induced hallucination, where explicit thinking introduces misleading or unsupported content and changes a correct direct answer into a hallucinatedone. This is particularly concerning because the correct answer is already accessible to the same model under the non-thinking mode, but is lost after generating an explicit thinking trace.  \nPreprint.  \n(a) Illustration of Thinking-Induced Hallucination  \n(b) Residual Value of Explicit Thinking  \nCorrect Answer  \nIncorrect Answer  \nResidual  \n(c) Standard GRPO vs. MARGO  \nFigure 1: Overview of MARGO. (a) Explicit thinking can induce hallucination by overturning a correct non-thinking answer into an incorrect prediction. (b) We view explicit thinking in LRMs as a residual over the model’s direct-answer tendency, which can be either helpful or harmful. (c) Unlike standard GRPO, w","cbCaietlVVxSl4Gr","https://ap.wps.com/l/cbCaietlVVxSl4Gr","pdf",1450087,4,1,19,"English","en",105,"# Abstract\n# Introduction\n## Thinking-Induced Hallucination in Factual QA\n## Residual Value View of Explicit Thinking\n## MARGO: Mixed-Mode Advantage Regularization","[{\"question\":\"What problem does MARGO address in large reasoning models?\",\"answer\":\"MARGO targets thinking-induced hallucination in factuality-oriented QA, where explicit thinking traces can overturn correct direct-answer predictions and introduce factual drift.\"},{\"question\":\"How does the paper explain why explicit thinking can be harmful?\",\"answer\":\"It models explicit thinking as a residual over the model’s direct-answer tendency: the residual can either recover missing knowledge or introduce unsupported associations that corrupt the final response.\"},{\"question\":\"What is the core idea behind MARGO’s training method?\",\"answer\":\"MARGO builds mixed-mode rollout groups containing both thinking and non-thinking trajectories for each question, using non-thinking rollouts as same-model references in advantage estimation to suppress harmful thinking while preserving beneficial behavior.\"}]",1784192050,48,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"mitigating-factual-hallucination-in-large-reasoning-models-via-mixed-mode-advantage-regularization","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/mitigating-factual-hallucination-in-large-reasoning-models-via-mixed-mode-advantage-regularization/84019/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-27","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does MARGO address in large reasoning models?","Question",{"text":75,"@type":76},"MARGO targets thinking-induced hallucination in factuality-oriented QA, where explicit thinking traces can overturn correct direct-answer predictions and introduce factual drift.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the paper explain why explicit thinking can be harmful?",{"text":80,"@type":76},"It models explicit thinking as a residual over the model’s direct-answer tendency: the residual can either recover missing knowledge or introduce unsupported associations that corrupt the final response.",{"name":82,"@type":73,"acceptedAnswer":83},"What is the core idea behind MARGO’s training method?",{"text":84,"@type":76},"MARGO builds mixed-mode rollout groups containing both thinking and non-thinking trajectories for each question, using non-thinking rollouts as same-model references in advantage estimation to suppress harmful thinking while preserving beneficial behavior.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},"General","general"]