[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86016-en":3,"doc-seo-86016-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86016,1099514067415,"Rowan","https://ap-avatar.wpscdn.com/avatar/100002539d78ffe74a7?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779092875211072502",8,"Research & Report","To Answer or to Abstain: Mitigating Search-Agent Hallucinations via Abstention-Aware Reinforcement Learning","Large language models equipped with search tools and outcome-reward reinforcement learning achieve strong open-domain QA performance, yet current training largely rewards correct outputs while not penalizing fabricated answers when retrieval fails, which intensifies hallucinations. The document introduces Abstention-Aware Reinforcement Learning (AWA-RL), dynamically shaping abstention rewards using query-specific prior capability and continuous on-policy observations. It proposes RA-F1 to quantify capability–reliability trade-offs, improving precision and overall RA-F1 with limited raw-accuracy loss.","To Answer or to Abstain: Mitigating Search-Agent Hallucinations via Abstention-Aware Reinforcement Learning  \nFengji Zhang1 , Tianyu Fan2 , Yuxiang Zheng2 , Xinyao Niu2 , Chengen Huang2 , Jacky Keung1 , Bei Chen2  \n1 City University of Hong Kong 2Alibaba Group  \narXiv :2607 . 10738v 1 [ cs .LG] 12 Jul 2026  \nAbstract  \nRecent advances in equipping Large Language Models (LLMs) with search tools and outcome-reward reinforcement learning (RL) have achieved new state-of-the-art results on open-domain QA tasks. However, we argue that current training paradigms harbor a critical vulnerability: they predominantly reward correct answers but fail to penalize fabricated ones when retrieval fails, thereby implicitly exacerbating hallucinations. To address this, we propose Abstention-Aware Reinforcement Learning (AWA-RL), which dynamically shapes the abstention reward utilizing the model’s queryspecific prior capabilities and continuous onpolicy training observations. We also introduce a novel metric, RA-F1, to measure the capability-reliability trade-off. Compared to non-abstaining baselines, AWA-RL boosts absolute precision by up to 10.3% and overall RA-F1 by 2.9%, with only marginal sacrifice in raw accuracy. These results confirm that AWA-RL successfully yields highly capable and reliable search agents. The code, data, and model weights are publicly available at [https://github.com/zfj1998/AWA-RL](https://github.com/zfj1998/AWA-RL).  \n1 Introduction  \nHallucination remains a long-standing and notoriously difficult challenge in LLMs. It is challenging to ensure the factual grounding of generated output. (Bang et al., 2025 ; Kalai et al., 2025) A prevailing approach to mitigating hallucinations is providing sufficient context to ground the answers. Progressing from early Retrieval-Augmented Generation methods,(Agrawal et al., 2024 ; Béchard and Ayala, 2024) recent advancements have introduced more advanced search agents. (Li et al., 2025b,a) By leveraging outcome-reward RL, these agents learn to interact with search environments iteratively to arrive at ground-truth answers.  \nHowever, this training paradigm harbors a critical vulnerability: current RL objectives predom-  \ninantly incentivize models to output the correct answer, but lack a mechanism to reward abstention. (Jin et al., 2025) Consequently, when an agent encounters queries beyond its capability boundaries or fails to retrieve sufficient evidence, it tends to fabricate an answer rather than safely abstain. This “over-answering” phenomenon could exacerbate hallucinations. (Song et al., 2025) The critical research question thus arises: How can we equip search agents with the self-awareness to recognize their capability boundaries and confidently abstain from ungrounded answers?  \nLearning to abstain is a non-trivial objective. Existing abstention-related studies (Xu et al., 2024b ; Song et al., 2025 ; Zhou et al., 2026 ; Zheng et al., 2025 ; Xu et al., 2024a) primarily suffer from three major limitations 1 : (1) Mismatched Scenarios: They predominantly focus on single-turn QA rather than multi-turn agentic settings, where capability boundaries dynamically shift based on tool-use and search results. (2) Degradation of Capabilities: Most approaches employ post-hoc pipeline training, which first trains for capability, then tunes for abstention. This disconnected process often severely compromises the model’s original search and reasoning abilities. (3) Reliance on Heuristics: They heavily rely on synthesizing new unanswerable datasets, which inevitably introduces human heuristics that fail to generalize across models with varying intrinsic capabilities.  \nTo address these challenges, we propose AWA-RL (Abstention-aWAre Reinforcement Learning) . The core intuition behind AWA-RL is that the incentive to abstain should not be a universal constant but rather dynamically calibrated to the model’s intrinsic problem-solving capability for each specific query. At a high level, we first e","cbCaikJzFTYO6ko2","https://ap.wps.com/l/cbCaikJzFTYO6ko2","pdf",852688,2,1,24,"English","en",105,"# Introduction\n# Methodology\n## Task Formulation","[{\"question\":\"为什么现有RL训练容易导致搜索智能体产生幻觉？\",\"answer\":\"现有目标主要奖励正确答案，却缺少在检索失败时对编造内容的惩罚机制，因此模型更倾向于在无法可靠获得证据时“硬答”，而不是安全地拒答。\"},{\"question\":\"AWA-RL通过哪些机制来提升“拒答”的可靠性？\",\"answer\":\"AWA-RL首先估计模型在给定问题上的先验成功率作为拒答基线；随后用可解释的“courage”因素进行非线性乐观映射，抑制过早拒答倾向；在RL训练中实时监控实际拒答与能力估计的偏差，动态更新拒答奖励。\"},{\"question\":\"RA-F1指标用于衡量什么权衡关系？\",\"answer\":\"RA-F1用于度量能力与可靠性之间的权衡：既关注模型的作答能力，也评估其在需要拒答时的可靠性表现。\"}]",1784207813,60,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"to-answer-or-to-abstain-mitigating-search-agent-hallucinations-via-abstention-aware-reinforcement-learning","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/to-answer-or-to-abstain-mitigating-search-agent-hallucinations-via-abstention-aware-reinforcement-learning/86016/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-26","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"为什么现有RL训练容易导致搜索智能体产生幻觉？","Question",{"text":75,"@type":76},"现有目标主要奖励正确答案，却缺少在检索失败时对编造内容的惩罚机制，因此模型更倾向于在无法可靠获得证据时“硬答”，而不是安全地拒答。","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"AWA-RL通过哪些机制来提升“拒答”的可靠性？",{"text":80,"@type":76},"AWA-RL首先估计模型在给定问题上的先验成功率作为拒答基线；随后用可解释的“courage”因素进行非线性乐观映射，抑制过早拒答倾向；在RL训练中实时监控实际拒答与能力估计的偏差，动态更新拒答奖励。",{"name":82,"@type":73,"acceptedAnswer":83},"RA-F1指标用于衡量什么权衡关系？",{"text":84,"@type":76},"RA-F1用于度量能力与可靠性之间的权衡：既关注模型的作答能力，也评估其在需要拒答时的可靠性表现。","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,109,114,119,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":29,"slug":108},5,"Comic","comic",{"id":110,"doc_module":4,"doc_module_name":46,"category_name":111,"show_sort_weight":112,"slug":113},6,"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]