[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86163-en":3,"doc-seo-86163-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":11,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},86163,962075114101,"Seraphina","https://ap-avatar.wpscdn.com/avatar/e000253a75eb197efd?x-image-process=image/resize,m_fixed,w_180,h_180&k=1780044092746381165",8,"Research & Report","STAMP: Provenance-Guided Credit Assignment for Deep Search Agents","Reinforcement learning for deep-search agents has focused on trajectory-level scoring such as outcome correctness, citation-aware rewards, and evidence coverage, leaving supporting-document exposing actions without targeted credit. STAMP introduces provenance-guided credit assignment using a reference-based verifier over training-time evidence graphs and first-exposure attribution to trace each supported citation back to its initial action. Step credit is injected via sign-preserving advantage modulation that preserves trajectory rewards and group rankings, improving GRPO on multiple benchmarks and complementing outcome and citation rubric rewards.","STAMP: Provenance-Guided Credit Assignment for Deep Search Agents  \nKe Xu1,2,* , Han Xu1,*,†, Xinran Chen1 , Yuqian Wang1 , Zhixuan Li1 , Xiaojian Liu1 , Changwo Wu1 , Jianqiang Xia1 , Yuchen Li1  \n1Baidu Inc., 2Peking University  \n[xuke59@stu.pku.edu.cn](xuke59@stu.pku.edu.cn), [xhbj66@gmail.com](xhbj66@gmail.com)  \narXiv :2607 . 1 1 172v 1 [ cs .AI] 13 Jul 2026  \nAbstract  \nReinforcement learning for deep-search agents has largely focused on trajectory-level scoring—outcome correctness, citation-aware rewards, and evidence coverage. Yet the actions that expose supporting documents receive no targeted credit, a gap we call the reward-credit mismatch. We propose STAMP, in which a reference-based verifier judges whether each cited document supports an entity or relation in a training-time evidence graph, and firstexposure attribution traces each supported citation back to the action that first surfaced it. This step credit is injected through signpreserving advantage modulation, which redistributes advantage across steps without changing the trajectory-level reward or the relative ranking of trajectories within each group. On BrowseComp, BrowseComp-ZH, and xbenchDS, STAMP improves the GRPO baseline by +2.0/+5.5/+3.0 points under matched SFT initialization, training data, and search tools, and composes with both outcome-only and citationrubric base rewards. Component ablations confirm that the provenance-based credit signal and the sign-preserving advantage modulation each contribute to the gains.  \n1 Introduction  \nDeep search agents discover hidden entities, verify relations across webpages, and assemble cited evidence to support a final answer (Yao et al., 2023 ; OpenAI, 2025a) . Reinforcement learning has become the dominant training paradigm for improving such agents, but long-horizon search exposesa distinction that trajectory-level training tends to blur: trajectory-level scoring of a rollout differs from step-level credit assignment to the actions that produced it (Zhang et al., 2026a) .  \nMost recent progress targets the scoring side: outcome correctness, or how to score final answers, citations, and evidence coverage at the trajectory level (Jin et al., 2025 ; Gao et al., 2025 ;  \nLi et al., 2025) . Richer trajectory-level rewards—including rubric-based scoring, citation-aware rewards, entity-or relation-level evidence rewards, and multi-turn outcome supervision (Zhang et al., 2026b ; Zhao et al., 2026)—have substantially improved deep-search agents, especially when paired with synthetic data and cold-start trajectories. Our analysis is consistent with this direction, but suggests that the value of a rollout is largely determined by the evidence its steps actually produce: Fig. 1a reports accuracy rising sharply from no target evidence (18.0%) to entityonly evidence (30.8%) to relation-verified evidence (69 . 1%) . Outcome-level rewards score whether such evidence ends up cited, but do not directly reward the actions that exposed it.  \nThe complementary side—step-level evidence, namely which search or read actions first expose target entities and verify relations—remains far less exploited. In standard outcome-supervised RL, the trajectory’s correctness verdict is the only signal that survives group-relative normalization, and the resulting advantage is broadcast uniformly to every action token (Shao et al., 2025 ; Guo et al., 2025) . This broadcast hides a strong mismatch between outcomes and evidence-producing actions: Fig. 1b shows that 83.5% of steps in correct rollouts produce no evidence at all, while 7.0% of steps in incorrect rollouts still expose useful entity-or relation-level evidence. We call this form of creditassignment gap the reward-credit mismatch: the outcome channel captures whether a trajectory succeeds, but the evidence channel—which actions expose the supporting documents—remains an underused supervision signal.  \nOne remedy is to inject step-level reward bonuses, but in standard outcome-supervise","cbCaig36FxFx93O1","https://ap.wps.com/l/cbCaig36FxFx93O1","pdf",515834,1,15,"English","en",105,"# Abstract\n# Introduction\n## Reward–Credit Mismatch\n## Reference-Based Provenance and Step Credit\n## STAMP Method Overview\n## Evaluation and Results","[{\"question\":\"What is the reward-credit mismatch in deep-search agent training?\",\"answer\":\"Trajectory-level rewards evaluate whether a rollout succeeds and whether evidence gets cited, but they do not localize credit to the specific actions that first exposed the supporting evidence. This mismatch means many steps in correct rollouts produce no evidence, while some steps in incorrect rollouts still expose useful evidence.\"},{\"question\":\"How does STAMP assign step-level credit to actions?\",\"answer\":\"STAMP uses provenance recorded by citations and a reference-based verifier to check whether each cited document supports an entity or relation in the training-time evidence graph. First-exposure attribution then traces each supported citation back to the action that first surfaced it, turning that into step credit.\"},{\"question\":\"What mechanism preserves trajectory-level reward while injecting step credit?\",\"answer\":\"STAMP injects step credit through sign-preserving advantage modulation, which redistributes advantage across steps without changing the trajectory-level reward or the relative ranking of trajectories within each group.\"}]",1784209024,38,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"stamp-provenance-guided-credit-assignment-for-deep-search-agents","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/stamp-provenance-guided-credit-assignment-for-deep-search-agents/86163/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-27","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":11},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What is the reward-credit mismatch in deep-search agent training?","Question",{"text":75,"@type":76},"Trajectory-level rewards evaluate whether a rollout succeeds and whether evidence gets cited, but they do not localize credit to the specific actions that first exposed the supporting evidence. This mismatch means many steps in correct rollouts produce no evidence, while some steps in incorrect rollouts still expose useful evidence.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does STAMP assign step-level credit to actions?",{"text":80,"@type":76},"STAMP uses provenance recorded by citations and a reference-based verifier to check whether each cited document supports an entity or relation in the training-time evidence graph. First-exposure attribution then traces each supported citation back to the action that first surfaced it, turning that into step credit.",{"name":82,"@type":73,"acceptedAnswer":83},"What mechanism preserves trajectory-level reward while injecting step credit?",{"text":84,"@type":76},"STAMP injects step credit through sign-preserving advantage modulation, which redistributes advantage across steps without changing the trajectory-level reward or the relative ranking of trajectories within each group.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":45,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":45,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":45,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]