[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86278-en":3,"doc-seo-86278-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86278,1099514068365,"Aurelia","https://ap-avatar.wpscdn.com/avatar/10000253d8d9f28188e?_k=1776742907772140068",8,"Research & Report","Auditing the Risk Claims of Distributional Reinforcement Learning","Distributional reinforcement learning agents output full return distributions that are often consumed as genuine risk information for interpretability, risk-sensitive control, and safety monitoring. This work tests whether the resulting risk claims are actually true. A decision-relevant audit uses an excess Wasserstein gap screening metric, Monte Carlo ground truth via snapshot-restart rollouts, and rigorous statistical refutation (permutation nulls, bootstrap, and FDR control). Across QR-DQN, C51, and IQN, most top claimed trade-offs are refuted, indicating learned “risk” largely reflects training artifacts rather than environment stochasticity.","Auditing the Risk Claims of Distributional Reinforcement Learning  \nHari Prasad  \n[h4ri.prasad@gmail.com](h4ri.prasad@gmail.com)  \narXiv :2607 . 1 1607v 1 [ cs .AI] 13 Jul 2026  \nAbstract  \nDistributional reinforcement learning agents learn full return distributions that are increasingly read at face value: for interpretability, risk-sensitive control, and safety monitoring. We ask a question theory anticipates but that has not been measured directly: are the risk claims of a trained distributional agent true? Our audit combines a decision-relevant screening metric (the excess Wasserstein gap between the top two actions, which equals the mass by which first-order stochastic dominance is violated), ground truth from snapshot-restart Monte Carlo, and a statistical harness (permutation nulls, bootstrap refutation, FDR control) without which the audit itself manufactures false conclusions. Across QR-DQN, C51, and IQN on MinAtar (33 runs), 40–95% of the strongest claimed risk trade-offs are refuted at 95% confidence, the placement of the strongest claims is statistically indistinguishable from truth-blind, and essentially no claim is confirmable: for these agents, the learned “risk” reflects a training artifact rather than environment stochasticity. The artifact is structural (fully formed early in training, uncorrelated with final score, idiosyncratic to each seed) and appears unchanged at full-Atari scale, with every top Breakout claim of a pretrained near-state-of-the-art QR-DQN refuted. Positive controls of known magnitude confirm 96–100% of real claims (correlation 0 .89–0.92): the reading measures the agents, not the audit. Acting on the heads’ CVaR advice at their mostflagged states ranges from beneficial to significantly worse than chance. Neither training for risk nor ensembling removes the artifact, and recalibration passes the audit only by nullifying the claims: the head is uninformative, not merely miscalibrated. We release the toolkit and document two silent pitfalls that produced convincing but wrong audits of our own.  \nIntroduction  \nDistributional reinforcement learning (RL) replaces the scalar value function with a full distribution over returns (Bellemare, Dabney, and Munos 2017; Dabney et al. 2018b,a) . The idea has been unusually successful: distributional heads drove state-of-the-art results on the Arcade Learning Environment (Bellemare, Dabney, and Munos 2017) and anchor a growing textbook theory (Bellemare, Dabney, and Rowland 2023) . Along the way, the learned distributions stopped being treated as an internal implementation detail: they are visualized to interpret agent behavior (Greydanus et al. 2018), thresholded to produce risk-  \nsensitive policies via conditional value-at-risk (CVaR) and related functionals (Keramati et al. 2020; Lim and Malik 2022), and monitored as safety signals from autonomous driving (Hoel, Wolff, and Laine 2021) to safe-RL pipelines broadly (Garc´ıa and Fernndez 2015) . All of these uses share one assumption: that where the learned distribution of one action differs in shape from another (promising, say, a safe small return versus a gamble of equal mean), the environment actually contains that risk structure.  \nTheory gives reasons for doubt. Quantile temporaldifference (QTD) learning converges to fixed points that need not coincide with the true return distribution (Rowland et al. 2023a), and the benefits of distributional methods appear even when only the mean is used (Rowland et al. 2023b; Lyle, Bellemare, and Castro 2019), suggesting the distributions can help without being correct. The uncertaintyestimation literature has long suspected the same flaw architecturally: the spread of a quantile head conflates aleatoricand epistemic uncertainty (Clements et al. 2019), motivating a series of ensemble and evidential repairs (Hoel, Wolff, and Laine 2021; Eriksson et al. 2022; Stutts et al. 2024) . What is missing from both threads is a measurement: how wrong are the distributions","cbCaiomBiz6lreFp","https://ap.wps.com/l/cbCaiomBiz6lreFp","pdf",2064509,4,1,25,"English","en",105,"# Abstract\n# Introduction\n## Background: Distributional RL and Assumptions\n## Motivation: Why Theory Suggests Distrust\n## Measurement: Decision-Level Audit Framework\n# Contributions","[{\"question\":\"What core question does the paper ask about distributional reinforcement learning risk claims?\",\"answer\":\"The paper asks whether the risk trade-offs claimed by a trained distributional agent are actually true, rather than artifacts of the learning procedure.\"},{\"question\":\"How does the proposed audit decide which states to test?\",\"answer\":\"It screens states using an excess Wasserstein gap between the top two actions, which is positive exactly when neither action first-order stochastically dominates the other—matching when a genuine risk-sensitive deviation would be expected.\"},{\"question\":\"What do the results show about the reliability of learned risk information?\",\"answer\":\"Across multiple algorithms and games, the strongest claimed risk trade-offs are largely refuted at high confidence, and the learned “risk” behaves like a structural training artifact that is not confirmable by environment ground truth.\"}]",1784209996,63,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"auditing-the-risk-claims-of-distributional-reinforcement-learning","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/auditing-the-risk-claims-of-distributional-reinforcement-learning/86278/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-26","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What core question does the paper ask about distributional reinforcement learning risk claims?","Question",{"text":75,"@type":76},"The paper asks whether the risk trade-offs claimed by a trained distributional agent are actually true, rather than artifacts of the learning procedure.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the proposed audit decide which states to test?",{"text":80,"@type":76},"It screens states using an excess Wasserstein gap between the top two actions, which is positive exactly when neither action first-order stochastically dominates the other—matching when a genuine risk-sensitive deviation would be expected.",{"name":82,"@type":73,"acceptedAnswer":83},"What do the results show about the reliability of learned risk information?",{"text":84,"@type":76},"Across multiple algorithms and games, the strongest claimed risk trade-offs are largely refuted at high confidence, and the learned “risk” behaves like a structural training artifact that is not confirmable by environment ground truth.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]