[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86289-en":3,"doc-seo-86289-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86289,1099514068365,"Aurelia","https://ap-avatar.wpscdn.com/avatar/10000253d8d9f28188e?_k=1776742907772140068",8,"Research & Report","Bet on Features Anytime-Valid and Feature-Aware Auditing of Conditional Quantile Forecasters","Black-box conditional quantile forecasts drive sequential decisions under asymmetric costs, yet deployment requires continuous monitoring as data drift and regimes change. Standard fixed-horizon backtests fail to provide reliable calibration checks, especially because calibration is information-dependent across auditors with different feature visibility. The framework introduces distribution-free, game-theoretic anytime-valid testing for conditional quantile forecasters with non-i.i.d. losses. Evidence is interpreted at feature level, and experiments show strong miscalibration for Chronos-2 w.r.t. relevant features.","Bet on Features: Anytime-Valid and Feature-Aware Auditing of  \nConditional Quantile Forecasters  \narXiv :2607 . 1 1653v 1 [ cs .LG] 13 Jul 2026  \nIvane Antonov∗  \nJulius-Maximilians-Universität Würzburg [ivane.antonov@uni-wuerzburg.de](ivane.antonov@uni-wuerzburg.de)  \nRichard Pibernik  \nJulius-Maximilians-Universität Würzburg Zaragoza Logistics Center [richard.pibernik@uni-wuerzburg.de](richard.pibernik@uni-wuerzburg.de)  \nSohom Mukherjee∗  \nJulius-Maximilians-Universität Würzburg [sohom.mukherjee@uni-wuerzburg.de](sohom.mukherjee@uni-wuerzburg.de)  \nYo Joong Choe  \nINSEAD [yojoong.choe@insead.edu](yojoong.choe@insead.edu)  \nAbstract  \nBlack-box conditional quantile forecasts are widely used for sequential decisions under asymmetric costs, such as inventory planning in supply chain management. Once deployed, such forecasters must be monitored continuously as data streams drift and regimes change; this invalidates standard, fixed-horizon backtests for calibration. Further, existing backtests do not take into account that the notion of calibration is, in fact, information-dependent: forecasts can look calibrated to an auditor with coarse information while being miscalibrated to an auditor with richer information. We develop a distribution-free and game-theoretic testing framework for continuously auditing black-box conditional quantile forecasters with non-i.i.d. losses, such that the resulting evidence process is powerful against predictably chosen alternatives specified by the features available to the auditor. We first formalize notions of conditional quantile calibration when different sets of features are available to the auditor, establishing that the coarseness of the auditor’s information set determines the hardness of the testing problem. We then identify the sets of alternatives for which the auditor can achieve power, and focusing on contextual bets linear in the features, we derive finite-time detection guarantees for such alternatives, all without an i.i.d. assumption. The resulting evidence processes are interpretable at the feature level, as they quantify fine-grained,“feature-aware” evidence for miscalibration. We empirically validate these methods on simulated and real data, finding that a popular time series forecaster (Chronos-2) is highly miscalibrated w.r.t. multiple relevant features.  \n1 Introduction  \n“Most of the literature implicitly assumes homoskedastic errors even when this is clearly violated, and proceed by merely testing for correct unconditional coverage. Consequently, I set out to build a consistent framework for conditional interval forecast evaluation.”  \n—Christoffersen (1998)  \nIn his seminal work, Christoffersen (1998) named the requirement that has informed forecast evaluation for nearly three decades: tests must be sensitive to conditional miscalibration, because  \n∗ Equal contribution.  \nfailures that average out marginally can be predictable and exploitable in context. This perspective is increasingly relevant as large black-box models (Ansari et al. , 2025 ; Liu et al. , 2025) are used for the probabilistic forecasting of covariate-rich time series in real-world applications (e.g. , Yang et al. , 2025) . Modern deployments add a second statistical difficulty: forecasts need to be monitored continuously. Practitioners inspect calibration as outcomes arrive, stop once evidence accumulates, or intervene after a suspected distribution shift (Hoga and Demetrescu, 2023) . Fixed-horizon backtests are not designed for this workflow and can lose their nominal Type-I guarantees under optional stopping. This necessitates incorporating ideas from the safe anytime-valid inference paradigm (Ville, 1939 ; Shafer et al. , 2011 ; Vovk and Wang, 2021 ; Ramdas et al. , 2023 ; Grünwald et al. , 2024) . An audit must, therefore, satisfy two requirements at once: it must be conditional in a sense rich enough for covariate-driven failures, and it must be anytime-valid. In particular, for decision-making pro","cbCaidikfkAbNzHf","https://ap.wps.com/l/cbCaidikfkAbNzHf","pdf",2358558,3,1,39,"English","en",105,"# Abstract\n# Introduction\n## Motivation: limits of fixed-horizon backtests\n## Information-aware conditional calibration\n## Sequential betting game framework","[{\"question\":\"Why do fixed-horizon backtests become unreliable after deployment?\",\"answer\":\"Continuous monitoring with optional stopping can invalidate nominal Type-I guarantees of fixed-horizon backtests, so evidence must remain valid over time.\"},{\"question\":\"What does “information-dependent” calibration mean in this work?\",\"answer\":\"Calibration can appear correct to an auditor with coarse information while being incorrect to an auditor with richer feature access; the testing notion and alternatives depend on what features the auditor observes.\"},{\"question\":\"How does the proposed method produce evidence for miscalibration?\",\"answer\":\"It formulates an anytime-valid, distribution-free testing framework as a sequential betting game and constructs evidence processes that are interpretable at the feature level, quantifying feature-aware miscalibration.\"}]",1784210110,98,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"bet-on-features-anytime-valid-and-feature-aware-auditing-of-conditional-quantile-forecasters","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/bet-on-features-anytime-valid-and-feature-aware-auditing-of-conditional-quantile-forecasters/86289/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-26","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why do fixed-horizon backtests become unreliable after deployment?","Question",{"text":75,"@type":76},"Continuous monitoring with optional stopping can invalidate nominal Type-I guarantees of fixed-horizon backtests, so evidence must remain valid over time.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What does “information-dependent” calibration mean in this work?",{"text":80,"@type":76},"Calibration can appear correct to an auditor with coarse information while being incorrect to an auditor with richer feature access; the testing notion and alternatives depend on what features the auditor observes.",{"name":82,"@type":73,"acceptedAnswer":83},"How does the proposed method produce evidence for miscalibration?",{"text":84,"@type":76},"It formulates an anytime-valid, distribution-free testing framework as a sequential betting game and constructs evidence processes that are interpretable at the feature level, quantifying feature-aware miscalibration.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]