[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-81615-en":3,"doc-seo-81615-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},81615,34359740700684,"Finn","https://ap-avatar.wpscdn.com/avatar/1f400023980c374ae676?_k=1777273430885731487",8,"Research & Report","From Cross-Validation to SURE: Asymptotic Risk of Tuned Regularized Estimators","Derives an asymptotic risk characterization for regularized empirical risk minimization estimators whose tuning parameters are selected via n-fold cross-validation. The out-of-sample prediction loss converges in distribution to squared-error risk of corresponding shrinkage estimators in the normal means model, with tuning driven by Stein’s unbiased risk estimate (SURE). The resulting risk function yields a parameter-dependent view of predictive performance beyond worst-case regret bounds, and is built on uniform CV-to-SURE convergence plus generic separation of SURE’s global minimum.","arXiv :2603 .20388v2 [math . ST] 10 Jul 2026  \nFrom Cross-Validation to SURE: Asymptotic Risk of Tuned Regularized Estimators  \nKarun Adusumilli∗ Maximilian Kasy† Ashia Wilson‡  \nJuly 13, 2026  \nAbstract  \nWe derive the asymptotic risk function of regularized empirical risk minimization (ERM) estimators tuned by n-fold cross-validation (CV) . The out-of-sample prediction loss of such estimators converges in distribution to the squared-error loss (risk function) of shrinkage estimatorsin the normal means model, tuned by Stein’s unbiased risk estimate (SURE) . This risk function provides a more fine-grained picture of predictive performance than uniform bounds on worst-case regret, which are common in learning theory: it quantifies how risk varies with the true parameter.  \nAs key intermediate steps, we show that (i) n-fold CV converges uniformly to SURE, and (ii) while SURE typically has multiple local minima, its global minimum is generically well separated. Well-separation ensures that uniform convergence of CV to SURE translates into convergence of the tuning parameter chosen by CV to that chosen by SURE.  \n∗ Department of Economics, University [of Pennsylvania. akarun@sas.upenn.edu](of Pennsylvania. akarun@sas.upenn.edu).  \n†Corresponding author. Department of Economics, University of Oxford. maximil[ian.kasy@economics.ox.ac.uk. Maximilian Kasy was](ian.kasy@economics.ox.ac.uk. Maximilian Kasy was) supported by the Alfred P. Sloan Foundation, under the grant “Social foundations for statistics and machine learning.”  \n‡Department of Electrical Engineering and Computer Science, MIT.  \n1 Introduction  \nBackground The goal of supervised learning is to produce good predictions for new observations.1 An important class of estimators for supervised learning are regularized empirical risk minimization (ERM) estimators that are tuned using cross-validation (CV) .  \nERM estimators minimize in-sample average prediction loss (empirical risk), among a given class of predictors. Examples include ordinary least squares and maximum likelihood. ERM estimators are prone to overfitting if the class of predictors is large. Such estimators achieve low in-sample loss, but can perform poorly for new observations. To counter overfitting, regularization is used. Regularization adds a penalty term to the ERM objective; common penalties include the L2 norm of parameters (in Ridge regression) and the L 1 norm (in Lasso regression) . Adding a penalty avoids overfitting, by reducing the estimator variance, at the cost of introducing some bias, which might result in underfitting.  \nTo achieve good performance, avoiding both overfitting and underfitting, the amount of penalization needs to be carefully tuned. This can be done by choosing weights for the penalty that minimize a cross-validation estimate of predictive loss. We focus on n-fold CV, where predictions are evaluated for one hold-out observation at a time, and predictive loss is estimated by averaging evaluations over each of the n observations.  \nRisk functions The present paper characterizes the behavior of estimators of this form by deriving an asymptotic approximation to their risk function. A large literature in learning theory characterizes such estimators by proving bounds on their worst-case regret—the supremum over data-generating processes (DGPs) of the difference between an estimator’s risk and the expected loss of the best predictor in the given class. Such bounds provide strong robustness guarantees, but they might not be informative about the behavior of predictive algorithms for realistic DGPs. By focusing on the risk function,  \n1We would like to dedicate this paper to Gary Chamberlain, whose conversations provided the original inspiration for this project.  \nFigure 1: Risk function for JS-shrinkage, dimension 10  \n1 2 3 4 5 6  \n∥θ∥  \nwe obtain a more fine-grained characterization: the risk function tells us how expected predictive performance depends on the DGP, while worst-case ","cbCaicJZwpE0ZMlu","https://ap.wps.com/l/cbCaicJZwpE0ZMlu","pdf",597051,4,1,70,"English","en",105,"# Abstract\n# Introduction\n## Background and problem setup\n## Risk functions and motivation\n## Main result\n## Key steps","[{\"question\":\"What does the paper derive about cross-validation-tuned regularized estimators?\",\"answer\":\"It derives an asymptotic risk function, showing that the out-of-sample prediction loss of n-fold CV-tuned regularized ERM estimators converges in distribution to the squared-error loss of a corresponding SURE-tuned shrinkage estimator in the normal means model.\"},{\"question\":\"How is SURE connected to the tuning performed by cross-validation?\",\"answer\":\"The paper proves that n-fold cross-validation converges uniformly to SURE and that the tuning parameter chosen by CV converges to the tuning parameter chosen by SURE, relying on well-separation of SURE’s global minimum.\"},{\"question\":\"Why do the authors focus on risk functions instead of worst-case regret bounds?\",\"answer\":\"Risk functions provide a finer, parameter-dependent characterization of expected predictive performance under realistic data-generating processes, whereas worst-case regret only describes behavior under the least favorable processes.\"}]",1784174793,176,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"from-cross-validation-to-sure-asymptotic-risk-of-tuned-regularized-estimators","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/from-cross-validation-to-sure-asymptotic-risk-of-tuned-regularized-estimators/81615/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-25","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What does the paper derive about cross-validation-tuned regularized estimators?","Question",{"text":75,"@type":76},"It derives an asymptotic risk function, showing that the out-of-sample prediction loss of n-fold CV-tuned regularized ERM estimators converges in distribution to the squared-error loss of a corresponding SURE-tuned shrinkage estimator in the normal means model.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How is SURE connected to the tuning performed by cross-validation?",{"text":80,"@type":76},"The paper proves that n-fold cross-validation converges uniformly to SURE and that the tuning parameter chosen by CV converges to the tuning parameter chosen by SURE, relying on well-separation of SURE’s global minimum.",{"name":82,"@type":73,"acceptedAnswer":83},"Why do the authors focus on risk functions instead of worst-case regret bounds?",{"text":84,"@type":76},"Risk functions provide a finer, parameter-dependent characterization of expected predictive performance under realistic data-generating processes, whereas worst-case regret only describes behavior under the least favorable processes.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,104,109,114,119,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":22,"slug":103},"Exam","exam",{"id":105,"doc_module":4,"doc_module_name":46,"category_name":106,"show_sort_weight":107,"slug":108},5,"Comic",60,"comic",{"id":110,"doc_module":4,"doc_module_name":46,"category_name":111,"show_sort_weight":112,"slug":113},6,"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":105,"slug":137},19,"General","general"]