[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-119634-en":3,"doc-seo-119634-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},119634,13056703020460,"Valentina","https://ap-avatar.wpscdn.com/avatar/be000253dac470eee5d?_k=1778207105932848923",8,"Research & Report","Moving from Machine Learning to Statistics - the case of Expected Points in American football","Expected points is a value function central to player evaluation and strategic in-game decision-making in sports analytics, with American football as a key setting. Estimating expected points often relies on machine learning tools that introduce selection bias, show artifacts consistent with overfitting, fail to report uncertainty in point predictions, and ignore the strong dependence structure of observational football data. The work analyzes these issues and proposes expected points models that address them, including a general approach to mitigate overfitting using a catalytic prior to smooth the learning process.","arXiv :2409 .04889v1 [ stat .AP] 7 Sep 2024  \nMoving from Machine Learning to Statistics: the case of Expected Points in American  \nfootball  \nRyan S. Brill∗, Ryan Yee†, Sameer K. Deshpande†, and Abraham J. Wyner‡  \nSeptember 10, 2024  \nAbstract  \nExpected points is a value function fundamental to player evaluation and strategic in-game decision-making across sports analytics, particularly in American football. To estimate expected points, football analysts use machine learning tools, which are not equipped to handle certain challenges. They suffer from selection bias, display counter-intuitive artifacts of overfitting, do not quantify uncertainty in point estimates, and do not account for the strong dependence structure of observational football data. These issues are not unique to American football or even sports analytics; they are general problems analysts encounter across various statistical applications, particularly when using machine learning in lieu of traditional statistical models. We explore these issues in detail and devise expected points models that account for them. We also introduce a widely applicable novel methodological approach to mitigate overfitting, using a catalytic prior to smooth our machine learning models.  \nKeywords: applications and case studies, machine learning, statistics in sports, catalytic prior  \n∗ Graduate Group in Applied Mathematics and Computational Science, University of Pennsylvania. Correspondence to: [ryguy123@sas.upenn.edu](ryguy123@sas.upenn.edu)  \n†Dept. of Statistics, University of Wisconsin–Madison  \n‡Dept. of Statistics and Data Science, The Wharton School, University of Pennsylvania  \n1 Introduction  \nSports analytics has become a multibillion-dollar industry. Each team in Major League Baseball (MLB), the National Basketball Association (NBA), and the National Football League (NFL) has at least one analytics staffer. Many of these teams have full-fledged sports analytics research groups and hire sports analytics consulting firms. The popularity of these firms, many of which consist of large teams of Ph.D.s in statistics, mathematics, and computer science, has exploded recently.  \nTwo fundamental areas of interest to these teams and firms are player evaluation and ingame strategic decision-making. The quantitative approach relies on a valuation function, typically an expected value, that measures the value of each game-state. Analysts can evaluate an individual player by the value added across each of his plays. They can also evaluate a coach by how often he makes decisions that maximize the value of the next game-state.  \nThe most prominent and widely used value function across all of sports analytics is expected points (EP) . In baseball, EP (or expected runs) is the expected number of runs scored from the current game-state through the end of the half-inning.1 In American football, which runs in continuous time, EP is the expected net number of points scored from the current game-state through the next scoring event (or the end of the half) (Yurko et al., 2018) . In soccer, EP (or expected possession value) is the likelihood that the team with possession of the ball scores the next goal minus the likelihood it concedes the next goal, given the current game-state (Fern´andez et al., 2021) . EP is defined similarly for any sport with scoring and time components.  \nIn this work, we focus on EP for American football as our primary case study for several reasons. First, the creation and development of expected points methodologies in football has been and continues to be an open source endeavor (Carter and Machol, 1971; Carroll et al., 1989; Romer, 2006; Burke, 2009) . State of the art EP models today are open source and reproducible from publicly available data (Yurko et al. , 2018; Baldwin, 2021a; Carl and Baldwin, 2022) . In contrast, soccer EP methodologies and data aren’t fully publicly available. All state of the art analyses are proprietary – fit by teams or by  \n1[https","cbCaitVbfnsp8yRQ","https://ap.wps.com/l/cbCaitVbfnsp8yRQ","pdf",5371350,1,32,"English","en",105,"# Abstract\n# Introduction\n## Sports analytics and valuation functions\n## Expected points across sports\n## Why focus on American football EP\n## Epochs and game-state modeling\n## Estimating EP from data: binning and averaging","[{\"question\":\"What is expected points and why is it important in sports analytics?\",\"answer\":\"Expected points is a value function that measures the expected net points from a given game-state to a future scoring event. It supports both player evaluation and strategic decision-making during games.\"},{\"question\":\"What limitations do analysts face when using machine learning to estimate expected points?\",\"answer\":\"Machine learning approaches can produce selection bias, overfitting-related artifacts, omit uncertainty quantification, and disregard the dependence structure typical of observational football data.\"},{\"question\":\"How does the paper address these problems for expected points in American football?\",\"answer\":\"It develops expected points models that account for the stated issues and introduces a widely applicable overfitting-mitigation method using a catalytic prior to smooth the machine learning models.\"}]","Moving from Machine Learning to Statistics - the case of Expected Points in American football | PDF",1785725399,81,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"moving-from-machine-learning-to-statistics-the-case-of-expected-points-in-american-football","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/moving-from-machine-learning-to-statistics-the-case-of-expected-points-in-american-football/119634/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-03",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What is expected points and why is it important in sports analytics?","Question",{"text":75,"@type":76},"Expected points is a value function that measures the expected net points from a given game-state to a future scoring event. It supports both player evaluation and strategic decision-making during games.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What limitations do analysts face when using machine learning to estimate expected points?",{"text":80,"@type":76},"Machine learning approaches can produce selection bias, overfitting-related artifacts, omit uncertainty quantification, and disregard the dependence structure typical of observational football data.",{"name":82,"@type":73,"acceptedAnswer":83},"How does the paper address these problems for expected points in American football?",{"text":84,"@type":76},"It develops expected points models that account for the stated issues and introduces a widely applicable overfitting-mitigation method using a catalytic prior to smooth the machine learning models.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]