[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-118186-en":3,"doc-seo-118186-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},118186,8796095461564,"Liam","https://ap-avatar.wpscdn.com/davatar_155a257f0dc6eb9ab79c44ca47cae57d",8,"Research & Report","Ecosystem-level Analysis of Deployed Machine Learning Reveals Homogeneous Outcomes","Machine learning is traditionally studied at the model level by evaluating accuracy, robustness, bias, efficiency, and related properties. In real deployments, outcomes depend on the context and the collection of deployed models, so the societal impact cannot be inferred from a single model alone. This work introduces ecosystem-level analysis across text, images, and speech, finding systemic failure where some users are misclassified by all available models. Even as individual models improve, systemic failures persist. Medical imaging for dermatology shows new racial disparities in model predictions not reflected in human predictions.","Ecosystem-level Analysis of Deployed Machine Learning Reveals Homogeneous Outcomes  \nConnor Toups∗ Stanford University  \nRishi Bommasani∗† Stanford University  \narXiv :2307 .05862v2 [ cs .LG] 3 Apr 2024  \nKathleen A. Creel  \nNortheastern University  \nSarah H. Bana  \nChapman University  \nDan Jurafsky  \nStanford University  \nPercy Liang  \nStanford University  \nAbstract  \nMachine learning is traditionally studied at the model level: researchers measure and improve the accuracy, robustness, bias, efficiency, and other dimensions of specific models. In practice, however, the societal impact of any machine learning model depends on the context into which it is deployed. To capture this, we introduce ecosystem-level analysis: rather than analyzing a single model, we consider the collection of models that are deployed in a given context. For example, ecosystem-level analysis in hiring recognizes that a job candidate’s outcomes are determined not only by a single hiring algorithm or firm but instead by the collective decisions of all the firms to which the candidate applied. Across three modalities (text, images, speech) and eleven datasets, we establish a clear trend: deployed machine learning is prone to systemic failure, meaning some users are exclusively misclassified by all models available. Even when individual models improve overtime, we find these improvements rarely reduce the prevalence of systemic failure.  \nInstead, the benefits of these improvements predominantly accrue to individuals who are already correctly classified by other models. In light of these trends, we analyze medical imaging for dermatology, a setting where the costs of systemic failure are especially high. While traditional analyses reveal that both models and humans exhibit racial performance disparities, ecosystem-level analysis reveals new forms of racial disparity in model predictions that do not present in human predictions. These examples demonstrate that ecosystem-level analysis has unique strengths in characterizing the societal impact of machine learning.1  \n1 Introduction  \nMachine learning (ML) is pervasively deployed. Systems based on ML mediate our communication and healthcare, influence where we shop or what we eat, and allocate opportunities like loans and jobs. Research on the societal impact of ML typically focuses on the behavior of individual models. If we center people, however, we recognize that the impact of ML on our lives depends on the aggregate result of our many interactions with ML models.  \nIn this work, we introduce ecosystem-level analysis to better characterize the societal impact of machine learning on people. Our insight is that when a ML model is deployed, the impact on users depends not only on its behavior but also on the behavior of other models and decision-makers (left of Figure 1) . For example, the decision of a single hiring algorithm to reject or accept a candidate does  \n∗Equal contribution.  \n†Corresponding author: [nlprishi@stanford.edu](nlprishi@stanford.edu).  \n1All code is available at [https://github.com/rishibommasani/EcosystemLevelAnalysis](https://github.com/rishibommasani/EcosystemLevelAnalysis).  \n37th Conference on Neural Information Processing Systems (NeurIPS 2023) .  \nFigure 1: Ecosystem-level analysis. Individuals interact with decision-makers (left), receiving outcomes that constitute the failure matrix (right) .  \nnot determine whether or not the candidate secures a job; the outcome of her search depends on the decisions made by all the firms to which she applied. Likewise, in selecting consumer products like voice assistants, users choose from options such as Amazon Alexa, Apple Siri, or Google Assistant. From the user’s perspective, what is important is that at least one product works.  \nIn both settings, there is a significant difference from the user’s perspective between systemic failure, in which zero systems correctly evaluate them or work for them, and any other state. The difference in ","cbCailQakfKfUMoD","https://ap.wps.com/l/cbCailQakfKfUMoD","pdf",5617126,1,24,"English","en",105,"# Abstract\n# Introduction\n## Motivation: from model-level metrics to people-centered outcomes\n## Failure matrix and systemic failure\n## Audit setting across modalities and datasets","[{\"question\":\"What is ecosystem-level analysis in deployed machine learning?\",\"answer\":\"It studies the collection of models deployed in a context rather than analyzing a single model. Outcomes for users depend on combined decisions from multiple decision-makers and models.\"},{\"question\":\"What pattern does the study find across datasets?\",\"answer\":\"It finds homogeneous outcomes, including systemic failure where some users receive exclusively negative outcomes from all available models. The study also finds that average model improvements rarely reduce systemic failure prevalence.\"},{\"question\":\"How does ecosystem-level analysis change conclusions about racial disparities in medical imaging?\",\"answer\":\"Traditional analyses show performance gaps for models and humans, but ecosystem-level analysis reveals additional forms of racial disparity in model predictions that do not appear in human predictions.\"}]","Ecosystem-level Analysis of Deployed Machine Learning Reveals Homogeneous Outcomes | PDF",1785682083,60,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"ecosystem-level-analysis-of-deployed-machine-learning-reveals-homogeneous-outcomes","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/ecosystem-level-analysis-of-deployed-machine-learning-reveals-homogeneous-outcomes/118186/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-02",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What is ecosystem-level analysis in deployed machine learning?","Question",{"text":75,"@type":76},"It studies the collection of models deployed in a context rather than analyzing a single model. Outcomes for users depend on combined decisions from multiple decision-makers and models.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What pattern does the study find across datasets?",{"text":80,"@type":76},"It finds homogeneous outcomes, including systemic failure where some users receive exclusively negative outcomes from all available models. The study also finds that average model improvements rarely reduce systemic failure prevalence.",{"name":82,"@type":73,"acceptedAnswer":83},"How does ecosystem-level analysis change conclusions about racial disparities in medical imaging?",{"text":84,"@type":76},"Traditional analyses show performance gaps for models and humans, but ecosystem-level analysis reveals additional forms of racial disparity in model predictions that do not appear in human predictions.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,109,114,119,122,127,130,134],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":29,"slug":108},5,"Comic","comic",{"id":110,"doc_module":4,"doc_module_name":46,"category_name":111,"show_sort_weight":112,"slug":113},6,"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]