[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-123982-en":3,"doc-seo-123982-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},123982,8796095462418,"Noah","https://ap-avatar.wpscdn.com/avatar/80000253c1241d02b47?x-image-process=image/resize,m_fixed,w_180,h_180&k=1778826106357471780",8,"Research & Report","What is it for a Machine Learning Model to Have a Capability?","The paper addresses what contemporary machine learning models can do, focusing on how to evaluate “capabilities” as models proliferate across society. It argues that the meaning of model ability is rarely examined, especially regarding what evidence supports claims that a model can perform an action. Using large language models as an example, it proposes a conditional analysis of model abilities (CAMA) and operationalizes it for LLM evaluation. It further uses CAMA to interpret evaluation practices and enable fair inter-model comparisons.","Defining Model Capabilities  \narXiv :2405 .08989v1 [ cs .AI] 14 May 2024  \nWhat is it for a Machine Learning Model to Have a  \nCapability?  \nJacqueline Harding􀀃 [hardingj@stanford.edu](hardingj@stanford.edu)  \nDepartment of Philosophy Stanford University  \nNathaniel Sharadin sharadin@hku.hk  \nDepartment of Philosophy University of Hong Kong  \nAbstract  \nWhat can contemporary machine learning (ML) models do? Given the proliferation of ML models in society, answering this question matters to a variety of stakeholders, both public and private. The evaluation of models’ capabilities is rapidly emerging as a key sub􀀌eld of modern ML, buoyed by regulatory attention and government grants. Despite this, the notion of an ML model possessing a capability has not been interrogated: what are we saying when we say that a model is able to do something? And what sorts of evidence bear upon this question?  \nIn this paper, we aim to answer these questions, using the capabilities of large language models (LLMs) as a running example. Drawing on the large philosophical literature on abilities, we develop an account of ML models’ capabilities which can be usefully applied to the nascent science of model evaluation. Our core proposal is a conditional analysis of model abilities (CAMA): crudely, a machine learning model has a capability to X just when it would reliably succeed at doing X if it ‘tried’. The main contribution of the paper is making this proposal precise in the context of ML, resulting in an operationalisation of CAMA applicable to LLMs. We then put CAMA to work, showing that it can help make sense of various features of ML model evaluation practice, as well as suggest procedures for performing fair inter-model comparisons.  \nKeywords: arti􀀌cial intelligence, machine learning, ability modals, capabilities, evaluation, benchmarks  \n1 Introduction  \nAs machine learning (ML) models proliferate, there is increased focus on their capabilities. This is especially true for general-purpose models such as large language models (LLMs) .1 Many di􀀋erent stakeholders, including policymakers and regulators, activists,  \n∗ . Correspondence to JH. JH and NS formulated the ideas in the paper together. JH wrote the bulk of the paper.  \n1. General-purpose models (sometimes called ‘foundation’ models (Bommasani et al., 2022)) are contrasted with single-purpose ML models, which have been trained to perform a single well-de􀀌ned task, e.g. , playing Go (Silver et al., 2017) . The distinction between single and general-purpose models is not precise. But it is su􀀎ciently well understood in this context, since there is consensus that LLMs are generalpurpose models, and LLMs are our focus in this paper. For readability’s sake, we drop the ‘generalpurpose’ pre􀀌x in what follows.  \nHarding and Sharadin  \nML researchers, and end-users themselves have an interest in understanding what, exactly, ML models are able to do.2  \nWe might be interested in evaluating whether models can pass the bar exam (Katz et al. ,  \n2023; OpenAI, 2023), produce photo-realistic images of Pope Francis wearing a pu􀀋er jacket (Vincent, 2023), defeat ninth-dan humans at games like Go (Silver et al., 2017), or deceive humans while playing the game Diplomacy (Meta, 2022; Meta Fundamental AI Research Team et al. ,  \n2022) . Many capabilities of interest relate to the safety of model deployment: we want to know if models can produce hate speech (Hacker et al., 2023), generate targeted misinformation (Benson, 2023), produce CSAM (Burgess, 2023), design novel toxic molecules (Urbina et al., 2022), enable bad actors to more easily develop novel pathogens (Lloyd et al. , 2023), or write e􀀋ective phishing emails (Hazell, 2023) . Furthermore, we are told that ML models’ capabilities may be dangerous, harmful, or bene􀀌cial (Shevlane et al. , 2023), emergent (Wei et al. , 2022a), autonomous (OpenAI, 2023), surprising (Lee et al. , 2023), novel (Sheynin et al., 2022), that they have abilities in chemistr","cbCaiuLIFQ3c0Ixb","https://ap.wps.com/l/cbCaiuLIFQ3c0Ixb","pdf",455880,1,38,"English","en",105,"# Introduction\n# Preliminaries: Ability Modals and the Science of Model Evaluation\n## Ability Modals","[{\"question\":\"What is the main question of the paper about model capabilities?\",\"answer\":\"It asks what it means to say an ML model has a capability—what the capability claim expresses and what kinds of evidence support it.\"},{\"question\":\"Why does the paper use large language models (LLMs) as an example?\",\"answer\":\"It uses LLM capabilities to develop and demonstrate an account that can be applied to the broader science of evaluating ML models.\"},{\"question\":\"What is the core proposal called in the paper?\",\"answer\":\"The paper proposes a conditional analysis of model abilities (CAMA), where a model has a capability to X when it would reliably succeed at doing X if it “tried”.\"}]","What is it for a Machine Learning Model to Have a Capability? | PDF",1785819596,96,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"what-is-it-for-a-machine-learning-model-to-have-a-capability","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/what-is-it-for-a-machine-learning-model-to-have-a-capability/123982/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-04",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What is the main question of the paper about model capabilities?","Question",{"text":75,"@type":76},"It asks what it means to say an ML model has a capability—what the capability claim expresses and what kinds of evidence support it.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"Why does the paper use large language models (LLMs) as an example?",{"text":80,"@type":76},"It uses LLM capabilities to develop and demonstrate an account that can be applied to the broader science of evaluating ML models.",{"name":82,"@type":73,"acceptedAnswer":83},"What is the core proposal called in the paper?",{"text":84,"@type":76},"The paper proposes a conditional analysis of model abilities (CAMA), where a model has a capability to X when it would reliably succeed at doing X if it “tried”.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]