[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86128-en":3,"doc-seo-86128-105":29,"detail-sidebar-cat-0-en-105":90},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},86128,962075114765,"Quinn","https://ap-avatar.wpscdn.com/davatar_a8503ba1806abce46bf441b54a3ca4cd",8,"Research & Report","Are LLMs Ready for Scientific Discovery Capability-Oriented Benchmark for AI Scientists","Existing benchmarks for scientific data analysis primarily measure LLM performance through code execution or workflow completion, overlooking that scientific analysis supports distinct claim types with different assumptions and validity rules. SDABench reorganizes evaluation around six capabilities—descriptive, exploratory, inferential, predictive, causal, and mechanistic—across Biology, Chemistry, Environment, Geography, and Physics. It includes 527 real-data instances and 6,000 synthetic instances, each in multiple-choice and open-ended formats. Results on 15 LLMs show strong descriptive performance but sharp degradation on tasks requiring assumption selection, latent-process modeling, and mechanistic reasoning.","Are LLMs Ready for Scientific Discovery? A Capability-Oriented Benchmark for  \nAI Scientists  \nChuhan Shi 1 , Xiaoquan Ren 1 , Sicheng Song2 , Haobo Li3 ,  \nRui Sheng3∗, Yushi Sun3∗  \n1 Southeast University  \n2East China Normal University  \n3Hong Kong University of Science and Technology  \n[chuhanshi@seu.edu.cn](chuhanshi@seu.edu.cn), [renxiaoquan987@gmail.com](renxiaoquan987@gmail.com), [scsong@dase.ecnu.edu.cn](scsong@dase.ecnu.edu.cn),  \n[rshengac@connect.ust.hk](rshengac@connect.ust.hk), [ysunbp@connect.ust.hk](ysunbp@connect.ust.hk)  \narXiv :2607 . 1 1079v 1 [ cs .AI] 13 Jul 2026  \nAbstract  \nExisting benchmarks for scientific data analysis evaluate LLMs primarily on code execution or workflow completion, overlooking that scientific analysis serves to support distinct types of scientific claims: hypothesis exploration, statistical inference, mechanistic explanation, each with different assumptions and validity criteria. We introduce SDABench, a benchmark that reorganizes evaluation around six capabilities (descriptive, exploratory, inferential, predictive, causal, and mechanistic) across five domains (Biology, Chemistry, Environment, Geography, Physics) . SDABench comprises 527 real-data instances (SDA-Real) and 6,000 synthetic instances (SDA-Synth), each in both multiple-choice and open-ended formats, constructed through an automated pipeline. Evaluating 15 representative LLMs, we find that models handle descriptive analysis well but degrade sharply on tasks requiring assumption selection, latent-process modeling, or mechanistic reasoning. SDABench further provides a five-stage error analysis framework that locates where LLMs fail: more advanced models more reliably identify the relevant scope and variables, but still struggle to select appropriate analytical procedures, model variable relationships, and draw valid conclusions.  \n1 Introduction  \nLarge language models (LLMs) are increasingly transforming scientific discovery by assisting researchers with tasks ranging from protein structure prediction (Jumper et al. 2021) to automated hypothesis generation (Kumbhar et al. 2025; Tang et al. 2025) . At the core of this discovery pipeline is scientific data analysis (Shi et al. 2026), the phase that converts raw experimental, simulated, or observational data into validated scientific knowledge (Majumder et al. 2024a,b) . Ensuring the reliability of AI-driven science therefore requires systematically evaluating how effectively LLMs analyze and reason with scientific data.  \nHowever, existing benchmarks such as DataSciBench (Zhang et al. 2024a), KramaBench (Lai et al. 2026), and ScienceAgentBench (Chen et al. 2025) primarily evaluate scientific data analysis from a computational perspective. They assess whether a system can generate and execute analysis code or pipelines, such as selecting features and building models, scoring the results against  \n∗Corresponding author.  \nground-truth answers or task-specific success criteria. While these benchmarks have expanded the realism and scope of data analysis evaluation, they carry two critical limitations. First, they can mistake successful execution for scientific validity: computing the requested result does not mean that the result supports the intended scientific claim. Second, they hide specific reasoning failures behind aggregate scores: a single workflow score cannot reveal whether a model fails at basic description or at complex causal inference. Scientific data analysis ultimately serves not merely to build models or maximize accuracy but to support the scientific claims drawn from data. These claims take different forms, including description, hypothesis exploration, statistical inference, prediction, causal estimation, and mechanistic explanation (Leek and Peng 2015) . They require different assumptions, evidence boundaries, and validity criteria, even when they rely on similar analytical procedures. Therefore, the evaluation for scientific data analysis should be organi","cbCaik7QWgOSXkv6","https://ap.wps.com/l/cbCaik7QWgOSXkv6","pdf",6504456,1,9,"English","en",105,"# Introduction\n## Capability-oriented framework for scientific discovery\n## Benchmark construction: SDA-Real and SDA-Synth\n## Evaluation setup and findings","[{\"question\":\"Why do existing scientific data analysis benchmarks fall short for evaluating AI scientific discovery?\",\"answer\":\"They often equate successful code execution or workflow completion with scientific validity and use aggregate scores that conceal where reasoning breaks down, such as failing basic description versus complex causal inference.\"},{\"question\":\"How does SDABench evaluate LLMs for scientific discovery?\",\"answer\":\"SDABench evaluates six discovery capabilities—descriptive, exploratory, inferential, predictive, causal, and mechanistic—across five domains, using both multiple-choice and open-ended question formats with real and synthetic data instances.\"},{\"question\":\"What performance pattern is observed when evaluating representative LLMs with SDABench?\",\"answer\":\"Models handle descriptive analysis well, but degrade sharply on tasks that require choosing appropriate assumptions, modeling latent processes, or performing mechanistic reasoning; error analysis further shows remaining difficulty in selecting analytical procedures and drawing valid conclusions.\"}]",1784208711,23,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":85,"head_meta":87,"extra_data":89,"updated_unix":27},"are-llms-ready-for-scientific-discovery-capability-oriented-benchmark-for-ai-scientists","",{"@graph":35,"@context":84},[36,53,67],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/are-llms-ready-for-scientific-discovery-capability-oriented-benchmark-for-ai-scientists/86128/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":61,"encodingFormat":60,"isAccessibleForFree":62,"interactionStatistic":63},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-16",true,{"@type":64,"interactionType":65,"userInteractionCount":4},"InteractionCounter",{"@type":66},"ViewAction",{"@type":68,"mainEntity":69},"FAQPage",[70,76,80],{"name":71,"@type":72,"acceptedAnswer":73},"Why do existing scientific data analysis benchmarks fall short for evaluating AI scientific discovery?","Question",{"text":74,"@type":75},"They often equate successful code execution or workflow completion with scientific validity and use aggregate scores that conceal where reasoning breaks down, such as failing basic description versus complex causal inference.","Answer",{"name":77,"@type":72,"acceptedAnswer":78},"How does SDABench evaluate LLMs for scientific discovery?",{"text":79,"@type":75},"SDABench evaluates six discovery capabilities—descriptive, exploratory, inferential, predictive, causal, and mechanistic—across five domains, using both multiple-choice and open-ended question formats with real and synthetic data instances.",{"name":81,"@type":72,"acceptedAnswer":82},"What performance pattern is observed when evaluating representative LLMs with SDABench?",{"text":83,"@type":75},"Models handle descriptive analysis well, but degrade sharply on tasks that require choosing appropriate assumptions, modeling latent processes, or performing mechanistic reasoning; error analysis further shows remaining difficulty in selecting analytical procedures and drawing valid conclusions.","https://schema.org",{"og:url":51,"og:type":86,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":88,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":91},[92,96,100,104,109,114,119,122,126,129,133],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":93,"show_sort_weight":94,"slug":95},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":97,"show_sort_weight":98,"slug":99},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":101,"show_sort_weight":102,"slug":103},"Exam",70,"exam",{"id":105,"doc_module":4,"doc_module_name":45,"category_name":106,"show_sort_weight":107,"slug":108},5,"Comic",60,"comic",{"id":110,"doc_module":4,"doc_module_name":45,"category_name":111,"show_sort_weight":112,"slug":113},6,"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":45,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":21,"doc_module":4,"doc_module_name":45,"category_name":123,"show_sort_weight":124,"slug":125},"Religion & Spirituality",20,"religion-spirituality",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":127,"show_sort_weight":124,"slug":128},"World Cup","world-cup",{"id":130,"doc_module":4,"doc_module_name":45,"category_name":131,"show_sort_weight":130,"slug":132},10,"Lifestyle","lifestyle",{"id":134,"doc_module":4,"doc_module_name":45,"category_name":135,"show_sort_weight":105,"slug":136},19,"General","general"]