[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83319-en":3,"doc-seo-83319-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83319,1374391974585,"Genevieve","https://ap-avatar.wpscdn.com/davatar_276721f389ce27ea32af1340a28f341c",8,"Research & Report","CausalDS Benchmarking Causal Reasoning in Data-Science Agents","Large language models increasingly serve as data-science agents, combining abstract reasoning with tool-based workflows. Existing benchmarks often separate symbolic causal reasoning from practical data analysis, and causal datasets frequently rely on curated examples that limit novelty. CausalDS provides scenes built from sampled structural causal models with generated observational data and a realistic natural-language story. It derives tasks across Pearl’s rungs, includes coding and uncertainty quantification under imperfect observations, and scores abstention when causal claims are not warranted.","CAUSALDS: BENCHMARKING CAUSAL REASONING IN DATA-SCIENCE AGENTS  \nAndrej Leban  \nDepartment of Statistics University of Michigan Ann Arbor, MI, United States [leban@umich.edu](leban@umich.edu)  \nYuekai Sun  \nDepartment of Statistics University of Michigan Ann Arbor, MI, United States [yuekai@umich.edu](yuekai@umich.edu)  \narXiv :2607 .08093v 1 [ cs .AI] 9 Jul 2026  \nAbstract  \nLarge language models (LLMs) increasingly act as integrated data-science agents, combining abstract reasoning with advanced tool use. Yet the relevant benchmark landscape largely divides into symbolic causal reasoning benchmarks without realistic data analysis or data analysis benchmarks without a principled causal data-generating structure. Furthermore, existing causal evaluation datasets are often restricted to curated examples from existing sources, with diversity coming from limited templatized variations rather than from systematic generation of novel synthetic causal structures. We introduce CausalDS, a benchmark for evaluating causal reasoning in agentic data-science workflows. Each benchmark instance is a scene consisting of a sampled structural causal model (SCM) with generated observational data and an accompanying synthetic natural-language story grounded in a realistic domain. We optionally ground the composition of the benchmark components in empirical distributions obtained from real-world datasets, thus retaining empirical structure while reducing the “causal parrot” risk through completely synthetic generation. From each scene, we then derive tasks spanning all three of Pearl’s rungs, with typical data-science prediction tasks appearing as Rung 1 . Most tasks include a data science coding component, where the model typically needs to use several tools to arrive at the final answer due to the frequent presence of imperfect observations, which are generated by an observation model. Additionally, recognizing when a question admits no warranted answer and abstaining is treated as a first-class scored outcome. The benchmark thus jointly evaluates symbolic causal reasoning, data science, uncertainty quantification, abstention, and tool use/coding.  \n1 Introduction  \nModern LLMs are increasingly powerful in agentic settings and are routinely used in data-science workflows (Jing et al., 2025; Chan et al., 2025; Gu et al., 2024; Majumder et al., 2025) . Their actual causal reasoning capabilities, however, remain contentious (Zecevic et al., 2023; Jin et al., 2023) . In realistic causal data science, the relevant task is not merely to answer a causal question in text. An analyst must interpret a domain description, reason about the implied causal structure, inspect observational data, decide what is identifiable, and then either estimate the target quantity or decline to answer when the available information is insufficient.  \nCausalDS 1 evaluates this setting along five axes that are usually tested separately. The first is symbolic causal reasoning: interpreting a causal scenario, reasoning over graph structure, and distinguishing associational, interventional, and counterfactual targets. The second is data-science execution: using tabular data and standard analysis tools to produce estimates and predictions. The third is uncertainty quantification: attaching calibrated uncertainty to those estimates. The fourth is epistemic abstention: recognizing when the requested causal claim is not warranted by the released data and assumptions. The fifth is tool use and coding: carrying the analysis out through code in an  \n[1](1github.com/andleb/causalds)[github.com/andleb/causalds](1github.com/andleb/causalds)  \nagentic, file-backed environment. These axes are separable in principle, but realistic causal analysis requires their interaction.  \nExisting evaluations tend to isolate parts of this problem. Data-science agent benchmarks stress coding and open-ended analysis but usually lack a hidden causal data-generating structure (Lai et al., 2022; Jing et al.,","cbCaiew8zPuUkQFn","https://ap.wps.com/l/cbCaiew8zPuUkQFn","pdf",846054,3,1,55,"English","en",105,"# Introduction\n## Evaluation axes and benchmark design\n## Main contributions","[{\"question\":\"What problem does CausalDS target in evaluating agentic data-science capabilities?\",\"answer\":\"CausalDS targets the gap where benchmarks either focus on symbolic causal reasoning without realistic data analysis or focus on data-science execution without a principled causal data-generating structure. It evaluates causal reasoning within agentic data-science workflows using synthetic scenes with ground truth.\"},{\"question\":\"How is a CausalDS benchmark instance constructed?\",\"answer\":\"Each instance is a scene that samples a structural causal model (SCM), generates observational tabular data, and pairs it with a synthetic natural-language story grounded in a realistic domain.\"},{\"question\":\"What kinds of tasks and scoring outcomes does CausalDS include?\",\"answer\":\"Tasks span Pearl’s rungs, with typical data-science prediction tasks appearing as Rung 1. Many tasks require coding and tool use under imperfect observations, and the benchmark treats abstaining when no warranted causal answer exists as a first-class scored outcome.\"}]",1784186710,139,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"causalds-benchmarking-causal-reasoning-in-data-science-agents","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/causalds-benchmarking-causal-reasoning-in-data-science-agents/83319/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-22","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does CausalDS target in evaluating agentic data-science capabilities?","Question",{"text":75,"@type":76},"CausalDS targets the gap where benchmarks either focus on symbolic causal reasoning without realistic data analysis or focus on data-science execution without a principled causal data-generating structure. It evaluates causal reasoning within agentic data-science workflows using synthetic scenes with ground truth.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How is a CausalDS benchmark instance constructed?",{"text":80,"@type":76},"Each instance is a scene that samples a structural causal model (SCM), generates observational tabular data, and pairs it with a synthetic natural-language story grounded in a realistic domain.",{"name":82,"@type":73,"acceptedAnswer":83},"What kinds of tasks and scoring outcomes does CausalDS include?",{"text":84,"@type":76},"Tasks span Pearl’s rungs, with typical data-science prediction tasks appearing as Rung 1. Many tasks require coding and tool use under imperfect observations, and the benchmark treats abstaining when no warranted causal answer exists as a first-class scored outcome.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]