[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86218-en":3,"doc-seo-86218-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86218,1374391974564,"Clementine","https://ap-avatar.wpscdn.com/avatar/14000253aa45c000a9e?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779874745381141002",8,"Research & Report","Predicting Program Comprehension with Foundation Models of Human Cognition","Software engineering relies on developers understanding code, yet predicting comprehension behavior remains difficult because existing solutions trade off accuracy, scalability, or experimental realism. Drawing on psychology, the paper models human actions through general cognitive regularities learned from large-scale behavioral data. It evaluates Centaur, trained on 160 psychological experiments, across nine program-comprehension studies, comparing predicted response distributions against humans and benchmarking against Llama 3.1.","Predicting Program Comprehension with Foundation Models of Human Cognition  \nYannick Lehmen, Marvin Wyrich, Anna-Maria Maurer, Norman Peitek, Sven Apel  \nSaarland Informatics Campus, Saarland University  \nSaarbr¨ucken, Germany  \narXiv :2607 . 1 1372v 1 [ cs . SE] 13 Jul 2026  \nAbstract—Software engineering depends on the ability of developers to understand code, yet predicting how they do so remains an open challenge despite decades of research. Existing approaches rely either on simplified proxy measures that limit accuracy or on non-trivial measurements requiring elaborate experimental setups that are difficult to scale and apply in practice. In contrast, recent work in psychology suggests an alternative perspective: Instead of modeling task-specific phenomena directly, human behavior can be captured through general cognitive regularities learned from large-scale behavioral data. This idea treats complex human behavior as the observable outcome of underlying cognitive processes that manifest consistently across tasks and domains.  \nIn this paper, we explore this perspective in the context of program comprehension. We evaluate Centaur, a foundation model trained on 160 general psychological experiments, on 9 previously published program-comprehension studies. We assess how well its predicted response distributions align with human response data and compare Centaur’s performance to its base model, Llama 3.1. To better understand the source of its performance, we subsequently conduct ablation studies to isolate the contribution of different sources of information, such as the code artifacts, task-related context, and prior trials and participant responses.  \nIn a nutshell, we find that Centaur more closely aligns with human response patterns than its base model, is significantly less reliant on information from prior trials and responses, and benefits more from task-related information. These findings suggest that behavioral patterns learned from general psychological data can transfer to complex software engineering tasks such as program comprehension. More broadly, they point toward foundation models of human cognition as a basis for modeling developer behavior in software engineering, opening a pathway toward a unified, data-driven perspective on human-centered software engineering grounded in cognitive science.  \nIndex Terms—program comprehension, foundation model, cognitive modeling, developer behavior  \nI. INTRODUCTION  \nUnderstanding how developers comprehend code is central to improving software quality and developer productivity [1]–[3] . A wide range of approaches has been proposed to assess and predict the process of program comprehension, including controlled experiments with developers [4], proxy measures such as complexity metrics [5], and neurocognitive methods that measure brain activity during programming tasks [6] . These approaches reflect different operationalizations of program comprehension, each capturing particular aspects of how humans engage with code.  \nYet, despite this diversity of approaches, reliably predicting how developers will respond in program comprehension tasks  \nremains an open challenge: Controlled experiments with programmers provide precise insights but do not scale beyond specific experimental settings, while automated proxy measurements such as complexity metrics offer scalability but show only weak correlations with observed human responses [5],[7], [8] . Even more direct approaches, such as neurocognitive methods that measure brain activity during programming tasks, provide rich signals but are expensive, difficult to generalize, and limited in scope [9], [10] . Across all these approaches, a common limitation is that they do not capture transferable regularities in how humans respond, making it difficult to build predictive models of program comprehension that generalize beyond specific study designs and individuals.  \nWe look towards psychology for a new perspective: Faced with i","cbCailtXEDuifVYJ","https://ap.wps.com/l/cbCailtXEDuifVYJ","pdf",613069,3,1,12,"English","en",105,"# Introduction\n## Background and motivation\n## Related work on program comprehension\n## Foundation models of human cognition\n## Research goal","[{\"question\":\"What problem does the paper target in software engineering research?\",\"answer\":\"Predicting how developers respond during program comprehension tasks remains an open challenge despite decades of work. The paper argues current approaches struggle with accuracy, scalability, or generalization.\"},{\"question\":\"How is Centaur trained and what is its intended modeling perspective?\",\"answer\":\"Centaur is a foundation model fine-tuned on data from 160 psychological experiments to reproduce human behavior. The approach emphasizes learning general cognitive regularities from large behavioral datasets rather than task-specific measurements.\"},{\"question\":\"How do the authors evaluate Centaur for program comprehension?\",\"answer\":\"The evaluation compares Centaur’s predicted response distributions to human response data across nine previously published program-comprehension studies. It also benchmarks against its base model, Llama 3.1, and uses ablation studies to separate contributions from code artifacts, task context, and prior trial/participant responses.\"}]",1784209534,30,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"predicting-program-comprehension-with-foundation-models-of-human-cognition","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/predicting-program-comprehension-with-foundation-models-of-human-cognition/86218/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-26","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the paper target in software engineering research?","Question",{"text":75,"@type":76},"Predicting how developers respond during program comprehension tasks remains an open challenge despite decades of work. The paper argues current approaches struggle with accuracy, scalability, or generalization.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How is Centaur trained and what is its intended modeling perspective?",{"text":80,"@type":76},"Centaur is a foundation model fine-tuned on data from 160 psychological experiments to reproduce human behavior. The approach emphasizes learning general cognitive regularities from large behavioral datasets rather than task-specific measurements.",{"name":82,"@type":73,"acceptedAnswer":83},"How do the authors evaluate Centaur for program comprehension?",{"text":84,"@type":76},"The evaluation compares Centaur’s predicted response distributions to human response data across nine previously published program-comprehension studies. It also benchmarks against its base model, Llama 3.1, and uses ablation studies to separate contributions from code artifacts, task context, and prior trial/participant responses.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":29,"slug":121},"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]