[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82841-en":3,"doc-seo-82841-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},82841,2336464648322,"Aria","https://ap-avatar.wpscdn.com/avatar/2200025388227c56fec?_k=1778556882303663488",8,"Research & Report","Correctness, Confidence, and Context: Framing Software Assurance in the AI Age","Software engineering’s relationship with “correctness” is complex, balancing the desire for formal proof against pragmatics, cost constraints, and the many properties beyond functional correctness. Confidence must reflect fitness for intended purpose and manage failure risk, not merely code-level claims. Generative AI shifts assurance toward statistical “probably approximately correct” outputs, limiting deterministic guarantees and excluding tacit or implicit context guiding expert judgment. A systematic framework is proposed to reason about confidence-building techniques and their cost-effective combinations.","Correctness, confidence, and context:  \nFraming software assurance in the AI age Narrative for keynote presentation at FSE 2026, July 2026  \nMary Shaw  \nSoftware & Societal Systems Dept Carnegie Mellon University  \nPittsburgh PA USA  \n[mary.shaw@cs.cmu.edu](mary.shaw@cs.cmu.edu)  \nABSTRACT  \nSoftware engineering has a complicated relationship with “correctness”. We recognize the challenges of full formal rigor as well as many required properties beyond functional correctness. Although we satisfice in practice, we are still stuck in the mindset that we could reason our way to correctness, if only we had enough information.  \nGenerative AI has introduced a new dimension to assurances: its foundation is statistical rather than formal. Traditional software engineering establishes confidence through rigorous reasoning, domain knowledge and expert judgment. In contrast, generative AI’s results are sophisticated predictions, in Valiant’s words“probably approximately correct” [32] . This inherently limits assurances about the results are to probabilistic assertions. Further, the nuances and implicit associations that guide human judgment are not accessible to its training sets, so that tacit knowledge cannot be incorporated in its models.  \nWe have many approaches for developing assurances that a software system does what it’s expected to do, though most of them focus on the specification of the code rather than the requirements for the system, let alone fitness for purpose. We have failed to develop a systematic understanding of the relative merits ofthe various approaches. I hope that generative AI will finally force us to tackle this.  \nTo that end, I will challenge us to think systematically about our assurance techniques. We need ways to make informed, reasoned choices about cost-effective combinations of approaches to developing confidence in our systems.  \nWe call ourselves software engineers. Let’s act like engineers.  \nCCS CONCEPTS  \n• Software and its engineering  \nKEYWORDS  \nSoftware correctness, software confidence, fitness for purpose, sufficient correctness,“good enough” software, software credentials, implicit context, tacit context, limitations of AI  \nThis work is licensed under a Creative Commons attribution 4.0 International License. Copyright © 2026 by Mary Shaw  \n1. Software Engineering and Correctness  \nSoftware engineering has a complicated relation with correctness. On the one hand, our roots are in symbolic, logical, precise reasoning, and we have a deep-seated urge to prove our software is correct. On the other hand, we recognize that this is impractical or even impossible for a variety of reasons ranging from intellectual complexity to the variety of properties we care about to engineering pragmatics about cost-effectiveness. The tension between formality and pragmatics is often over context—the extent to which a piece of software should be evaluated in isolation or in its complete context of use. Practically, we should seek confidence that it is fit for its intended purpose. This, of course, requires us to address how fit software for a task needs to be and how to manage the risk that it will fail [8] .  \nThe tension is heightened by modern AI, which couples great power with great uncertainty, all sitting on a statistical predictive basis rather than a rigorous mathematical one. It produces results that are, in Leslie Valiant’s description [32] probably approximately correct.  \nI’ve been pondering this tension for some time, hoping that modern AI will finally force SE to be more systematic. I’ll lay out some of the tensions about correctness within SE, discuss some inherent problems with AI, and propose a framework for reasoning about assurances. This will recognize that any software project will be required to satisfy only a subset of the large set of properties that are sometimes desirable. It will also recognize the many techniques available to us for evaluating the properties, and it will place techniqu","cbCaihR68dBZYe2l","https://ap.wps.com/l/cbCaihR68dBZYe2l","pdf",1588203,1,9,"English","en",105,"# Software Engineering and Correctness\n## Correctness versus pragmatics and context\n## The challenge raised by modern AI\n# The Landscape of “Correctness”\n## Limits of natural-language requirements\n## Implicit and tacit knowledge","[{\"question\":\"Why is proving full software correctness often impractical in software engineering?\",\"answer\":\"Software roots emphasize logical proof, but practical limits arise from intellectual complexity, the variety of desired properties, and cost-effectiveness constraints. As a result, projects often pursue confidence that software is fit for its intended purpose rather than exhaustive formal correctness.\"},{\"question\":\"How does generative AI change what software assurance can guarantee?\",\"answer\":\"Generative AI produces sophisticated statistical predictions rather than results grounded in rigorous formality. This confines assurances to probabilistic statements and prevents incorporation of tacit or implicit associations that human judgment uses.\"},{\"question\":\"What does “fitness for purpose” mean for building confidence in software systems?\",\"answer\":\"Confidence should be tied to how the software performs in its intended context and the risk it may fail. This requires addressing what fit software for the task entails and how to manage failure risk, not only whether code meets a narrow specification.\"}]",1784183363,23,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"correctness-confidence-and-context-framing-software-assurance-in-the-ai-age","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/correctness-confidence-and-context-framing-software-assurance-in-the-ai-age/82841/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why is proving full software correctness often impractical in software engineering?","Question",{"text":75,"@type":76},"Software roots emphasize logical proof, but practical limits arise from intellectual complexity, the variety of desired properties, and cost-effectiveness constraints. As a result, projects often pursue confidence that software is fit for its intended purpose rather than exhaustive formal correctness.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does generative AI change what software assurance can guarantee?",{"text":80,"@type":76},"Generative AI produces sophisticated statistical predictions rather than results grounded in rigorous formality. This confines assurances to probabilistic statements and prevents incorporation of tacit or implicit associations that human judgment uses.",{"name":82,"@type":73,"acceptedAnswer":83},"What does “fitness for purpose” mean for building confidence in software systems?",{"text":84,"@type":76},"Confidence should be tied to how the software performs in its intended context and the risk it may fail. This requires addressing what fit software for the task entails and how to manage failure risk, not only whether code meets a narrow specification.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,127,130,134],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":21,"doc_module":4,"doc_module_name":45,"category_name":124,"show_sort_weight":125,"slug":126},"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":45,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":45,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":45,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]