[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82078-en":3,"doc-seo-82078-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},82078,13056703019404,"Miles","https://ap-avatar.wpscdn.com/davatar_29158cc5080c5b710cf443261637dec0",8,"Research & Report","Programmers Are Poor and Overconfident Judges of LLM-Generated Assertions","Code comprehension and code review remain critical software engineering activities, and AI code generation increases their importance. Generative AI can augment code with assertions and natural-language explanations, yet effectiveness for evaluating such support is unclear. A controlled experiment with 86 Python programmers and a follow-up think-aloud study examine accuracy in judging correctness and completeness of generated assertions of varying quality and assess how explanations affect those judgments. Participants judge correct assertions reasonably well (74%) but incorrect ones poorly (49%) with statistically significant accuracy differences. Natural-language explanations show no overall benefit: low-quality explanations reduce assessment accuracy while increasing confidence, suggesting reliability artifacts may require careful evaluation rather than blind adoption.","Programmers Are Poor and Overconfident Judges of LLM-Generated Assertions  \nZhanna Kaufman  \n[zhannakaufma@umass.edu](zhannakaufma@umass.edu)[ ](zhannakaufma@umass.edu)University of Massachusetts Amherst Amherst, MA, USA  \nYuriy Brun  \n[brun@cs.umass.edu](brun@cs.umass.edu)[ ](brun@cs.umass.edu)University of Massachusetts Amherst Amherst, MA, USA  \nAdithya Murali  \n[adithyamurali@cs.wisc.edu](adithyamurali@cs.wisc.edu)[ ](adithyamurali@cs.wisc.edu)University of Wisconsin–Madison Madison, WI, USA  \nMadeline Endres  \n[mendres@umass.edu](mendres@umass.edu)[ ](mendres@umass.edu)University of Massachusetts Amherst Amherst, MA, USA  \narXiv :2607 .08885v 1 [ cs . SE] 9 Jul 2026  \nAbstract  \nCode comprehension and code review are already critically important software engineering tasks, and the rising use of AI code generation tools is only increasing that importance. Generative AI has the possibility of supporting these activities, for example by augmenting code with assertions and natural-language explanations describing code behavior. However, little is known about how effective such support may be. We conduct a controlled experiment with 86 Python programmers and a follow-up think-aloud study to examine developers’ ability to assess the correctness and completeness of generated assertions of varying quality, and to investigate how natural-language explanations influence these assessments. While programmers can somewhat accurately judge correct assertions (74% accuracy), they perform poorly when shown incorrect assertions (49% accuracy), despite reporting similar levels of confidence in both judgments. This difference in judgment accuracy is statistically significant (􀀿 \u003C 0. 001): the odds of a developer accurately judging a correct assertion was nearly three times higher than the odds of accurately judging an incorrect assertion (OR = 2.94) . Surprisingly, natural-language explanations of assertions provided no overall benefit. Furthermore, low-quality explanations could impair specification assessment accuracy (􀀿 = 0.037, 􀀤􀀧 = 0. 58) while simultaneously increasing developer confidence (􀀿 = 0.005, 3. 99/5 vs. 4. 25/5) . Our findings suggest that, contrary to common assumptions, AI assistance may not improve the reliability of code comprehension and review. More broadly, our findings highlight the importance of helping developers evaluate machine-generated reliability artifacts, in addition to generating them.  \n1 Introduction  \nGenerative AI tools have transformed software engineering, and are now a routine part of virtually every aspect of development including code generation [63], architectural design [34], debugging [63], and verification [20] . This makes code comprehension and code review, already important parts of the software engineering lifecycle, even more important, as humans are asked to be part of the loop, reviewing, adapting, and adopting automatically generated code [55, 78]. However, reviewing code is difficult, time-consuming, and cognitively demanding [5, 6] . Programmers often struggle to understand non-self-authored code [74], including LLM-generated code [69], as such code may be incorrect or misaligned with user intent [17, 77], which can be difficult to fix in practice [80] . To reduce  \nthis burden, researchers have proposed shifting developer review toward smaller reliability-oriented artifacts such as tests [17, 41], specifications [40, 81], or assertions [16] . For example, Endres et al. propose using LLMs to generate executable assertions that allow developers to review intended program behavior rather than the full implementation, reducing the amount of code that must be inspected while enabling run-time validation of generated code. Similarly, Lahiri envisions a future in which LLMs help formalize user intent through behavioral artifacts such as tests, assertions, and formal specifications [40] . By formalizing intended program behavior, such artifacts have the potential to help developers determine","cbCaivhue2nramyA","https://ap.wps.com/l/cbCaivhue2nramyA","pdf",643106,2,1,12,"English","en",105,"# Abstract\n# 1 Introduction\n## Background: AI in software engineering\n## Reliability artifacts and key assumption\n## Research questions and study focus","[{\"question\":\"How accurately can programmers judge the correctness and completeness of LLM-generated assertions?\",\"answer\":\"Programmers show relatively good accuracy for correct assertions (74%), but accuracy drops substantially for incorrect assertions (49%). They also consider both correctness and completeness when evaluating generated assertions.\"},{\"question\":\"What effect do natural-language explanations have on programmers’ judgment of assertions?\",\"answer\":\"Natural-language explanations provide no overall benefit. Low-quality explanations can impair specification assessment accuracy while simultaneously increasing developer confidence.\"},{\"question\":\"Why does the study focus on postcondition assertions?\",\"answer\":\"Postconditions have well-defined formal semantics, enabling precise measurement of correctness and completeness. The setup also reflects a realistic human-in-the-loop reliability workflow where LLMs can generate useful assertions but still require human review.\"}]",1784178096,30,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"programmers-are-poor-and-overconfident-judges-of-llm-generated-assertions","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/programmers-are-poor-and-overconfident-judges-of-llm-generated-assertions/82078/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-21","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"How accurately can programmers judge the correctness and completeness of LLM-generated assertions?","Question",{"text":75,"@type":76},"Programmers show relatively good accuracy for correct assertions (74%), but accuracy drops substantially for incorrect assertions (49%). They also consider both correctness and completeness when evaluating generated assertions.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What effect do natural-language explanations have on programmers’ judgment of assertions?",{"text":80,"@type":76},"Natural-language explanations provide no overall benefit. Low-quality explanations can impair specification assessment accuracy while simultaneously increasing developer confidence.",{"name":82,"@type":73,"acceptedAnswer":83},"Why does the study focus on postcondition assertions?",{"text":84,"@type":76},"Postconditions have well-defined formal semantics, enabling precise measurement of correctness and completeness. The setup also reflects a realistic human-in-the-loop reliability workflow where LLMs can generate useful assertions but still require human review.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":29,"slug":121},"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]