[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84907-en":3,"doc-seo-84907-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},84907,1099514068035,"Ezra","https://ap-avatar.wpscdn.com/davatar_276721f389ce27ea32af1340a28f341c",8,"Research & Report","Harnessing Code Agents for Automatic Software Verification","Formal verification guarantees software correctness but does not scale because interactive theorem provers like Coq require substantial expert effort to construct proofs. Prior LLM methods hard-code or constrain proof strategies and still verify only a limited share of target theorems. This work removes those strategy constraints by using a general LLM code agent within a verification harness that enforces soundness, completeness, and termination. Evaluations show fully automatic, full coverage results on Iris modules and strong gains versus prior LLM provers across multiple provers.","Harnessing Code Agents for Automatic Software  \nVerification  \nShuangxiang Kan  \nSingapore Management University Singapore [sxkan@smu.edu.sg](sxkan@smu.edu.sg)  \nShuanglong Kan  \nBarkhausen Institut  \nDresden, Germany[shuanglong.kan@barkhauseninstitut.org](shuanglong.kan@barkhauseninstitut.org)  \nSebastian Ertel Barkhausen Institut Dresden, Germany  \n[Sebastian.Ertel@barkhauseninstitut.org](Sebastian.Ertel@barkhauseninstitut.org)  \narXiv :2607 .0634 1v 1 [ cs .FL] 7 Jul 2026  \nAbstract—Formal verification offers the strongest available guarantee of software correctness, but it does not scale: the proofs demanded by interactive theorem provers such as Coq require enormous expert effort. Large language models (LLMs) promise to generate these proofs automatically. Existing approaches wire a fixed, human-designed proof strategy into the system and constrain the model to follow it—retrieving relevant premises and predicting tactics one step at a time, or recursively splitting a goal into subgoals by divide-and-conquer—yet still prove only a fraction of their target theorems.  \nWe show that imposing such a strategy is unnecessary, and limiting: handing the whole lemma to a general LLM code agent—for example, Claude Code [1]—free to choose its own approach, and wrapping it in a verification harness is both simpler and more effective, achieving surprisingly full coverage—every targeted lemma proved, with no failures and no Coq expert intervention. The agent writes the proofs with effective feedback and hard constraints from the harness that keep each one sound (accepted only when the prover’s kernel closes it), complete (no obligation left unproved or silently dropped), and terminating (no divergent, non-terminating tactics). We evaluate the harness + code agent for verified software development along three dimensions. (1) Core logic for software verification: on Iris, the state-of-the-art separation logic for concurrent and memorymanipulating programs, Aria proves all 4,257 lemmas of the four core modules and the 217 lemmas verifying Rust’s standard libraries (Arc, Mutex, RwLock, RefCell) built on it—both in full and fully automatically. (2) Comparison with prior LLMprovers: on reglang, where prior LLM provers manage barely one in eight, Aria proves all 318. (3) Generality across provers: on iris-lean, the unfinished Lean 4 port of Iris, it proves 72 notyet-ported lemmas, showing the approach is not specific to Coq. We conclude that a state-of-the-art model—Claude Opus 4.7—is capable of writing proofs for state-of-the-art verified software development fully and automatically.  \nIndex Terms—Empirical study, formal verification, interactive theorem proving, Coq, separation logic, Iris, LLMs, agents, automated proof synthesis  \nI. INTRODUCTION  \nSoftware defects cost the global economy trillions of dollars each year, and the systems where a single bug is most catastrophic—compilers, operating-system kernels, cryptographic libraries—are precisely the ones that most need strong correctness guarantees. Formal verification provides the strongest such guarantee available: a machine-checked proof, constructed in an interactive theorem prover (ITP) such as Coq [2], [3] or Isabelle [4], that a program meets its specification. Landmark efforts such as the CompCert verified  \ncompiler (developed roughly 20 years as of 2026) demonstrate that this level of assurance is attainable in practice. Yet verified software remains the exception rather than the rule, for one stubborn reason: the proofs must be written by hand.  \nThe obstacle is the cost of those proofs. Constructing an ITP proof requires a highly trained expert to compose tactics one step at a time, reasoning simultaneously about the program, its specification, and the prover’s underlying type theory. Nowhere is this harder than in separation logic [5] and its modern realization, Iris [6], [7], where concurrency, higher-order state, and ghost resources—auxiliary, proof-only bookkeeping, invi","cbCaiscWxW2ayRyn","https://ap.wps.com/l/cbCaiscWxW2ayRyn","pdf",296532,1,12,"English","en",105,"# Abstract\n# Introduction","[{\"question\":\"What problem does the document address in software verification?\",\"answer\":\"Formal verification offers strong guarantees, but proof construction in interactive theorem provers is costly and manual, which prevents verified software from scaling to everyday engineering.\"},{\"question\":\"How does the proposed approach differ from prior LLM-based theorem proving methods?\",\"answer\":\"Instead of imposing a fixed, human-designed proof strategy or constraining the model to follow it step-by-step, the approach lets a general LLM code agent choose its own method, while a harness enforces correctness requirements.\"},{\"question\":\"What results does the document report for the harness plus code agent?\",\"answer\":\"The evaluation reports full coverage where every targeted lemma is proved automatically, including all 4,257 lemmas in Iris core modules and verification of 217 Rust library lemmas; it also shows major improvements over prior LLM provers on other benchmarks and demonstrates generality across provers.\"}]",1784199283,30,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"harnessing-code-agents-for-automatic-software-verification","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/harnessing-code-agents-for-automatic-software-verification/84907/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the document address in software verification?","Question",{"text":75,"@type":76},"Formal verification offers strong guarantees, but proof construction in interactive theorem provers is costly and manual, which prevents verified software from scaling to everyday engineering.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the proposed approach differ from prior LLM-based theorem proving methods?",{"text":80,"@type":76},"Instead of imposing a fixed, human-designed proof strategy or constraining the model to follow it step-by-step, the approach lets a general LLM code agent choose its own method, while a harness enforces correctness requirements.",{"name":82,"@type":73,"acceptedAnswer":83},"What results does the document report for the harness plus code agent?",{"text":84,"@type":76},"The evaluation reports full coverage where every targeted lemma is proved automatically, including all 4,257 lemmas in Iris core modules and verification of 217 Rust library lemmas; it also shows major improvements over prior LLM provers on other benchmarks and demonstrates generality across provers.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,122,127,130,134],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":28,"slug":121},"research-report",{"id":123,"doc_module":4,"doc_module_name":45,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":45,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":45,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":45,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]