[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83892-en":3,"doc-seo-83892-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83892,8796095461564,"Liam","https://ap-avatar.wpscdn.com/davatar_155a257f0dc6eb9ab79c44ca47cae57d",8,"Research & Report","Can Code Specify a System Precisely Enough to Formally Verify It","Formal verification is rarely used in everyday production software because modeling costs historically outweigh benefits. This paper evaluates a lower-cost specification approach on production code: the payment workflow of an operational restaurant point-of-sale system that must keep register, terminal, and processor consistent. Three results follow: protocol correctness relative to a stated failure model, an emulator audit failure via correlated-oracle response divergence, and replication showing contract structure—not specification language—governs reliable LLM specifications.","arXiv :2607 .05076v 1 [ cs . SE] 6 Jul 2026  \nCan Code Specify a System Precisely Enough to  \nFormally Verify It?  \nJean-Jacques Dubray  \n[jdubray@gmail. com](jdubray@gmail. com)  \nJuly 4, 2026  \nAbstract  \nFormal verification is seldom applied to everyday production software, because the cost of writing and maintaining a model has historically exceeded the benefit. A companion study [1] developed a lower-cost alternative and evaluated it on a benchmark: specifications are graded against traces captured from the running system. It found that when large language models write the specifications, their reliability is governed by the structure of the specification contract rather than by the specification language. This paper evaluates both the alternative and that finding on production software: the payment workflow of an operational restaurant point-of-sale system, which must keep the register, the payment terminal, and the payment processor in agreement. We report three results. First, the core protocol is correct relative to a hand-built, line-cited model under a failure model that we state precisely (§4.1) . The audit identified seven failure-handling gaps, nearly all with a common root cause; three were reproduced as real executions, and a patch closing them was re-checked with all failure gates enabled, after which a follow-up patch closed a further defect that the re-check itself exposed. Systematic extensions of the failure model (crash–restart, stale reads, and two attempts) each identified the windows they were designed to probe. Second, a single probe of the production payment sandbox exposed a response-shape divergence that renders an entire recovery ladder unreachable against the live API. The emulator-based audit was structurally unable to detect this divergence, because the code and the emulator share the same misreading; we characterize this as a correlated-oracle failure. Third, the companion study’s central finding replicates across seven models from two vendors: contract structure, not language, governs what LLMs specify reliably. The replication concerns the ordering of contracts and the failure taxonomy rather than the absolute level: only the strongest models reached the corpus ceiling, and the harder task restores discriminating power that the benchmark had lost.  \n1 Introduction  \nThe companion study [1] extended SysMoBench [4, 5] with its first non-formal specification backend (JavaScript in the SAM pattern, contributed upstream in pull request [2] and documented in a from-scratch walkthrough [3]) and arrived, across four experiments and successive adversarial audits, at the following ranking: on the tasks tested, LLM specification quality was governed first by contract structure, second by the prompt, and not at all by the specification language. This ranking carried an explicit caveat. The tasks were a kernel spinlock with three observable states and a small lock service, and the result was conditional on replication against a system whose transition relation cannot be inferred from its action names.  \nThis paper provides that replication, and extends it. The target is not a benchmark task but production code: the payment state-alignment workflow of a live point-of-sale system, in daily use with real cards and real transactions. The change of setting introduces a question that a benchmark cannot pose. The question is not “can models specify this system?” but “is this system correct?”, with consequences (lost charges, double charges, and short-charged orders) measured in currency rather than in benchmark points.  \nThis paper is self-contained. The three specification contracts under comparison are reproduced in Appendix A, every scoring definition is stated in §3, and §4.1 states exactly what the model checker verified, under which failure modes, with which instantiation.  \nThe system. The workflow keeps three agents in agreement; they share no memory and communicate only through narrow, failure-prone ch","cbCaiqkpmo0zwuIg","https://ap.wps.com/l/cbCaiqkpmo0zwuIg","pdf",549661,2,1,25,"English","en",105,"# Abstract\n# Introduction\n## The companion study and its ranking\n## Replication on production payment code\n## System overview and workflow constraints\n## Provenance and audit setup","[{\"question\":\"Why is formal verification uncommon in everyday production software?\",\"answer\":\"Because creating and maintaining the specification model has historically cost more than the expected benefit.\"},{\"question\":\"What does the paper evaluate beyond benchmark tasks?\",\"answer\":\"It replicates prior findings on real production payment workflow code, where correctness affects outcomes like lost or double charges rather than benchmark points.\"},{\"question\":\"What governs how reliably LLMs specify the system in this study?\",\"answer\":\"Reliability is governed by the structure of the specification contracts rather than by the specification language.\"}]",1784191269,63,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"can-code-specify-a-system-precisely-enough-to-formally-verify-it","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/can-code-specify-a-system-precisely-enough-to-formally-verify-it/83892/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-22","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why is formal verification uncommon in everyday production software?","Question",{"text":75,"@type":76},"Because creating and maintaining the specification model has historically cost more than the expected benefit.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What does the paper evaluate beyond benchmark tasks?",{"text":80,"@type":76},"It replicates prior findings on real production payment workflow code, where correctness affects outcomes like lost or double charges rather than benchmark points.",{"name":82,"@type":73,"acceptedAnswer":83},"What governs how reliably LLMs specify the system in this study?",{"text":84,"@type":76},"Reliability is governed by the structure of the specification contracts rather than by the specification language.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]