[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82590-en":3,"doc-seo-82590-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},82590,34359740700684,"Finn","https://ap-avatar.wpscdn.com/avatar/1f400023980c374ae676?_k=1777273430885731487",8,"Research & Report","Antaeus: Hunting Repository-Level Logic Vulnerabilities via Context-Grounded LLM Reasoning","LLM-based vulnerability detectors have strong results for memory-safety and other established bug classes, but logic vulnerabilities require inferring application-specific security invariants and handling implicit assumptions about intended behavior. Frontier agentic models struggle in large repositories because relevant invariants are buried among unrelated code. ANTAEUS introduces a repository-level pipeline that prioritizes functions, grounds LLM reasoning in explicit repository evidence, validates findings comparatively, and outputs traceable, actionable sink/condition/evidence reports. On 28 real repositories, it detects and explains 15 vulnerabilities with strong performance and comparable token cost.","Antaeus: Hunting Repository-Level Logic Vulnerabilities via Context-Grounded  \nLLM Reasoning  \nMichele Armillotta  \nUniversity College London University of Bologna  \nNicol Romandini University of Bologna  \nRebecca Montanari University of Bologna  \nLorenzo Cavallaro University College London  \narXiv :2607 .0 1 138v 1 [ cs .CR] 1 Jul 2026  \nAbstract—LLM-based vulnerability detectors have shown promising results in identifying memory-safety bugs and wellestablished vulnerability classes, where violations can often be expressed in terms of established security properties such as unsafe data propagation, bounds violations, or invalid memory accesses. Logic vulnerabilities, however, pose a fundamentally different challenge, as their identification requires inferring application-specific security invariants and often relies on implicit assumptions about intended behavior. Even frontier agentic models struggle in this setting, despite their ability to inspect and traverse large repositories, because the security invariants relevant to these vulnerabilities are often implicit and buried among large amounts of unrelated code. Motivated by this gap, we present ANTAEUS, a framework for detecting logic vulnerabilities that grounds LLM reasoning in repository-level code context. ANTAEUS follows a repositoryscale pipeline that combines function prioritization, contextgrounded reasoning, comparative validation, and structured reporting. First, it ranks functions using lightweight repositorywide security signals, directing costly LLM analysis toward the most relevant code regions and reducing model calls, cost, and triage effort. For each prioritized function, ANTAEUS grounds the model in explicit repository evidence, combining local code context with a repository-level view of the application’s functionality, security-relevant resources, and trust boundaries. This grounded input enables the model to reason about how the function is executed within the broader application rather than as an isolated snippet. ANTAEUS then identifies securitysensitive sinks, derives the safety conditions required for safe execution, and checks whether those conditions are locally satisfied. Candidate findings are subjected to comparative validation, which prunes concerns that reflect project-wide norms rather than distinctive violations. Finally, ANTAEUS reports the sinks, the violated safety conditions, and the supporting evidence, making findings specific, actionable, and traceable. We evaluate ANTAEUS on 28 real-world repositories with confirmed logic vulnerabilities and compare it against both function-level LLM analysis and frontier agentic models, including Opus 4.8 Agentic and Codex 5.4. ANTAEUS detects and explains 15 vulnerabilities, substantially outperforming stateof-the-art baselines while maintaining a comparable token usage and cost budget.  \n1. Introduction  \nLogic vulnerabilities arise from failures to enforce the application-specific security invariants that govern intended program behavior [1] . Unlike memory-safety bugs, which often manifest through invalid memory accesses, or dataflow vulnerabilities, which can often be expressed as propagation from sources to sinks, logic vulnerabilities stem from missing, misplaced, or incorrectly composed enforcement of application-specific security invariants [2],[3] . These invariants define who may perform an operation, which resources may be accessed, what information may be exposed, and under which workflow conditions an action is safe. Their detection depends on reconstructing the intended security invariants implicitly encoded by the surrounding codebase, rather than matching local code against known dangerous patterns. Logic vulnerabilities have long been recognized as an important but comparatively underexplored class of software defects [4]. Most automated vulnerability-detection techniques have focused on bug classes with clearer syntactic or semantic anchors, such as memory-safety errors, inject","cbCaifkbqh5BoHK0","https://ap.wps.com/l/cbCaifkbqh5BoHK0","pdf",888257,1,16,"English","en",105,"# Introduction\n## Logic vs. other vulnerability classes\n## Limits of existing automated approaches\n## Repository-level challenge in C/C++","[{\"question\":\"What makes logic vulnerabilities different from memory-safety or dataflow vulnerabilities?\",\"answer\":\"Logic vulnerabilities come from missing, misplaced, or incorrectly composed enforcement of application-specific security invariants. Unlike memory-safety bugs or source-to-sink dataflow issues, they often lack universal sources, canonical sinks, or obvious values to track.\"},{\"question\":\"Why do frontier agentic models perform poorly for repository-level logic vulnerability detection?\",\"answer\":\"Relevant invariants are often implicit and buried among large amounts of unrelated code. As a result, agentic models can inspect large repositories but still fail to reliably ground reasoning in the specific evidence that defines the correct security behavior.\"},{\"question\":\"How does ANTAEUS detect logic vulnerabilities end to end?\",\"answer\":\"ANTAUS ranks functions with lightweight repository-wide signals, grounds LLM reasoning in explicit repository evidence combining local context and repository-level functionality, derives required safety conditions and checks them locally, performs comparative validation to prune non-distinctive concerns, and reports sinks, violated conditions, and supporting evidence.\"}]",1784181679,40,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"antaeus-hunting-repository-level-logic-vulnerabilities-via-context-grounded-llm-reasoning","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/antaeus-hunting-repository-level-logic-vulnerabilities-via-context-grounded-llm-reasoning/82590/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What makes logic vulnerabilities different from memory-safety or dataflow vulnerabilities?","Question",{"text":75,"@type":76},"Logic vulnerabilities come from missing, misplaced, or incorrectly composed enforcement of application-specific security invariants. Unlike memory-safety bugs or source-to-sink dataflow issues, they often lack universal sources, canonical sinks, or obvious values to track.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"Why do frontier agentic models perform poorly for repository-level logic vulnerability detection?",{"text":80,"@type":76},"Relevant invariants are often implicit and buried among large amounts of unrelated code. As a result, agentic models can inspect large repositories but still fail to reliably ground reasoning in the specific evidence that defines the correct security behavior.",{"name":82,"@type":73,"acceptedAnswer":83},"How does ANTAEUS detect logic vulnerabilities end to end?",{"text":84,"@type":76},"ANTAUS ranks functions with lightweight repository-wide signals, grounds LLM reasoning in explicit repository evidence combining local context and repository-level functionality, derives required safety conditions and checks them locally, performs comparative validation to prune non-distinctive concerns, and reports sinks, violated conditions, and supporting evidence.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,119,122,127,130,134],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":28,"slug":118},7,"Healthcare","healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":45,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":45,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":45,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":45,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]