[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86206-en":3,"doc-seo-86206-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86206,1374391974564,"Clementine","https://ap-avatar.wpscdn.com/avatar/14000253aa45c000a9e?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779874745381141002",8,"Research & Report","Compile Then Page: Executable SOP Programs and a Capability-Gated Runtime for Procedural LLM Agents","Enterprise agents must execute long-horizon, conditional, safety-critical SOPs reliably. The work compiles machine-readable SOP constraints into executable pseudo-code and runs them with a program-guided (PG) stack machine that pages only the active frame while an LLM performs semantic execution. A three-arm SOPBench study across six models shows compiled text never significantly hurts and yields up to 16.0 points over official prose. Runtime guidance is capability-gated: strong models benefit, weak models degrade. Mechanism probes attribute the effect to spontaneous state discipline, and targeted cursor paging largely recovers refusal gains. Practical guidance: compile first, then enable active-frame paging only after a model-level discipline check.","Compile, Then Page: Executable SOP Programs and a Capability-Gated Runtime for Procedural LLM Agents  \nChenglin Yu 1 , Li Yin2 , Ying Yu3 , Hongxia Yang4 , Ming Li 1 ,5  \n1Department of Industrial and Systems Engineering, The Hong Kong Polytechnic University  \n2Department of Data and Systems Engineering, The University of Hong Kong  \n3College of Economics and Management, Zhejiang Normal University  \n4Department of Computing, The Hong Kong Polytechnic University  \n5Research Institute for Generative AI, The Hong Kong Polytechnic University  \n[Correspondence: ming.li@polyu.edu.hk](Correspondence: ming.li@polyu.edu.hk)  \narXiv :2607 . 1 1346v 1 [ cs .AI] 13 Jul 2026  \nAbstract  \nEnterprise agents must follow long-horizon, conditional, safetycritical standard operating procedures (SOPs) . We compile machine-readable SOP constraints into executable pseudo-code and run them with a program-guided (PG) stack machine that pages the active frame while an LLM performs semantic execution. A three-arm SOPBench study across six models separates representation from runtime: compiled text never significantly hurts and gains up to 16.0 points where official prose underperforms. Runtime guidance is capability-gated. Two strong models independently show positive seven-domain PG contrasts (58:19 and 75:31 discordant pairs), whereas weak models are harmed. A full-program cursor ablation—active frame first, complete program retained—recovers much of the strong-model refusal gain; selective visibility adds a smaller improvement. Paired probe and audit measurements track this divide to spontaneous state discipline rather than reconstruction ability. On Bank the three primary arms rise 70.4→86.4→92 . 8, with 100% refusal correctness. Practical guidance: compile first; enable active-frame paging only after a model-level discipline check.  \n1 Introduction  \nDeploying language agents in customer-facing operations is, toa first approximation, a procedure-following problem. Banks, clinics, and service desks encode their obligations as standard operating procedures (SOPs): long-horizon workflows with conditional branches and safety-critical refusal rules—what todo, what to verify first, when to decline. On SOPBench [Liet al., 2025], the best official configuration passes 82.4% of Bank tasks on the subset we study, with the best refusal accuracy reaching 90.7%—an error rate a compliance office would not accept.  \nThe dominant practice puts the entire SOP into the prompt as text. Resident prose (∼ 104 characters per turn on Bank) competes with the dialog for attention [Liu et al., 2024a], and verbalization strips executable structure: alternative order, gate  \nshort-circuits, the boundary between what must hold and how to verify it. The symptoms: skipped checks, premature actions, incorrect refusals.  \nWe treat the SOP as a program rather than a text. An offline, deterministic compiler translates the benchmark’s machinereadable dependency constraints—the same source its verbalizer renders as the baseline prose—into a two-layer pseudo-code program: process functions, one per user goal, and rule subroutines pairing each constraint with its verification recipe and an evidence-bearing return. An online program-guided (PG) runtime executes it as a two-layer virtual machine: a symbolic stack machine does the bookkeeping (stack, cursor, variables, recovery) and pages only the active frame into context, while the LLM converses, calls tools, and judges evidence. Enforcement is deliberately soft—attention, not permissions—at the price that correctness depends on faithful local-frame execution: a design bet we measure.  \nA three-arm design (official text; compiled text; compiled program plus runtime) separates the two layers on identical tasks, and six models spanning capability tiers chart where each layer pays. The compiled representation never significantly hurts and pays where official prose underperforms; a content ×format factorial shows the code-format package is ","cbCaibjMw0gGiE2x","https://ap.wps.com/l/cbCaibjMw0gGiE2x","pdf",745749,7,1,9,"English","en",105,"# Abstract\n# Introduction\n# Related Work","[{\"question\":\"文中如何将SOP约束转化为可执行形式？\",\"answer\":\"通过离线的、确定性的编译器，将基准提供的机器可读依赖约束翻译为两层伪代码程序：以用户目标为粒度的进程函数，以及将每条约束与其验证配方和可提供证据的返回绑定的规则子程序。\"},{\"question\":\"程序引导（PG）运行时如何与LLM协同工作？\",\"answer\":\"运行时以两层虚拟机方式执行：符号堆栈机负责栈、游标、变量与恢复等账务，并仅将“当前活动帧”分页到上下文；LLM则进行对话、工具调用与基于证据的判断，从而实现程序约束的软执行。\"},{\"question\":\"“能力门控（capability-gated）”对不同模型的效果有什么差异？\",\"answer\":\"结果显示，运行时指导对强模型带来独立的正向提升，而弱模型在Bank七个领域上会出现显著下降（文中提到约14–26分的损失）。机制验证表明差异与“自发的状态纪律”有关，而非仅由重建能力导致。\"}]",1784209458,23,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"compile-then-page-executable-sop-programs-and-a-capability-gated-runtime-for-procedural-llm-agents","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/compile-then-page-executable-sop-programs-and-a-capability-gated-runtime-for-procedural-llm-agents/86206/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-26","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"文中如何将SOP约束转化为可执行形式？","Question",{"text":76,"@type":77},"通过离线的、确定性的编译器，将基准提供的机器可读依赖约束翻译为两层伪代码程序：以用户目标为粒度的进程函数，以及将每条约束与其验证配方和可提供证据的返回绑定的规则子程序。","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"程序引导（PG）运行时如何与LLM协同工作？",{"text":81,"@type":77},"运行时以两层虚拟机方式执行：符号堆栈机负责栈、游标、变量与恢复等账务，并仅将“当前活动帧”分页到上下文；LLM则进行对话、工具调用与基于证据的判断，从而实现程序约束的软执行。",{"name":83,"@type":74,"acceptedAnswer":84},"“能力门控（capability-gated）”对不同模型的效果有什么差异？",{"text":85,"@type":77},"结果显示，运行时指导对强模型带来独立的正向提升，而弱模型在Bank七个领域上会出现显著下降（文中提到约14–26分的损失）。机制验证表明差异与“自发的状态纪律”有关，而非仅由重建能力导致。","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,120,123,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":46,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":107,"slug":137},19,"General","general"]