[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86435-en":3,"doc-seo-86435-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86435,137441390410,"Hazel","https://ap-avatar.wpscdn.com/avatar/2000252f4ab5702993?_k=1776741390130283984",8,"Research & Report","Closed Loop Control with Rule Aligned Small Language Models and Multi Agent Self Correction","A key step toward autonomous industrial operation is the creation and reconfiguration of control policies from natural-language requirements with minimal manual redesign. This work explores an approach where a compact Small Language Model is retrained for control reasoning and embedded in a plant-aware validator guided correction loop, using a digital-twin style checker. In randomized thermal-control simulations, the method delivers 91.5% average action alignment at 3.84 s mean latency and maintains a 95% in-range rate under symbolic remapping, supporting edge deployable reconfigurable control.","Closed-Loop Control with Rule-Aligned Small Language Models and  \nMulti-Agent Self-Correction  \nYuchen Wang  \nSchool of Computer Science University of Sheffield Sheffield, UK  \n[ywang1016@sheffield.ac.uk](ywang1016@sheffield.ac.uk)  \nTong Liu  \nSchool of Computer Science University of Sheffield Sheffield, UK [t.liu@sheffield.ac.uk](t.liu@sheffield.ac.uk)  \nJaval Vyas  \nDepartment of Chemical Engineering Imperial College London London, UK [j.vyas24@imperial.ac.uk](j.vyas24@imperial.ac.uk)  \nMehmet Mercangöz  \nDepartment of Chemical Engineering Imperial College London London, UK[m.mercangoz@imperial.ac.uk](m.mercangoz@imperial.ac.uk)  \narXiv :2607 .097 13v 1 [ cs .AI] 24 Jun 2026  \nAbstract—A key step toward autonomous industrial operation is the ability to create and reconfigure control policies from natural-language requirement specifications, with minimal or no manual redesign. In this setting, policy generation by AI agents can be a credible path when paired with a plant-aware validator (e.g., a digital twin) that can check generated candidate actions before execution. However, practical deployment is constrained by inference latency and compute footprint: large cloud-based models are often too slow, opaque, or data-sensitive for edge closed-loop use. This work investigates whether a compact Small Language Model (SLM) can be retrained for control reasoning and embedded in a validator-guided correction loop. We use a Qwen2.5-1.5B model aligned via Group Relative Policy Optimization (GRPO), combined with (i) an action agent, (ii) a symbolic/digital-twin-style validation layer, and (iii) a reprompting agent that iteratively steers outputs toward valid actions. In randomized thermal-control simulations (30 experiments with 500 steps each), the framework achieves 91.5% average action-alignment accuracy (86.3%–100% across cases) at 3.84 s mean inference latency. Under symbolic re-mapping, it maintains a 95% in-range rate, indicating robust physical regulation despite reduced token-level agreement. These results support SLM+validator architectures as a practical path toward reconfigurable autonomous control at the edge.  \nKeywords: Autonomous systems, Small language models, Edge AI, Generative AI  \nI. INTRODUCTION  \nA central requirement for autonomous industrial operation is the ability to create and reconfigure control policies from high-level, natural-language requirement specifications with minimal manual redesign. This is particularly important in environments where objectives and constraints evolve overtime (e.g., energy priorities, safety envelopes, and operating policies) . In this setting, policy generation by AI agents can bea credible path when paired with a plant-aware validator (e.g., a digital twin) that can check generated candidate actions before execution [1] .  \nConventional industrial control architectures are highly effective for deterministic numerical regulation, but they are less flexible when symbolic operating intent must be incorporated quickly. Translating new expert directives into  \ndeployed logic typically requires manual re-engineering of rules, supervisory layers, or interlocks, which slows reconfiguration and increases engineering overhead. This creates a persistent gap between high-level operational intent and low-level control execution.  \nLanguage models provide a promising interface for closing this gap because they can parse and operationalize naturallanguage instructions and support iterative self-correction during inference [2], [3] . However, deploying cloud-hosted Large Language Models (LLMs) in closed-loop industrial workflows remains challenging: API dependence introduces variable end-to-end latency, while external data transfer raises governance and sovereignty concerns for sensitive operational telemetry [4], [5] . These issues are especially limiting for edge and offline autonomy.  \nThis work investigates whether a compact Small Language Model (SLM) can be trained to perform control-or","cbCairlk2GWbhZXq","https://ap.wps.com/l/cbCairlk2GWbhZXq","pdf",1131652,4,1,6,"English","en",105,"# Introduction\n## Motivation and problem gap\n## Challenges of cloud LLM deployment\n## Proposed SLM + validator correction framework\n# Related Work\n## LLM-driven autonomous control and advantages","[{\"question\":\"How does the framework generate control policies from natural-language requirements?\",\"answer\":\"It trains a compact Small Language Model for control reasoning and uses a validator-guided correction loop so candidate actions are checked before execution.\"},{\"question\":\"Why is a Small Language Model used instead of a large cloud LLM?\",\"answer\":\"The approach targets edge deployment constraints, reducing problems such as variable latency, compute footprint, and data governance issues from external telemetry transfer.\"},{\"question\":\"What are the reported simulation results for action alignment and regulation quality?\",\"answer\":\"Across 30 thermal-control simulation experiments, it achieves 91.5% average action-alignment accuracy and 3.84 s mean inference latency, while maintaining a 95% in-range rate under symbolic re-mapping.\"}]",1784211729,15,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"closed-loop-control-with-rule-aligned-small-language-models-and-multi-agent-self-correction","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/closed-loop-control-with-rule-aligned-small-language-models-and-multi-agent-self-correction/86435/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-25","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"How does the framework generate control policies from natural-language requirements?","Question",{"text":75,"@type":76},"It trains a compact Small Language Model for control reasoning and uses a validator-guided correction loop so candidate actions are checked before execution.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"Why is a Small Language Model used instead of a large cloud LLM?",{"text":80,"@type":76},"The approach targets edge deployment constraints, reducing problems such as variable latency, compute footprint, and data governance issues from external telemetry transfer.",{"name":82,"@type":73,"acceptedAnswer":83},"What are the reported simulation results for action alignment and regulation quality?",{"text":84,"@type":76},"Across 30 thermal-control simulation experiments, it achieves 91.5% average action-alignment accuracy and 3.84 s mean inference latency, while maintaining a 95% in-range rate under symbolic re-mapping.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,114,119,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":111,"show_sort_weight":112,"slug":113},"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]