[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85198-en":3,"doc-seo-85198-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},85198,962075114101,"Seraphina","https://ap-avatar.wpscdn.com/avatar/e000253a75eb197efd?x-image-process=image/resize,m_fixed,w_180,h_180&k=1780044092746381165",8,"Research & Report","Measure the Sim-to-Real Gap Designing an Affordable Real-World Benchmark Platform for Reinforcement Learning in AIoT Systems","Reinforcement learning (RL) improves autonomous AIoT systems, yet real-world trial-and-error can be costly and hazardous, making most research rely on simulation. This creates a Sim-to-Real transfer problem where algorithmic robustness and the Sim-to-Real gap must be evaluated before deploying to real environments. Since no universal AIoT Sim-to-Real benchmark platform existed, the work builds an affordable real-world AIoT platform using \u003CUSD 400 components and edge-side vision-guided gaming. A human-score maximizing objective reduces safety risk, showing 1160% degradation after deployment from simulation and ~38% human-level performance after DQN training.","Measure the Sim-to-Real Gap: Designing an Affordable Real-World Benchmark Platform for Reinforcement Learning in AIoT Systems  \nRongping Zhou, Omid Tavallaie, Shuaijun Chen, Albert Y. Zomaya  \narXiv :2607 . 10309v 1 [ cs .AI] 11 Jul 2026  \nAbstract—Reinforcement learning (RL) is commonly employed to enhance the performance of autonomous systems, including the Autonomous Internet of Things (AIoT). However, the trial-anderror nature of RL, when conducted in real-world environments, is costly and hazardous in some scenarios. Consequently, the majority of RL research is conducted in simulation. This reliance introduces challenges related to the Sim-to-Real transferability. Evaluating the Sim-to-Real algorithmic robustness and the Simto-Real gap is a critical prerequisite for research aimed at improving RL performance in the real world. Therefore, industries such as robotics have developed concurrent simulation and physical platforms to facilitate this research. However, a universal Sim-to-Real benchmark platform for AIoT does not currently exist. To address these concerns, we developed a real-world AIoT platform for studying RL in AIoT. On this platform, an agent deployed on an edge device plays video games on a separate host computer via a hardware-emulated keyboard, guided by vision input. This platform uses commercially available components costing less than USD 400, together with two computers. Because the system’s objective is game score maximization, it inherently mitigates safety risks associated with real-world RL deployments. Experimental results show the simulation-trained agent suffers a 1160% performance degradation relative to the human-level performance after real-world deployment, indicating a significant Sim-to-Real gap. Direct real-world training using the deep Q-network (DQN) RL algorithm achieves approximately 38% of human-level performance after 10 million training steps, demonstrating the feasibility of RL under real-world conditions. These results suggest that the proposed Sim-to-Real benchmark platform provides a substantial foundation for qualitative and quantitative evaluations of RL in real-world AIoT systems.  \nIndex Terms—reinforcement learning, AIoT, Sim-to-Real, video game.  \nI. INTRODUCTION  \nREINFORCEMENT learning (RL) [1], when combined  \nwith deep learning [2], serves as an effective methodology for improving the performance of autonomous systems, including robotic platforms [3], [4], game-playing agents [5], artificial intelligence systems incorporating Large Language Models (LLMs) [6], specialized systems such as tokamak equipment [7], and the AIoT [8], [9], [11], [12] . The iterative trial-and-error mechanism inherent in RL can require  \nRongping Zhou, Shuaijun Chen, Albert Y. Zomaya are with the School of Computer Science, The University of Sydney, Camperdown NSW 2050, Australia (e-mail: [rzho0616@uni.sydney.edu.au](rzho0616@uni.sydney.edu.au); [shuaijun.chen@sydney.edu.au](shuaijun.chen@sydney.edu.au); [albert.zomaya@sydney.edu.au](albert.zomaya@sydney.edu.au)).  \nOmid Tavallaie is with the Department of Engineering Science, University of Oxford, Wellington Square, Oxford OX1 2JD, United Kingdom, and the Department of Computer Science, The University of Western Australia, 241, 35 Stirling Hwy, Crawley WA 6009 (email: [omid.tavallaie@eng.ox.ac.uk](omid.tavallaie@eng.ox.ac.uk); [omid.tavallaie@uwa.edu.au](omid.tavallaie@uwa.edu.au)) .  \nsubstantial resources and may introduce safety hazards that can reach a life-threatening level in certain real-world scenarios. As a result, research in these domains employs simulated environments, which provide cost-effective, controlled, and repeatable conditions for RL algorithms, agent development, and evaluation. RL algorithms and agents are typically designed, trained, and tested in simulation before being deployed to real-world environments. RL training is expensive in terms of computing resources, regardless of whether it is conducted in simulation or t","cbCaiaf5AdgVx0Cz","https://ap.wps.com/l/cbCaiaf5AdgVx0Cz","pdf",2198099,3,1,17,"English","en",105,"# Introduction\n## Sim-to-Real challenge and RL safety concerns\n## Limitations of existing benchmarks","[{\"question\":\"Why does reinforcement learning in real-world environments often rely on simulation first?\",\"answer\":\"RL’s iterative trial-and-error can require substantial resources and may introduce serious safety hazards, so researchers use simulated environments that are controlled, repeatable, and cost-effective before real-world deployment.\"},{\"question\":\"What gap does the document focus on?\",\"answer\":\"The document targets the Sim-to-Real gap, emphasizing how simulation-trained agents may lose robustness when transferred to real-world settings and how to measure that gap for RL algorithms.\"},{\"question\":\"How is the proposed real-world benchmark platform designed and why is it safer?\",\"answer\":\"An edge-deployed agent plays video games on a separate host via a hardware-emulated keyboard, guided by vision input, using commercially available components costing less than USD 400. Because the task objective is maximizing game score, the platform mitigates safety risks associated with real-world RL deployments.\"}]",1784201685,43,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"measure-the-sim-to-real-gap-designing-an-affordable-real-world-benchmark-platform-for-reinforcement-learning-in-aiot-systems","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/measure-the-sim-to-real-gap-designing-an-affordable-real-world-benchmark-platform-for-reinforcement-learning-in-aiot-systems/85198/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-23","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why does reinforcement learning in real-world environments often rely on simulation first?","Question",{"text":75,"@type":76},"RL’s iterative trial-and-error can require substantial resources and may introduce serious safety hazards, so researchers use simulated environments that are controlled, repeatable, and cost-effective before real-world deployment.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What gap does the document focus on?",{"text":80,"@type":76},"The document targets the Sim-to-Real gap, emphasizing how simulation-trained agents may lose robustness when transferred to real-world settings and how to measure that gap for RL algorithms.",{"name":82,"@type":73,"acceptedAnswer":83},"How is the proposed real-world benchmark platform designed and why is it safer?",{"text":84,"@type":76},"An edge-deployed agent plays video games on a separate host via a hardware-emulated keyboard, guided by vision input, using commercially available components costing less than USD 400. Because the task objective is maximizing game score, the platform mitigates safety risks associated with real-world RL deployments.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]