[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85493-en":3,"doc-seo-85493-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},85493,962075006959,"Anda","https://ap-avatar.wpscdn.com/avatar/e0002397efbe92a78e?_k=1776741047341049297",8,"Research & Report","Performance Isolation for Inference Processes in Edge GPU Systems","This work analyzes key isolation mechanisms in modern NVIDIA GPUs—MPS, MIG, and Green Contexts—to achieve predictable inference timing for safety-critical deep learning applications. The methodology combines performance testing, evaluation of partitioning effects, and analysis of temporal isolation across processes on both NVIDIA A100 and Jetson Orin. Findings show MIG delivers strong isolation, while Green Contexts enable fine-grained SM allocation with low overhead, though without memory isolation. Limitations and future research directions for improving temporal predictability in shared GPUs are identified.","Performance Isolation for Inference Processes in  \nEdge GPU Systems  \n1st Juan Jos Mart´ın  \nDISCA  \nUniversitat Polite`cnica de Vale`ncia Valncia, Spain [juamaros@upvnet.upv.es](juamaros@upvnet.upv.es)  \n2nd Jos Flich DISCA  \nUniversitat Polite`cnica de Vale`ncia Valncia, Spain [jflich@disca.upv.es](jflich@disca.upv.es)  \n3rd Carles Hernndez DISCA  \nUniversitat Polite`cnica de Vale`ncia Valncia, Spain [carherlu@upv.es](carherlu@upv.es)  \narXiv :2601 .07600v 3 [ cs .OS] 13 Jul 2026  \nAbstract—This work analyzes the main isolation mechanisms available in modern NVIDIA GPUs: MPS, MIG, and the recent Green Contexts, to ensure predictable inference time in safetycritical applications using deep learning models. The experimental methodology includes performance tests, evaluation of partitioning impact, and analysis of temporal isolation between processes, considering both the NVIDIA A100 and Jetson Orin platforms. It is observed that MIG provides a high level of isolation. At the same time, Green Contexts represent a promising alternative for edge devices by enabling fine-grained SM allocation with low overhead, albeit without memory isolation. The study also identifies current limitations and outlines potential research directions to improve temporal predictability in shared GPUs.  \nIndex Terms—Deep learning models, Ensemble of neural networks, Process isolation, GPU, Multi-Process Service (MPS), Multi-Instance GPUs (MIG), CUDA, Inferences, Green Contexts (GC)  \nI. INTRODUCTION  \nArtificial intelligence based on deep learning models (DLMs) is becoming increasingly common across various industries, including e-commerce, image and video analysis, and finance, among many others. However, in applications where malfunctions can result in the loss of human life or environmental damage, this technology is subject to strict compliance with certification standards (e.g., ISO 26262 [1] in the automotive domain) . This article focuses on meeting the predictability requirements of tasks executed on a graphics processing unit (GPU) in the context of safety-related applications.  \nTo enable the use of DLMs in the context of functional  \nsafety applications, ISO/IEC TR5469 [2] proposes using prediction strategies based on diverse neural network ensembles [3], [4] . This technique avoids relying on the output of a single neural network as the final prediction. Instead, the exact inference is performed across multiple models, and the final result is determined by a voting mechanism among all models. This approach enhances the overall robustness [5] and reliability [6] of the system.  \nHowever, in high-criticality tasks, not only is the accuracy of the prediction important, but the timing guarantees provided by the system are also necessary. In other words, a DLM must produce a correct prediction and do so within the required time  \nconstraints. Meeting this requirement is particularly challenging when GPUs are concurrently executing multiple tasks.  \nWith these two premises in mind, this study aims to analyze the alternatives modern GPUs offer to enable temporal isolation between multiple concurrently running tasks.  \nII. BACKGROUND  \nNowadays, numerous GPU-based systems exist for neural network model inference, ranging from large computing clusters to embedded systems with power consumption below 10 W [7], including general-purpose GPUs.  \nDespite this, in almost all cases, computational resources are underutilized, especially when performing inferences with small batch sizes, such as in real-time processing systems or autonomous driving platforms.  \nThis underutilization opens the possibility of parallelizing inference processes on the GPU, a strategy widely employed on CPUs. This approach addresses two issues simultaneously: on one hand, it enables full utilization of available computational resources, and on the other, it facilitates the use of neural network ensemble strategies.  \nNeural network ensembles is a technique to improve system accu","cbCaingxBWicSlxZ","https://ap.wps.com/l/cbCaingxBWicSlxZ","pdf",770947,3,1,10,"English","en",105,"# Introduction\n# Background\n# Isolation Mechanisms in Modern GPUs","[{\"question\":\"What problem does the paper address in edge GPU inference systems?\",\"answer\":\"The paper targets the need for predictable inference time in safety-critical applications when GPUs execute multiple tasks concurrently.\"},{\"question\":\"Which NVIDIA isolation mechanisms are analyzed?\",\"answer\":\"The study evaluates MPS, MIG, and Green Contexts to understand their ability to isolate concurrent inference processes.\"},{\"question\":\"What are the main observations about MIG and Green Contexts?\",\"answer\":\"MIG provides a high level of isolation, while Green Contexts offer fine-grained SM allocation with low overhead, but they do not provide memory isolation.\"}]",1784203989,25,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"performance-isolation-for-inference-processes-in-edge-gpu-systems","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/performance-isolation-for-inference-processes-in-edge-gpu-systems/85493/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-25","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the paper address in edge GPU inference systems?","Question",{"text":75,"@type":76},"The paper targets the need for predictable inference time in safety-critical applications when GPUs execute multiple tasks concurrently.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"Which NVIDIA isolation mechanisms are analyzed?",{"text":80,"@type":76},"The study evaluates MPS, MIG, and Green Contexts to understand their ability to isolate concurrent inference processes.",{"name":82,"@type":73,"acceptedAnswer":83},"What are the main observations about MIG and Green Contexts?",{"text":84,"@type":76},"MIG provides a high level of isolation, while Green Contexts offer fine-grained SM allocation with low overhead, but they do not provide memory isolation.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":22,"slug":133},"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]