[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84433-en":3,"doc-seo-84433-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},84433,1099513958607,"Jiven","https://ap-avatar.wpscdn.com/avatar/100002390cf8733938c?x-image-process=image/resize,m_fixed,w_180,h_180&k=1778829742770036399",8,"Research & Report","Scheduling Techniques of AI Models on Modern Heterogeneous Edge GPU A Critical Review","Scheduling Techniques of AI Models on Modern Heterogeneous Edge GPU presents a critical review of deep neural network (DNN) schedulers for NVIDIA Jetson edge platforms. The work addresses the need for automated, optimal execution of complex DNN workloads while maximizing utilization of heterogeneous accelerators, including CPU, GPU, DLA, PVA, and VIC. It analyzes scheduler methodologies, performance, and effectiveness, summarizes the current research state, and outlines future research directions to improve edge AI capabilities under constrained resources.","Scheduling Techniques of AI Models on Modern Heterogeneous Edge GPU -A Critical Review  \nAshiyana Abdul Majeed, Mahmoud Meribout, Senior Member, IEEE, and Safa Mohammed Sali  \narXiv :2506 .01377v2 [ cs .DC] 12 Jul 2026  \nAbstract—In recent years, the development of specialized edge computing devices has significantly increased, driven by the growing demand for AI models. These devices, such as the NVIDIA Jetson series, must efficiently handle increased data processing and storage requirements. However, despite these advancements, there remains a lack of frameworks that automate the optimal execution of deep neural network (DNN). Therefore, efforts have been made to create schedulers that can manage complex data processing needs while ensuring the efficient utilization of all available accelerators within these devices, including the CPU, GPU, deep learning accelerator (DLA), programmable vision accelerator (PVA), and video image compositor (VIC). Such schedulers would maximize the performance of edge computing systems, which is crucial in resource-constrained environments. This paper aims to comprehensively review the various DNN schedulers implemented on NVIDIA Jetson devices. It examines their methodologies, performance, and effectiveness in addressing the demands of modern AI workloads. By analyzing these schedulers, this review highlights the current state of the research in the field. It identifies future research and development areas, further enhancing edge computing devices’ capabilities.  \nIndex Terms—accelerator, DLA, neural network, performance, scheduler.  \nI. INTRODUCTION AS AI-System-on-Chip (AI-SoC) become increasingly  \npopular and widely adopted in various embedded devices, it is vital to allocate the tasks in such a way that ensures high performance while operating under low power, particularly in edge devices [1] . One such powerful edge device is the NVIDIA Jetson series, which contains hardware accelerators that are well-adapted for AI and graphics applications. They comprise several hardware engines dedicated to various parallel computation models and interfaced with highspeed and low-power synchronous dynamic random-access memory (SDRAM) . However, in most cases, these devices balance heavy workloads, including multiple DNNs, that can impair their performance. Some works consider optimizing the DNN execution in their GPUs, such as [2], but often overlook other accelerators, leading to suboptimal performance. Proper task scheduling and partitioning into their hardware engines can significantly enhance efficiency, thereby using the full potential of edge computing. Building schedulers remains challenging due to the heterogeneous nature of their hardware architecture and the unpredictable memory contention between different hardware engines. The contention arises from using a shared bus that connects all hardware engines to the shared  \nAshiyana Abdul Majeed, Dr Mahmoud Meribout, and Safa Mohammed Sali are with the Department of Computer and Information Engineering, Khalifa University, Abu Dhabi, UAE (email: [100059454@ku.ac.ae](100059454@ku.ac.ae), mah[moud.meribout@ku.ac.ae](moud.meribout@ku.ac.ae), [safa.msali@ku.ac.ae](safa.msali@ku.ac.ae)) .  \nSDRAM. Hence, extensive work is necessary to develop these devices’ scheduling and hardware partitioning algorithms.  \nWith the rising demand for Jetson devices, it is of the utmost importance that a framework is developed to guarantee efficient and effective performance. For instance, the Jetson series has been used in autonomous driving. [3] evaluates the use of the Xavier and Orin device for pedestrian detection using an MM-Net model. Another example is a delivery robot that uses the Xavier device to navigate and reach the destination. Here, a modified single shot detector (SSD) is utilized for object detection [4] . Apart from robotics, the Jetson series has also been utilized for monitoring purposes, particularly in oceanography. Employing such devices allow","cbCaiqiKkvGkRf0I","https://ap.wps.com/l/cbCaiqiKkvGkRf0I","pdf",933637,2,1,12,"English","en",105,"# Introduction\n## NVIDIA Jetson heterogeneity and motivation\n## Goals and challenges in DNN scheduling\n## Example scheduler approaches","[{\"question\":\"Why is task scheduling important for NVIDIA Jetson edge devices?\",\"answer\":\"Jetson devices balance heavy workloads across multiple hardware accelerators, and improper execution can degrade performance. Efficient scheduling and partitioning across hardware engines helps improve overall efficiency and system utilization under low-power constraints.\"},{\"question\":\"What makes DNN scheduling challenging on heterogeneous Jetson hardware?\",\"answer\":\"The hardware architecture is heterogeneous and shared resources like SDRAM and buses introduce unpredictable memory contention between engines. This contention complicates achieving both high throughput and low latency with controlled power usage.\"},{\"question\":\"What does the paper aim to contribute regarding DNN schedulers?\",\"answer\":\"The paper comprehensively reviews DNN schedulers implemented on NVIDIA Jetson devices by examining their methodologies and effectiveness. The analysis identifies gaps and future research and development areas to enhance edge computing capabilities.\"}]",1784195606,30,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"scheduling-techniques-of-ai-models-on-modern-heterogeneous-edge-gpu-a-critical-review","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/scheduling-techniques-of-ai-models-on-modern-heterogeneous-edge-gpu-a-critical-review/84433/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-18","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why is task scheduling important for NVIDIA Jetson edge devices?","Question",{"text":75,"@type":76},"Jetson devices balance heavy workloads across multiple hardware accelerators, and improper execution can degrade performance. Efficient scheduling and partitioning across hardware engines helps improve overall efficiency and system utilization under low-power constraints.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What makes DNN scheduling challenging on heterogeneous Jetson hardware?",{"text":80,"@type":76},"The hardware architecture is heterogeneous and shared resources like SDRAM and buses introduce unpredictable memory contention between engines. This contention complicates achieving both high throughput and low latency with controlled power usage.",{"name":82,"@type":73,"acceptedAnswer":83},"What does the paper aim to contribute regarding DNN schedulers?",{"text":84,"@type":76},"The paper comprehensively reviews DNN schedulers implemented on NVIDIA Jetson devices by examining their methodologies and effectiveness. The analysis identifies gaps and future research and development areas to enhance edge computing capabilities.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":29,"slug":121},"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]