[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-119290-en":3,"doc-seo-119290-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},119290,7971461740909,"Levi","https://ap-avatar.wpscdn.com/davatar_155a257f0dc6eb9ab79c44ca47cae57d",8,"Research & Report","ML-NIC - accelerating machine learning inference using smart network interface cards","Low-latency machine-learning inference is essential for mission-critical applications such as autonomous driving, defense target recognition, and network traffic analysis. To reduce inference delay, existing approaches offload computation to specialized hardware like GPUs, but mapping models onto programmable network devices remains limited. This work presents ML-NIC, a framework that deploys trained models directly into programmable network devices’ data-plane computational cores, leveraging device parallelism. Experiments show at least 6x lower latency and 16x higher throughput with minimal impact on model effectiveness, along with tighter latency bounds under concurrent traffic and reduced CPU and RAM usage.","TYPE Original Research PUBLISHED 06 January 2025  \nDOI 10. 3389/fcomp.2024.1493399  \nOPEN ACCESS  \nEDITED BY  \nYong Wang,  \nGuilin University of Electronic Technology, China  \nREVIEWED BY  \nLuca Deri,  \nUniversity of Pisa, Italy Sabina Rossi,  \nCa’ Foscari University of Venice, Italy  \n*CORRESPONDENCE  \nSean Choi  \n [sean.choi@scu.edu](sean.choi@scu.edu)  \nRECEIVED 09 September 2024  \nACCEPTED 04 December 2024  \nPUBLISHED 06 January 2025  \nCITATION  \nKapoor R, Anastasiu DC and Choi S (2025) ML-NIC: accelerating machine learning inference using smart network interface cards. Front. Comput. Sci. 6:1493399 .  \ndoi: 10.3389/fcomp.2024.1493399  \nCOPYRIGHT  \n© 2025 Kapoor, Anastasiu and Choi. This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY) . The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.  \nML-NIC: accelerating machine learning inference using smart network interface cards  \nRaghav Kapoor1 , David C. Anastasiu2 and Sean Choi1*  \n1 Cloud Lab, Department of Computer Science and Engineering, Santa Clara University, Santa Clara, CA, United States, 2Anastasiu Lab, Department of Computer Science and Engineering, Santa Clara University, Santa Clara, CA, United States  \nLow-latency inference for machine learning models is increasingly becoming a necessary requirement, as these models are used in mission-critical applications such as autonomous driving, military defense (e.g., target recognition), and network tra􀀈c analysis. A widely studied and used technique to overcome this challenge is to o􀀉oad some or all parts of the inference tasks onto specialized hardware such as graphic processing units. More recently, o􀀉oading machine learning inference onto programmable network devices, such as programmable network interface cards or a programmable switch, is gaining interest from both industry and academia, especially due to the latency reduction and computational beneﬁts of performing inference directly on the data plane where the network packets are processed. Yet, current approaches are relatively limited in scope, and there is a need to develop more general approaches for mapping o􀀉oading machine learning models onto programmable network devices. Tofulﬁll such a need, this work introduces a novel framework, called ML-NIC, for deploying trained machine learning models onto programmable network devices’ data planes. ML-NIC deploys models directly into the computational cores of the devices to e􀀈ciently leverage the inherent parallelism capabilities of network devices, thus providing huge latency and throughput gains. Our experiments show that ML-NIC reduced inference latency by at least 6􀀂 on average and in the 99th percentile and increased throughput by at least 16x with little to no degradation in model e􀀀ectiveness compared to the existing CPU solutions. In addition, ML-NIC can provide tighter guaranteed latency boundsin the presence of other network tra􀀈c with shorter tail latencies. Furthermore, ML-NIC reduces CPU and host server RAM utilization by 6 .65% and 320 .80 MB. Finally, ML-NIC can handle machine learning models that are 2 .25 􀀂 larger than the current state-of-the-art network device o􀀉oading approaches.  \nKEYWORDS  \nmachine learning, SmartNIC, Netronome, data plane, inference  \n1 Introduction  \nMachine learning (ML) permeates a vast amount of everyday life, from personalized recommendations to stock market analysis and novel drug synthesis. While the machine learning models created to solve problems in these various 􀀂elds are proven to be highly e􀀓ective, these models often need large amount of time to make predictions (also referred to as model inference) on data instances. Often ti","cbCaieDmd1EEYVbY","https://ap.wps.com/l/cbCaieDmd1EEYVbY","pdf",2525537,1,17,"English","en",105,"# Introduction\n## Motivation for low-latency ML inference\n## Offloading inference to specialized hardware and its limits\n## Programmable data planes and the ML-NIC approach","[{\"question\":\"What problem does ML-NIC address?\",\"answer\":\"ML-NIC targets the high latency of machine-learning inference for latency-critical applications by moving inference closer to where data packets are processed.\"},{\"question\":\"How does ML-NIC deploy machine learning models?\",\"answer\":\"It deploys trained models directly into the computational cores of programmable network devices’ data planes to exploit inherent parallelism.\"},{\"question\":\"What performance improvements does ML-NIC report?\",\"answer\":\"Experiments show at least 6x average latency reduction and latency improvement in the 99th percentile, at least 16x throughput increase, and reduced CPU and host server RAM usage, with little to no degradation in model effectiveness.\"}]","ML-NIC - accelerating machine learning inference using smart network interface cards | PDF",1785723541,43,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"ml-nic-accelerating-machine-learning-inference-using-smart-network-interface-cards","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/ml-nic-accelerating-machine-learning-inference-using-smart-network-interface-cards/119290/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-06","2026-08-03",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What problem does ML-NIC address?","Question",{"text":76,"@type":77},"ML-NIC targets the high latency of machine-learning inference for latency-critical applications by moving inference closer to where data packets are processed.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How does ML-NIC deploy machine learning models?",{"text":81,"@type":77},"It deploys trained models directly into the computational cores of programmable network devices’ data planes to exploit inherent parallelism.",{"name":83,"@type":74,"acceptedAnswer":84},"What performance improvements does ML-NIC report?",{"text":85,"@type":77},"Experiments show at least 6x average latency reduction and latency improvement in the 99th percentile, at least 16x throughput increase, and reduced CPU and host server RAM usage, with little to no degradation in model effectiveness.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,121,124,129,132,136],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":46,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":46,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":46,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":133,"doc_module":4,"doc_module_name":46,"category_name":134,"show_sort_weight":133,"slug":135},10,"Lifestyle","lifestyle",{"id":137,"doc_module":4,"doc_module_name":46,"category_name":138,"show_sort_weight":107,"slug":139},19,"General","general"]