[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82170-en":3,"doc-seo-82170-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},82170,687197207057,"Sage","https://ap-avatar.wpscdn.com/davatar_29158cc5080c5b710cf443261637dec0",8,"Research & Report","EvoLP Self-Evolving Latency Predictor for Model Compression in Real-Time Edge Systems","Edge devices increasingly run deep learning for embedded deployment, where real-time response and tight resource budgets demand latency-targeted model compression. Direct latency measurement on physical devices is expensive and slow, making it impractical for large search spaces. EvoLP introduces an efficient framework that predicts inference latency on edge hardware and evolves during compression to improve accuracy. Evaluations on multiple edge devices and model variants show superior performance and effective guidance toward higher accuracy under strict latency constraints.","EvoLP: Self-Evolving Latency Predictor for Model Compression in Real-Time Edge Systems  \nShuo Huai, Hao Kong, Shiqing Li, Xiangzhong Luo, Ravi Subramaniam, Christian Makaya, Qian Lin and Weichen Liu  \narXiv :2607 .09063v 1 [ cs .LG] 10 Jul 2026  \nAbstract—Edge devices are increasingly utilized for deploying deep learning applications on embedded systems. The real-time nature of many applications and the limited resources of edge devices necessitate latency-targeted neural network compression. However, measuring latency on real devices is challenging and expensive. Therefore, this letter presents a novel and efficient framework, named EvoLP, to accurately predict the inference latency of models on edge devices. This predictor can evolve to achieve higher latency prediction precision during the network compression process. Experimental results demonstrate that EvoLP outperforms previous state-of-the-art approaches by being evaluated on three edge devices and four model variants. Moreover, when incorporated into a model compression framework, it effectively guides the compression process for higher model accuracy while satisfying strict latency constraints. We open source EvoLP at [https://github.com/ntuliuteam/EvoLP](https://github.com/ntuliuteam/EvoLP).  \nI. INTRODUCTION  \nDEEP Neural Networks (DNNs) have been widely de  \nployed on edge devices to eliminate the issues of data confidentiality breaches and unstable network bandwidth in accessing cloud servers [1] . However, current models grow in complexity to improve accuracy, rendering them unsuitable for resource-limited edge devices. Neural architecture search (NAS) can design customized DNN models for edge devices, but it is time-consuming and requires heavy engineering efforts [2] . Model compression is promising to improve DNNs’efficiency on edge devices. Many DNN applications, such as virtual/augmented reality, and autonomous driving, demand strict response latency, thus, the inference latency of DNNs should be treated as a hard constraint in model compression. Model compression based on real latency provides additional benefits from better exploration of hardware features [3] . However, traditionally, considering latency as a compression objective requires interaction with the physical device to obtain the real latency, which is time-consuming. Each latency measurement process may take several minutes [4],  \n© 2024 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works. This is the author’s accepted version of the article published in IEEE Embedded Systems Letters, vol. 16, no. 2, pp. 174-177, June 2024, DOI: 10. 1109/LES.2023.3321599.  \nS. Huai, H. Kong, S. Li, X. Luo, and W. Liu (Corresponding author, [liu@ntu.edu.sg](liu@ntu.edu.sg)) are with the School of Computer Science and Engineering, Nanyang Technological University, Singapore; S. Huai and H. Kong are also with the HP-NTU Digital Manufacturing Corporate Lab, Nanyang Technological University, Singapore; R. Subramaniam, C. Makaya, and Q. Lin are with HP Inc., Palo Alto, California, USA.  \nThis work is partially supported under the RIE2020 Industry Alignment Fund – Industry Collaboration Projects (IAF-ICP) Funding Initiative, as well as cash and in-kind contribution from the industry partner, HP Inc., through the HP-NTU Digital Manufacturing Corporate Lab (I1801E0028), and partially supported by the Ministry of Education, Singapore, under its Academic Research Fund Tier 2 (MOE2019-T2-1-071), and Nanyang Technological University, Singapore, under its NAP  \nmaking it prohibitively expensive in large design spaces. Some approaches are proposed to avoid the high cost of measuring latency, including using","cbCaiaxznw78Hq0z","https://ap.wps.com/l/cbCaiaxznw78Hq0z","pdf",730468,1,5,"English","en",105,"# Introduction\n## Motivation and Challenges\n## Related Work\n## EvoLP Contributions","[{\"question\":\"Why is latency a hard constraint in edge model compression?\",\"answer\":\"Many real-time edge applications require strict response latency, so inference latency must be treated as a non-negotiable constraint during compression.\"},{\"question\":\"What makes traditional latency measurement impractical for large design spaces?\",\"answer\":\"Measuring latency on physical devices is time-consuming and expensive, and building data sets through sampling can take days, making it infeasible for extensive model variants.\"},{\"question\":\"How does EvoLP improve latency prediction during compression?\",\"answer\":\"EvoLP uses a neural latency predictor with an evolution scheme that can update and improve prediction precision as the model compression process proceeds, without interrupting it.\"}]",1784178576,13,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"evolp-self-evolving-latency-predictor-for-model-compression-in-real-time-edge-systems","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/evolp-self-evolving-latency-predictor-for-model-compression-in-real-time-edge-systems/82170/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why is latency a hard constraint in edge model compression?","Question",{"text":75,"@type":76},"Many real-time edge applications require strict response latency, so inference latency must be treated as a non-negotiable constraint during compression.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What makes traditional latency measurement impractical for large design spaces?",{"text":80,"@type":76},"Measuring latency on physical devices is time-consuming and expensive, and building data sets through sampling can take days, making it infeasible for extensive model variants.",{"name":82,"@type":73,"acceptedAnswer":83},"How does EvoLP improve latency prediction during compression?",{"text":84,"@type":76},"EvoLP uses a neural latency predictor with an evolution scheme that can update and improve prediction precision as the model compression process proceeds, without interrupting it.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,109,114,119,122,127,130,134],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":21,"doc_module":4,"doc_module_name":45,"category_name":106,"show_sort_weight":107,"slug":108},"Comic",60,"comic",{"id":110,"doc_module":4,"doc_module_name":45,"category_name":111,"show_sort_weight":112,"slug":113},6,"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":45,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":45,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":45,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":45,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":45,"category_name":136,"show_sort_weight":21,"slug":137},19,"General","general"]