[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83012-en":3,"doc-seo-83012-105":30,"detail-sidebar-cat-0-en-105":90},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83012,7971461740909,"Levi","https://ap-avatar.wpscdn.com/davatar_155a257f0dc6eb9ab79c44ca47cae57d",8,"Research & Report","Energy Efficient GPU DVFS for Fine Tuning of SLMs on Resource constrained Embedded Devices","Dynamic Voltage Frequency Scaling (DVFS) on resource-constrained embedded GPU platforms is crucial for energy-efficient small language model (SLM) fine-tuning, since privacy- and personalization-driven adaptation runs locally and repeats forward-backward optimization across many mini-batches. The work first characterizes fine-tuning behavior of representative encoder-only BERT variants and autoregressive decoder-only Pythia variants on GLUE benchmarks, then proposes an ML-based model selection to choose energy-optimal DVFS settings. Results on NVIDIA Jetson AGX Orin show 13.11% average energy savings, up to 26.73% versus MAXN Mode 0.","Energy-Efficient GPU DVFS for Fine-Tuning of SLMs on Resource-constrained Embedded Devices  \nJurn-Gyu Park, Sanzhar Zholdybayev, Aidar Amangeldi, and Ademi Zhanuzakova  \narXiv :2607 .05933v 1 [ cs .PF] 7 Jul 2026  \nAbstract—Dynamic Voltage Frequency Scaling (DVFS) on resource-constrained embedded GPU platforms is essential for energy-efficient small language model (SLM) fine-tuning, as privacy- and personalization-driven adaptation increasingly requires local execution and involves repeated forward-backward optimization over many mini-batches, making it substantially more time- and energy-intensive than single-pass inference. To this end, 1) we first characterize the fine-tuning behavior of representative encoder-only SLMs of BERT variants, and autoregressive decoder-only SLMs of Pythia variants on GLUE benchmarks. In addition to the characterizations, 2) we propose a simple yet effective ML-based model selection that selects energyoptimal GPU DVFS settings on resource-constrained embedded platforms. Our results on NVIDIA Jetson AGX Orin demonstrate average 13.11% energy savings (up to 26.73%) over MAXN Mode 0, which has no explicit power cap.  \nIndex Terms—Transformer fine-tuning, small language models (SLMs), dynamic voltage frequency scaling (DVFS), embedded systems.  \nI. INTRODUCTION  \nT RANbeddSFedORMERand mobimleodpelstfoonrmtegarertenGPU-basinglyasedadoemptedfor on-device SLM fine-tuning in healthcare, finance, and autonomous systems, where privacy, latency, and offline operation requirements forbid cloud offload [1] [2] . These batterybased and resource-constrained devices require crucial yet challenging dynamic power management (DPM) techniques for power and energy savings without accuracy degradation, using power modes and/or dynamic voltage frequency scaling (DVFS) on heterogeneous processors.  \nA line of studies focuses on DVFS for large language model (LLM) inference workloads [3]–[5] and datacenter training [6]–[8]; moreover, the off-the-shelf Jetson AGX Orin series support four different power modes: 15W, 30W (Default/Mode 2), 50W, and the unconstrained MAXN performance mode (Mode 0) . To the best of our knowledge, our work firstly aims energy-efficient GPU DVFS for small language models (SLMs) fine-tuning using encoder-based and decoderbased SLM models on Jetson AGX Orin (64GB) unconstrained MAXN Mode, with the help of comprehensive applicationspecific workloads characterization. The contributions of the paper are as follows:  \n• We characterize how GPU frequency scaling affects finetuning latency, power consumption, and total energy on Transformer-based SLM workloads.  \nThis manuscript is currently under review for publication; (Corresponding author: [jurn.park@nu.edu.kz](jurn.park@nu.edu.kz).)  \nThe authors gratefully acknowledge Prof. Atakan Varol for providing access to the computational resources of the Institute of Smart Systems and Artificial Intelligence (ISSAI), and Yerbol Absalyamov and Vladimir Albrekht for their assistance.  \nThe authors are with the School of Engineering and Digital Sciences, Nazarbayev University, 010000 Astana, Kazakhstan (e-mails: jurn.park,sanzhar.zholdybayev,aidar.amangeldi,[ademi.zhanuzakova@nu.edu.kz](ademi.zhanuzakova@nu.edu.kz)).  \nFig. 1. Motivating example: Normalized fine-tuning energy versus GPU frequency for BERT-base and Pythia-70M on SST-2 using Jetson AGX Orin.  \n• We propose lightweight GPU frequency governors for energy-efficient SLM fine-tuning based on interpretable ML-based DT model policy.  \n• We demonstrate up to 26 .73% energy savings on NVIDIA Jetson AGX Orin accross benchmark fine-tuning workloads.  \nII. MOTIVATION AND RELATED WORK  \nA. Motivation  \nNVIDIA Jetson platforms [9] provide predefined power modes [10] for power-aware embedded deployment. However, these coarse-grained profiles do not guarantee energy-efficient fine-tuning, since total energy depends on both power draw and execution time. This issue is critical for on-device SLM adaptation, where","cbCaipn8JH4OP4aQ","https://ap.wps.com/l/cbCaipn8JH4OP4aQ","pdf",1537133,3,1,4,"English","en",105,"# Introduction\n# Motivation and Related Work\n## Motivation\n## Related Work","[{\"question\":\"Why is DVFS important for SLM fine-tuning on embedded GPUs?\",\"answer\":\"DVFS enables dynamic power management during repeated forward-backward optimization, reducing energy use under limited compute, memory, and battery constraints while preserving accuracy.\"},{\"question\":\"What models and benchmarks are used in the characterization study?\",\"answer\":\"The paper characterizes encoder-only BERT-variant SLMs and decoder-only Pythia-variant SLMs on GLUE benchmarks, analyzing how GPU frequency impacts latency, power, and total energy.\"},{\"question\":\"How does the proposed method select energy-optimal DVFS settings?\",\"answer\":\"It uses a lightweight ML-based model selection/policy to choose GPU frequency governor settings that minimize energy on resource-constrained embedded platforms.\"}]",1784184665,10,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":85,"head_meta":87,"extra_data":89,"updated_unix":28},"energy-efficient-gpu-dvfs-for-fine-tuning-of-slms-on-resource-constrained-embedded-devices","",{"@graph":36,"@context":84},[37,52,67],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":22},"https://docshare.wps.com/document/energy-efficient-gpu-dvfs-for-fine-tuning-of-slms-on-resource-constrained-embedded-devices/83012/",{"url":51,"name":13,"@type":53,"author":54,"headline":13,"publisher":56,"fileFormat":59,"inLanguage":24,"description":14,"dateModified":60,"datePublished":61,"encodingFormat":59,"isAccessibleForFree":62,"interactionStatistic":63},"DigitalDocument",{"name":9,"@type":55},"Person",{"url":41,"name":57,"@type":58},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":64,"interactionType":65,"userInteractionCount":20},"InteractionCounter",{"@type":66},"ViewAction",{"@type":68,"mainEntity":69},"FAQPage",[70,76,80],{"name":71,"@type":72,"acceptedAnswer":73},"Why is DVFS important for SLM fine-tuning on embedded GPUs?","Question",{"text":74,"@type":75},"DVFS enables dynamic power management during repeated forward-backward optimization, reducing energy use under limited compute, memory, and battery constraints while preserving accuracy.","Answer",{"name":77,"@type":72,"acceptedAnswer":78},"What models and benchmarks are used in the characterization study?",{"text":79,"@type":75},"The paper characterizes encoder-only BERT-variant SLMs and decoder-only Pythia-variant SLMs on GLUE benchmarks, analyzing how GPU frequency impacts latency, power, and total energy.",{"name":81,"@type":72,"acceptedAnswer":82},"How does the proposed method select energy-optimal DVFS settings?",{"text":83,"@type":75},"It uses a lightweight ML-based model selection/policy to choose GPU frequency governor settings that minimize energy on resource-constrained embedded platforms.","https://schema.org",{"og:url":51,"og:type":86,"og:title":13,"og:site_name":57,"og:description":14},"article",{"robots":88,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":91},[92,96,100,104,109,114,119,122,127,130,133],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":93,"show_sort_weight":94,"slug":95},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":97,"show_sort_weight":98,"slug":99},"Literature",80,"literature",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":101,"show_sort_weight":102,"slug":103},"Exam",70,"exam",{"id":105,"doc_module":4,"doc_module_name":46,"category_name":106,"show_sort_weight":107,"slug":108},5,"Comic",60,"comic",{"id":110,"doc_module":4,"doc_module_name":46,"category_name":111,"show_sort_weight":112,"slug":113},6,"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":29,"doc_module":4,"doc_module_name":46,"category_name":131,"show_sort_weight":29,"slug":132},"Lifestyle","lifestyle",{"id":134,"doc_module":4,"doc_module_name":46,"category_name":135,"show_sort_weight":105,"slug":136},19,"General","general"]