[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-124588-en":3,"doc-seo-124588-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},124588,1099513958762,"Logic","https://ap-avatar.wpscdn.com/avatar/1000023916a998db790?x-image-process=image/resize,m_fixed,w_180,h_180&k=1784791008015729253",6,"Technology","Green Carbon Footprint for Model Inference Serving via Exploiting Mixed-Quality Models and GPU Partitioning","Large-scale high performance computing (HPC) systems and datacenters run machine learning (ML) inference services that drive substantial compute cycles and carbon emissions. The work introduces Clover, a carbon-aware ML inference serving runtime that balances performance, accuracy, and carbon output by combining mixed-quality model variants with GPU resource partitioning. Experiments show substantial carbon reductions while maintaining high accuracy and meeting service level agreement (SLA) targets, positioning Clover as a practical step toward carbon neutrality for datacenters and HPC.","Green Carbon Footprint for Model Inference Serving via Exploiting Mixed-Quality Models and GPU Partitioning  \nBaolin Li Northeastern University  \nSiddharth Samsi  \nMIT  \nVijay Gadepally Devesh Tiwari  \nMIT Northeastern University  \narXiv :2304 .09781v1 [ cs .DC] 19 Apr 2023  \nABSTRACT  \nThis paper presents a solution to the challenge of mitigating carbon emissions from large-scale high performance computing (HPC) systems and datacenters that host machine learning (ML) inference services. ML inference is critical to modern technology products, but it is also a significant contributor to datacenter compute cyclesand carbon emissions. We introduce Clover, a carbon-friendly ML inference service runtime system that balances performance, accuracy, and carbon emissions through mixed-quality models and GPU resource partitioning. Our experimental results demonstrate that Clover is effective in substantially reducing carbon emissions while maintaining high accuracy and meeting service level agreement (SLA) targets. Therefore, it is a promising solution toward achieving carbon neutrality in HPC systems and datacenters.  \n1 INTRODUCTION  \nReducing carbon emissions is of critical importance to combat the growing threat of climate change, as noted by United Nation and other agencies [1]. The large-scale high performance computing (HPC) systems and datacenters that host information technology services account for 2% of the global carbon emission [2] – the amount of datacenter workload has grown by 260% in the past 6 years and is expected to keep growing, and, its contribution to the global carbon emission is likely to increase by many folds [3, 4] .  \nA relevant and contributing trend is that technology companies are increasingly incorporating artificial intelligence (AI) into their products and hosting trained machine learning (ML) models in datacenters GPUs to offer ML inference services to customers. Inadvertently, these inference services have exacerbated the carbon emission challenge because they account for a large proportion of the datacenter compute cycles. For example, many of Google’s billion-user services are empowered by AI and their inference represents 60% of the AI infrastructure emissions [5]; Meta has expanded their infrastructure capacity by 2.5× to meet the ML inference demand [6]; AWS and NVIDIA have estimated that inference accounts for 90% of the ML workloads in HPC and cloud datacenters [7, 8] .  \nTherefore, the goal of this paper is to design a novel carbonfriendly ML inference service runtime system. Unfortunately, carbon-friendliness is often at odds with other desirable properties, including performance (low inference latency) and inference accuracy (high accuracy requires complex models and more computation-intensive operations, increasing the carbon footprint) . But currently, we do not have the tools to effectively and automatically navigate this trade-off space and make ML inference services carbon-friendly. Nevertheless, given the growing importance of carbon-free operation in HPC systems and datacenters, finding  \nsolutions to this problem is becoming increasingly critical [9–12] .  \nClover Key Ideas and Contributions. The following summarizes the key insights behind Clover, the challenges in exploiting the observed opportunities, and Clover’s contributions.  \nOpportunity Space of Mixed-Quality Models and GPU Partitioning for Carbon Saving. This is the first study to present experimental evidence to demonstrate the opportunities and trade-offs in mixed-quality models and GPU partitioning for carbon savings. Our experiments reveal that creating a mixture of model variants (i.e., a mixture of low-and high-quality models) can result in significant carbon savings while maintaining a high level of accuracy. Additionally, GPU partitioning can also contribute to carbon reduction by optimizing resource utilization, although it can lead to increased latency and potential violations of SLA targets. Unfortunately, navig","cbCaisdB14eQZ8R2","https://ap.wps.com/l/cbCaisdB14eQZ8R2","pdf",1702868,1,13,"English","en",105,"# Introduction\n## Key problem and motivation\n## Clover key ideas and contributions\n# Background\n## Carbon intensity and carbon emission\n# System framework (inferred)","[{\"question\":\"What challenge does Clover address in ML inference serving?\",\"answer\":\"Clover targets the growing carbon emissions caused by large-scale datacenter and HPC ML inference workloads while still meeting performance and accuracy needs.\"},{\"question\":\"How do mixed-quality models and GPU partitioning reduce carbon emissions?\",\"answer\":\"Mixed-quality model variants enable carbon-accuracy trade-offs, while GPU partitioning improves resource utilization; together they reduce emissions but can affect latency and SLA, which Clover mitigates.\"},{\"question\":\"How does Clover ensure SLA targets while optimizing for carbon?\",\"answer\":\"Clover’s optimization engine dynamically adapts to carbon intensity and selects mixed-quality model variants and GPU partitioning choices to reduce carbon and maintain high accuracy without violating SLA constraints.\"}]","Green Carbon Footprint for Model Inference Serving via Exploiting Mixed-Quality Models and GPU Partitioning | PDF",1785893190,33,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"green-carbon-footprint-for-model-inference-serving-via-exploiting-mixed-quality-models-and-gpu-partitioning","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/technology/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/green-carbon-footprint-for-model-inference-serving-via-exploiting-mixed-quality-models-and-gpu-partitioning/124588/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-05",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What challenge does Clover address in ML inference serving?","Question",{"text":75,"@type":76},"Clover targets the growing carbon emissions caused by large-scale datacenter and HPC ML inference workloads while still meeting performance and accuracy needs.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How do mixed-quality models and GPU partitioning reduce carbon emissions?",{"text":80,"@type":76},"Mixed-quality model variants enable carbon-accuracy trade-offs, while GPU partitioning improves resource utilization; together they reduce emissions but can affect latency and SLA, which Clover mitigates.",{"name":82,"@type":73,"acceptedAnswer":83},"How does Clover ensure SLA targets while optimizing for carbon?",{"text":84,"@type":76},"Clover’s optimization engine dynamically adapts to carbon intensity and selects mixed-quality model variants and GPU partitioning choices to reduce carbon and maintain high accuracy without violating SLA constraints.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,113,118,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":111,"slug":112},50,"technology",{"id":114,"doc_module":4,"doc_module_name":46,"category_name":115,"show_sort_weight":116,"slug":117},7,"Healthcare",40,"healthcare",{"id":119,"doc_module":4,"doc_module_name":46,"category_name":120,"show_sort_weight":121,"slug":122},8,"Research & Report",30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]