[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84748-en":3,"doc-seo-84748-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},84748,137441390410,"Hazel","https://ap-avatar.wpscdn.com/avatar/2000252f4ab5702993?_k=1776741390130283984",8,"Research & Report","Lights, Camera, Carbon: Architectural Scaling Laws for Video Generation Energy Consumption","A bidirectional framework estimates energy consumption of text-to-video (T2V) and text-to-video-audio (T2VA) models from observable generation parameters and architectural first principles, without requiring weights, model size, or implementation access. Forward prediction derives energy from resolution, duration, and batch size; backward recovery infers architectural scaling from measured inference times, using prediction accuracy to validate architectural consistency. Energy profiles of compute-bound diffusion models decompose into quadratic and linear scaling terms with coefficients tied to architectural complexity, validated on open models and GPU setups for standardized sustainability benchmarking.","arXiv :2607 .04553v 1 [ cs .MM] 5 Jul 2026  \nLights, Camera, Carbon: Architectural Scaling Laws for Video Generation Energy Consumption  \nNidhal Jegham  \nUniversity of Rhode Island, Sustainable AI Group Rhode Island, United States [nidhal@sustainableaigroup.com](nidhal@sustainableaigroup.com)  \nBoris Gamazaychikov  \nSustainable AI Group Paris, France  \nSasha Luccioni  \nSustainable AI Group Montreal, Canada  \nFigure 1: Measured open-model energy consumption versus estimated proprietary-model energy consumption for an 8-second 720p generation.  \nAbstract  \nWe present a bidirectional framework for estimating the energy consumption of text-to-video (T2V) and text-to-video-audio (T2VA) models from architectural first principles and observable generation parameters such as resolution and duration, requiring no access to weights, model size, or implementation details. Forward, it predicts energy from generation parameters and architectural principles; backward, it recovers architectural scaling behavior from observed inference times, with accuracy serving as a criterion for architectural validity. Building on the established compute-bound nature of video diffusion models, we demonstrate that each model’s energy profile obeys theoretically derived scaling laws, decomposing into quadratic and linear terms whose coefficients directly reflect the underlying architectural complexity. Validated across six open-source models spanning 8.3B–27B parameters and three GPU configurations, this decomposition achieves below 3% MAPE across all architectures. This approach offers a standardized, empirically and theoretically grounded framework for sustainability benchmarking across T2V models and architectures.  \nPreprint.  \n1 Introduction  \nRecent years have seen a rise in the widescale deployment of machine learning (ML) systems ina variety of user-facing applications and tasks, generating responses to user queries in modalities ranging from text [1] to audio [2] and images [3] . However, the rapidly growing environmental impacts associated with this growth are an increasingly important topic to examine and factor into the development and deployment of ML systems [4] . As such, the sustainability and, more specifically, the energy demands of different ML systems have become a nascent but rapidly developing field of study. Starting with the seminal work of Strubell et al in 2019 [5], the subsequent years of scholarship have shed more light on different ML tasks, the relative contribution of the different stages of the ML life cycle, as well as the factors that influence them, mainly focusing on textbased models [6, 7, 8, 9, 10, 11] . Most recently, the topic has been broadened to also include other modalities such as image generation and speech-to-text [12, 13] .  \nGiven that text-to-video (T2V) generation is a relatively novel ML task, its energy consumption has yet to be analyzed in depth in an empirical way. The few existing studies that examine this modality have focused on specific models [14] or a few open model families [15] – these studies have found that video generation is not simply a scaled-up version of image generation, mainly because it requires iterative denoising across both spatial and temporal dimensions, which generate hundreds of frames per output (as opposed to a single image, in the case of image generation) . The most complete study to date, by Delavande et al. [15], found that AI video generation operates in a compute-bound regime, with its energy requirements and latency scaling near-quadratically as the resolution and length of output videos increases. Comparing video generation with other modalities, they found that generating a single short video can consume approximately 90 Wh of energy, making it 30 times more costly than image generation and over 2,000 times more costly than text generation.  \nIn the current study, we go beyond previous analyses to theoretically derive and empirically validate the scaling laws governi","cbCaioV9yz41aoEZ","https://ap.wps.com/l/cbCaioV9yz41aoEZ","pdf",758738,2,1,17,"English","en",105,"# Introduction\n## Background\n# Methodology\n## Bidirectional estimation framework\n# Results\n## Energy scaling factors and validation\n# Discussion\n## Significance and future work","[{\"question\":\"How does the framework estimate energy consumption without accessing model weights or implementation details?\",\"answer\":\"It uses a bidirectional approach: forward prediction estimates energy from observable generation parameters and architectural first principles, while backward analysis recovers architectural scaling from measured inference times. Accuracy then serves as a criterion for architectural validity.\"},{\"question\":\"What scaling behavior is expected for video diffusion model energy profiles?\",\"answer\":\"The document states that, under the compute-bound regime, each model’s energy profile obeys theoretically derived scaling laws. The energy decomposes into quadratic and linear terms whose coefficients reflect underlying architectural complexity.\"},{\"question\":\"How was the approach validated across models and hardware?\",\"answer\":\"Validation uses six open-source models spanning 8.3B–27B parameters and three GPU configurations. The decomposition achieves below 3% MAPE across all architectures, and the framework is also applied to estimate energy for proprietary systems where direct measurement is infeasible.\"}]",1784198018,43,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"lights-camera-carbon-architectural-scaling-laws-for-video-generation-energy-consumption","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/lights-camera-carbon-architectural-scaling-laws-for-video-generation-energy-consumption/84748/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-23","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"How does the framework estimate energy consumption without accessing model weights or implementation details?","Question",{"text":75,"@type":76},"It uses a bidirectional approach: forward prediction estimates energy from observable generation parameters and architectural first principles, while backward analysis recovers architectural scaling from measured inference times. Accuracy then serves as a criterion for architectural validity.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What scaling behavior is expected for video diffusion model energy profiles?",{"text":80,"@type":76},"The document states that, under the compute-bound regime, each model’s energy profile obeys theoretically derived scaling laws. The energy decomposes into quadratic and linear terms whose coefficients reflect underlying architectural complexity.",{"name":82,"@type":73,"acceptedAnswer":83},"How was the approach validated across models and hardware?",{"text":84,"@type":76},"Validation uses six open-source models spanning 8.3B–27B parameters and three GPU configurations. The decomposition achieves below 3% MAPE across all architectures, and the framework is also applied to estimate energy for proprietary systems where direct measurement is infeasible.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]