[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84214-en":3,"doc-seo-84214-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},84214,962075114101,"Seraphina","https://ap-avatar.wpscdn.com/avatar/e000253a75eb197efd?x-image-process=image/resize,m_fixed,w_180,h_180&k=1780044092746381165",8,"Research & Report","CarbonCLIP: Enhance Carbon Prediction from Satellite Imagery via Integrated Street View Semantics and Temporal Context Training","Accurately estimating urban carbon emissions is essential for sustainable urban planning, but current satellite-based approaches are hard to apply across cities due to data heterogeneity and limited fine-grained semantic-temporal context. CarbonCLIP introduces a task-oriented multimodal distillation framework that transfers contextual knowledge into a unified satellite representation using dual-branch contrastive learning. Spatial priors come from street-view text descriptions generated by large multimodal models, and temporal priors are encoded by a month encoder. Pretraining uses multimodal data, while inference uses only satellite imagery for scalable deployment; experiments on Beijing and Singapore validate improved prediction performance.","arXiv :2607 .07292v 1 [ cs .CV] 8 Jul 2026  \nCARBONCLIP: ENHANCE CARBON PREDICTION FROM SATELLITE IMAGERY VIA INTEGRATED STREET-VIEW SEMANTICS AND TEMPORAL CONTEXT TRAINING  \nZeru Yang 1 ,2 , Fang-Ying Gong3 , Steve H.L. Yim4 ,5 , Chau Yuen2 ,5  \n1Energy Research Institute at NTU, Interdisciplinary Graduate Programme, Nanyang Technological University, Singapore  \n2 School of Electrical and Electronic Engineering, Nanyang Technological University, Singapore  \n3 School of Public Administration and Policy, Renmin University of China, Beijing, China  \n4Asian School of the Environment, Nanyang Technological University, Singapore  \n5 Center for Climate Change and Environmental Health, Nanyang Technological University, Singapore [zeru001@e.ntu.edu.sg](zeru001@e.ntu.edu.sg) , [fangying.gong@ruc.edu.cn](fangying.gong@ruc.edu.cn) , [steve.yim@ntu.edu.sg](steve.yim@ntu.edu.sg) , [chau.yuen@ntu.edu.sg](chau.yuen@ntu.edu.sg)  \nABSTRACT  \nAccurately estimating urban carbon emissions is critical for sustainable urban planning, yet many  \nexisting approaches remain difficult to apply consistently across cities due to data-source hetero  \ngeneity and the lack of fine-grained semantic-temporal context in remote sensing data. We propose  \nCarbonCLIP, a task-oriented multimodal distillation framework that improves satellite-based car  \nbon emission prediction by transferring contextual knowledge into a unified satellite representation  \nthrough dual-branch contrastive learning. Unlike conventional methods that rely on static visual  \nfeatures, CarbonCLIP explicitly bridges the gap between top-down satellite views and ground-level  \nhuman activities. Specifically, the spatial branch uses fine-grained textual descriptions automatically  \ngenerated from street-view images by Large Multimodal Models (LMMs) to provide semantic priors  \nreflecting building functions, infrastructure, and urban activities, while the temporal branch employs  \na month encoder to encode temporal priors associated with monthly emission variation. CarbonCLIP  \nrequires multimodal data only during the pretraining phase; during inference, it relies solely on  \nsatellite imagery, thereby supporting scalable deployment when ground-level data are unavailable  \nat inference. Experiments on Beijing and Singapore demonstrate that CarbonCLIP outperforms  \nbaselines in both study cities. The results validate that our method effectively transfers multimodal  \nknowledge into satellite representations, offering a robust solution for satellite-based urban carbon  \nmodeling.  \n1 Introduction  \nWith the rapid urbanization occurring worldwide, cities have become the dominant sources of anthropogenic carbon emissions, accounting for more than one-third of global totals [1, 2] . Accurately quantifying and estimating urban carbon emissions is essential for developing sustainable and climate-resilient cities [3] . Urban sustainability requires not only reducing carbon footprints but also integrating intelligent monitoring tools that can inform policy-making and promote equitable environmental outcomes. A core challenge in achieving this goal lies in how we perceive and analyze cities at scale. Carbon emissions in urban areas are deeply intertwined with spatial patterns such as land use and infrastructure, as well as human-scale factors like greenery, density, and activity levels.  \nIn recent years, visual data have become a cornerstone of urban analytics and environmental monitoring, providing unprecedented means of observing cities at multiple scales. Satellite imagery offers a comprehensive overview of urban environments, capturing spatial structure and physical layout [4] . As a scalable and consistent data source, it has been widely applied in air pollution mapping [5], agricultural monitoring [6], and carbon stock estimation [7] . Advancesin remote sensing, including high-resolution Planet imagery at 3 m resolution [8], enable detailed detection of urban expansion and environmental change [9","cbCaij0UrWOLUFUM","https://ap.wps.com/l/cbCaij0UrWOLUFUM","pdf",1799564,4,1,21,"English","en",105,"# Introduction\n## Challenges in Satellite-Based Carbon Estimation\n## Limits of Static Visual Features\n## Motivation for Human-Centric and Temporal Context\n## Framework Overview: CarbonCLIP","[{\"question\":\"What problem does CarbonCLIP address in urban carbon emission prediction?\",\"answer\":\"It tackles the difficulty of applying existing satellite-based methods across cities caused by heterogeneous data sources and the lack of fine-grained semantic-temporal context in remote sensing inputs.\"},{\"question\":\"How does CarbonCLIP integrate street-view information into satellite representation learning?\",\"answer\":\"It uses a spatial branch where street-view images are converted into fine-grained textual descriptions by large multimodal models, providing semantic priors tied to building functions, infrastructure, and activities.\"},{\"question\":\"When is street-view or other multimodal data required for CarbonCLIP?\",\"answer\":\"Multimodal data are used only during pretraining; during inference, the method relies solely on satellite imagery, enabling scalable deployment when ground-level data are unavailable.\"}]",1784194044,53,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"carbonclip-enhance-carbon-prediction-from-satellite-imagery-via-integrated-street-view-semantics-and-temporal-context-training","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/carbonclip-enhance-carbon-prediction-from-satellite-imagery-via-integrated-street-view-semantics-and-temporal-context-training/84214/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-25","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does CarbonCLIP address in urban carbon emission prediction?","Question",{"text":75,"@type":76},"It tackles the difficulty of applying existing satellite-based methods across cities caused by heterogeneous data sources and the lack of fine-grained semantic-temporal context in remote sensing inputs.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does CarbonCLIP integrate street-view information into satellite representation learning?",{"text":80,"@type":76},"It uses a spatial branch where street-view images are converted into fine-grained textual descriptions by large multimodal models, providing semantic priors tied to building functions, infrastructure, and activities.",{"name":82,"@type":73,"acceptedAnswer":83},"When is street-view or other multimodal data required for CarbonCLIP?",{"text":84,"@type":76},"Multimodal data are used only during pretraining; during inference, the method relies solely on satellite imagery, enabling scalable deployment when ground-level data are unavailable.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]