[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-118342-en":3,"doc-seo-118342-105":30,"detail-sidebar-cat-0-en-105":95},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},118342,7971461740909,"Levi","https://ap-avatar.wpscdn.com/davatar_155a257f0dc6eb9ab79c44ca47cae57d",8,"Research & Report","Better, Not Just More - Data-centric machine learning for Earth observation","Remote sensing data connect pressing global challenges with machine learning opportunities, but achieving full potential requires deep understanding of data and their roles across the ML pipeline. The work contrasts a predominantly model-centric research focus with the need to treat data quality, acquisition, and curation as critical cycle components. It argues that benchmark-driven, data-static development leads to limited robustness in real deployment, motivating criteria for data quality. Emphasis is placed on enabling adaptive solutions through quality and quantity considerations for specific domains.","Better, Not Just More  \nData-centric machine learning for Earth observation  \nRIBANA ROSCHER, MARC RUSSWURM, CAROLINE GEVAERT,  \nMICHAEL KAMPFFMEYER,  \nJEFERSSON A. DOS SANTOS,  \nMARIA VAKALOPOULOU, RONNY HÄNSCH, STINE HANSEN, KEILLER NOGUEIRA, JONATHAN PREXL, AND DEVIS TUIA  \nRemote sensing data are a central link between press  \ning global challenges and the possibilities of machine learning (ML) methods. To realize their full potential, a deep understanding of the data and their utilization possibilities is essential. Generally, data play a fundamental role in the ML cycle, with steps from 1) informing the problem definition, 2) data creation, 3) data curation, 4) enabling model training, and 5) evaluation to 6) the eventual model deployment that feeds back to modifying the problem definition. We show this cycle in Figure 1, where each node is colored by its focus of being problem-centric (darkgray), data-centric (blue), or model-centric (dark green). Yet, current research in ML is predominantly model-centric and focuses on model design and evaluation (step 4). This primarily emphasizes optimizing the accuracy and efficiency of the models themselves [1], [2], [3] and considers the dataset a static benchmark rather than a dynamic representation of the application. Even in applied ML areas, such as geospatial data analysis, the research toward integrating new ML methods has recently largely focused on refining the algorithms and fine-tuning the model parameters [4]. Data-centric aspects like data acquisition (step 1) and curation (step 2) are done manually to a large degree, and data quality is rarely considered. At the same time, most domain expertise remains concentrated in classic feature engineering.  \nINTRODUCTION  \nThe rapid development of ML methods has been primarily driven by the availability of large-scale datasets and advancements in computing power [5], [6]. This has allowed researchers to focus on developing complex models capable  \nDigital Object Identifier 10.1109/MGRS.2024.3470986  \nDate of publication 31 October 2024; date of current version 12 December 2024.  \n©SHUTTERSTOCK.COM/ELENABSL  \nof capturing patterns and relationships within the data. However, many developments have been pursued in controlled benchmark settings only, where the natural variability and real-world characteristics of the data are bypassed. In the geospatial domain, for example, this is achieved through one-time predefined data location sampling, as seen in the predefined regions of the Sen12MS dataset [7], or through class balancing techniques commonly used in crop type mapping [8]. Although the great importance of these benchmarks should not be diminished, as a result, many methods developed are often detached from reality and seldom robust in real-world deployment. This requires considering all steps in the ML pipeline and acknowledging them as a cycle (Figure 1) as well as a high degree of robustness for acquisition specification and generalization to diverse shifts.  \nDespite the availability of enormous amounts of data [6], including well-curated benchmarks, and significant progress made through methodological advances, thereis currently a saturation point where many established architectures and training methods achieve comparable accuracy ([9], [139]) . This is exemplified by foundation models or general-purpose computing architectures for multiple modalities and tasks [10], [11], [12], [13] . These large-scale models, trained on massive but mainly domain-agnostic datasets, increasingly replace modality-specific recurrent and convolutional models . This  \nDECEMBER 2024 IEEE GEOSCIENCE AND REMOTE SENSING MAGAZINE 2473-2397/24©2024IEEE 335  \nAuthorized licensed use limited to: UNIVERSITY OF TWENTE.. Downloaded on January 13,2025 at 08:39:05 UTC from IEEE Xplore. Restrictions apply.  \n\n| \u003Cbr>development comes with the prevailing belief that “more substantial scientific value. Therefore, while foundation |  |\n| --- | --- |\n| data ar","cbCaisJRt3qj08zx","https://ap.wps.com/l/cbCaisJRt3qj08zx","pdf",3548152,1,21,"English","en",105,"# Introduction\n## Data-centric vs model-centric ML pipeline\n## Limits of benchmark-driven development\n## Need for data quality criteria\n## Towards automated data interaction","[{\"question\":\"为什么文章强调数据中心（data-centric）的机器学习？\",\"answer\":\"文章指出当前研究多偏向模型设计与评估，而数据在问题定义、数据创建、数据整理、训练与部署反馈中扮演基础作用。要释放遥感数据的潜力，必须理解数据并提升其可用性。\"},{\"question\":\"文章如何描述机器学习流程在数据上的关键环节？\",\"answer\":\"文章将ML周期视为循环，包含问题定义、数据创建、数据整理、模型训练、评估以及最终部署，并指出数据获取与整理在很大程度上仍依赖手工，同时数据质量常被忽视。\"},{\"question\":\"为什么基于受控基准的研究可能难以在真实场景中落地？\",\"answer\":\"文章认为许多方法在受控基准设置下追求模式捕获，忽略了真实数据的自然变异与特征，因此在真实部署时往往缺乏鲁棒性。\"},{\"question\":\"文章提出了哪些与数据质量相关的方向或标准？\",\"answer\":\"文章指出与地理空间领域相关、对ML尤其重要的五项数据质量标准，包括多样性与完整性、准确性、一致性、无偏性以及任务相关性。\"}]","Better, Not Just More - Data-centric machine learning for Earth observation | PDF",1785683175,53,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":90,"head_meta":92,"extra_data":94,"updated_unix":28},"better-not-just-more-data-centric-machine-learning-for-earth-observation","",{"@graph":36,"@context":89},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/better-not-just-more-data-centric-machine-learning-for-earth-observation/118342/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-02",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81,85],{"name":72,"@type":73,"acceptedAnswer":74},"为什么文章强调数据中心（data-centric）的机器学习？","Question",{"text":75,"@type":76},"文章指出当前研究多偏向模型设计与评估，而数据在问题定义、数据创建、数据整理、训练与部署反馈中扮演基础作用。要释放遥感数据的潜力，必须理解数据并提升其可用性。","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"文章如何描述机器学习流程在数据上的关键环节？",{"text":80,"@type":76},"文章将ML周期视为循环，包含问题定义、数据创建、数据整理、模型训练、评估以及最终部署，并指出数据获取与整理在很大程度上仍依赖手工，同时数据质量常被忽视。",{"name":82,"@type":73,"acceptedAnswer":83},"为什么基于受控基准的研究可能难以在真实场景中落地？",{"text":84,"@type":76},"文章认为许多方法在受控基准设置下追求模式捕获，忽略了真实数据的自然变异与特征，因此在真实部署时往往缺乏鲁棒性。",{"name":86,"@type":73,"acceptedAnswer":87},"文章提出了哪些与数据质量相关的方向或标准？",{"text":88,"@type":76},"文章指出与地理空间领域相关、对ML尤其重要的五项数据质量标准，包括多样性与完整性、准确性、一致性、无偏性以及任务相关性。","https://schema.org",{"og:url":52,"og:type":91,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":93,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":96},[97,101,105,109,114,119,124,127,132,135,139],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":106,"show_sort_weight":107,"slug":108},"Exam",70,"exam",{"id":110,"doc_module":4,"doc_module_name":46,"category_name":111,"show_sort_weight":112,"slug":113},5,"Comic",60,"comic",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},6,"Technology",50,"technology",{"id":120,"doc_module":4,"doc_module_name":46,"category_name":121,"show_sort_weight":122,"slug":123},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":125,"slug":126},30,"research-report",{"id":128,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":130,"slug":131},9,"Religion & Spirituality",20,"religion-spirituality",{"id":130,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":130,"slug":134},"World Cup","world-cup",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":136,"slug":138},10,"Lifestyle","lifestyle",{"id":140,"doc_module":4,"doc_module_name":46,"category_name":141,"show_sort_weight":110,"slug":142},19,"General","general"]