[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82046-en":3,"doc-seo-82046-105":29,"detail-sidebar-cat-0-en-105":90},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},82046,7971461740909,"Levi","https://ap-avatar.wpscdn.com/davatar_155a257f0dc6eb9ab79c44ca47cae57d",8,"Research & Report","HERO: A Heterogeneity-Aware Benchmark Library for Federated Continual Learning","Federated continual learning (FCL) must measure how distributed clients learn from evolving data while preserving prior knowledge. Existing FCL evaluations are difficult to compare because they simultaneously alter datasets, task splits, client splits, task orders, models, memory assumptions, and reporting rules. HERO is introduced as a heterogeneity-aware benchmark library that disentangles task split, client data split, and client task sequence. HERO-Core uses α (client skew) and ρ (task-order mismatch) to evaluate methods, including image-based and graph-based DomainIL portability results.","arXiv :2607 .08784v1 [ cs .LG] 13 Jun 2026  \nHERO: A Heterogeneity-Aware Benchmark Library for Federated Continual Learning  \nThinh T. H. NguyenLe-Tuan Nguyen∗ , Minh-Duong Nguyen, Nhi Trinh, Anh Tran Nam Nguyet, Dung D. Le†, Kok-Seng Wong†  \nVinUniversity, Hanoi, Vietnam  \n{thinh.nth,[tuan.nl](tuan.nl),duong.nm2,23nhi.ttt,23anh.tnn2,dung.ld,[wong.ks}@vinuni.edu.vn](wong.ks}@vinuni.edu.vn)  \nAbstract  \nFederated continual learning (FCL) evaluates how distributed clients learn from changing data streams while retaining previously learned knowledge. Existing evaluations are difficult to compare because they often change datasets, task splits, client data splits, task orders, backbones, memory assumptions, and reporting rules simultaneously. We introduce HERO, a heterogeneity-aware benchmark library for FCL. HERO builds benchmark streams by separating three choices that are often coupled, namely the task split, the client data split, and the client task sequence. In HERO-Core, the main comparable benchmark, α controls client data skew and ρ controls task-order mismatch. We evaluate representative FCL methods on CIFAR-100 and TinyImageNet using final average accuracy, average forgetting, and bottom-10% client accuracy. We also include a graph-based DomainIL portability case study on OGB-MolPCBA, where scaffold-domain granularity changes the input distribution while the prediction task remains fixed. Our results show that method behavior changes across easy and heterogeneous settings, that average accuracy can hide weak bottom-client performance, that task-order mismatch favors different strategies from synchronized evaluation, and that the same HERO interface can expose domain-shift difficulty beyond image-based FCIL.  \nHERO releases benchmark streams, configurations, method implementations, and reporting scripts to support reproducible and setting-aware FCL evaluation.  \n1 Motivation and Benchmark Gap  \nFederated learning (FL) allows many clients to train a shared model without centralizing raw data, but standard FL evaluation usually treats the learning problem as fixed over time [1, 2] . This assumption is often too narrow in practice. Clients may collect data continuously, new classes may appear, and the model may need to adapt to new distributions without losing earlier knowledge. Federated continual learning (FCL) studies this setting by combining federated optimization with continual learning (CL) [3–5] . This combination makes evaluation difficult. On the FL side, methods must handle non-IID client data, client drift, and partial participation [6–10] . On the CL side, methods must handle non-stationary streams, growing task sequences, and forgetting [11–16] . A common FCL setting is federated class-incremental learning (FCIL), where clients and the server learn an expanding label space without task identity at test time [17–19] .  \nRecent FCL methods use replay, generative replay, distillation, task tracing, prompt tuning, personalization, class balancing, gradient correction, and resource-aware training [18–25] . These techniques address important problems, but their reported results are still hard to compare. Different  \npapers often change the dataset, task split, client data split, task sequence, backbone, memory bud-∗ Co-first Authors.  \n†Co-corresponding Authors: [dung.ld@vinuni.edu.vn](dung.ld@vinuni.edu.vn and wong.ks@vinuni.edu.vn)[ and](dung.ld@vinuni.edu.vn and wong.ks@vinuni.edu.vn)[ wong.ks@vinuni.edu.vn](dung.ld@vinuni.edu.vn and wong.ks@vinuni.edu.vn).  \nPreprint.  \nget, and reporting rule at the same time. As a result, a method may look stronger because it is more robust, but it may also look stronger because it is evaluated under an easier stream or a different set of assumptions. This creates a benchmark gap. FCL needs a reusable library that can construct comparable streams, expose the main sources of heterogeneity, document method assumptions, and report more than one average score. Without such a libra","cbCaijtUrYQtJpAY","https://ap.wps.com/l/cbCaijtUrYQtJpAY","pdf",2888902,1,30,"English","en",105,"# Motivation and Benchmark Gap\n## Federated continual learning setting and evaluation challenges\n## Benchmark gap and need for a reusable library\n# HERO: Design and Benchmark Construction\n## Separating task split, client data split, and client task sequence\n## HERO-Core comparable benchmark\n# Experimental Evaluation and Findings\n## Image-based FCIL benchmarks\n## Graph-based DomainIL portability case study","[{\"question\":\"Why are current federated continual learning (FCL) results hard to compare across papers?\",\"answer\":\"Because evaluations often change multiple ingredients at once, including datasets, task splits, client splits, task orders, model backbones, memory assumptions, and reporting rules. This makes differences in performance ambiguous.\"},{\"question\":\"How does HERO isolate key sources of heterogeneity in FCL evaluation?\",\"answer\":\"HERO separates three often-coupled choices: the task split, the client data split, and the client task sequence. HERO-Core then uses α to control client data skew and ρ to control task-order mismatch.\"},{\"question\":\"What metrics and benchmark settings does HERO-Core use to assess methods?\",\"answer\":\"Methods on CIFAR-100 and TinyImageNet are evaluated using final average accuracy, average forgetting, and bottom-10% client accuracy, highlighting both overall performance and weak client behavior. A portability case study extends evaluation via a graph-based DomainIL setting on OGB-MolPCBA.\"}]",1784177790,76,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":85,"head_meta":87,"extra_data":89,"updated_unix":27},"hero-a-heterogeneity-aware-benchmark-library-for-federated-continual-learning","",{"@graph":35,"@context":84},[36,53,67],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/hero-a-heterogeneity-aware-benchmark-library-for-federated-continual-learning/82046/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":61,"encodingFormat":60,"isAccessibleForFree":62,"interactionStatistic":63},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-16",true,{"@type":64,"interactionType":65,"userInteractionCount":4},"InteractionCounter",{"@type":66},"ViewAction",{"@type":68,"mainEntity":69},"FAQPage",[70,76,80],{"name":71,"@type":72,"acceptedAnswer":73},"Why are current federated continual learning (FCL) results hard to compare across papers?","Question",{"text":74,"@type":75},"Because evaluations often change multiple ingredients at once, including datasets, task splits, client splits, task orders, model backbones, memory assumptions, and reporting rules. This makes differences in performance ambiguous.","Answer",{"name":77,"@type":72,"acceptedAnswer":78},"How does HERO isolate key sources of heterogeneity in FCL evaluation?",{"text":79,"@type":75},"HERO separates three often-coupled choices: the task split, the client data split, and the client task sequence. HERO-Core then uses α to control client data skew and ρ to control task-order mismatch.",{"name":81,"@type":72,"acceptedAnswer":82},"What metrics and benchmark settings does HERO-Core use to assess methods?",{"text":83,"@type":75},"Methods on CIFAR-100 and TinyImageNet are evaluated using final average accuracy, average forgetting, and bottom-10% client accuracy, highlighting both overall performance and weak client behavior. A portability case study extends evaluation via a graph-based DomainIL setting on OGB-MolPCBA.","https://schema.org",{"og:url":51,"og:type":86,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":88,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":91},[92,96,100,104,109,114,119,121,126,129,133],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":93,"show_sort_weight":94,"slug":95},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":97,"show_sort_weight":98,"slug":99},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":101,"show_sort_weight":102,"slug":103},"Exam",70,"exam",{"id":105,"doc_module":4,"doc_module_name":45,"category_name":106,"show_sort_weight":107,"slug":108},5,"Comic",60,"comic",{"id":110,"doc_module":4,"doc_module_name":45,"category_name":111,"show_sort_weight":112,"slug":113},6,"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":45,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":21,"slug":120},"research-report",{"id":122,"doc_module":4,"doc_module_name":45,"category_name":123,"show_sort_weight":124,"slug":125},9,"Religion & Spirituality",20,"religion-spirituality",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":127,"show_sort_weight":124,"slug":128},"World Cup","world-cup",{"id":130,"doc_module":4,"doc_module_name":45,"category_name":131,"show_sort_weight":130,"slug":132},10,"Lifestyle","lifestyle",{"id":134,"doc_module":4,"doc_module_name":45,"category_name":135,"show_sort_weight":105,"slug":136},19,"General","general"]