[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86318-en":3,"doc-seo-86318-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86318,13056703020460,"Valentina","https://ap-avatar.wpscdn.com/avatar/be000253dac470eee5d?_k=1778207105932848923",8,"Research & Report","From Global to Factor-Wise Expert Composition in Discrete Diffusion Models","Discrete diffusion models support compositional generation by combining multiple pre-trained experts to generalize beyond any single expert’s training data. Prior theoretical corrections add time-dependent mixing weights, yet they treat each generated state monolithically with global scalar weighting per expert. FactorDiff introduces factor-wise decomposition and dynamic per-factor expert routing, instantiated for spatial/pixel-level compositions. Experiments on the ARC-AGI benchmark show factor-specific routing consistently outperforms complex global weighting on tasks demanding logical consistency and spatial disentanglement.","From Global to Factor-Wise Expert Composition in Discrete Diffusion Models  \nHaozhe Huang 1,2 Yudong W. Xu2,3 Abhijoy Mandal 1,2 Alán Aspuru-Guzik 1,2,4  \n1Department of Computer Science, University of Toronto  \n2 Vector Institute for Artificial Intelligence  \n3 Department of Mechanical & Industrial Engineering, University of Toronto  \n4 Senior Fellow, Canadian Institute for Advanced Research (CIFAR)  \narXiv :2607 . 1 1758v 1 [ cs .LG] 13 Jul 2026  \nAbstract  \nDiscrete diffusion models offer a powerful framework for solving complex reasoning tasks, particularly through compositional generation, which combines multiple pre-trained experts to generalize beyond their individual training data. Recent theoretical corrections introduce time-dependent mixing weights to better align composed diffusion dynamics with the intended target. However, these methods are fundamentally limited by working on a per-sample basis, treating each generated state monolithically and ignoring the potential spatial or functional specializations of different experts. In this work, we address this limitation by proposing FactorDiff– a factor-wise composition framework for diffusion models. We posit that samples can be further decomposed into smaller factors, and propose a sampling process that dynamically routes each factor to the most relevant expert. We instantiate this framework with spatial/pixel-level compositions and validate it on the ARC-AGI benchmark, demonstrating that simple factor-specific routing consistently outperforms complex global scalar weighting schemes on tasks that require logical consistency and spatial disentanglement.  \n1 INTRODUCTION  \nDiscrete diffusion models have recently emerged as a powerful paradigm for solving complex reasoning and generative tasks by learning the underlying data distribution directly from samples [Austin et al., 2021, Lou et al., 2024, Ye et al., 2025] . A key advantage of this framework is its potential for compositional generation—the ability to combine multiple pre-trained models (or experts) to solve tasks that lie outside the training distribution of any single model [Du et al., 2020] . This capability is critical for real-world applications  \nwhere collecting data for every possible combination of desired properties is intractable.  \nRecent works have sought to improve compositional sampling by introducing time-dependent correction weights for each expert. Methods such as SuperDiff [Skreta et al., 2025b], RNE [He et al., 2026], and Feynman-Kac Correctors [Skreta et al., 2025a, Hasan et al., 2025] derive these weights mathematically to ensure the combined score field approximates the true product distribution. While theoretically grounded, these approaches share a fundamental limitation: they assign a single global scalar weight to each expert at every timestep. This implicitly assumes that each expert is equally knowledgeable across the entire set of variables in the state space x.  \nThis \"monolithic\" assumption is problematic for complex generative tasks where experts have specialized domains or conflicting objectives. For instance, in a spatial reasoning task, one expert may be valid only for a specific region, while another governs the boundary conditions. Assigning a single weight forces a compromise that dilutes the signal of the correct expert in its region of competence.  \nIn this work, we argue that effective composition requires a more granular approach. We propose a factor-wise composition framework in which the state is decomposed into user-defined factors, and expert contributions are routed atthe factor level rather than applied uniformly to the entire sample. Instead of assigning a single scalar weight for each expert, our method dynamically assigns weights per factor during sampling, enabling heterogeneous expert contributions within a single generated state. We instantiate this framework with position-level routing for grid-structured reasoning and demonstrate the benefits of this int","cbCaioZDJDLyVelh","https://ap.wps.com/l/cbCaioZDJDLyVelh","pdf",1378373,7,1,21,"English","en",105,"# Abstract\n# Introduction\n# Preliminaries\n## ARC-AGI\n## Concrete Scores\n## FactorDi","[{\"question\":\"What limitation do existing compositional sampling methods have in discrete diffusion models?\",\"answer\":\"They use a single global scalar weight for each expert at every timestep, assuming experts are equally valid across the entire state space. This monolithic weighting prevents experts from specializing for different regions or functional components.\"},{\"question\":\"How does FactorDiff change expert composition during sampling?\",\"answer\":\"It decomposes a sample into user-defined factors and dynamically routes each factor to the most relevant expert. This assigns weights at the factor level rather than uniformly across the whole generated state.\"},{\"question\":\"How is FactorDiff instantiated and evaluated in the paper?\",\"answer\":\"The framework is instantiated with spatial/pixel-level compositions using a 2D masked diffusion architecture to preserve grid topology. It is validated on the ARC-AGI benchmark and yields gains over per-sample scalar-weight baselines, especially when experts are highly specialized and complementary.\"}]",1784210452,53,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"from-global-to-factor-wise-expert-composition-in-discrete-diffusion-models","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/from-global-to-factor-wise-expert-composition-in-discrete-diffusion-models/86318/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-27","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What limitation do existing compositional sampling methods have in discrete diffusion models?","Question",{"text":76,"@type":77},"They use a single global scalar weight for each expert at every timestep, assuming experts are equally valid across the entire state space. This monolithic weighting prevents experts from specializing for different regions or functional components.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How does FactorDiff change expert composition during sampling?",{"text":81,"@type":77},"It decomposes a sample into user-defined factors and dynamically routes each factor to the most relevant expert. This assigns weights at the factor level rather than uniformly across the whole generated state.",{"name":83,"@type":74,"acceptedAnswer":84},"How is FactorDiff instantiated and evaluated in the paper?",{"text":85,"@type":77},"The framework is instantiated with spatial/pixel-level compositions using a 2D masked diffusion architecture to preserve grid topology. It is validated on the ARC-AGI benchmark and yields gains over per-sample scalar-weight baselines, especially when experts are highly specialized and complementary.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":46,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":107,"slug":138},19,"General","general"]