[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84507-en":3,"doc-seo-84507-105":29,"detail-sidebar-cat-0-en-105":83},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},84507,549758146520,"Patrick","https://ap-avatar.wpscdn.com/avatar/80002397d8c0411e94?_k=1775819394049821470",8,"Research & Report","Diffusion Language Model Parallel Decoding via Product-of-Experts Bridge","Diffusion language models enable parallel decoding and can greatly reduce generation latency, yet their token independence limits output quality versus autoregressive models. Recent bridging through importance sampling suffers from expensive large-particle requirements due to distribution mismatch. PoE-Bridge introduces an intermediate Product-of-Experts distribution built from the DLM proposal and AR target. It drafts candidates in parallel, uses rejection sampling to move them toward the PoE, then importance-samples to correct toward the AR target. Mixed-temperature sampling and elastic rejection windows improve diversity and reduce wasted verification, yielding up to 5× speedup while recovering at least 95% of AR performance.","Diffusion Language Model Parallel Decoding via Product-of-Experts Bridge  \nJuntong Shi 1 Brian L. Trippe 1 Jure Leskovec 1 Stefano Ermon 1 Minkai Xu 1  \narXiv :2606 .08048v 1 [ cs .CL] 6 Jun 2026  \nAbstract  \nDiffusion language models (DLMs) offer substantial speed advantages through parallel decoding, but the lack of token dependencies limits generation quality compared to autoregressive (AR) models. Recent progress attempts to bridge the gap via importance sampling, with DLM being the proposal and AR being the target. However, due to the huge gap between their distributions, the sampling requires a large number of particles and is thus expensive to compute. In this paper, we introduce PoE-Bridge, a novel decoding framework that drastically improves generation speed and accuracy by introducing an intermediate distribution to bridge the gap. The distribution is constructed as a Product-of-Experts (PoE) of the DLM proposal and the AR target. With the intermediate distribution, we first use the DLM to draft multiple continuations in parallel, then apply rejection sampling to verify the drafted tokens and move the resulting candidates toward the PoE. We then use importance sampling to further correct the PoE-aligned candidates toward the AR target. We further propose several improved techniques, including mixed-temperature sampling for enhanced diversity and elastic rejection windows for reducing wasted verification. Empirically, PoE-Bridge achieves significantly improved accuracy with 5 × speedup over the standard DLM decoding approach, and recovers at least 95% of the target AR model’s performance, efficiently advancing most of the quality gap on challenging mathematical reasoning and coding tasks. Our code is available at [https://github.com/](https://github.com/)[ ](https://github.com/)juntongshi48/poe-bridge.  \nBT, JL, and SE co-supervised the project, secured resources, and contributed to the conceptualization and critical revision of the manuscript. 1 Stanford University. Correspondence to: Juntong Shi  \n\u003C[juntong@stanford.edu](juntong@stanford.edu)>, Minkai Xu \u003C[minkai@cs.stanford.edu](minkai@cs.stanford.edu)>.  \nProceedings of the 43 rd International Conference on Machine Learning, Seoul, South Korea. PMLR 306, 2026 . Copyright 2026 by the author(s) .  \n1. Introduction  \nAutoregressive (AR) language models remain the dominant approach to text generation (Vaswani et al., 2017 ; Wolf et al., 2020 ; Touvron et al., 2023), achieving strong performance on challenging tasks such as mathematical reasoning (Wei et al., 2022 ; Trinh et al., 2024 ; Yang et al., 2024) and code synthesis (Roziere et al., 2023 ; Hui et al., 2024) . However, AR decoding is inherently sequential—generating tokens one by one in a strict left-to-right order. This sequential dependency inherently limits inference-time parallelism and makes latency and throughput a major bottleneck.  \nDiffusion language models (DLMs) (Sahoo et al., 2024 ; Shi et al., 2024 ; Ye et al., 2025 ; Nie et al., 2025), derived from the broader discrete diffusion framework (Austin et al., 2021b ; Campbell et al., 2022 ; Lou et al., 2024), provide an appealing alternative by enabling parallel decoding. By generating and refining multiple tokens simultaneously, DLM can, in principle, produce entire sequences in a fraction of the sequential steps required by AR models. In practice, however, DLMs still lag behind strong AR models in generation quality. This gap largely stems from the conditional independence assumptions used to enable parallel decoding, under which tokens generated within the same step are modeled independently rather than jointly. As a result, the anticipated speedups from parallel generation are difficult to realize without incurring substantial degradation in output quality.  \nThis gap suggests that effective parallel decoding requires a mechanism to inject inter-token dependencies into DLM drafts while preserving their parallel sampling efficiency. Notably, while sa","cbCaibODgsmsZvtp","https://ap.wps.com/l/cbCaibODgsmsZvtp","pdf",928208,1,17,"English","en",105,"# Introduction\n# PoE-Bridge Method\n## Intermediate Product-of-Experts Distribution\n## Parallel Drafting and Rejection Sampling\n## Importance Sampling Toward Autoregressive Target\n# Experimental Results","[{\"question\":\"Which enhancements does PoE-Bridge add to improve diversity and efficiency?\",\"answer\":\"The paper proposes mixed-temperature sampling to enhance diversity and elastic rejection windows to reduce wasted verification, improving both generation quality and computational efficiency.\"}]",1784196192,43,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":78,"head_meta":80,"extra_data":82,"updated_unix":27},"diffusion-language-model-parallel-decoding-via-product-of-experts-bridge","",{"@graph":35,"@context":77},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/diffusion-language-model-parallel-decoding-via-product-of-experts-bridge/84507/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71],{"name":72,"@type":73,"acceptedAnswer":74},"Which enhancements does PoE-Bridge add to improve diversity and efficiency?","Question",{"text":75,"@type":76},"The paper proposes mixed-temperature sampling to enhance diversity and elastic rejection windows to reduce wasted verification, improving both generation quality and computational efficiency.","Answer","https://schema.org",{"og:url":51,"og:type":79,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":81,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":84},[85,89,93,97,102,107,112,115,120,123,127],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":86,"show_sort_weight":87,"slug":88},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":90,"show_sort_weight":91,"slug":92},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Exam",70,"exam",{"id":98,"doc_module":4,"doc_module_name":45,"category_name":99,"show_sort_weight":100,"slug":101},5,"Comic",60,"comic",{"id":103,"doc_module":4,"doc_module_name":45,"category_name":104,"show_sort_weight":105,"slug":106},6,"Technology",50,"technology",{"id":108,"doc_module":4,"doc_module_name":45,"category_name":109,"show_sort_weight":110,"slug":111},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":113,"slug":114},30,"research-report",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},9,"Religion & Spirituality",20,"religion-spirituality",{"id":118,"doc_module":4,"doc_module_name":45,"category_name":121,"show_sort_weight":118,"slug":122},"World Cup","world-cup",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":124,"slug":126},10,"Lifestyle","lifestyle",{"id":128,"doc_module":4,"doc_module_name":45,"category_name":129,"show_sort_weight":98,"slug":130},19,"General","general"]