[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82099-en":3,"doc-seo-82099-105":30,"detail-sidebar-cat-0-en-105":83},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},82099,1374391975076,"Riley","https://ap-avatar.wpscdn.com/avatar/14000253ca4ec9f6853?x-image-process=image/resize,m_fixed,w_180,h_180&k=1783305029341752051",8,"Research & Report","BlockServe Block-Grained Continuous Batching for High-Throughput Diffusion LLM Serving","Efficient serving of diffusion large language models (dLLMs) is constrained by convergence heterogeneity: in batched workloads, different request sequences complete denoising at different block boundaries, so faster requests wait for slower stragglers, creating compute bubbles and high tail latency. BlockServe introduces block-grained continuous batching with block boundary eviction, mixed-state execution via gather-scatter indexing, and a compute-aware admission controller using a token budgeted refill strategy. On Dream and LLaDA across five benchmarks, BlockServe delivers 1.9–10.6× throughput over Fast-dLLM with comparable generation quality.","BlockServe: Block-Grained Continuous Batching for High-Throughput Diffusion LLM Serving  \nYuanjie Zhu, Liangwei Yang, Ke Xu, Weizhi Zhang, Shanghao Li, Zihe Song, and Philip S. Yu  \nUniversity of Illinois Chicago, USA  \n{yzhu224, lyang84, kxu25, wzhan42, sli261, zsong29, [psyu](psyu}@uic.edu)[}](psyu}@uic.edu)[@uic.edu](psyu}@uic.edu)  \narXiv :2607 .08930v 1 [ cs .LG] 9 Jul 2026  \nAbstract—Efficient serving of diffusion large language models (dLLMs) is hindered by convergence heterogeneity: when batching multiple requests, different sequences converge at different rates, causing faster requests to stall behind slower stragglers and introducing compute bubbles and tail latency. We present BlockServe, a continuous batching framework that integrates blockgrained scheduling—immediately evicting completed requests at block boundaries—with mixed-state execution that extends dual cache and parallel decoding to heterogeneous batches via gatherscatter indexing. Furthermore, a compute-aware admission controller expands effective batch capacity through token-budgeted refill. On Dream and LLaDA across five benchmarks, BlockServe achieves 1.9–10.6× throughput over Fast-dLLM with comparable generation quality, establishing block-grained scheduling as a foundation for high-throughput offline dLLM inference.  \nIndex Terms—diffusion large language models, continuous batching, model serving, block-grained scheduling, throughput  \nI. INTRODUCTION  \nThe deployment of Large Language Models (LLMs) has shifted focus from model architecture to efficient serving systems [1] . In the autoregressive (AR) paradigm, continuous batching [2], [3] has become the industry standard, scheduling at the granularity of a single token to maintain high GPU utilization. However, the emergence of Diffusion LLMs (dLLMs) [4], [5] introduces a fundamental paradigm shift. Unlike AR models that generate tokens sequentially, dLLMs employ a parallel denoising process, generating entire blocks of tokens iteratively. While recent works like Fast-dLLM [6] have successfully optimized dLLMs for low-latency singlerequest inference, efficient batched serving for these models remains challenging when requests converge at different rates across block boundaries.  \nWe term this phenomenon convergence heterogeneity. In autoregressive serving, output length variation is the primary source of scheduling heterogeneity, which token-level continuous batching handles by preempting at every generation step [7], [8] . dLLMs, however, generate in discrete blocks, and requests reach completion at different block boundaries even under the same generation budget. As a result, some requests can finish early while others continue denoising, creating mixed completion states within a single batched serving iteration. In current dLLM batching, the entire batch is gated by the slowest request (the “straggler”), forcing alreadycompleted requests to remain idle while occupying resources in the active batch. The resulting “compute bubbles”—idle  \ntime where GPU resources are underutilized—lead to severe long-tail latency, calling for a dedicated block-level batch scheduling framework.  \nTo address these challenges, we propose BlockServe, a specialized serving framework that enables high-throughput continuous batching for dLLMs without altering the underlying model architecture. BlockServe integrates a block-grained scheduler with a mixed-state memory manager, preventing stragglers from stalling the batch while enabling the execution of requests at different block indices within a unified dense tensor. A compute-aware admission controller under a token budget further adapts concurrency to workload geometry. These components form a block-centric execution loop: completed requests are reclaimed at block boundaries, heterogeneous in-flight requests remain executable in a shared dense batch, and the recovered capacity is then converted into additional admissions. Our contributions are summarized as follows:  \n•","cbCaihaC4vL4Arkx","https://ap.wps.com/l/cbCaihaC4vL4Arkx","pdf",574535,3,1,10,"English","en",105,"# Introduction\n# Method\n## Problem formulation\n## BlockServe scheduling loop\n## Mixed-state memory and execution","[{\"question\":\"What mechanisms enable mixed-state execution across heterogeneous batches?\",\"answer\":\"BlockServe uses a mixed-state memory manager that aligns states in unified dense tensors and extends dual cache and parallel decoding via gather-scatter indexing so requests at different block indices can execute together without changing the model architecture.\"}]",1784178206,25,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":78,"head_meta":80,"extra_data":82,"updated_unix":28},"blockserve-block-grained-continuous-batching-for-high-throughput-diffusion-llm-serving","",{"@graph":36,"@context":77},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/blockserve-block-grained-continuous-batching-for-high-throughput-diffusion-llm-serving/82099/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-19","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71],{"name":72,"@type":73,"acceptedAnswer":74},"What mechanisms enable mixed-state execution across heterogeneous batches?","Question",{"text":75,"@type":76},"BlockServe uses a mixed-state memory manager that aligns states in unified dense tensors and extends dual cache and parallel decoding via gather-scatter indexing so requests at different block indices can execute together without changing the model architecture.","Answer","https://schema.org",{"og:url":51,"og:type":79,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":81,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":84},[85,89,93,97,102,107,112,115,120,123,126],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":86,"show_sort_weight":87,"slug":88},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":90,"show_sort_weight":91,"slug":92},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Exam",70,"exam",{"id":98,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},5,"Comic",60,"comic",{"id":103,"doc_module":4,"doc_module_name":46,"category_name":104,"show_sort_weight":105,"slug":106},6,"Technology",50,"technology",{"id":108,"doc_module":4,"doc_module_name":46,"category_name":109,"show_sort_weight":110,"slug":111},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":113,"slug":114},30,"research-report",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},9,"Religion & Spirituality",20,"religion-spirituality",{"id":118,"doc_module":4,"doc_module_name":46,"category_name":121,"show_sort_weight":118,"slug":122},"World Cup","world-cup",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":22,"slug":125},"Lifestyle","lifestyle",{"id":127,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":98,"slug":129},19,"General","general"]