[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83909-en":3,"doc-seo-83909-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83909,8796095461610,"Oliver","https://ap-avatar.wpscdn.com/davatar_276721f389ce27ea32af1340a28f341c",8,"Research & Report","DSpark: Confidence-Scheduled Speculative Decoding with Semi-Autoregressive Generation","DSpark presents a speculative decoding framework that accelerates Large Language Model inference by improving both draft quality and serving efficiency. It combines high-throughput parallel generation with adaptive, load-aware verification to avoid throughput collapse caused by high rejection-risk tokens. DSpark uses a semi-autoregressive architecture to model intra-block dependencies and mitigate suffix decay, then applies confidence-scheduled verification that selects verification length from estimated prefix survival probabilities and engine throughput profiles. Offline benchmarks and DeepSeek-V4 live deployment show 60%–85% per-user speedups with reduced verification waste.","arXiv :2607 .05 147v 1 [ cs .AI] 6 Jul 2026  \nDSpark: Confidence-Scheduled Speculative Decoding with Semi-Autoregressive Generation  \nXin Cheng1,2,∗, Xingkai Yu2,∗, Chenze Shao2,∗, Jiashi Li2,∗, Yunfan Xiong2,∗  \nYi Qian2, Jiaqi Zhu2, Shirong Ma2, Xiaokang Zhang2, Jiasheng Ye2, Qinyu Chen2, Chengqi Deng2, Jiping Yu2, Damai Dai2, Zhengyan Zhang2, Yixuan Wei2, Yixuan Tan2, Wenkai Yang2, Runxin Xu2, Yu Wu2, Zhean Xu2, Xuanyu Wang2, Muyang Chen2, Rui Tian2, Xiao Bi2, Zhewen Hao2, Shaoyuan Chen2, Huanqi Cao2, Wentao Zhang2, Anyi Xu2, Huishuai Zhang1, Dongyan Zhao1, Wenfeng Liang2  \n1 Peking University 2 DeepSeek-AI  \n{chengxin, xingkai, shaochenze, [js.li](js.li), [yunfanxiong}@deepseek.com](yunfanxiong}@deepseek.com)  \nAbstract  \nSpeculative decoding accelerates Large Language Model (LLM) inference by decoupling draft generation from target verification. While recent parallel drafters efficiently propose long token sequences in a single forward pass, they suffer from rapid acceptance decay due to a lack of inter-token dependencies. Furthermore, indiscriminately verifying these extended blocks wastes critical batch capacity on tokens with high rejection risks, severely degrading throughputin high-concurrency serving systems. We introduce DSpark, a speculative decoding framework that unifies high-throughput parallel generation with adaptive, load-aware verification. To maintain draft quality, DSpark utilizes a semi-autoregressive architecture—coupling a parallel backbone with a lightweight sequential module—to introduce intra-block dependency modeling and mitigate suffix decay. To optimize system efficiency, DSpark employs confidence-scheduled verification, dynamically tailoring the verification length for each request based on estimated prefix survival probabilities and engine-specific throughput profiles. On offline benchmarks across diverse domains, DSpark substantially improves the accepted length over state-of-the-art autoregressive and parallel drafters. When deployed within the DeepSeek-V4 serving system under live user traffic, DSpark successfully mitigates verification waste. Compared to the established production baseline (MTP-1), DSpark accelerates per-user generation speeds by 60%–85% at matched throughput levels. More importantly, by preventing severe throughput degradation under strict interactivity constraints, it enables performance tiers that were previously unattainable, shifting the Pareto frontier of our serving system. To facilitate community progress, we open-source the DSpark checkpoints alongside DeepSpec, an algorithm-driven training repository for speculative decoding.  \n1. Introduction  \nLarge Language Models (LLMs) generate text autoregressively: each new token requires a full forward pass conditioned on all preceding tokens, making inference latency proportional to the output length. The resulting low GPU utilization and high user-perceived waiting time constitute  \n*  \nEqual contribution.  \na primary bottleneck in production LLM serving, particularly for latency-sensitive scenarios such as real-time conversational assistants and multi-turn agentic workflows. Speculative decoding (Chen et al., 2023; Leviathan et al., 2023) offers a principled solution: a lightweight draft model proposes a block of candidate tokens, and the full-size target model verifies the entire block in a single forward pass via rejection sampling, accepting the longest prefix consistent with the target distribution and appending one bonus token. Because verification is parallel and the acceptance rule preserves the target distribution exactly, speculative decoding accelerates generation without any quality loss.  \nThe design of the draft model governs the trade-off between drafting latency and acceptance rate. Early drafters are autoregressive (Cheng et al., 2024; Li et al., 2024b), conditioning each position on previously sampled tokens. However, their drafting latency grows linearly with the block size, forcing these methods to us","cbCaip07XnAGPwki","https://ap.wps.com/l/cbCaip07XnAGPwki","pdf",1055890,4,1,33,"English","en",105,"# Introduction\n## Background: Autoregressive Inference and Speculative Decoding\n## Bottlenecks: Quality Decay and Verification Waste\n## DSpark: Confidence-Scheduled Semi-Autoregressive Framework","[{\"question\":\"What problem does DSpark address in speculative decoding for LLM serving?\",\"answer\":\"DSpark targets two bottlenecks: rapid acceptance decay from independent parallel drafting within a block, and inefficient verification that wastes batch capacity on tokens with high rejection risk under high concurrency.\"},{\"question\":\"How does DSpark improve draft quality compared with fully parallel drafters?\",\"answer\":\"DSpark adopts a semi-autoregressive design that couples a parallel backbone with a lightweight sequential module, enabling intra-block dependency modeling and reducing suffix decay.\"},{\"question\":\"How does DSpark choose how many tokens to verify for each request?\",\"answer\":\"DSpark uses confidence-scheduled verification, dynamically setting verification length based on estimated prefix survival probabilities and engine-specific throughput profiles to balance speed and batch usage.\"}]",1784191379,83,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"dspark-confidence-scheduled-speculative-decoding-with-semi-autoregressive-generation","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/dspark-confidence-scheduled-speculative-decoding-with-semi-autoregressive-generation/83909/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-27","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does DSpark address in speculative decoding for LLM serving?","Question",{"text":75,"@type":76},"DSpark targets two bottlenecks: rapid acceptance decay from independent parallel drafting within a block, and inefficient verification that wastes batch capacity on tokens with high rejection risk under high concurrency.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does DSpark improve draft quality compared with fully parallel drafters?",{"text":80,"@type":76},"DSpark adopts a semi-autoregressive design that couples a parallel backbone with a lightweight sequential module, enabling intra-block dependency modeling and reducing suffix decay.",{"name":82,"@type":73,"acceptedAnswer":83},"How does DSpark choose how many tokens to verify for each request?",{"text":84,"@type":76},"DSpark uses confidence-scheduled verification, dynamically setting verification length based on estimated prefix survival probabilities and engine-specific throughput profiles to balance speed and batch usage.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]