[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85293-en":3,"doc-seo-85293-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},85293,687197207919,"Theodora","https://ap-avatar.wpscdn.com/avatar/a000253d6f5f7c60be?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779446848396160552",8,"Research & Report","TIGER Text-conditioned Visual Gated Routing with Acceptance Alignment for Multimodal Speculative Decoding","Speculative decoding speeds up autoregressive generation by having a lightweight drafter propose tokens that a larger verifier checks and accepts. While effective for text-only LLMs, gains in vision-language models are limited because drafting often diverges on vision-critical content, and existing methods do not target irrelevant visual evidence or optimize the verifier-accepted prefix length that determines speedup. This work proposes TIGER, which selects sparse context-relevant visual tokens from the drafter’s current textual state and trains the drafter via acceptance-aligned group-based policy optimization using verifier-derived prefix-length rewards.","TIGER: Text-conditioned Visual Gated Routing with Acceptance Alignment for Multimodal Speculative Decoding  \nQuynh Vo Cong-Duy Nguyen Ponhvoan Srey Luu Anh Tuan Thong Nguyen*  \nNational University of Singapore Nanyang Technological University Center of AI Research, VinUniversity  \n[thong.nguyen@u.nus.edu](thong.nguyen@u.nus.edu)  \narXiv :2607 . 1 1 13 1v 1 [ cs .CL] 13 Jul 2026  \nAbstract  \nSpeculative decoding accelerates autoregressive generation by letting a lightweight drafter propose multiple tokens that are verified by a larger target model. Although effective for text-only LLMs, speculative decoding yields limited gains in VLMs because drafters often diverge on vision-critical content, while existing multimodal acceleration methods do not directly address irrelevant visual evidence or optimize the verifier-accepted prefix length that governs speedup. We propose TIGER, a Textconditioned vIsual GatEd Routing framework for multimodal speculative decoding. TIGER dynamically selects a sparse set of contextrelevant visual tokens based on the drafter’s current textual state, rather than expose the full visual token set or a fixed compressed interface. To better align training with inferencetime efficiency, we optimize the drafter with acceptance-aligned group-based policy training using verifier-derived rewards based on accepted prefix length, built on top of distillation warm start with KL anchoring. This encourages the drafter not only to imitate the target model, but also to produce speculative continuations that survive verification for longer prefixes. Experiments show that TIGER yields consistent gains in accepted prefix length and speculative speedup under exact verifier-side speculative decoding, while achieving favorable qualitylatency trade-offs with comparable downstream accuracy in visual-routing analyses.  \n1 Introduction  \nVision-language models (VLMs) have rapidly become the backbone of multimodal assistants (Liu et al., 2023 ; Li et al., 2023a), enabling instruction following over images and videos for tasks such as visual question answering, document understanding, chart reasoning, and grounded dialogue (Laurençon et al., 2024) . Despite steady  \n* Corresponding author  \nprogress in model quality, real-world deployment remains constrained by autoregressive decoding cost: generating outputs token-by-token requires repeated passes through large decoders and extensive key–value (KV) caching (Sadhukhan et al., 2025) . As VLMs scale to higher resolutions and richer visual contexts, decoding latency and serving cost increasingly dominate practical deployment.  \nSpeculative decoding offers a principled, lossless route to accelerate generation. A lightweight drafter proposes multiple tokens, and a large verifier checks and accepts a prefix of these proposals, thereby reducing the number of expensive verifier decoding steps while preserving the target distribution through a modified rejectionsampling test (Wang et al., 2025b) . In text-only LLMs, speculative decoding can provide substantial speedups when the drafter closely matches the target model, and recent work further improves efficiency through better drafting structures such as dynamic draft trees (Li et al., 2024) . However, directly transferring this paradigm to VLMs often yields limited gains.  \nThe main challenge is that multimodal speculative decoding is highly sensitive to vision-critical tokens—for example, OCR strings in TextVQAstyle queries (Singh et al., 2019), numerical answers in chart or science reasoning (Lu et al., 2022), compositional relations in GQA (Hudson and Manning, 2019), and grounded object descriptions where hallucination must be avoided (Li et al., 2023b) . On such tokens, a small drafter often diverges early from the verifier, causing short accepted prefixes and sharply reducing speculative speedup. We argue that this failure stems from two coupled mismatches. First, the drafter is often exposed to either the full visual token set ora ","cbCaiaNOZC1USot2","https://ap.wps.com/l/cbCaiaNOZC1USot2","pdf",8696457,2,1,31,"English","en",105,"# Abstract\n# 1 Introduction","[{\"question\":\"What problem does TIGER address in multimodal speculative decoding?\",\"answer\":\"TIGER addresses the limited speedup in vision-language models caused by early divergence between a lightweight drafter and the verifier on vision-critical content, which leads to short accepted prefixes.\"},{\"question\":\"How does TIGER improve the visual input used by the drafter?\",\"answer\":\"TIGER dynamically selects a sparse subset of context-relevant visual tokens conditioned on the drafter’s current textual state, instead of exposing the full visual token set or a fixed compressed interface.\"},{\"question\":\"How is the drafter trained to improve runtime efficiency?\",\"answer\":\"TIGER trains the drafter with acceptance-aligned group-based policy optimization, using verifier-derived rewards based on the verifier-accepted prefix length, and uses a distillation warm start with KL anchoring for stable training.\"}]",1784202294,78,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"tiger-text-conditioned-visual-gated-routing-with-acceptance-alignment-for-multimodal-speculative-decoding","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/tiger-text-conditioned-visual-gated-routing-with-acceptance-alignment-for-multimodal-speculative-decoding/85293/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-23","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does TIGER address in multimodal speculative decoding?","Question",{"text":75,"@type":76},"TIGER addresses the limited speedup in vision-language models caused by early divergence between a lightweight drafter and the verifier on vision-critical content, which leads to short accepted prefixes.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does TIGER improve the visual input used by the drafter?",{"text":80,"@type":76},"TIGER dynamically selects a sparse subset of context-relevant visual tokens conditioned on the drafter’s current textual state, instead of exposing the full visual token set or a fixed compressed interface.",{"name":82,"@type":73,"acceptedAnswer":83},"How is the drafter trained to improve runtime efficiency?",{"text":84,"@type":76},"TIGER trains the drafter with acceptance-aligned group-based policy optimization, using verifier-derived rewards based on the verifier-accepted prefix length, and uses a distillation warm start with KL anchoring for stable training.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]