[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85573-en":3,"doc-seo-85573-105":29,"detail-sidebar-cat-0-en-105":90},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},85573,1649267921044,"Ava Thompson","https://us-avatar.wpscdn.com/avatar/1800007509477c92dfb?_k=1782875107921204101",8,"Research & Report","BARD Bridging AutoRegressive and Diffusion Vision Language Models Via Highly Efficient Progressive Block Merging and Stage Wise Distillation","Autoregressive vision-language models deliver strong multimodal capability, but token-by-token decoding creates a hard inference bottleneck. Diffusion vision-language models enable more parallel refinement, yet naive conversion of pretrained autoregressive VLMs into large-block dVLMs causes significant quality loss. BARD bridges this gap by progressively merging blocks, then recovering performance with stage-wise intra-dVLM distillation from a small-block diffusion anchor, plus mixed noise scheduling and memory-friendly training for long multimodal sequences.","arXiv :2604 . 16514v5 [ cs .CV] 12 Jul 2026  \nBARD: Bridging AutoRegressive and Diffusion Vision-Language Models Via Highly Efficient Progressive Block Merging and  \nStage-Wise Distillation  \nBaoyou Chen 1,3,★ HanChen Xia 1,★ Peng Tu 1,★ Haojun Shi 1 Liwei Zhang 1 Yuxuan Yao2,3  \nWeihao Yuan4 Siyu Zhu 1,2,3,†∗  \n1 Shanghai Academy of AI for Science, 2 Shanghai Innovation Institute, 3Fudan University  \n4Nanjing University  \nFigure 1: Quality–efficiency comparison of BARD-VL and representative open dVLMs. Left: radar chart on seven multimodal benchmarks, where BARD-VL at 2B/4B/8B is compared with prior open diffusion VLMs. BARD-VL 4B and 8B exhibit the strongest overall performance among the compared dVLMs, showing that the proposed bridge substantially narrows the capability gap of existing diffusion VLMs. Right: OCRBench accuracy versus decoding throughput. Even though it is smaller than the 7B/8B baselines, BARD-VL 4B traces a clearly better accuracy–throughput trade-off, retaining higher accuracy across abroad range of decoding speeds.  \nAbstract  \nAutoregressive vision-language models (VLMs) deliver strong multimodal capability, but their token-by-token decoding imposes a fundamental inference bottleneck. Diffusion VLMs offer a more parallel decoding paradigm, yet directly converting a pretrained autoregressive VLM into a large-block diffusion VLM (dVLM) often leads to substantial quality degradation. In this work, we present BARD, a simple and effective bridging framework that converts apretrained autoregressive VLM into a same-architecture, decodingefficient dVLM. Our approach combines progressive supervised block merging, which gradually enlarges the decoding block size, with stage-wise intra-dVLM distillation from a fixed small-block diffusion anchor to recover performance lost at larger blocks. We further incorporate a mixed noise scheduler to improve robustness and token revision during denoising, and memory-friendly training  \n∗★ Equal contribution.† Corresponding author.  \nto enable efficient training on long multimodal sequences. A key empirical finding is that direct autoregressive-to-diffusion distillation is poorly aligned and can even hurt performance, whereas distillation within the diffusion regime is consistently effective. Experimental results show that, with ≤ 4.4􀀢 data, BARD-VL transfers strong multimodal capability from Qwen3-VL to a large-block dVLM. Remarkably, BARD-VL establishes a new SOTA among comparable-scale open dVLMs on our evaluation suite at both 4Band 8B scales. At the same time, BARD-VL achieves up to 3× decoding throughput speedup compared to the source model. Code is available at: [https://github.com/fudan-generative-vision/Bard-VL](https://github.com/fudan-generative-vision/Bard-VL).  \nCCS Concepts  \n• Computing methodologies → Natural language generation; Neural networks; Computer vision.  \nConference acronym’34, November 10–14, 2026, Rio de Janeiro, Brazil  \nKeywords  \nvision-language models, multimodal understanding, discrete diffusion, parallel decoding, multimodal reasoning  \n1 Introduction  \nAutoregressive vision-language models (VLMs) have become the dominant foundation for multimodal understanding and agentic interaction. Their success spans visual reasoning, document understanding, grounded question answering, and emerging multimodal agents. Yet their token-by-token causal decoding remains inherently sequential, imposing a hard inference bottleneck for practical deployment. Diffusion-based decoding offers a compelling alternative: by refining a partially generated response block by block, a diffusion VLM (dVLM) can update multiple tokens in parallel and therefore expose a fundamentally different quality–efficiency trade-off [2, 11, 29] .  \nRecent progress has established diffusion models as a viable direction for both language and vision-language modeling. Broadly speaking, existing methods follow two main paradigms. The first builds diffusion-native models from scratch or from ","cbCailW504uBEydC","https://ap.wps.com/l/cbCailW504uBEydC","pdf",5697739,1,9,"English","en",105,"# Abstract\n# Keywords\n# Introduction\n## Autoregressive VLM bottleneck and diffusion decoding\n## Two diffusion modeling paradigms\n## The BARD bridging framework","[{\"question\":\"What problem does BARD address when converting autoregressive VLMs to diffusion VLMs?\",\"answer\":\"Directly converting a pretrained autoregressive VLM into a large-block diffusion VLM often causes substantial quality degradation due to misalignment in the distillation process.\"},{\"question\":\"How does BARD enable efficient large-block diffusion decoding?\",\"answer\":\"BARD first converts an autoregressive checkpoint into a stable small-block diffusion anchor, then gradually increases decoding parallelism via progressive supervised block merging rather than an abrupt transition.\"},{\"question\":\"What training improvements does BARD add beyond the bridging procedure?\",\"answer\":\"It introduces a mixed noise scheduler to improve robustness and token revision during denoising, and memory-friendly training that packs clean and noisy responses with shared multimodal context to reduce overhead.\"}]",1784204681,23,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":85,"head_meta":87,"extra_data":89,"updated_unix":27},"bard-bridging-autoregressive-and-diffusion-vision-language-models-via-highly-efficient-progressive-block-merging-and-stage-wise-distillation","",{"@graph":35,"@context":84},[36,53,67],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/bard-bridging-autoregressive-and-diffusion-vision-language-models-via-highly-efficient-progressive-block-merging-and-stage-wise-distillation/85573/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":61,"encodingFormat":60,"isAccessibleForFree":62,"interactionStatistic":63},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-16",true,{"@type":64,"interactionType":65,"userInteractionCount":4},"InteractionCounter",{"@type":66},"ViewAction",{"@type":68,"mainEntity":69},"FAQPage",[70,76,80],{"name":71,"@type":72,"acceptedAnswer":73},"What problem does BARD address when converting autoregressive VLMs to diffusion VLMs?","Question",{"text":74,"@type":75},"Directly converting a pretrained autoregressive VLM into a large-block diffusion VLM often causes substantial quality degradation due to misalignment in the distillation process.","Answer",{"name":77,"@type":72,"acceptedAnswer":78},"How does BARD enable efficient large-block diffusion decoding?",{"text":79,"@type":75},"BARD first converts an autoregressive checkpoint into a stable small-block diffusion anchor, then gradually increases decoding parallelism via progressive supervised block merging rather than an abrupt transition.",{"name":81,"@type":72,"acceptedAnswer":82},"What training improvements does BARD add beyond the bridging procedure?",{"text":83,"@type":75},"It introduces a mixed noise scheduler to improve robustness and token revision during denoising, and memory-friendly training that packs clean and noisy responses with shared multimodal context to reduce overhead.","https://schema.org",{"og:url":51,"og:type":86,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":88,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":91},[92,96,100,104,109,114,119,122,126,129,133],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":93,"show_sort_weight":94,"slug":95},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":97,"show_sort_weight":98,"slug":99},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":101,"show_sort_weight":102,"slug":103},"Exam",70,"exam",{"id":105,"doc_module":4,"doc_module_name":45,"category_name":106,"show_sort_weight":107,"slug":108},5,"Comic",60,"comic",{"id":110,"doc_module":4,"doc_module_name":45,"category_name":111,"show_sort_weight":112,"slug":113},6,"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":45,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":21,"doc_module":4,"doc_module_name":45,"category_name":123,"show_sort_weight":124,"slug":125},"Religion & Spirituality",20,"religion-spirituality",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":127,"show_sort_weight":124,"slug":128},"World Cup","world-cup",{"id":130,"doc_module":4,"doc_module_name":45,"category_name":131,"show_sort_weight":130,"slug":132},10,"Lifestyle","lifestyle",{"id":134,"doc_module":4,"doc_module_name":45,"category_name":135,"show_sort_weight":105,"slug":136},19,"General","general"]