[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-81815-en":3,"doc-seo-81815-105":31,"detail-sidebar-cat-0-en-105":93},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":28,"seo_description":14,"update_tm":29,"read_time":30},81815,137441390410,"Hazel","https://ap-avatar.wpscdn.com/avatar/2000252f4ab5702993?_k=1776741390130283984",8,"Research & Report","Discrete Diffusion Language Models for Interactive Radiology Report Drafting","Discrete diffusion language models generate text by denoising a token canvas bidirectionally rather than emitting tokens left to right, but most medical foundation models still follow autoregressive decoding. This work adapts DiffusionGemma-26B, a mixture-of-experts diffusion model, and benchmarks it against its same-size autoregressive sibling Gemma-4-26B using an identical LoRA recipe on medical visual question answering datasets. Diffusion matches or exceeds autoregressive performance, runs 3.5–4.4× faster, and adds interactive any-order infill for drafting radiology reports.","arXiv :2607 .0 1436v 1 [ cs .AI] 1 Jul 2026  \nDiscrete Diffusion Language Models for Interactive Radiology Report Drafting  \nMax Van Puyvelde* 1,2 [maxvpuyv@stanford.edu](maxvpuyv@stanford.edu)  \nWim Van Criekinge† 2 [wim.vancriekinge@ugent.be](wim.vancriekinge@ugent.be)  \nH. Ibrahim Gulluk* 3  \n[gulluk@stanford.edu](gulluk@stanford.edu)[ ](gulluk@stanford.edu)Olivier Gevaert† 1 [ogevaert@stanford.edu](ogevaert@stanford.edu)  \n1Department of Biomedical Data Science, Stanford University School of Medicine  \n2Department of Mathematical Modelling, Statistics & Bioinformatics, Ghent University  \n3Department of Electrical Engineering, Stanford University  \nAbstract  \nDiffusion language models, which generate text by denoising a token canvas bidirectionally instead of emitting tokens left to right, have become competitive with autoregressive (AR) generation. Medical foundation models, however, remain almost entirely autoregressive. We adapt a mixture-of-experts diffusion language model, DiffusionGemma-26B, and benchmark it against its same-size AR sibling Gemma-4-26B under an identical LoRA recipe on medical visual question answering datasets, scored by a verbosity-robust LLM judge. Diffusion matches or exceeds AR on all of them, and the finetuned model (3 .8B active) is competitive with frontier vision-language models; its decoding is also 3.5–4.4× faster. Beyond this parity, the diffusion model offers a drafting capability AR lacks: any-order infill. Because the canvas is denoised bidirectionally, a radiologist can fix report fragments and have the model fill the text between them, an operation inherent to diffusion but not to autoregression, which is subpar at it. This suits real reports, which are often terse or inconsistent across clinicians and institutions.  \n1 Introduction  \nAutoregressive (AR) generation, which produces text one token at a time from left to right, underlies nearly all large language and vision-language models. Discrete diffusion language models [1, 18, 19] are a recent alternative: they generate a sequence by iteratively denoising a fixed token canvas, with each position attending to the entire canvas rather than only to preceding tokens. On general text these models are competitive with autoregressive models of comparable size [18, 22], which makes them a plausible backbone for domains that have so far relied on autoregression. One open instance, DiffusionGemma-26B [6], couples this denoising decoder with a native multimodal encoder, and belongs to a model family that also includes a same-size autoregressive model, Gemma-4-26B [5]; the two share size, family, and lineage, and differ chiefly in their generative paradigm.  \nExisting medical foundation models, however, are almost exclusively autoregressive. Radiology report generation (RRG), the task of drafting a report from an image, is dominated by AR models [2, 7– 9, 12, 25], as are medical vision-language assistants [15] . Whether a diffusion language model is viable as a medical foundation model, both accurate enough and useful in the clinical workflow, is largely untested. A few diffusion models already generate CXR reports [4, 17, 23], but produce complete reports only and do not address interactive drafting.  \n*Joint first authors. †Joint senior authors.  \nWe finetune both the diffusion model and its autoregressive sibling on paired image-text data from medical visual-question-answering datasets, under an identical LoRA recipe that varies only the generative paradigm (same backbone size, vision tower, LoRA targets, and data), and benchmark them against each other and frontier vision-language models with a verbosity-robust LLM judge.  \nBeyond accuracy, the two paradigms differ in what they can be conditioned on. Reporting practice varies: negative and normal findings are stated explicitly in some settings and omitted in others, and section conventions differ across institutions. A tool that completes or normalizes a report around content the radiologi","cbCaicxbSDtNLx24","https://ap.wps.com/l/cbCaicxbSDtNLx24","pdf",6851973,5,1,16,"English","en",105,"# Abstract\n# Introduction\n# Related Work","[{\"question\":\"How do diffusion language models differ from autoregressive models for text generation in this study?\",\"answer\":\"Diffusion models denoise a fixed token canvas bidirectionally, attending to the entire canvas, while autoregressive models generate tokens sequentially left to right based only on preceding text.\"},{\"question\":\"What is the main advantage of diffusion for radiology report drafting proposed in the paper?\",\"answer\":\"Diffusion enables any-order infill: radiologists can fix report fragments at arbitrary positions and the model fills the text between them, leveraging context on both sides of the gap.\"},{\"question\":\"How does DiffusionGemma-26B compare with Gemma-4-26B in the reported benchmarks?\",\"answer\":\"Using the same-size backbones and an identical LoRA setup, Diffusion matches or exceeds autoregressive results across medical VQA datasets and decodes 3.5–4.4× faster.\"}]","Discrete Diffusion Language Models for Interactive Radiology Report Drafting | PDF",1784176322,40,{"code":4,"msg":32,"data":33},"ok",{"site_id":25,"language":24,"slug":34,"title":13,"keywords":35,"description":14,"schema_data":36,"social_meta":88,"head_meta":90,"extra_data":92,"updated_unix":29},"discrete-diffusion-language-models-for-interactive-radiology-report-drafting","",{"@graph":37,"@context":87},[38,55,70],{"@type":39,"itemListElement":40},"BreadcrumbList",[41,45,49,52],{"item":42,"name":43,"@type":44,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":46,"name":47,"@type":44,"position":48},"https://docshare.wps.com/document/","Document",2,{"item":50,"name":12,"@type":44,"position":51},"https://docshare.wps.com/document/research-report/",3,{"item":53,"name":13,"@type":44,"position":54},"https://docshare.wps.com/document/discrete-diffusion-language-models-for-interactive-radiology-report-drafting/81815/",4,{"url":53,"name":13,"@type":56,"author":57,"headline":13,"publisher":59,"fileFormat":62,"inLanguage":24,"description":14,"dateModified":63,"datePublished":64,"encodingFormat":62,"isAccessibleForFree":65,"interactionStatistic":66},"DigitalDocument",{"name":9,"@type":58},"Person",{"url":42,"name":60,"@type":61},"DocShare","Organization","application/pdf","2026-08-01","2026-07-16",true,{"@type":67,"interactionType":68,"userInteractionCount":20},"InteractionCounter",{"@type":69},"ViewAction",{"@type":71,"mainEntity":72},"FAQPage",[73,79,83],{"name":74,"@type":75,"acceptedAnswer":76},"How do diffusion language models differ from autoregressive models for text generation in this study?","Question",{"text":77,"@type":78},"Diffusion models denoise a fixed token canvas bidirectionally, attending to the entire canvas, while autoregressive models generate tokens sequentially left to right based only on preceding text.","Answer",{"name":80,"@type":75,"acceptedAnswer":81},"What is the main advantage of diffusion for radiology report drafting proposed in the paper?",{"text":82,"@type":78},"Diffusion enables any-order infill: radiologists can fix report fragments at arbitrary positions and the model fills the text between them, leveraging context on both sides of the gap.",{"name":84,"@type":75,"acceptedAnswer":85},"How does DiffusionGemma-26B compare with Gemma-4-26B in the reported benchmarks?",{"text":86,"@type":78},"Using the same-size backbones and an identical LoRA setup, Diffusion matches or exceeds autoregressive results across medical VQA datasets and decodes 3.5–4.4× faster.","https://schema.org",{"og:url":53,"og:type":89,"og:title":13,"og:site_name":60,"og:description":14},"article",{"robots":91,"canonical":53},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":94},[95,99,103,107,111,116,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":47,"category_name":96,"show_sort_weight":97,"slug":98},"Story & Novel",90,"story-novel",{"id":48,"doc_module":4,"doc_module_name":47,"category_name":100,"show_sort_weight":101,"slug":102},"Literature",80,"literature",{"id":54,"doc_module":4,"doc_module_name":47,"category_name":104,"show_sort_weight":105,"slug":106},"Exam",70,"exam",{"id":20,"doc_module":4,"doc_module_name":47,"category_name":108,"show_sort_weight":109,"slug":110},"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":47,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":47,"category_name":118,"show_sort_weight":30,"slug":119},7,"Healthcare","healthcare",{"id":11,"doc_module":4,"doc_module_name":47,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":47,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":47,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":47,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":47,"category_name":137,"show_sort_weight":20,"slug":138},19,"General","general"]