[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85960-en":3,"doc-seo-85960-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},85960,13056703019404,"Miles","https://ap-avatar.wpscdn.com/davatar_29158cc5080c5b710cf443261637dec0",8,"Research & Report","Improving Sample Diversity in Autoregressive Text-to-Image Generation via Cluster Truncation","Diffusion models deliver top-tier image quality for text-to-image (T2I) generation, yet they can suffer from sample diversity collapse. This study examines whether autoregressive (AR) T2I models can improve the quality–diversity trade-off and push the Pareto frontier. The work shows that persistent high token entropy and redundancy in visual token spaces limit existing token-level decoding methods. To address this, a cluster-level entropy truncation strategy, p-less cluster, is proposed and evaluated across multiple AR models and datasets, achieving maximal diversity while preserving image quality and prompt alignment.","arXiv :2607 . 10535v1 [ cs .CV] 12 Jul 2026  \nImproving Sample Diversity in Autoregressive Text-to-Image Generation via Cluster Truncation  \nTrang Nguyen 1 ,2 , Shuang Wu2 , Runyan Tan2 , Phillip Howard2  \n1University of Massachusetts Amherst, 2Thoughtworks  \n[tramnguyen@umass.edu](tramnguyen@umass.edu) , {shuang.wu, runyan.tan, [phillip.howard}@thoughtworks.com](phillip.howard}@thoughtworks.com)  \nAbstract  \nWhile diffusion models achieve state-of-the-art image quality for text-to-image  \n(T2I) generation, recent work has demonstrated that they suffer from sample di  \nversity collapse. In this work, we investigate whether autoregressive (AR) image  \ngeneration models can push the Pareto frontier between image quality and sample  \ndiversity. With recent advances in quality and efficiency, AR models have emerged as a viable alternative to diffusion-based image generation. Beyond enabling new use cases such as interleaved image-text generation, their sequential generation process makes them compatible with a wide range of token-based decoding strategies originally developed to improve diversity in text generation. Motivated by the potential of a better diversity-quality tradeoff in the AR paradigm, we present  \nthe first systematic study of sample diversity in AR image generation models. We  \nshow that two key properties of AR image generation, persistently high token-level entropy and substantial redundancy in visual token spaces, limit the effectiveness of existing token-level decoding methods for diversity enhancement. We therefore propose p-less cluster, a new decoding strategy that performs entropy-based truncation sampling at cluster level rather than at token level. We evaluate our  \napproach and baseline decoding methods across four autoregressive T2I models  \nand two datasets using a comprehensive suite of metrics spanning image quality, prompt alignment, and diversity. Our results show that p-less cluster unlocks the  \ngreatest diversity across most evaluated autoregressive T2I models and datasets  \nwhile maintaining image quality and prompt alignment.  \n1 Introduction  \nText-to-image (T2I) generation has witnessed remarkable progress in recent years, with diffusionbased models such as Flux[15], Stable Diffusion [1], and Imagen[25] producing photorealistic outputs that closely follow complex text prompts. Despite this progress, diversity collapse has remained a persistent limitation with T2I generation. Diversity collapse manifests at two levels: distributional diversity (the model ignores rare or underrepresented modes of the data distribution) and sample diversity (repeated sampling for the same prompt yields little variation in composition, style, or content) [8] . Critically, these two axes of diversity are distinct and do not always improve together:  \nfor example, SDXL-DMD2 [37], a distilled diffusion model, achieves superior distributional diversity as measured by FID [12] compared to the base model, yet exhibits markedly lower per-prompt sample diversity [8] . In this work, we focus on the latter: sample diversity, i.e., generating a diverse set of images for the same prompt. This capability is particularly important in creative applications, where users are often presented with multiple generated candidates to choose from [22, 2] .  \nRecently, autoregressive (AR) image generation [26, 34, 4, 3] has emerged as a compelling alternative paradigm to diffusion-style T2I generation. Such models have enabled new use cases such as interleaved image and text generation by unifying language and image generation in a single transformer-based architecture. This has important implications for how sampling is conducted in T2I  \nPreprint.  \nAutoregressive  \nJanus Pro Emu3  \nFigure 1: p-less cluster enables diverse i.i.d. sampling across autoregressive (AR) image generative models while maintaining acceptable image quality, avoiding the issue of diversity collapse.  \ngeneration: by formulating image synthesis as sequential to","cbCaimY32Lr1ks6X","https://ap.wps.com/l/cbCaimY32Lr1ks6X","pdf",11654544,3,1,26,"English","en",105,"# Abstract\n# Introduction\n## Diversity collapse in T2I\n## Autoregressive models and decoding\n## Proposed p-less cluster method\n## Experimental evaluation plan","[{\"question\":\"What problem does the paper address in text-to-image generation?\",\"answer\":\"It addresses sample diversity collapse, where repeated sampling for the same prompt produces limited variation in generated images.\"},{\"question\":\"Why are existing token-level decoding methods less effective for AR image generation?\",\"answer\":\"The paper attributes the limitation to persistently high token-level entropy and substantial redundancy in visual token spaces, which constrain the impact of token-level truncation strategies.\"},{\"question\":\"What is p-less cluster and how does it differ from token-level truncation?\",\"answer\":\"p-less cluster truncates based on entropy at the cluster level rather than at individual tokens, encouraging exploration across semantically distinct modes while filtering low-probability regions that can cause artifacts.\"}]",1784207377,66,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"improving-sample-diversity-in-autoregressive-text-to-image-generation-via-cluster-truncation","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/improving-sample-diversity-in-autoregressive-text-to-image-generation-via-cluster-truncation/85960/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the paper address in text-to-image generation?","Question",{"text":75,"@type":76},"It addresses sample diversity collapse, where repeated sampling for the same prompt produces limited variation in generated images.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"Why are existing token-level decoding methods less effective for AR image generation?",{"text":80,"@type":76},"The paper attributes the limitation to persistently high token-level entropy and substantial redundancy in visual token spaces, which constrain the impact of token-level truncation strategies.",{"name":82,"@type":73,"acceptedAnswer":83},"What is p-less cluster and how does it differ from token-level truncation?",{"text":84,"@type":76},"p-less cluster truncates based on entropy at the cluster level rather than at individual tokens, encouraging exploration across semantically distinct modes while filtering low-probability regions that can cause artifacts.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]