[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82533-en":3,"doc-seo-82533-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},82533,549758146520,"Patrick","https://ap-avatar.wpscdn.com/avatar/80002397d8c0411e94?_k=1775819394049821470",8,"Research & Report","Low Perplexity is Repetition: A One-Dimensional Self-Conditioning Attractor in Continuous Diffusion","Continuous diffusion language models report record-low generative perplexity (Gen-PPL), but the metric can be misleading: these models repeat substantially more than human text, and Gen-PPL rewards that repetition. Removing the repeated component increases Gen-PPL markedly and reveals systematic repetition loops. The cause is traced to a one-dimensional contractive attractor in self-conditioning feedback, where a single direction steers samples into a repetition basin. ACE (Attractor-Contrast-Escape) subtracts this direction, reducing repetition near human level while keeping quality competitive and transferring across model sizes and samplers.","arXiv :2607 .00588v 1 [ cs .CL] 1 Jul 2026  \nLOW PERPLEXITY IS REPETITION: A ONEDIMENSIONAL SELF-CONDITIONING ATTRACTOR IN CONTINUOUS DIFFUSION LMS  \nShuai Zhang 1 ,2 , Zijie Chen2 , Hongliang He2 , Lun Du3 ,†, Zhenzhong Lan2 ,†  \n1Zhejiang University 2Westlake University 3Ant Group [zhangshuai@westlake.edu.cn](zhangshuai@westlake.edu.cn) [lanzhenzhong@westlake.edu.cn](lanzhenzhong@westlake.edu.cn)[ ](lanzhenzhong@westlake.edu.cn)†Corresponding authors  \nABSTRACT  \nContinuous diffusion language models such as ELF report record-low generative perplexity (Gen-PPL) . We find a catch: these models repeat far more than human text, and Gen-PPL rewards rather than penalizes that repetition, so its low scores overstate quality. Strip the repetition and ELF-B’s Gen-PPL rises from 19.5 to  \n27.7; the smallest model even posts the best Gen-PPL because it repeats most. Wetrace the repetition to its source: a contractive attractor along a single direction in the self-conditioning feedback loop, the loop that feeds each step’s clean estimate into the next. Because the failure is one-dimensional, a one-dimensional fix suffices, and we propose one. ACE (Attractor-Contrast-Escape) subtracts that single, label-free direction from the feedback at each step. Estimated once on the 105M model, the direction cuts repetition to near the human level while keeping quality competitive, and transfers near-unchanged to the 342M and 652M models and across samplers; the same recipe recovers useful directions on other architectures.  \nSince Gen-PPL itself rewards repetition, we instead measure the compute each fix needs to produce human-clean text, where ACE is 1.5–5 × cheaper.  \n1 INTRODUCTION  \nContinuous diffusion language models (DLMs) are a promising non-autoregressive route to text generation: they denoise a whole sequence in parallel within a differentiable embedding space, steerable by gradients and guidance. Self-conditioning (Chen et al., 2022) improves their sample quality by feeding the model’s own clean estimate back into each step to refine the next. Recent models such as ELF (Hu et al., 2026) report low generative perplexity (Gen-PPL), the number the field reads as generation quality. We find that this headline hides a defect: ELF’s samples repeat far more than human text, and Gen-PPL rewards the repetition instead of penalizing it. Strip therepetition and ELF-B’s Gen-PPL rises from 19.5 to 27.7, enough for the larger ELF-M to overtake it; the smallest model posts the best Gen-PPL only because it repeats the most (Table 1) .  \nThe defect is heavy and systematic. A large share of ELF samples lock onto a few repeated 4-grams and loop them for hundreds of words (Table 17), which human text essentially never does. The link is not only across models but within one: at a fixed setting, sample-level repetition correlates with the GPT-2 (Radford et al., 2019) PPL the samples are scored by (Table 16) . The defect stays hidden because the certifying metric is blind to it: repeated text is highly probable under the scorer, so it earns a flatteringly low Gen-PPL, analogous to likelihood-based degeneration in autoregressive generation (Holtzman et al., 2020; Welleck et al., 2020) .  \nWe trace the defect to its mechanism rather than stop at the symptom. Like audio feedback, this self-conditioning loop settles on whatever is most self-predictable, which is repeated content. Two probes pin this down. Turning the feedback strength up, with nothing else changed, drives repetition up and Gen-PPL down together (§3): the loop creates the repetition the metric then rewards. And linearizing the loop, its Jacobian has a single slowest-contracting mode, so the repeated state is a one-dimensional contractive attractor along one direction d (Fig. 1; §4), a basin  \n~~ ~~ 2 ~~ ~~ 1 0 1 2 3 4  \nu projection onto d  \nFigure 1: Repetition is a basin; ACE escapes it. Even as ELF denoises toward a clean sample, self-conditioning drags its representation u along one direction d","cbCaisWUEaGBUSUW","https://ap.wps.com/l/cbCaisWUEaGBUSUW","pdf",704115,1,25,"English","en",105,"# Abstract\n# Introduction\n## Defect: Gen-PPL rewards repetition\n## Mechanism: one-dimensional self-conditioning attractor\n## Fix: ACE (Attractor-Contrast-Escape)\n# Evaluation considerations","[{\"question\":\"Why do continuous diffusion language models show very low Gen-PPL despite poor human-likeness?\",\"answer\":\"Gen-PPL rewards the repetition produced by the models. Because repeated text is highly probable under the scorer, the metric decreases even as human-like quality worsens.\"},{\"question\":\"What mechanism causes the excessive repetition in self-conditioned continuous diffusion models?\",\"answer\":\"The self-conditioning feedback loop contains a one-dimensional contractive attractor. As denoising proceeds, representations drift along a single direction into a repetition basin.\"},{\"question\":\"How does ACE reduce repetition while maintaining competitive generation quality?\",\"answer\":\"ACE subtracts the single high-minus-low repetition direction from the self-conditioning feedback at every step. The direction is recovered label-free and no retraining is required, leading to near-human repetition rates with competitive quality.\"}]",1784181317,63,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"low-perplexity-is-repetition-a-one-dimensional-self-conditioning-attractor-in-continuous-diffusion","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/low-perplexity-is-repetition-a-one-dimensional-self-conditioning-attractor-in-continuous-diffusion/82533/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why do continuous diffusion language models show very low Gen-PPL despite poor human-likeness?","Question",{"text":75,"@type":76},"Gen-PPL rewards the repetition produced by the models. Because repeated text is highly probable under the scorer, the metric decreases even as human-like quality worsens.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What mechanism causes the excessive repetition in self-conditioned continuous diffusion models?",{"text":80,"@type":76},"The self-conditioning feedback loop contains a one-dimensional contractive attractor. As denoising proceeds, representations drift along a single direction into a repetition basin.",{"name":82,"@type":73,"acceptedAnswer":83},"How does ACE reduce repetition while maintaining competitive generation quality?",{"text":84,"@type":76},"ACE subtracts the single high-minus-low repetition direction from the self-conditioning feedback at every step. The direction is recovered label-free and no retraining is required, leading to near-human repetition rates with competitive quality.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":45,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":45,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":45,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]