[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86543-en":3,"doc-seo-86543-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86543,962075114101,"Seraphina","https://ap-avatar.wpscdn.com/avatar/e000253a75eb197efd?x-image-process=image/resize,m_fixed,w_180,h_180&k=1780044092746381165",8,"Research & Report","DynEval: Holistic Evaluations of T2I Generative Models in the Wild","Recent progress in text-to-image (T2I) generation enables highly realistic outputs, but evaluating them reliably at scale remains difficult. Existing automatic evaluators often rely on static prompt sets and miss subtle failure modes such as partial prompt misalignment, compositional errors, or visually plausible yet semantically incorrect generations. DynEval introduces a dynamic evaluation framework that jointly assesses text-image alignment and image quality, training on GenDB (500K pairs) and DynEvalInstruct (250K triplets). Fine-tuned DynEval-2B/4B achieves stronger correlation with human judgments across 11 benchmarks and detailed failure-mode analysis.","arXiv :2607 . 11199v1 [ cs .CV] 13 Jul 2026  \nDynEval: Holistic Evaluations of T2I Generative Models in the Wild  \nShyam Marjit 1∗, Dheeraj Baiju 1∗, Anuj Shikarkhane 1∗†, Akhil Sakthieswaran 1†, Sayak Paul2 , and Anirban Chakraborty 1  \n1 Indian Institute of Science 2 Hugging Face  \n[shyammarjit@iisc.ac.in](shyammarjit@iisc.ac.in) , [anirban@iisc.ac.in](anirban@iisc.ac.in)[ ](anirban@iisc.ac.in)Project Page: [https://vcl-iisc.github.io/dyneval](https://vcl-iisc.github.io/dyneval)  \nAbstract. Recent advances in text-to-image (T2I) generation have led to models capable of producing highly realistic images. Yet, reliably evaluating their outputs remains challenging, especially at scale. Existing automatic evaluators, often relying on a static prompt set, struggle to capture subtle failure modes such as partial prompt misalignment, compositional errors or visually plausible but semantically incorrect generations.  \nIn this work, we introduce DynEval, a Dynamic Evaluation framework designed to jointly assess text-to-image alignment and image quality of T2I models. To support scalable training beyond limited humanannotated data, we construct two large datasets. First, we build GenDB, a collection of 500K prompt-image pairs generated from human-written prompts drawn from DiffusionDB using a tiered prompt-model generation strategy. Second, building upon GenDB, we construct DynEvalInstruct, a 250K instruction dataset comprising prompt-image-response triplets distilled from a structured evaluation pipeline that decomposes evaluation into text-image alignment and visual quality reasoning. Using this dataset, we perform full fine-tuning of a compact evaluator through a curriculum learning strategy to effectively distill the superior evaluation capabilities of a larger teacher vision-language model, resulting in DynEval-2B and DynEval-4B. In extensive comparisons against existing evaluators across 11 benchmarks, our evaluator achieves a higher overall correlation with human judgments. Furthermore, it provides finegrained analysis of the capabilities and failure modes of 36 T2I models across 42 subcategories and 9 semantic dimensions.  \n1 Introduction  \nText-to-image (T2I) generation has evolved from multi-stage U-Net [61] based diffusion models such as SDXL [58] to diffusion transformers [57] that scale efficiently to high-fidelity outputs [20, 40 , 42] . In parallel, efficiency-focused diffusion designs have reduced training and deployment costs while enabling highresolution synthesis on modest hardware [12–14, 83 , 84] . Alongside diffusion,  \n* Equal contribution.  \n† Work done during internship at VCL, IISc.  \n2 S. Marjit et al.  \n\n|  |  |  |  |\n| --- | --- | --- | --- |\n| EvalMuse | \"Four leaves on a branch. \"\u003Cbr>\u003Cbr>\"A small dog in a cozy orange sweater sitting \"A bright blue sky where planes have beside a cat wearing a stylish blue bow tie. \" drawn 'Sky's the Limit' with contrails. \"\u003Cbr>\u003Cbr>Model: DeepFloyd IF-I-XL\u003Cbr>FGA-BLIP2 Score: 0.53\u003Cbr>DynEval Score: 1.00\u003Cbr>Human Score: 0.87\u003Cbr>Model: SDXL 2.1\u003Cbr>FGA-BLIP2 Score: 0.75\u003Cbr>DynEval Score: 0.65\u003Cbr>Human Score: 0.40\u003Cbr>Model: Midjourney 6\u003Cbr>FGA-BLIP2 Score: 0.75\u003Cbr>DynEval Score: 1.00\u003Cbr>Human Score: 0.93\u003Cbr>Model: DALLE 3\u003Cbr>FGA-BLIP2 Score: 0.87\u003Cbr>DynEval Score: 0.78\u003Cbr>Human Score: 0.60\u003Cbr>Model: DALLE 3\u003Cbr>FGA-BLIP2 Score: 0.50\u003Cbr>DynEval Score: 0.77\u003Cbr>Human Score: 0.73\u003Cbr>Model: SDXL 2.1\u003Cbr>FGA-BLIP2 Score: 0.61\u003Cbr>DynEval Score: 0.40\u003Cbr>Human Score: 0.40 | \"A shoe with no laces, standing alone. \" |  |\n|  |  |  |  |\n|  |  |  |  |\n|  |  |  |  |\n\nFig. 1: Qualitative comparison of DynEval-4B with four representative text-to-image evaluation methods [24 , 28 , 34 , 35] across multiple T2I models [4, 7 , 13 , 15 , 19 , 20 , 25 , 32 , 37 , 41 , 52 , 58 , 60 , 64 , 66 , 80 , 83 , 85] . All scores are normalized to the [0 , 1] range. The first GenEval example additionally includes a real reference image, illustrating the limitations of detector-based evaluation. Compared with ex","cbCaidVwFOfKkjwU","https://ap.wps.com/l/cbCaidVwFOfKkjwU","pdf",6315337,7,1,40,"English","en",105,"# Introduction\n# DynEval: Dynamic Evaluation Framework\n## Datasets: GenDB and DynEvalInstruct\n## Curriculum fine-tuning of DynEval-2B and DynEval-4B\n# Experimental Results and Benchmarks\n## Correlation with human judgments\n## Failure modes and capability analysis","[{\"question\":\"Why is evaluating text-to-image (T2I) outputs challenging at scale?\",\"answer\":\"Automatic evaluators often depend on fixed prompt sets, which makes them weak at capturing subtle issues like partial prompt misalignment, compositional errors, and semantically wrong but visually plausible generations.\"},{\"question\":\"What is DynEval designed to measure in T2I models?\",\"answer\":\"DynEval jointly evaluates text-image alignment and image quality, using a dynamic framework rather than a static prompt-only evaluator.\"},{\"question\":\"How were the datasets for training DynEval constructed?\",\"answer\":\"GenDB is built from 500K prompt-image pairs generated from human-written prompts using a tiered prompt-model generation strategy. DynEvalInstruct extends GenDB with 250K instruction triplets distilled from a structured evaluation pipeline that decomposes evaluation into alignment and visual quality reasoning.\"}]",1784212525,101,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"dyneval-holistic-evaluations-of-t2i-generative-models-in-the-wild","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/dyneval-holistic-evaluations-of-t2i-generative-models-in-the-wild/86543/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-27","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"Why is evaluating text-to-image (T2I) outputs challenging at scale?","Question",{"text":76,"@type":77},"Automatic evaluators often depend on fixed prompt sets, which makes them weak at capturing subtle issues like partial prompt misalignment, compositional errors, and semantically wrong but visually plausible generations.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"What is DynEval designed to measure in T2I models?",{"text":81,"@type":77},"DynEval jointly evaluates text-image alignment and image quality, using a dynamic framework rather than a static prompt-only evaluator.",{"name":83,"@type":74,"acceptedAnswer":84},"How were the datasets for training DynEval constructed?",{"text":85,"@type":77},"GenDB is built from 500K prompt-image pairs generated from human-written prompts using a tiered prompt-model generation strategy. DynEvalInstruct extends GenDB with 250K instruction triplets distilled from a structured evaluation pipeline that decomposes evaluation into alignment and visual quality reasoning.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,119,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":46,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":22,"slug":118},"Healthcare","healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":107,"slug":137},19,"General","general"]