[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84747-en":3,"doc-seo-84747-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},84747,137441390410,"Hazel","https://ap-avatar.wpscdn.com/avatar/2000252f4ab5702993?_k=1776741390130283984",8,"Research & Report","Fidelity-Diversity Metrics for Text","Language modeling maturity increases attention on how datasets are composed and curated for training. Yet data augmentation decisions require assessments beyond generic quality scores. The work introduces two embedding-based metrics—fidelity to measure how closely candidate text matches reference data and diversity to measure coverage of reference modes—derived from optimal transport divergence between discrete text summaries. Experiments on M2D2 and GSM8K-style synthetic datasets separate fidelity deficits from diversity deficits and link diversity shortfalls to reduced downstream fine-tuning accuracy.","arXiv :2607 .04563v 1 [ cs .CL] 6 Jul 2026  \nFidelity-Diversity Metrics for Text  \nAmanda Wang∗ Tudor Manole† Florentina Bunea‡ John Thickstun∗  \n[arw274@cornell.edu](arw274@cornell.edu) , [tmanole@mit.edu](tmanole@mit.edu) , [fb238@cornell.edu](fb238@cornell.edu) , [jthickstun@cornell.edu](jthickstun@cornell.edu)  \nAbstract  \nAs language modeling technology matures, there is an increasing research focus on the composition and curation of datasets used to train these models. For instance, practitioners commonly seek to augment high-quality datasets with additional text to enhance the performance of models trained on that data. However, informed decisions about data augmentation require more nuanced assessments about data quality. We build on work measuring the precision and recall of generative models to develop a pair of metrics that quantify (1) fidelity, capturing how closely candidate text resembles reference data, and (2) diversity, capturing how well it covers the modes of the reference dataset. Our metrics are based on optimal transport divergence functionals between discrete text summaries. In experiments on M2D2 text datasets, we show that these metrics are able to disentangle a lack of fidelity from a lack of diversity in deficient candidate text. In further experiments, our metrics detect diversity deficits in synthetic GSM8K-style math datasets, which correlate with degradations in downstream accuracy of language models finetuned on this synthetic data.  \n1 Introduction  \nIn tandem with the development of large language models (LMs), the research community has developed quantitative evaluation metrics to assess the quality of text generated by these models. In this paper we take a broader perspective on the evaluation of text, reimagining this style of quantitative evaluation as a tool to improve our understanding of the datasets used for training and finetuning LMs. We are particularly interested in synthetic training data, which bridges model evaluation and dataset analysis: in this setting, generated text is both the training data input and sampled output of language models [11 , 47] .  \nThe predominant metrics for evaluating LM-generated text compare model outputs to reference text using a scalar numerical score. These metrics can be computed either on the instance level, comparing one model output to one reference [51], or between i.i.d. samples from distributions of generated text and reference text [33] . We study the latter scenario, comparing an evaluation dataset collected from one distribution of text (possibly LM-generated, but not necessarily) to a reference dataset. We seek to develop more nuanced  \n∗ Department of Computer Science, Cornell University  \n†Statistics and Data Science Center, Massachusetts Institute of Technology ‡Department of Statistics, Cornell University  \nmetrics for insight into data quality, disaggregating our measure of data quality along two axes: fidelity, how similar evaluation text is to the reference text, and diversity, how broadly the evaluation text covers the modes of the reference text distribution.  \nWe quantify fidelity and diversity through two separate notions of transport between discrete measures. Our metrics operate on pretrained text embeddings; each dataset is summarized by a discrete measure consisting of point masses on a finite set of embeddings that represent modes of the data distribution. Fidelity is defined through the cost of agreedy nearest neighbor assignment between evaluation modes and reference modes. We arrive at this notion through an optimal transport argument, providing an alternative motivation for nearest-neighbor metrics that have previously proposed as a measure of data fidelity [38] . Diversity is defined as the gap between this greedy assignment and the global cost of optimal transport between the two measures. Taken together, these metrics decompose the Wasserstein transport cost between dataset representations into two components: costs","cbCaiovlHfMPyj2E","https://ap.wps.com/l/cbCaiovlHfMPyj2E","pdf",3864983,1,32,"English","en",105,"# Introduction\n# Related work","[{\"question\":\"What problem does the document address for language-model training datasets?\",\"answer\":\"It addresses how to make informed decisions about dataset composition and augmentation by evaluating candidate text beyond surface-level similarity.\"},{\"question\":\"How do the proposed metrics distinguish fidelity from diversity?\",\"answer\":\"Fidelity measures the cost of greedy nearest-neighbor assignment between evaluation and reference modes, while diversity measures the gap to the global optimal transport cost, reflecting how broadly modes are covered.\"},{\"question\":\"What evidence do the experiments provide about the metrics’ usefulness?\",\"answer\":\"On M2D2, the metrics disentangle missing fidelity from missing diversity. On synthetic GSM8K-style math datasets, diversity deficits correlate with downstream accuracy degradation after fine-tuning language models.\"}]",1784198018,81,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"fidelity-diversity-metrics-for-text","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/fidelity-diversity-metrics-for-text/84747/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the document address for language-model training datasets?","Question",{"text":75,"@type":76},"It addresses how to make informed decisions about dataset composition and augmentation by evaluating candidate text beyond surface-level similarity.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How do the proposed metrics distinguish fidelity from diversity?",{"text":80,"@type":76},"Fidelity measures the cost of greedy nearest-neighbor assignment between evaluation and reference modes, while diversity measures the gap to the global optimal transport cost, reflecting how broadly modes are covered.",{"name":82,"@type":73,"acceptedAnswer":83},"What evidence do the experiments provide about the metrics’ usefulness?",{"text":84,"@type":76},"On M2D2, the metrics disentangle missing fidelity from missing diversity. On synthetic GSM8K-style math datasets, diversity deficits correlate with downstream accuracy degradation after fine-tuning language models.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":45,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":45,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":45,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]