[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86043-en":3,"doc-seo-86043-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},86043,1099514067438,"River Wang","https://ap-avatar.wpscdn.com/avatar/100002539ee87300030?x-image-process=image/resize,m_fixed,w_180,h_180&k=1780474512215547542",8,"Research & Report","Abstractiveness Metrics for Evaluating Text Summarization: A Refined Formulation with Empirical Validation","Quantifying abstractiveness in generated summaries enables evaluation beyond surface overlap metrics such as ROUGE. The paper introduces Reference Abstraction (RA), Summary Abstraction (SA), and Abstraction Ratio (AR), heuristic measures capturing divergence from extractive copying of the source text. Using a harmonic-mean document-length formulation modulated by a cubic non-overlap factor, the metrics provide bounded, dimensionally consistent outputs with nonlinear sensitivity at extractive–abstractive boundaries. Experiments on 100 XSUM documents across four models show strong separation between extractive and abstractive systems and AR signals summaries needing manual review for hallucination.","arXiv :2607 . 10806v 1 [ cs .CL] 12 Jul 2026  \nAbstractiveness Metrics for Evaluating Text  \nSummarization:  \nA Refined Formulation with Empirical Validation  \nPraveenkumar Katwe 1 , Rakesh Chandra Balabantaray 1 , Kali Prasad Vittala2  \n1 Department of Computer Science and Engineering,  \nInternational Institute of Information Technology, Bhubaneswar, India  \n2 Salesforce India Pvt Ltd, Bengaluru, India  \n[c121007@iiit-bh. ac. in](c121007@iiit-bh. ac. in)  \nAbstract  \nQuantifying abstractiveness in generated summaries is essential for evaluating summarization models beyond surface-level metrics like ROUGE. We introduce Reference Abstraction (RA), Summary Abstraction (SA), and Abstraction Ratio (AR)—a set of principled heuristic metrics that measure how much a summary diverges from extractive copying of the sourcetext. The formulation uses the harmonic mean of document lengths modulated by a cubic non-overlap factor, yielding dimensionally consistent, bounded output with non-linear sensitivity to the extractive-abstractive boundary. Evaluation on 100 XSUM documents across four summarization models (BART-large-cnn, Pegasus-xsum, DistilBart, MT5-small) demonstrates that the metrics successfully discriminate between extractive models (SA ≈ 0.12–0.26) and abstractive models (SA ≈ 0.96–1.77), and that the Abstraction Ratio identifies summaries requiring manual evaluation for potential hallucination. Code and results are available at [https://github.com/katweNLP/AbstractionStudy](https://github.com/katweNLP/AbstractionStudy).  \n1 Introduction  \nAbstractive summarization generates novel text that captures the gist of a source document. Unlike extractive methods that copy sentences verbatim, abstractive models paraphrase, generalize, and sometimes hallucinate—introducing content not grounded in the source. While fluency and informativeness have well-established metrics (ROUGE [1], BERTScore [2]), abstractiveness itself—the degree to which a summary diverges from extractive copying—has lacked a principled quantitative measure.  \nPrior work has used novel n-gram ratio [3] as a proxy, but this binary token-level measure does not account for document length asymmetry or provide graduated discrimination at different overlap levels. We propose three complementary metrics—RA, SA, and AR—that address these limitations through a formulation combining harmonic-mean length weighting with cubic nonoverlap amplification.  \nThese metrics serve as a screening tool: high RA indicates that the reference itself is highly abstractive (and potentially hallucinated); mismatched AR values flag summaries requiring manual evaluation. The metrics motivate and precede the Entity Hallucination Index (EHI) [4], which provides fine-grained hallucination decomposition.  \n2 Related Work  \n2.1 Reference-Based Quality Metrics  \nThe dominant paradigm for summarization evaluation relies on comparing generated text against human-written references. ROUGE [1] measures n-gram overlap between generated and reference summaries, providing recall-oriented quality scores. While ROUGE remains the de facto standard for benchmarking, it fundamentally rewards extractive copying—a model that reproduces the reference verbatim achieves a perfect score regardless of whether it faithfully represents the source document. BERTScore [2] addresses surface-level matching limitations by computing semantic similarity via contextual embeddings, but remains reference-based: it measures how well the generated summary matches human expectations, not whether it is grounded in the source.  \nA critical limitation shared by all reference-based metrics is their inability to detect faithfulness violations. A summary can achieve high ROUGE by capturing the same content choices as the reference while simultaneously introducing hallucinated details not present in the source. This disconnect between quality and faithfulness was demonstrated empirically by Fabbri et al. [5], who showed poor correlation betwe","cbCaimhgHzbDCBLu","https://ap.wps.com/l/cbCaimhgHzbDCBLu","pdf",1755447,1,13,"English","en",105,"# Abstractiveness Metrics for Evaluating Text Summarization\n## Introduction\n## Related Work\n### Reference-Based Quality Metrics\n### Faithfulness and Factuality\n### Abstractiveness Proxies\n### Positioning of Our Work","[{\"question\":\"What problem does the paper address in summarization evaluation?\",\"answer\":\"It addresses the lack of a principled way to quantify how abstractive a generated summary is, rather than relying only on surface-level quality metrics.\"},{\"question\":\"How do RA, SA, and AR measure abstractiveness?\",\"answer\":\"They quantify divergence from extractive copying by introducing Reference Abstraction, Summary Abstraction, and an Abstraction Ratio based on a harmonic-mean length formulation with a cubic non-overlap factor.\"},{\"question\":\"What do the experiments on XSUM show about these metrics?\",\"answer\":\"Results on 100 XSUM documents across four summarization models show that SA can discriminate extractive versus abstractive models, and that AR can flag cases that likely require manual evaluation for hallucination.\"}]",1784208043,33,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"abstractiveness-metrics-for-evaluating-text-summarization-a-refined-formulation-with-empirical-validation","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/abstractiveness-metrics-for-evaluating-text-summarization-a-refined-formulation-with-empirical-validation/86043/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-21","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the paper address in summarization evaluation?","Question",{"text":75,"@type":76},"It addresses the lack of a principled way to quantify how abstractive a generated summary is, rather than relying only on surface-level quality metrics.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How do RA, SA, and AR measure abstractiveness?",{"text":80,"@type":76},"They quantify divergence from extractive copying by introducing Reference Abstraction, Summary Abstraction, and an Abstraction Ratio based on a harmonic-mean length formulation with a cubic non-overlap factor.",{"name":82,"@type":73,"acceptedAnswer":83},"What do the experiments on XSUM show about these metrics?",{"text":84,"@type":76},"Results on 100 XSUM documents across four summarization models show that SA can discriminate extractive versus abstractive models, and that AR can flag cases that likely require manual evaluation for hallucination.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":45,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":45,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":45,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]