[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83995-en":3,"doc-seo-83995-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83995,7971461740909,"Levi","https://ap-avatar.wpscdn.com/davatar_155a257f0dc6eb9ab79c44ca47cae57d",8,"Research & Report","SpanUQ: Span-Level Uncertainty Quantification for Large Language Model Generation","Uncertainty estimation is vital for trustworthy large language model (LLM) deployment and for self-refinement during generation. Existing methods use token-level or sequence-level granularities, which either lack semantic coherence or cannot localize erroneous parts. The work introduces Span-Level Uncertainty Estimation (SLUE) and SPANUQ, a lightweight ∼25M-parameter span probe that distills uncertainty knowledge into a single forward pass using a DETR-style span decoder and a Mixture of Beta uncertainty model. Experiments across five LLMs validate SPANUQ-BENCH and show superior quality and faster inference.","arXiv :2607 .0572 1v 1 [ cs .CL] 7 Jul 2026  \nSPANUQ: Span-Level Uncertainty Quantification for Large Language Model Generation  \nYimeng Zhang 1  Yingying Zhuang 1 , Ziyi Wang2 , Yuxuan Lu2 , Pei Chen 1 , Aman Gupta 1 , Zhe Su 1 , Ming Tan 1 , Zhilin Zhang 1 , Qun Liu 1 , Manikandarajan Ramanathan 1 , Rajashekar Maragoud 1 , Edward Vul 1 , Jing Huang 1 , Dakuo Wang2  \n1Amazon, 2Northeastern University  \nAbstract  \nUncertainty estimation is essential not only for the trustworthy deployment of  \nlarge language models (LLMs) but also as a foundation for self-refinement in  \nLLM generation. However, existing approaches operate at suboptimal granularities:  \ntoken-level scores lack semantic coherence, while sequence-level scores fail to  \nlocalize errors. We formalize Span-Level Uncertainty Estimation (SLUE), a new  \ntask that targets the natural granularity for uncertainty: semantically coherent  \ntext spans, each conveying a single assessable unit of meaning. To address this  \ntask, we introduce SPANUQ, a lightweight (∼25M parameter) probe that distills  \nthe uncertainty knowledge from expensive multi-sample inference into a single  \nforward pass over LLM hidden states. SPANUQ employs a DETR-style span  \ndecoder to simultaneously detect spans and estimate their uncertainty via a Mixture  \nof Beta distribution, trained with a principled combination of Beta NLL regression  \nand contrastive ranking objectives. We construct SPANUQ-BENCH, the first  \nspan-level uncertainty benchmark comprising 20K prompts, ∼293K annotated  \nspans, and continuous soft labels derived from multi-sample claim verification.  \nExperiments on five LLM backbones show that SPANUQ consistently achieves  \nthe best span-level uncertainty quality (AUROC 0.908–0.944, MAE 0.110–0.129),  \noutperforming the strongest probe baseline and all sampling-based methods while  \nbeing 10–20× faster. Its DETR-based span detector attains 0.910 F1, surpassing  \nthe best heuristic by 39.4%, enabling precise error localization that sequence-level  \nmethods cannot provide. The framework generalizes across five LLMs spanning  \ntwo model families (AUROC 0.908–0.944), and we additionally observe that  \nsequence-level uncertainty is partially decomposable: the learned importance  \nweighted span composition achieves ρseq = 0 .839, suggesting that span-level  \nestimation subsumes sequence-level as a special case. The project page is available  \nat [https://damon-demon.github.io/SpanUQ.html](https://damon-demon.github.io/SpanUQ.html).  \n1 Introduction  \nLarge language models (LLMs) generate fluent and coherent text across a wide range of tasks, yet their propensity to produce factually incorrect statements, commonly termed hallucinations, remains a critical barrier to deployment in high-stakes domains such as healthcare, legal analysis, and scientific research [1] . A fundamental step toward trustworthy LLM deployment is the ability to estimate how uncertain the model is about each piece of information it generates. Existing approaches to uncertainty estimation operate at two extremes of granularity. Token-level methods compute scores for individual tokens using predictive entropy or learned probes [2–4] . While computationally efficient, these scores are semantically incomplete: a single token rarely constitutes a verifiable  \n∗ Corresponding author  \nPreprint.  \nfact, and function words receive scores that carry no informational value. Sequence-level methods produce a single score for the entire response by sampling multiple outputs and measuring their semantic dispersion [5–7] . These methods capture meaningful semantic uncertainty but require  \n10–20× inference cost and cannot localize which parts of a response are unreliable. Consider the example in Fig. 1: an LLM generates “Marie Curie was a Polish physicist who won the Nobel Prize in 1901 .” A token-level method (a) assigns separate scores to every token; function words like “was” and “a” receive low but meaningless scores, while “p","cbCaimE0qwjijVZK","https://ap.wps.com/l/cbCaimE0qwjijVZK","pdf",483379,3,1,27,"English","en",105,"# Abstract\n# Introduction","[{\"question\":\"What problem does SpanUQ address in uncertainty estimation for LLMs?\",\"answer\":\"It targets the mismatch between uncertainty granularity and meaningful semantic units, where token-level scores lack coherence and sequence-level scores cannot localize which text parts are unreliable.\"},{\"question\":\"How does SPANUQ estimate uncertainty at the span level?\",\"answer\":\"SPANUQ uses a lightweight ∼25M-parameter DETR-style span decoder to detect semantically coherent spans and estimate each span’s uncertainty via a Mixture of Beta distribution.\"},{\"question\":\"What is SPANUQ-BENCH and how is it constructed?\",\"answer\":\"SPANUQ-BENCH is the first span-level uncertainty benchmark with 20K prompts and about 293K annotated spans, using continuous soft labels derived from multi-sample claim verification.\"}]",1784191920,68,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"spanuq-span-level-uncertainty-quantification-for-large-language-model-generation","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/spanuq-span-level-uncertainty-quantification-for-large-language-model-generation/83995/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does SpanUQ address in uncertainty estimation for LLMs?","Question",{"text":75,"@type":76},"It targets the mismatch between uncertainty granularity and meaningful semantic units, where token-level scores lack coherence and sequence-level scores cannot localize which text parts are unreliable.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does SPANUQ estimate uncertainty at the span level?",{"text":80,"@type":76},"SPANUQ uses a lightweight ∼25M-parameter DETR-style span decoder to detect semantically coherent spans and estimate each span’s uncertainty via a Mixture of Beta distribution.",{"name":82,"@type":73,"acceptedAnswer":83},"What is SPANUQ-BENCH and how is it constructed?",{"text":84,"@type":76},"SPANUQ-BENCH is the first span-level uncertainty benchmark with 20K prompts and about 293K annotated spans, using continuous soft labels derived from multi-sample claim verification.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]