[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83551-en":3,"doc-seo-83551-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83551,962075006959,"Anda","https://ap-avatar.wpscdn.com/avatar/e0002397efbe92a78e?_k=1776741047341049297",8,"Research & Report","Beyond Document Grounding Span-Level Hallucination Detection","Hallucination detection for retrieval-augmented generation (RAG) is commonly measured on natural-language document evidence, but real grounded systems increasingly consume structured inputs such as source code, developer-tool output, markdown documents, tables, and repository metadata. A unified benchmark is introduced for span-level hallucination detection across code and structured documents. Grounded correct answers are modified by injecting localized hallucinations with exact character labels, then validated via code test splits and evidence-based review. A fine-tuned Qwen3.5-2B detector achieves 0.689 span-F1 on the unified test set and 0.60 on the code-agent source. ","Beyond Document Grounding: Span-Level Hallucination Detection over  \nCode, Tool Output, and Documents  \nÁdám Kovács1 , Bowei He2,3 , Xue Liu2,3 , István Boros1 , Szilveszter Tóth1 , Gábor Recski1,4  \n1 KR Labs, 2MBZUAI, 3McGill University, 4TU Wien  \nCorrespondence: [kovacs@krlabs.eu](kovacs@krlabs.eu)  \narXiv :2607 .00895v 1 [ cs .CL] 1 Jul 2026  \nAbstract  \nHallucination detection for retrieval-augmented generation (RAG) is usually evaluated on natural-language document evidence. However, grounded generation systems increasingly rely on structured inputs: source code, developertool output, markdown documents, tables, and repository metadata. We introduce a unified benchmark for span-level hallucination detection over code, tool output, structured documents, and existing natural-language RAG datasets. The benchmark is built by starting from grounded correct answers, injecting localized hallucinations with exact character labels, and validating the code test split with evidencebased review. Our fine-tuned Qwen3 .5-2B detector reaches 0.689 span-F1 on the unified test set and 0.60 on the code-agent source, where it substantially outperforms LettuceDetect-large (0 . 17) and the strongest zero-shot LLM judges we evaluated (at most 0 .22) . The same model remains competitive on established naturallanguage benchmarks, with 81.8 RAGTruth example-F1 and 0.724 English PsiloQA IoU.  \n1 Introduction  \nRetrieval-augmented generation (RAG) grounds model outputs in external evidence (Lewis et al., 2020), but it does not remove the need for verification. A generated answer can still contradict the retrieved context, introduce unsupported information, or cite a reference that is not present in the evidence. Hallucination detection methods therefore ask whether an answer is supported by the supplied context, often at the level of examples, sentences, tokens, or spans (Niu et al., 2024 ; Rykov et al., 2025 ; Vazquez et al., 2025) .  \nMost existing benchmarks and detectors study this problem in natural-language RAG, where both  \nthe evidence and the answer are usually document text (Niu et al., 2024 ; Belyi et al., 2025 ; Tanget al., 2024 ; Kovács and Recski, 2025 ; Rykov et al., 2025 ; Vazquez et al., 2025) . Real groundedgeneration systems are broader: coding agents work over repository files, git history, and test output (Jimenez et al., 2024); developer assistants summarize command output and tool observations (Kovacs, 2026); and research or documentation systems retrieve markdown pages, tables, citations, and structured documents (Recski et al., 2026 ; Index, 2026) . These settings are not well covered by current training data or evaluations: there is no shared span-level benchmark that covers generated code, tool observations, and structured documents under the same verification task.  \nWe study post-generation verification for this structured setting: given an answer that has already been produced, together with its request and context, a detector should flag the parts that the evidence does not support. We frame this at the span level rather than as a whole-answer accept/reject decision, because in code and tool output a single unsupported substring, such as a wrong field, a fabricated method name, a misreported value, or an invented section reference, can change program behavior or mislead a user while leaving the rest of the answer correct. A verifier should therefore point to the unsupported substring, not just reject the answer.  \nExisting hallucination-detection benchmarksand models leave this setting only partially covered. RAGTruth (Niu et al., 2024), Luna (Belyi et al., 2025), MiniCheck (Tang et al., 2024), and LettuceDetect (Kovács and Recski, 2025) verify generated text against retrieved documents. Code hallucination work studies generated snippets (Tian  \net al., 2025 ; Agarwal et al., 2025), generation-time divergence (Jiang et al., 2024), or agent trajectories (Liu et al., 2026) . These resources are useful, but they do not","cbCaiumAFfpk8HU8","https://ap.wps.com/l/cbCaiumAFfpk8HU8","pdf",282493,4,1,12,"English","en",105,"# Abstract\n# Introduction\n# Contributions","[{\"question\":\"What problem does the document address in hallucination detection for RAG?\",\"answer\":\"It targets hallucination detection when evidence is not only natural-language documents, but also structured sources like source code, developer-tool output, and markdown/table documents. The goal is to verify whether generated parts are supported by the provided evidence.\"},{\"question\":\"How is the span-level benchmark constructed?\",\"answer\":\"The benchmark starts from grounded correct answers and injects localized hallucinations, preserving exact character-level labels via an edit-based process. The dataset is then split by grounding source, with code test-split validation using evidence-based review.\"},{\"question\":\"How does the proposed detector perform compared with other methods?\",\"answer\":\"The fine-tuned Qwen3.5-2B detector reaches 0.689 span-F1 on the unified test set and 0.60 on the code-agent source. It substantially outperforms LettuceDetect-large and the evaluated strongest zero-shot LLM judges, while remaining competitive on established natural-language benchmarks.\"}]",1784188763,30,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"beyond-document-grounding-span-level-hallucination-detection","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/beyond-document-grounding-span-level-hallucination-detection/83551/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-25","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the document address in hallucination detection for RAG?","Question",{"text":75,"@type":76},"It targets hallucination detection when evidence is not only natural-language documents, but also structured sources like source code, developer-tool output, and markdown/table documents. The goal is to verify whether generated parts are supported by the provided evidence.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How is the span-level benchmark constructed?",{"text":80,"@type":76},"The benchmark starts from grounded correct answers and injects localized hallucinations, preserving exact character-level labels via an edit-based process. The dataset is then split by grounding source, with code test-split validation using evidence-based review.",{"name":82,"@type":73,"acceptedAnswer":83},"How does the proposed detector perform compared with other methods?",{"text":84,"@type":76},"The fine-tuned Qwen3.5-2B detector reaches 0.689 span-F1 on the unified test set and 0.60 on the code-agent source. It substantially outperforms LettuceDetect-large and the evaluated strongest zero-shot LLM judges, while remaining competitive on established natural-language benchmarks.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":29,"slug":121},"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]