[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82555-en":3,"doc-seo-82555-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},82555,962075006959,"Anda","https://ap-avatar.wpscdn.com/avatar/e0002397efbe92a78e?_k=1776741047341049297",8,"Research & Report","Phantom References Hallucinated Citations That Survive Peer Review at Top-Tier Conferences","Large language models can generate polished scientific writing that lacks real evidence, and once such text enters scholarly workflows, hallucinations may persist in the archival record. Measuring this risk via prose is hard because technical claims are context-dependent and require expert judgment. This work targets citations as a narrower, verifiable surface: a reference either resolves to a real work with matching authorship or it does not. It proposes RefChecker, an auditing pipeline across major bibliographic sources plus web re-verification.","Phantom References: Hallucinated Citations That Survive Peer Review  \nat Top-Tier Conferences  \nMark Russinovich∗ , Ram Shankar Siva Kumar§ , Ahmed Salem§  \n∗ Microsoft Azure  \n§ Microsoft  \narXiv :2607 .00738v2 [ cs .DL] 6 Jul 2026  \nAbstract—Large language models make it easy to produce scientific text that is polished, confident, and unsupported by real evidence. When such text enters scholarly workflows, hallucination can become part of the archival record rather than merely a transient model error. Measuring this risk through the prose of published papers is difficult: technical claims are contextual and often require expert judgment. However, citations expose a narrower, more auditable surface. A reference either resolves to a real scholarly work with compatible authorship, or it does not.  \nThis paper measures citation hallucination in peer-reviewed proceedings. We define a conservative notion of hallucinated citation that counts identity-level failures: non-existent works and substantial author-list mismatches. We explicitly exclude ordinary bibliographic drift, such as venue changes, year changes, publication-status updates, and minor name variants. To audit citations at scale, we build RefChecker, a referenceverification pipeline that resolves bibliography entries against multiple bibliographic sources and escalates unresolved cases to web-search re-verification. We apply RefChecker to accepted camera-ready papers from ICLR, ICML, NeurIPS, and USENIX Security.  \nOur results show that hallucinated citations have entered the archival record. Reference-level rates are usually below one percent, but proceedings contain enough papers and references for these failures to become visible at the paper level. In 2025, roughly one in twenty NeurIPS and USENIX Security papers contains at least two likely hallucinated academic-paper-like references under our strict definition. We also observe postChatGPT increases in several venues highlighted by a highcount tail of papers with multiple (5+) failures in the same bibliography, and likely hallucinated citations even among award-winning papers.  \nThese findings show that citation integrity is not reliably enforced by peer review alone, yet the problem is tractable: in one venue-scale scan, conference-scale auditing cost roughly four cents per paper. We open-source RefChecker to support routine, reproducible citation verification before publication ([https://github.com/markrussinovich/refchecker](https://github.com/markrussinovich/refchecker)).  \n1. Introduction  \nLarge language models hallucinate [1], [2], [3] . They fabricate quotations, invent sources, and produce confident,  \nwell-formed claims that have no base in the world. What began as a curiosity of early chatbots is now a recurring hazard in production: hallucinated content has surfaced in legal filings [4], news reporting, code repositories [5], and technical documents, often in settings where downstream readers have neither the time nor the expertise to verify every claim. As LLMs are folded into the workflows that produce written knowledge, hallucination stops being only a model failure and starts becoming part of the public record.  \nScientific writing is one of the workflows absorbing these tools the fastest. Researchers use LLMs to rephrase paragraphs, restructure arguments, summarize related work, and assemble bibliographies [6] . A rough title, a partial memory of a paper, or a URL can become a polished BibTex entry in seconds [7], [8] . On the surface, the resulting manuscript still looks like a paper: claims are supported, references are alphabetized, and the bibliography has the familiar shape of scholarly care. What has changed is the chain of evidence underneath, and with it the cost of placing a fabricated or misattributed reference into the scientific record.  \nAt first glance this looks hard to measure. Checking an arbitrary technical claim in the body of a paper is essentially an open problem in automated f","cbCaisJNT88tTMJX","https://ap.wps.com/l/cbCaisJNT88tTMJX","pdf",661551,1,14,"English","en",105,"# Abstract\n# Introduction\n## Measuring citation hallucination\n## RefChecker auditing pipeline\n## Findings on archival impact","[{\"question\":\"What problem does the paper focus on regarding large language models and scientific writing?\",\"answer\":\"It focuses on citation hallucinations—when LLM-generated manuscripts include references that look credible but fail identity-level resolution or authorship matching. These errors can become part of the archival record after peer review.\"},{\"question\":\"How does the paper define a “hallucinated citation”?\",\"answer\":\"A hallucinated citation is defined conservatively as an identity-level failure: nonexistent works and substantial author-list mismatches. The paper excludes ordinary bibliographic drift such as venue/year changes or minor name variants.\"},{\"question\":\"How does RefChecker verify citations at scale?\",\"answer\":\"RefChecker resolves bibliography entries against multiple bibliographic sources (including arXiv, Semantic Scholar, OpenAlex, CrossRef, DBLP, and the ACL Anthology). Unresolved or suspicious cases are escalated to an LLM-driven deep web search to reach a verdict.\"}]",1784181499,35,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"phantom-references-hallucinated-citations-that-survive-peer-review-at-top-tier-conferences","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/phantom-references-hallucinated-citations-that-survive-peer-review-at-top-tier-conferences/82555/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-23","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the paper focus on regarding large language models and scientific writing?","Question",{"text":75,"@type":76},"It focuses on citation hallucinations—when LLM-generated manuscripts include references that look credible but fail identity-level resolution or authorship matching. These errors can become part of the archival record after peer review.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the paper define a “hallucinated citation”?",{"text":80,"@type":76},"A hallucinated citation is defined conservatively as an identity-level failure: nonexistent works and substantial author-list mismatches. The paper excludes ordinary bibliographic drift such as venue/year changes or minor name variants.",{"name":82,"@type":73,"acceptedAnswer":83},"How does RefChecker verify citations at scale?",{"text":84,"@type":76},"RefChecker resolves bibliography entries against multiple bibliographic sources (including arXiv, Semantic Scholar, OpenAlex, CrossRef, DBLP, and the ACL Anthology). Unresolved or suspicious cases are escalated to an LLM-driven deep web search to reach a verdict.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":45,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":45,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":45,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]