[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85288-en":3,"doc-seo-85288-105":30,"detail-sidebar-cat-0-en-105":83},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},85288,687197207919,"Theodora","https://ap-avatar.wpscdn.com/avatar/a000253d6f5f7c60be?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779446848396160552",8,"Research & Report","Do LLMs Fabricate Legal Citations A Bilingual Benchmark on Saudi Data Protection Law and the GDPR","Organizations and regulators increasingly consult large language models for regulatory compliance, but incorrect statutory citations can silently enter legal advice and policy documents. The work introduces a bilingual benchmark of 120 questions assessing whether freely accessible LLMs fabricate article citations for the GDPR and the Saudi PDPL. It combines direct citation retrieval, false-premise verification, and unanswerable trap probes including repealed articles and deadlines in implementing regulations. Results show strong GDPR accuracy alongside substantial Saudi fabrication and high-confidence errors, indicating that verbatim verification gates are essential.","Do LLMs Fabricate Legal Citations? A Bilingual Benchmark on Saudi Data Protection Law and the  \nGDPR  \nNoura Suliman Alrajeh  \n[nalrajeh0002@stu.kau.edu.sa](nalrajeh0002@stu.kau.edu.sa)  \nIT Department, King Abdulaziz University, Jeddah, Saudi Arabia  \narXiv :2607 . 1 1 127v 1 [ cs .CL] 13 Jul 2026  \nAbstract—Organizations and regulators increasingly consult large language models (LLMs) for regulatory-compliance questions, yet a wrong statutory citation can silently propagate into legal advice, compliance documentation, and policy decisions. We introduce a bilingual benchmark of 120 questions probing whether freely accessible LLMs fabricate article citations for two data-protection instruments: the EU General Data Protection Regulation (GDPR) and the Saudi Personal Data Protection Law (PDPL). The benchmark pairs direct citationretrieval questions with false-premise verification probes and deliberately unanswerable “trap” questions – including questions about a repealed article and about deadlines that exist only in implementing regulations, not in the law itself. Every question is posed in both Arabic and English, and all scoring is fully automatic against a manually verified gold reference. Evaluating three freely accessible models (Gemini 2.5 Flash, GPT-OSS- 120B, Nemotron-3-Super-120B), we find a dramatic jurisdiction gap: near-ceiling citation accuracy on the GDPR (94-100% on direct retrieval) against majority fabrication on the Saudi PDPL (60-77%), invariant to query language; the highest fabrication rates (67%) arise from statute-vs-regulations confusion, and 91% of fabricated citations are asserted with confidence ≥ 0.8. Fabrication tracks the jurisdiction of the law, not the language of the query, and model confidence provides no protection —indicating that verbatim-verification safeguards – rather than model self-confidence – must gate any institutional reliance on LLMs for compliance screening.  \nThe benchmark, gold article index, and raw model outputs will be made publicly available upon publication.  \nIndex Terms—trustworthy AI, large language models, hallucination, legal NLP, data protection, GDPR, regulatory compliance, Arabic NLP  \nI. INTRODUCTION  \nLarge language models are rapidly becoming a first point of contact for legal and regulatory questions. Employees ask chat assistants whether their processing of customer data requires consent; startups ask which article of a data-protection law governs cross-border transfer; and compliance teams draft policies with LLM assistance. In this setting, a fabricated citation – a confident reference to an article number that does not support the claim, or does not exist at all – is uniquely harmful: unlike a vague answer, it carries the surface form of verifiability and is therefore likely to be copied into documents that downstream readers trust [3], [4] .  \nThe risk is amplified in two directions that prior work has largely left unexamined. First, language: most legalhallucination evaluations target English-language, U.S.-centric  \nlaw [3], [4], while hundreds of millions of users interact with LLMs in Arabic about legal systems whose authoritative texts are Arabic. Second, jurisdiction: the statutes most heavily represented in web training data (such as the GDPR [1]) may enjoy far better factual grounding than recently enacted laws from other regions, such as the Saudi Personal Data Protection Law (PDPL), issued by Royal Decree M/19 of 1443H and amended by Royal Decree M/148 of 1444H [2] . If LLM reliability degrades precisely where users have the fewest alternative resources, the trustworthiness gap becomesan equity problem for secure digital ecosystems.  \nThis paper asks a deliberately narrow, fully automatable question: when asked which article of a data-protection law governs a given matter, do freely accessible LLMs answer correctly, abstain honestly, or fabricate? We make four contributions:  \n1) A bilingual citation-fabrication benchmark. 120 questions (60 per la","cbCaieDC65mDSVrY","https://ap.wps.com/l/cbCaieDC65mDSVrY","pdf",163387,4,1,5,"English","en",105,"# Abstract\n# Introduction\n# Contributions","[{\"question\":\"What is the main empirical finding regarding jurisdiction and query language?\",\"answer\":\"Citation reliability depends on the jurisdiction of the law rather than the query language. Models show near-ceiling accuracy for GDPR, but majority fabrication for the Saudi PDPL, with many fabricated citations asserted at high confidence.\"}]",1784202274,13,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":78,"head_meta":80,"extra_data":82,"updated_unix":28},"do-llms-fabricate-legal-citations-a-bilingual-benchmark-on-saudi-data-protection-law-and-the-gdpr","",{"@graph":36,"@context":77},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/do-llms-fabricate-legal-citations-a-bilingual-benchmark-on-saudi-data-protection-law-and-the-gdpr/85288/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71],{"name":72,"@type":73,"acceptedAnswer":74},"What is the main empirical finding regarding jurisdiction and query language?","Question",{"text":75,"@type":76},"Citation reliability depends on the jurisdiction of the law rather than the query language. Models show near-ceiling accuracy for GDPR, but majority fabrication for the Saudi PDPL, with many fabricated citations asserted at high confidence.","Answer","https://schema.org",{"og:url":52,"og:type":79,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":81,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":84},[85,89,93,97,101,106,111,114,119,122,126],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":86,"show_sort_weight":87,"slug":88},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":90,"show_sort_weight":91,"slug":92},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Exam",70,"exam",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Comic",60,"comic",{"id":102,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},6,"Technology",50,"technology",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":112,"slug":113},30,"research-report",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},9,"Religion & Spirituality",20,"religion-spirituality",{"id":117,"doc_module":4,"doc_module_name":46,"category_name":120,"show_sort_weight":117,"slug":121},"World Cup","world-cup",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":123,"slug":125},10,"Lifestyle","lifestyle",{"id":127,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":22,"slug":129},19,"General","general"]