[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82730-en":3,"doc-seo-82730-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},82730,4398048949847,"Eliana","https://ap-avatar.wpscdn.com/avatar/400002536579ef2da7f?_k=1778318612642679267",8,"Research & Report","From Judgments to Issues: Structured Extraction of Legal Reasoning with Citation-Hallucination Control","An automated pipeline transforms Italian tax-court judgments into individual legal issues, producing for each issue a structured XML representation grounded in IRAC and the legal syllogism. The approach targets a large corpus of about 330,000 first- and second-instance decisions and relies on a cost-efficient general-purpose model (DeepSeek V3) to support large-scale processing. To mitigate citation hallucinations, an automatic filter compares model-generated references with references parsed from the judgment text using Linkoln normalized to URN-NIR, ECLI, and CELEX. Validation uses 50 expert-annotated judgments and evaluates both issue extraction and citation quality, including the hallucination filter, enabling downstream retrieval, citation-network analysis, and dataset construction.","FROM JUDGMENTS TO ISSUES: STRUCTURED EXTRACTION OF LEGAL REASONING WITH CITATION-HALLUCINATION CONTROL  \narXiv :2607 .03325v 1 [ cs .CL] 3 Jul 2026  \nGiovanni Piccioli  \nQuantitative and Digital Law Laboratory King’s College London Strand, London WC2R 2LS [giovannipiccioli@gmail.com](giovannipiccioli@gmail.com)  \nAlessia Fidelangeli  \nCIRSFID-Alma AI, Faculty of Law University of Bologna Bologna, Via Zamboni 27/29  \nPiera Santin  \nRobert Schuman Centre European University Institute Fiesole, Badia Fiesolana-Via dei Roccettini 9  \nPierpaolo Vivo  \nQuantitative and Digital Law Laboratory King’s College London Strand, London WC2R 2LS  \nJuly 7, 2026  \nABSTRACT  \nWe present an automated pipeline that decomposes Italian tax-court judgments into individual legal issues and extracts, for each issue, a structured XML representation grounded in the IRAC framework and the legal syllogism. The pipeline targets a corpus of approximately 330 ,000 first-and second-instance decisions of the Italian tax courts and is built around a capable yet cost-efficient general-purpose model (DeepSeek V3), a choice driven by the need to process several hundred thousand documents at a sustainable cost. To address the well-documented unreliability of large language models on legal citations, we couple the extraction step with an automatic hallucinationdetection filter that compares the references produced by the model with those identified in the judgment text by a dedicated parser (Linkoln), normalised to standard identifiers (URN-NIR, ECLI, CELEX). We validate the pipeline on 50 judgments annotated by two PhDs in tax law, computing inter-annotator agreement and LLM-vs-expert agreement on both issue extraction and legal citations, together with a stand-alone evaluation of the hallucination filter. To the best of our knowledge, this is the first issue-level, expert-validated structured extraction pipeline with hallucination control for Italian tax-court decisions, and it provides a concrete starting point for downstream applications such as issue-level retrieval, citation-network analysis, and the construction of large-scale datasets of legal reasoning.  \n1 Introduction  \nModern judicial systems generate court decisions at a scale that has long outpaced the capacity of legal professionals to read them individually. In Italy, the database of first-and second-instance tax-court decisions alone contains almost a million judgments, and grows by the hundreds of thousands every year. The practical value of this volume lies almost entirely in secondary use: precedent search by practitioners, citation-network analysis, evaluation of judicial behaviour, and the construction of large datasets for training and benchmarking legal NLP systems. Each of these uses, however, requires a representation of the decisions that is more compact and structured, than the original PDFs the courts produce. The central question this paper addresses is how to obtain such a representation automatically, at scale, and with quantified reliability.  \nItalian tax-court judgments are a particularly informative testbed for this question. On the one hand, the domain stresses any extraction pipeline that claims to be general. Tax disputes range from a few tens of euros to tens of millions, and routinely draw on civil, commercial, criminal, procedural, and European Union law in resolving a single case. The first  \nA PREPRINT-JULY 7, 2026  \ntwo instances are decided by panels that include non-career honorary judges, whose drafting style is markedly less standardized than that of higher courts: even the names and ordering of the structural sections of a judgment vary from one decision to the next. On the other hand, the domain also makes the problem tractable: an official open database is available, the volume is large enough that statistical evaluation is meaningful, and the parties are always the same pair (taxpayer and tax authority) . A pipeline that performs well here is therefore informative ","cbCaiqnlQqA1sZqr","https://ap.wps.com/l/cbCaiqnlQqA1sZqr","pdf",683627,2,1,33,"English","en",105,"# Introduction\n## Research problem and motivation\n## Unit of representation: legal issues\n## Structured representation: IRAC and legal syllogism\n## Citation-hallucination control and evaluation","[{\"question\":\"What is the main goal of the presented pipeline?\",\"answer\":\"Decompose Italian tax-court judgments into individual legal issues and extract, for each issue, a structured XML representation grounded in IRAC and the legal syllogism.\"},{\"question\":\"How does the method control citation hallucinations?\",\"answer\":\"It adds a hallucination-detection filter that compares model-produced references against references parsed from the judgment text by Linkoln, normalized to standard identifiers such as URN-NIR, ECLI, and CELEX.\"},{\"question\":\"What was used to validate the pipeline and what was evaluated?\",\"answer\":\"The pipeline was validated on 50 judgments annotated by two PhDs in tax law, computing agreement for issue extraction and citations between annotators and between LLM and experts, plus a standalone evaluation of the hallucination filter.\"}]",1784182550,83,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"from-judgments-to-issues-structured-extraction-of-legal-reasoning-with-citation-hallucination-control","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/from-judgments-to-issues-structured-extraction-of-legal-reasoning-with-citation-hallucination-control/82730/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-21","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What is the main goal of the presented pipeline?","Question",{"text":75,"@type":76},"Decompose Italian tax-court judgments into individual legal issues and extract, for each issue, a structured XML representation grounded in IRAC and the legal syllogism.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the method control citation hallucinations?",{"text":80,"@type":76},"It adds a hallucination-detection filter that compares model-produced references against references parsed from the judgment text by Linkoln, normalized to standard identifiers such as URN-NIR, ECLI, and CELEX.",{"name":82,"@type":73,"acceptedAnswer":83},"What was used to validate the pipeline and what was evaluated?",{"text":84,"@type":76},"The pipeline was validated on 50 judgments annotated by two PhDs in tax law, computing agreement for issue extraction and citations between annotators and between LLM and experts, plus a standalone evaluation of the hallucination filter.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]