[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84474-en":3,"doc-seo-84474-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},84474,687197100911,"Himbo","https://ap-avatar.wpscdn.com/avatar/a000239b6f1da00475?x-image-process=image/resize,m_fixed,w_180,h_180&k=1782698725881665579",8,"Research & Report","Where Experts Disagree, Models Fail: Detecting Implicit Legal Citations in French Court Decisions","Applying computational methods to law at scale depends on separating genuine legal reasoning from surface similarity. The work focuses on detecting implicit citations of the French Civil Code, where courts apply a statutory rule without naming it. A benchmark of 1,015 annotated passage–article pairs is introduced from three legal experts. Expert disagreement is shown to predict model failure: a supervised ensemble reaches F1=0.70 overall, while false positives concentrate on disputed cases. Using top-k ranking and multi-model consensus enables 76% precision for top-200 candidates without supervision.","Where Experts Disagree, Models Fail: Detecting Implicit Legal Citations in French Court Decisions  \nAvrile Floro1  \n[avrile.floro@ip-paris.fr](avrile.floro@ip-paris.fr)  \nSoline Pellez2  \n[soline.pellez@uphf.fr](soline.pellez@uphf.fr)  \nTamara Dhorasoo2  \n[dhorasoo.tamara@uphf.fr](dhorasoo.tamara@uphf.fr)  \nNils Holzenberger1  \n[nils.holzenberger@telecom-paris.fr](nils.holzenberger@telecom-paris.fr)  \n1Télécom Paris, Institut Polytechnique de Paris 2Université Polytechnique Hauts-de-France  \narXiv :2603 .22973v2 [ cs .AI] 13 Jul 2026  \nAbstract  \nApplying computational methods to law at scale requires separating genuine legal reasoning from surface similarity. We study this through a concrete task: detecting implicit citations of the French Civil Code, where a court applies a statutory rule without naming it: a post-hoc question about the reasoning a court actually used. We release a benchmark of 1,015 passage–article pairs annotated by three legal experts. Our central finding is that their disagreement is itself informative: the third of cases the experts dispute are where models fail. Our best ensemble reaches an F1 score of 0.70 overall. Yet, two-thirds of its false positives fallon those disputed cases, a concentration that holds across all ten models we evaluate. Disagreement is a signal of intrinsic difficulty, not annotation noise. This should not block useful tools, however: reframed as top-k ranking with multi-model consensus, the same signals reach 76% precision for the top-200 candidates without supervision.  \n1 Introduction  \nA lawyer researching case law faces an asymmetry: explicit citations are trivial to find through keyword search. However, implicit applications, where a court applies a legal rule without naming it, are hidden. Consider a practitioner seeking examples of how article 2274 of the French Civil Code (the presumption of good faith) is applied in practice. Searching for “article 2274” retrieves decisions that explicitly cite this provision. Yet many decisions apply the same legal reasoning without numerical reference, using formulations such as “the mere observation of the increase in rental debt is not sufficient to establish bad faith.” This blind spot is also relevant to quantitative legal scholarship. Take a researcher studying whether French courts have expanded the scope of the good-faith presumption over the past decade. If the analysis captures only decisions that explicitly cite article 2274, it misses  \nArticle 1192 (Contract interpretation)  \n“Clear and unambiguous clauses may not be interpreted, as this would amount to distortion.”  \n✓ Found by search × Invisible to search  \nFigure 1: Explicit and implicit statutory citations. While both excerpts apply article 1192, only the left one can be found by keyword search.  \ncases where the same provision is applied implicitly, and could potentially skew conclusions about jurisprudential trends. Our work begins to address this problem by evaluating the reliability of automatic detection. Specifically, this paper tackles the task of detecting implicit statutory citations (Figure 1): given a passage from a court decision anda candidate Civil Code article, determine whether the passage applies that article’s legal rule without explicitly mentioning it.  \nThis task is both practically important and methodologically challenging. Indeed, it requires distinguishing genuine legal reasoning from semantic similarity. But how difficult is this task, and where do current methods fail? We make four contributions 1 that characterize both the limits and the practical potential of computational approaches to this problem. First, we introduce an adversarial benchmark for implicit citation detection in French civil law (§3) . We train a bi-encoder on explicit citations, use it to retrieve semantically similar candidates, then perform adversarial filtering using o3 with a conservative prompt. The final dataset comprises 1,015 pairs. Second, we conduct an  \n1Data a","cbCaieaYHK6WZ8Ky","https://ap.wps.com/l/cbCaieaYHK6WZ8Ky","pdf",363003,1,19,"English","en",105,"# Abstract\n# Introduction\n# Related Work","[{\"question\":\"What problem does the document address in legal NLP?\",\"answer\":\"It addresses detecting implicit statutory citations in French court decisions, where a court applies a Civil Code rule without explicitly naming the article.\"},{\"question\":\"How is the benchmark dataset constructed and what is its size?\",\"answer\":\"The benchmark contains 1,015 passage–article pairs annotated by three legal experts. It is built for implicit citation detection using a retrieval-and-adversarial filtering pipeline.\"},{\"question\":\"What relationship is found between expert disagreement and model performance?\",\"answer\":\"Expert disagreement signals intrinsic difficulty: roughly one third of disputed cases correspond to where models fail, and most false positives occur in those disputed cases across evaluated models.\"}]",1784195883,48,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"where-experts-disagree-models-fail-detecting-implicit-legal-citations-in-french-court-decisions","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/where-experts-disagree-models-fail-detecting-implicit-legal-citations-in-french-court-decisions/84474/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the document address in legal NLP?","Question",{"text":75,"@type":76},"It addresses detecting implicit statutory citations in French court decisions, where a court applies a Civil Code rule without explicitly naming the article.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How is the benchmark dataset constructed and what is its size?",{"text":80,"@type":76},"The benchmark contains 1,015 passage–article pairs annotated by three legal experts. It is built for implicit citation detection using a retrieval-and-adversarial filtering pipeline.",{"name":82,"@type":73,"acceptedAnswer":83},"What relationship is found between expert disagreement and model performance?",{"text":84,"@type":76},"Expert disagreement signals intrinsic difficulty: roughly one third of disputed cases correspond to where models fail, and most false positives occur in those disputed cases across evaluated models.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":45,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":45,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":21,"doc_module":4,"doc_module_name":45,"category_name":136,"show_sort_weight":106,"slug":137},"General","general"]