[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-125870-en":3,"doc-seo-125870-105":31,"detail-sidebar-cat-0-en-105":93},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":28,"seo_description":14,"update_tm":29,"read_time":30},125870,1099523885336,"Violet","https://ap-avatar.wpscdn.com/davatar_276721f389ce27ea32af1340a28f341c",8,"Research & Report","Textual Similarity for Legal Precedents Discovery - Assessing the Performance of Machine Learning Techniques in an Administrative Court","Legal precedents ensure consistent jurisprudence, but escalating case volumes create a manual-identification bottleneck. This study applies natural language processing and machine learning to automate the discovery of similar administrative-court cases that may indicate precedents. Using a substantial Brazil-based legal dataset, it evaluates over one hundred combinations of document representations and text vectorizations with a statistically significant, expert-vetted sample. Results show granular text representations and concept-relation extraction as especially effective, while feature refinement remains critical for performance. These findings support decision-support systems for judicial contexts and inform future research integrating technology into legal decision-making.","International Journal of Information Management Data Insights 4 (2024) 100247  \nContents lists available at ScienceDirect  \nInternational Journal of Information Management Data Insights  \njournal [homepage:](homepage: www.elsevier.com/locate/jjimei)[ www.elsevier.com/locate/jjimei](homepage: www.elsevier.com/locate/jjimei)  \n| Textual similarity for legal precedents discovery: Assessing the performance   of machine learning techniques in an administrative court\u003Cbr>Hugo Mentzingena, *, Nuno Ant´onio a, Fernando Bacao a, Marcio Cunha b\u003Cbr>a NOVA Information Management School, Lisbon, Portugal\u003Cbr>b Minist´erio Público do Rio Grande do Sul, Rio Grande do Sul, Brazil |  |  |\n| --- | --- | --- |\n| A R T I C L E I N F O |  | A B S T R A C T |\n| Keywords:\u003Cbr>Language processing Court automation Case similarity Imbalanced data |  | The importance of legal precedents in ensuring consistent jurisprudence is undisputed. Particularly in jurisdictions following the Common law, but even in Civil law systems, uniformity in case law requires adherence to precedents. However, with the growing volume of cases, manual identification becomes a bottleneck, prompting the need for automation. Leveraging the capabilities of natural language processing (NLP) and machine learning (ML), our study delves into the potential of automation in identifying similar cases indicative of precedents. Drawing from a unique, substantial dataset of legal cases from an administrative court in Brazil, we extensively evaluated over one hundred combinations of document representations and text vectorizations. Contrary to earlier studies that relied on minimal validation samples, ours employed a statistically significant sample vetted by legal experts. Our findings reveal that models focusing on granular text representations perform optimally, especially when extracting concepts and relations. Notably, while intricate models may not always guarantee superior outcomes, the importance of refining textual features cannot be understated. These findings pave the way for creating efficient decision support systems in judicial contexts and set a direction for future research aiming to integrate technology in legal decision-making. |\n\n1. Introduction  \nAdministrative courts specialize in administrative law, a branch of public law focused on public administration (Amaral-Garcia, 2021). It encompasses the set of statutes and legal principles ruling the administration and regulation of government agencies (Cornell University Law School, 2022). In many countries, these courts deal with more cases than criminal or private civil justice, as their role is closely linked to providing public services such as licensing, residence permits, or granting social benefits (Nason, 2018).  \nThe administrative courts’ actions positively impact the quality and efficiency of public administration (Batalli & Pepaj, 2022). Many aspects, such as legislative changes, migratory flows, and economic activity, influence their workload. In any way, their intervention must result in prompt and consistent judgments (Gomez, 2021; Rhode, 2004). These institutions, however, deal with limited resources and strive to keep up with the caseload (Popova et al., 2021; Susskind, 2020).  \nThe relevance of data-driven decision-making is burgeoning within all management spheres (Kushwaha et al., 2021). Administrative law courts are no exception, as they are integrating technological  \nadvancements to enhance efficiency. Henkel et al. (2017) examined the potential of language technologies in public organizations, concluding that these technologies support the notion that automation, including AI-driven systems, can streamline case processing and decision-making in judicial settings. In the context of vast volumes of data, automation is a crucial factor in increasing the efficiency of a judicial court. It may be characterized as using technology to facilitate or minimize human involvement in case processing. When effectively i","cbCaikeknfy6KpZi","https://ap.wps.com/l/cbCaikeknfy6KpZi","pdf",6614476,7,1,21,"English","en",105,"# Introduction\n## Administrative courts and the need for consistent judgments\n## Automation and language technologies in judicial settings\n## Precedents as the basis for reasoning\n# (Subsequent sections not provided in the extracted text)","[{\"question\":\"Why is automation needed for identifying similar legal precedents in administrative courts?\",\"answer\":\"The growing volume of cases makes manual identification a bottleneck. Automation helps discover similar past cases more efficiently and supports consistent judging.\"},{\"question\":\"What data and evaluation approach does the study use?\",\"answer\":\"The study uses a substantial dataset of administrative legal cases from Brazil and evaluates over one hundred combinations of document representations and text vectorizations. A statistically significant sample vetted by legal experts is used for evaluation.\"},{\"question\":\"Which modeling choices perform best according to the findings?\",\"answer\":\"Models using granular text representations perform optimally, especially when extracting concepts and relations. More complex models do not always guarantee better outcomes, so refining textual features is emphasized.\"}]","Textual Similarity for Legal Precedents Discovery - Assessing the Performance of Machine Learning Techniques in an Administrative Court | PDF",1785901736,53,{"code":4,"msg":32,"data":33},"ok",{"site_id":25,"language":24,"slug":34,"title":13,"keywords":35,"description":14,"schema_data":36,"social_meta":88,"head_meta":90,"extra_data":92,"updated_unix":29},"textual-similarity-for-legal-precedents-discovery-assessing-the-performance-of-machine-learning-techniques-in-an-administrative-court","",{"@graph":37,"@context":87},[38,55,70],{"@type":39,"itemListElement":40},"BreadcrumbList",[41,45,49,52],{"item":42,"name":43,"@type":44,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":46,"name":47,"@type":44,"position":48},"https://docshare.wps.com/document/","Document",2,{"item":50,"name":12,"@type":44,"position":51},"https://docshare.wps.com/document/research-report/",3,{"item":53,"name":13,"@type":44,"position":54},"https://docshare.wps.com/document/textual-similarity-for-legal-precedents-discovery-assessing-the-performance-of-machine-learning-techniques-in-an-administrative-court/125870/",4,{"url":53,"name":13,"@type":56,"author":57,"headline":13,"publisher":59,"fileFormat":62,"inLanguage":24,"description":14,"dateModified":63,"datePublished":64,"encodingFormat":62,"isAccessibleForFree":65,"interactionStatistic":66},"DigitalDocument",{"name":9,"@type":58},"Person",{"url":42,"name":60,"@type":61},"DocShare","Organization","application/pdf","2026-08-23","2026-08-05",true,{"@type":67,"interactionType":68,"userInteractionCount":20},"InteractionCounter",{"@type":69},"ViewAction",{"@type":71,"mainEntity":72},"FAQPage",[73,79,83],{"name":74,"@type":75,"acceptedAnswer":76},"Why is automation needed for identifying similar legal precedents in administrative courts?","Question",{"text":77,"@type":78},"The growing volume of cases makes manual identification a bottleneck. Automation helps discover similar past cases more efficiently and supports consistent judging.","Answer",{"name":80,"@type":75,"acceptedAnswer":81},"What data and evaluation approach does the study use?",{"text":82,"@type":78},"The study uses a substantial dataset of administrative legal cases from Brazil and evaluates over one hundred combinations of document representations and text vectorizations. A statistically significant sample vetted by legal experts is used for evaluation.",{"name":84,"@type":75,"acceptedAnswer":85},"Which modeling choices perform best according to the findings?",{"text":86,"@type":78},"Models using granular text representations perform optimally, especially when extracting concepts and relations. More complex models do not always guarantee better outcomes, so refining textual features is emphasized.","https://schema.org",{"og:url":53,"og:type":89,"og:title":13,"og:site_name":60,"og:description":14},"article",{"robots":91,"canonical":53},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":94},[95,99,103,107,112,117,121,124,129,132,136],{"id":21,"doc_module":4,"doc_module_name":47,"category_name":96,"show_sort_weight":97,"slug":98},"Story & Novel",90,"story-novel",{"id":48,"doc_module":4,"doc_module_name":47,"category_name":100,"show_sort_weight":101,"slug":102},"Literature",80,"literature",{"id":54,"doc_module":4,"doc_module_name":47,"category_name":104,"show_sort_weight":105,"slug":106},"Exam",70,"exam",{"id":108,"doc_module":4,"doc_module_name":47,"category_name":109,"show_sort_weight":110,"slug":111},5,"Comic",60,"comic",{"id":113,"doc_module":4,"doc_module_name":47,"category_name":114,"show_sort_weight":115,"slug":116},6,"Technology",50,"technology",{"id":20,"doc_module":4,"doc_module_name":47,"category_name":118,"show_sort_weight":119,"slug":120},"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":47,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":125,"doc_module":4,"doc_module_name":47,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":47,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":133,"doc_module":4,"doc_module_name":47,"category_name":134,"show_sort_weight":133,"slug":135},10,"Lifestyle","lifestyle",{"id":137,"doc_module":4,"doc_module_name":47,"category_name":138,"show_sort_weight":108,"slug":139},19,"General","general"]