[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-117921-en":3,"doc-seo-117921-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},117921,16904993612988,"Olivia Brown","https://ap-avatar.wpscdn.com/davatar_a8503ba1806abce46bf441b54a3ca4cd",8,"Research & Report","Taking the Law More Seriously - Investigating Design Choices in Machine Learning Prediction Research","Machine learning approaches to court case prediction vary in success and legal reasonableness, partly because core legal requirements such as justification are difficult for learning systems. This work examines how specific research design choices influence both predictive effectiveness and legal alignment. Four models are trained on European Court of Human Rights cases, evaluating the impact of performance metrics, input coverage of case elements, degree of legal specialization, and temporal effects from past decisions, with experiments reporting accuracy and MCC.","Taking the Law More Seriously by Investigating Design Choices in Machine Learning Prediction Research  \nCor Steging1, * , Silja Renooij2 and Bart Verheij1  \n1 Bernoulli Institute of Mathematics, Computer Science and Artificial Intelligence, University of Groningen  \n2 Department of Information and Computing Sciences, Utrecht University  \nAbstract  \nApproaches to court case prediction using machine learning differ widely with varying levels of success and legal reasonableness. In part this is due to some aspects of law, such as justification, being inherently difficult for machine learning approaches. Another aspect is the effect of design choices and the extent to which these are legally reasonable, which has not yet been extensively studied. We create four machine learning models tasked with predicting cases from the European Court of Human Rights and we perform experiments in order to measure the role of the following four design choices and effects: the choice of performance metric; the effect of including different parts of the legal case; the effect of a more or less specialized legal focus; and the temporal effects of the available past legal decisions. Through this research, we aim to study design decisions and their limitations and how they affect the performance of machine learning models.  \nKeywords  \nCourt case prediction, design choices, machine learning  \n1. Introduction  \nRecently, much work has been done in the field of court case predictions. While automatically determining the outcome of court cases remains an academic exercise, the large variation in the ways that previous research has tackled the problem makes it nearly impossible to compare the approaches [1] . The law has unique characteristics, making it difficult to apply machine learning in the legal domain: machine learning is retrospective, assumes normally distributed, homogeneous data that is largely free of errors, and it often cannot explain its decisionmaking [2] . The law on the other hand is prospective, changes over time, contains wrong decisions, and demands arguments for the decisions made. These unique characteristics of the law are not always taken into account. To take the law more seriously, we must consider these when doing machine learning research in the field of AI & Law.  \nSome requirements of the law, such as justification, are inherently difficult for machine learning systems, and machine learning systems have been shown to use unsound reasoning [3] . However, despite their importance, our focus in this paper will not be on justification, responsibility or explainibility. Moreover, our goal is not to create a machine learning system that obtains a better  \nProceedings of the Sixth Workshop on Automated Semantic Analysis of Information in Legal Text (ASAIL 2023), June 23, 2023, Braga, Portugal.  \n* Corresponding author.  \n$ [c.c.steging@rug.nl](c.c.steging@rug.nl) (C. Steging); [s.renooij@uu.nl](s.renooij@uu.nl) (S. Renooij);  \n[bart.verheij@rug.nl](bart.verheij@rug.nl) (B. Verheij)  \n􀀚 0000-0001-6887-1687 (C. Steging); 0000-0003-4339-8146  \n(S. Renooij); 0000-0001-8927-8751 (B. Verheij)  \n© 2023 Copyright for this paper by its authors. Use permitted under Creative Commons License  \n\n|  | CEUR Workshop Proceedings |\n| --- | --- |\n\nAttribution 4 .0 International (CC BY 4 .0) .  \nCEUR Workshop Proceedings ([CEUR-WS.org](CEUR-WS.org))  \n[http://ceur-ws.org](http://ceur-ws.org)  \n[ISSN 1613-0073](ISSN 1613-0073)  \nperformance, or has a better alignment with legal experts [4] . Instead, we investigate the effect of specific design choices and effects in machine learning research, in order to better analyze performance and alignment with characteristics of the legal domain.  \nWe focus on research involving cases from the European Court of Human Rights (ECHR), which has been used as a benchmark in a number of studies. ECHR data is included in the LexGLUE benchmark datasets [5], and forms the basis of the ECHR-OD repository [6] . Previous ","cbCaigQ15RYyQGUQ","https://ap.wps.com/l/cbCaigQ15RYyQGUQ","pdf",1455970,1,11,"English","en",105,"# Introduction\n## Background\n## Experimental Setup\n## Experiments and Results\n## Conclusion","[{\"question\":\"Why is court case prediction in machine learning difficult in legal domains?\",\"answer\":\"Legal characteristics such as justification are hard for learning systems, and learning models may not produce sound reasoning or explain their decisions adequately. The law’s prospective, time-varying nature also differs from typical machine learning assumptions.\"},{\"question\":\"Which design choices does the research investigate?\",\"answer\":\"The study evaluates the choice of performance metric, the effect of including different parts of each legal case, the impact of using more or less specialized legal focus, and temporal effects based on available past decisions.\"},{\"question\":\"How are model performances measured in the experiments?\",\"answer\":\"Four model types are trained, and results are reported using accuracy and Matthew’s Correlation Coefficient (MCC) across tasks, models, and datasets, including both replication/expansion and generalist versus ensemble comparisons.\"}]","Taking the Law More Seriously - Investigating Design Choices in Machine Learning Prediction Research | PDF",1785680386,28,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"taking-the-law-more-seriously-investigating-design-choices-in-machine-learning-prediction-research","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/taking-the-law-more-seriously-investigating-design-choices-in-machine-learning-prediction-research/117921/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-02",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why is court case prediction in machine learning difficult in legal domains?","Question",{"text":75,"@type":76},"Legal characteristics such as justification are hard for learning systems, and learning models may not produce sound reasoning or explain their decisions adequately. The law’s prospective, time-varying nature also differs from typical machine learning assumptions.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"Which design choices does the research investigate?",{"text":80,"@type":76},"The study evaluates the choice of performance metric, the effect of including different parts of each legal case, the impact of using more or less specialized legal focus, and temporal effects based on available past decisions.",{"name":82,"@type":73,"acceptedAnswer":83},"How are model performances measured in the experiments?",{"text":84,"@type":76},"Four model types are trained, and results are reported using accuracy and Matthew’s Correlation Coefficient (MCC) across tasks, models, and datasets, including both replication/expansion and generalist versus ensemble comparisons.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]