[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-81894-en":3,"doc-seo-81894-105":31,"detail-sidebar-cat-0-en-105":93},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":28,"seo_description":14,"update_tm":29,"read_time":30},81894,8796095462418,"Noah","https://ap-avatar.wpscdn.com/avatar/80000253c1241d02b47?x-image-process=image/resize,m_fixed,w_180,h_180&k=1778826106357471780",8,"Research & Report","Shortcut Learning in Legal Judgment Prediction Empirical Evidence from the UK Employment Tribunal","Current Legal Judgment Prediction (LJP) is limited by reliance on post-hoc judicial materials, which can shift systems toward retrospective classification rather than true forecasting. This paper empirically examines shortcut learning in a claim-level setting using 33,158 UK Employment Tribunal claims and compares TF-IDF interpretable models with black-box LLMs. Performance increases when test items contain outcome-revealing leakage cues, and masking these features yields only negligible Macro-F1 loss.","1  \nShortcut Learning in Legal Judgment Prediction: Empirical Evidence from the UK Employment  \nTribunal  \nJoe Watson1,2*, Joana Ribeiro de Faria1, Marcus Tomalin3, Måns Magnusson4, Huiyuan Xie5, Hao Tian Yeung6, Christine Carter1, Jonathan Rutherford1, Felix Steffek1  \n1. Faculty of Law, University of Cambridge, Cambridge, United Kingdom  \n2. The Psychometrics Centre, Cambridge Judge Business School, University of Cambridge, Cambridge, United Kingdom  \n3. Faculty of English, University of Cambridge, Cambridge, United Kingdom  \n4. Department of Statistics, Uppsala University, Uppsala, Sweden  \n5. Department of Computer Science and Technology, Tsinghua University, Beijing, China  \n6. Department of Engineering, University of Cambridge, Cambridge, United Kingdom  \n* Corresponding author: Joe Watson, [jmw239@cam.ac.uk](jmw239@cam.ac.uk)  \n2  \nAbstract  \nCurrent Legal Judgment Prediction (LJP) is constrained by its reliance on post-hoc judicial materials, increasing the likelihood that models perform retrospective classification rather than true forecasting. This paper empirically investigates shortcut learning in this context by studying claim-level outcome prediction in UK Employment Tribunal (UKET) decisions. Using a corpus of 33,158 individual claims, we predict outcomes from claim texts and LLMextracted case summaries, evaluating models ranging from interpretable TF-IDF-based classifiers to black-box LLMs. While headline predictive performance figures appear strong, we demonstrate that such performance in LJP systems trained on post-hoc judicial text can be driven by the retrospective nature of the source material. Stratifying the test data by human judgments of leakage reveals that performance increases where outcome-revealing cues are embedded in the narrative. Moreover, a model trained on just the 4% of features identified as leakage achieves high performance, outperforming human experts. These findings substantiate concerns that LJP performance may be exaggerated by linguistic artefacts. Yet this vulnerability is not fatal to the research agenda. Instead, post-hoc judgments might be treated as potentially contaminated texts, requiring active auditing. Retraining models after masking leakage features results in only a negligible reduction in Macro-F1 . Hence, while models will opportunistically exploit shortcuts when available, they remain capable of extracting useful predictive signals when these artefacts are removed.  \nKeywords  \nLegal Judgment Prediction; shortcut learning; information leakage; UK Employment Tribunal  \n3  \nMain Text  \n1 Introduction  \nThe practical promise of Legal Judgment Prediction (LJP) rests on the possibility of forecasting court outcomes from information available before a decision is made, estimating a dispute’s likely resolution from ex ante case materials. In its strongest form, therefore, LJPis a temporally constrained task: a model should not rely on information that would only become available after the relevant legal decision has been reached. This requirement matters because LJP is often motivated by its potential to help litigants, lawyers, or policymakers assess likely outcomes before proceedings. For that promise to hold, predictions must be based on information available at that point, rather than on facts, findings, or reasoning later used to justify the result.  \nThis ideal is difficult to satisfy in practice because most LJP work relies on post-hoc judicial texts: decisions published after the court outcome has been reached. This reliance largely arises from data constraints. Pre-decisional materials, such as claim forms, response forms, comprehensive case files, and other pre-trial documents, are rarely available to researchers (Medvedeva et al. 2023), including in the context of the UK Employment Tribunal (UKET) . As a result, many LJP studies risk performing outcome identification or outcome classification rather than true forecasting (Medvedeva et al. 2023) . The problem is tha","cbCaitgwsKa5IJ8J","https://ap.wps.com/l/cbCaitgwsKa5IJ8J","pdf",1183223,5,1,25,"English","en",105,"# Introduction\n## Legal Judgment Prediction as a temporally constrained task\n## Data limitations and post-hoc judicial texts\n## Shortcut learning and validity concerns","[{\"question\":\"What problem does the paper identify in current Legal Judgment Prediction systems?\",\"answer\":\"It highlights that reliance on post-hoc judicial materials can cause models to perform retrospective outcome identification or classification instead of forecasting from information available before a decision.\"},{\"question\":\"How does the study test for shortcut learning in the UK Employment Tribunal context?\",\"answer\":\"It uses a corpus of 33,158 individual claims to predict outcomes from claim texts and LLM-extracted case summaries, and it evaluates whether performance depends on leakage cues revealed by human judgments.\"},{\"question\":\"What happens to model performance when leakage features are masked or removed?\",\"answer\":\"Retraining after masking leakage features results in only a negligible reduction in Macro-F1, indicating models can still extract useful predictive signals even when artefacts are removed.\"}]","Shortcut Learning in Legal Judgment Prediction Empirical Evidence from the UK Employment Tribunal | PDF",1784176912,63,{"code":4,"msg":32,"data":33},"ok",{"site_id":25,"language":24,"slug":34,"title":13,"keywords":35,"description":14,"schema_data":36,"social_meta":88,"head_meta":90,"extra_data":92,"updated_unix":29},"shortcut-learning-in-legal-judgment-prediction-empirical-evidence-from-the-uk-employment-tribunal","",{"@graph":37,"@context":87},[38,55,70],{"@type":39,"itemListElement":40},"BreadcrumbList",[41,45,49,52],{"item":42,"name":43,"@type":44,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":46,"name":47,"@type":44,"position":48},"https://docshare.wps.com/document/","Document",2,{"item":50,"name":12,"@type":44,"position":51},"https://docshare.wps.com/document/research-report/",3,{"item":53,"name":13,"@type":44,"position":54},"https://docshare.wps.com/document/shortcut-learning-in-legal-judgment-prediction-empirical-evidence-from-the-uk-employment-tribunal/81894/",4,{"url":53,"name":13,"@type":56,"author":57,"headline":13,"publisher":59,"fileFormat":62,"inLanguage":24,"description":14,"dateModified":63,"datePublished":64,"encodingFormat":62,"isAccessibleForFree":65,"interactionStatistic":66},"DigitalDocument",{"name":9,"@type":58},"Person",{"url":42,"name":60,"@type":61},"DocShare","Organization","application/pdf","2026-08-05","2026-07-16",true,{"@type":67,"interactionType":68,"userInteractionCount":20},"InteractionCounter",{"@type":69},"ViewAction",{"@type":71,"mainEntity":72},"FAQPage",[73,79,83],{"name":74,"@type":75,"acceptedAnswer":76},"What problem does the paper identify in current Legal Judgment Prediction systems?","Question",{"text":77,"@type":78},"It highlights that reliance on post-hoc judicial materials can cause models to perform retrospective outcome identification or classification instead of forecasting from information available before a decision.","Answer",{"name":80,"@type":75,"acceptedAnswer":81},"How does the study test for shortcut learning in the UK Employment Tribunal context?",{"text":82,"@type":78},"It uses a corpus of 33,158 individual claims to predict outcomes from claim texts and LLM-extracted case summaries, and it evaluates whether performance depends on leakage cues revealed by human judgments.",{"name":84,"@type":75,"acceptedAnswer":85},"What happens to model performance when leakage features are masked or removed?",{"text":86,"@type":78},"Retraining after masking leakage features results in only a negligible reduction in Macro-F1, indicating models can still extract useful predictive signals even when artefacts are removed.","https://schema.org",{"og:url":53,"og:type":89,"og:title":13,"og:site_name":60,"og:description":14},"article",{"robots":91,"canonical":53},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":94},[95,99,103,107,111,116,121,124,129,132,136],{"id":21,"doc_module":4,"doc_module_name":47,"category_name":96,"show_sort_weight":97,"slug":98},"Story & Novel",90,"story-novel",{"id":48,"doc_module":4,"doc_module_name":47,"category_name":100,"show_sort_weight":101,"slug":102},"Literature",80,"literature",{"id":54,"doc_module":4,"doc_module_name":47,"category_name":104,"show_sort_weight":105,"slug":106},"Exam",70,"exam",{"id":20,"doc_module":4,"doc_module_name":47,"category_name":108,"show_sort_weight":109,"slug":110},"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":47,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":47,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":47,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":125,"doc_module":4,"doc_module_name":47,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":47,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":133,"doc_module":4,"doc_module_name":47,"category_name":134,"show_sort_weight":133,"slug":135},10,"Lifestyle","lifestyle",{"id":137,"doc_module":4,"doc_module_name":47,"category_name":138,"show_sort_weight":20,"slug":139},19,"General","general"]