[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82966-en":3,"doc-seo-82966-105":28,"detail-sidebar-cat-0-en-105":90},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":11,"language":21,"language_code":22,"site_id":23,"html_lang":22,"table_of_contents":24,"faqs":25,"seo_title":13,"seo_description":14,"update_tm":26,"read_time":27},82966,687197207639,"Asher","https://ap-avatar.wpscdn.com/davatar_a8503ba1806abce46bf441b54a3ca4cd",8,"Research & Report","Prompting Beats Fine-Tuning: Generative Expected Value Scoring for Statutory Term Retrieval","Legal concepts in statutes are often conveyed through vague, open-textured terms, making interpretation dependent on identifying the right case-law explanations. This work ranks case-law sentences by usefulness for explaining a target statutory term using a dataset of 26,959 sentences covering 42 U.S. Code concepts and four explanatory-value categories. Encoder-only supervised fine-tuning (ModernBERT) is compared with zero-shot prompting of decoder-only models. Prompting yields the strongest overall results, surpassing prior state of the art under common NDCG cutoffs.","Prompting Beats Fine-Tuning: Generative Expected Value Scoring for Statutory Term Retrieval  \nAlvin Wang1 , Jaromir Savelka1  \n1 Carnegie Mellon University, 5000 Forbes Ave, Pittsburgh, PA 15213, USA  \nAbstract  \nLegal concepts in statutes are often expressed using vague terms, and practitioners frequently turn to case law to interpret them. We study the task ofranking case-law sentences by their usefulness for explaining a concept or target statutory term, using an established dataset of 26,959 sentences covering 42 U.S. Code concepts labeled into four explanatory-value categories. We compare two families of methods: (i) supervised fine-tuning of encoder-only models (ModernBERT) and (ii) zero-shot prompting of decoder-only models. We show that across all concepts and standard NDCG cutoffs, ModernBERT largely matches earlier BERT-family baselines. In contrast, prompting decoder-only models achieves the strongest overall effectiveness, with our best system surpassing all previously reported state-of-the-art results on this task.  \nKeywords  \nInformation retrieval, statutory interpretation, case-law analysis, relevant sentences  \n1. Introduction  \nUnderstanding laws may be challenging because they need to communicate general standards and refer to classes of persons, acts, things, and circumstances [1, p. 124] . Therefore, legislators must use vague [2], open textured [1] terms, abstract standards [3], principles, and values [4] . Understanding of any provision of law may depend on understanding the meaning of a term mentioned within. Potential doubts about its meaning may be removed by explanation or interpretation [5] . When searching a database of legal documents, a lawyer may retrieve many short text snippets (sentences) that mention a particular term. Some of the sentences are most likely useful for explaining the term but others may have very little value. Manually reviewing all the snippets is labor intensive because they may often be retrieved in thousands or more.  \nIn this work, we focus on the task ofranking text snippets (sentences) more highly that are useful for interpretation or explanation of a selected term as defined in [6] . These may include definitional sentences, sentences that state explicitly in a different way what the term means or state what it does not mean, sentences that provide an example, instance, or counterexample, and sentences that show how a court determines whether something is such an example.  \nThis paper substantially extends our previously published demonstration paper [7], which focused on the ranking task and offered the initial empirical comparison of encoder-only and decoder-only language models. That study produced two initial findings: (i) more modern fine-tuned encoder models perform comparably to earlier BERT-family baselines, and (ii) decoder-only prompting, particularly with GPT models, achieves competitive performance without any task-specific fine-tuning.  \nBuilding on those findings, the present work makes the following contributions: we (i) extend the model comparison beyond proprietary OpenAI systems to include open-weight models,(ii) provide amore detailed experimental setup and evaluation protocol,(iii) analyze performance across multiple query regimes (small/large, sparse/dense), and (iv) offer a deeper investigation into the role of context and model architecture in explanatory sentence ranking. Results reported in this paper are state-of-the-art.  \nProceedings of the Eighth International Workshop on Automated Semantic Analysis of Information in Legal Text (ASAIL 2026), 12 June 2026, Singapore  \n$ [alvinw2@andrew.cmu.edu](alvinw2@andrew.cmu.edu) (A. Wang); [jsavelka@andrew.cmu.edu](jsavelka@andrew.cmu.edu) (J. Savelka)  \n􀀚 0000-0002-3674-5456 (J. Savelka)  \n © 2026 Copyright for this paper by its authors. Use permitted under Creative Commons License Attribution 4.0 International (CC BY 4.0) .  \nFigure 1: The graph on the left shows the distribution of the labels. The g","cbCaiqINksxtm0uo","https://ap.wps.com/l/cbCaiqINksxtm0uo","pdf",771508,1,"English","en",105,"# Introduction\n# Related Work\n# Dataset","[{\"question\":\"What is the core task addressed in the paper?\",\"answer\":\"The paper studies ranking case-law sentences by how useful they are for explaining a selected statutory term, using labeled usefulness categories.\"},{\"question\":\"Which model families are compared in the experiments?\",\"answer\":\"It compares supervised fine-tuning of encoder-only models (ModernBERT) with zero-shot prompting of decoder-only models.\"},{\"question\":\"What main performance conclusion does the paper report?\",\"answer\":\"Prompting decoder-only models achieves the strongest overall effectiveness and its best system surpasses previously reported state-of-the-art results on this task under standard NDCG cutoffs.\"}]",1784184366,20,{"code":4,"msg":29,"data":30},"ok",{"site_id":23,"language":22,"slug":31,"title":13,"keywords":32,"description":14,"schema_data":33,"social_meta":85,"head_meta":87,"extra_data":89,"updated_unix":26},"prompting-beats-fine-tuning-generative-expected-value-scoring-for-statutory-term-retrieval","",{"@graph":34,"@context":84},[35,52,67],{"@type":36,"itemListElement":37},"BreadcrumbList",[38,42,46,49],{"item":39,"name":40,"@type":41,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":43,"name":44,"@type":41,"position":45},"https://docshare.wps.com/document/","Document",2,{"item":47,"name":12,"@type":41,"position":48},"https://docshare.wps.com/document/research-report/",3,{"item":50,"name":13,"@type":41,"position":51},"https://docshare.wps.com/document/prompting-beats-fine-tuning-generative-expected-value-scoring-for-statutory-term-retrieval/82966/",4,{"url":50,"name":13,"@type":53,"author":54,"headline":13,"publisher":56,"fileFormat":59,"inLanguage":22,"description":14,"dateModified":60,"datePublished":61,"encodingFormat":59,"isAccessibleForFree":62,"interactionStatistic":63},"DigitalDocument",{"name":9,"@type":55},"Person",{"url":39,"name":57,"@type":58},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":64,"interactionType":65,"userInteractionCount":20},"InteractionCounter",{"@type":66},"ViewAction",{"@type":68,"mainEntity":69},"FAQPage",[70,76,80],{"name":71,"@type":72,"acceptedAnswer":73},"What is the core task addressed in the paper?","Question",{"text":74,"@type":75},"The paper studies ranking case-law sentences by how useful they are for explaining a selected statutory term, using labeled usefulness categories.","Answer",{"name":77,"@type":72,"acceptedAnswer":78},"Which model families are compared in the experiments?",{"text":79,"@type":75},"It compares supervised fine-tuning of encoder-only models (ModernBERT) with zero-shot prompting of decoder-only models.",{"name":81,"@type":72,"acceptedAnswer":82},"What main performance conclusion does the paper report?",{"text":83,"@type":75},"Prompting decoder-only models achieves the strongest overall effectiveness and its best system surpasses previously reported state-of-the-art results on this task under standard NDCG cutoffs.","https://schema.org",{"og:url":50,"og:type":86,"og:title":13,"og:site_name":57,"og:description":14},"article",{"robots":88,"canonical":50},"index,follow",{"doc_id":7,"site_id":23},{"code":4,"msg":5,"data":91},[92,96,100,104,109,114,119,122,126,129,133],{"id":20,"doc_module":4,"doc_module_name":44,"category_name":93,"show_sort_weight":94,"slug":95},"Story & Novel",90,"story-novel",{"id":45,"doc_module":4,"doc_module_name":44,"category_name":97,"show_sort_weight":98,"slug":99},"Literature",80,"literature",{"id":51,"doc_module":4,"doc_module_name":44,"category_name":101,"show_sort_weight":102,"slug":103},"Exam",70,"exam",{"id":105,"doc_module":4,"doc_module_name":44,"category_name":106,"show_sort_weight":107,"slug":108},5,"Comic",60,"comic",{"id":110,"doc_module":4,"doc_module_name":44,"category_name":111,"show_sort_weight":112,"slug":113},6,"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":44,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":44,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":44,"category_name":124,"show_sort_weight":27,"slug":125},9,"Religion & Spirituality","religion-spirituality",{"id":27,"doc_module":4,"doc_module_name":44,"category_name":127,"show_sort_weight":27,"slug":128},"World Cup","world-cup",{"id":130,"doc_module":4,"doc_module_name":44,"category_name":131,"show_sort_weight":130,"slug":132},10,"Lifestyle","lifestyle",{"id":134,"doc_module":4,"doc_module_name":44,"category_name":135,"show_sort_weight":105,"slug":136},19,"General","general"]