[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-121665-en":3,"doc-seo-121665-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},121665,687197100911,"Himbo","https://ap-avatar.wpscdn.com/avatar/a000239b6f1da00475?x-image-process=image/resize,m_fixed,w_180,h_180&k=1785132997149421697",8,"Research & Report","I beg to differ - how disagreement is handled in the annotation of legal machine learning data sets","Legal documents such as contracts and laws invite interpretation, so disagreement persists across legal actors and can remain unresolved even after costly settlements. This work analyzes the state of the art in annotating legal machine learning datasets and how annotator disagreement is managed and reported. Results show most published datasets eliminate traces of disagreement rather than leveraging information in conflicting labels, and many provide limited detail on how a “gold standard” is derived. Based on the findings, the article proposes implementable ways to improve handling and reporting of disagreement.","Artificial Intelligence and Law  \n[https://doi.org/10.1007/s10506-023-09369-4](https://doi.org/10.1007/s10506-023-09369-4)  \nORIGINAL RESEARCH  \nI beg to differ: how disagreement is handled  \nin the annotation of legal machine learning data sets  \nDaniel Braun1  \nAccepted: 12 June 2023 © The Author(s) 2023  \nAbstract  \nLegal documents, like contracts or laws, are subject to interpretation. Different people can have different interpretations of the very same document. Large parts of judicial branches all over the world are concerned with settling disagreements that arise, in part, from these different interpretations. In this context, it only seems natural that during the annotation of legal machine learning data sets, disagreement, how to report it, and how to handle it should play an important role. This article presentsan analysis of the current state-of-the-art in the annotation of legal machine learning data sets. The results of the analysis show that all of the analysed data sets remove all traces of disagreement, instead of trying to utilise the information that might be contained in conflicting annotations. Additionally, the publications introducing the data sets often do provide little information about the process that derives the “gold standard” from the initial annotations, often making it difficult to judge the reliability of the annotation process. Based on the state-of-the-art, the article provides easily implementable suggestions on how to improve the handling and reporting of disagreement in the annotation of legal machine learning data sets.  \nKeywords Data annotation · Legal corpora · Annotator agreement  \n1 Introduction  \nDisagreement is the default state in legal proceedings. While it is often their goal to settle a disagreement, e.g. by a court decision, the state of disagreement can prevail for a long time and in some cases, a disagreement might never be settled. Parties ina lawsuit can disagree, legal scholars can disagree, courts can disagree with eachother, and even judges in the same court can disagree. Sometimes, the state of disagreement is so valuable to one of the involved parties, that they are willing to pay  \n* Daniel Braun  \n[d.braun@utwente.nl](d.braun@utwente.nl)  \n1 Department of High-Tech Business and Entrepreneurship, University of Twente, Hallenweg 17, 7522 NH Enschede, The Netherlands  \n1 3  \nlarge sums in out-of-court settlements, to prevent the official resolution of the underlying disagreement.  \nThe process and methods surrounding artificial intelligence (AI) and machine learning (ML), on the other hand, are optimised towards finding a single “truth”, a gold standard. From the annotation of data sets, where “outliers” are often eliminated by majority vote and a high inter-annotator agreement is seen as a sign of quality, to the presentation of the predictions models make, where often, only the most probable output will be shown, disagreement is systematically eradicated in favour of a single “truth”.  \nWhen legal machine learning data sets are annotated, these opposing worlds have to be combined. Annotations have to be correct from a legal perspective and usable from a technical perspective. This article presents an analysis of how disagreement between annotators is handled in the annotation of legal data sets. A review of 29 manually annotated machine learning data sets and the corresponding publications describing their annotation shows that most data sets are annotated by multiple annotators and that the majority of the accompanying papers describe how disagreement is handled, however, often only in little detail. None of the data sets we investigated provides the raw data including disagreeing annotations in the published corpus.  \nIn their paper “Analyzing Disagreements”, Beigman Klebanov et al. (2008) differentiate two types of disagreement in annotations: mistakes due to a lack of attention and “genuine subjectivity”. Disagreement that originates from a lack of attention does","cbCaiec5rLFnAZpu","https://ap.wps.com/l/cbCaiec5rLFnAZpu","pdf",862270,1,24,"English","en",105,"# Introduction\n## Disagreement in legal proceedings\n## Disagreement and “gold standard” in AI/ML annotation\n## Research scope and dataset review\n## Types and sources of disagreement\n# State-of-the-art analysis and recommendations","[{\"question\":\"Why is disagreement central to legal proceedings and legal ML annotation?\",\"answer\":\"Disagreement is the default state in legal proceedings because interpretations can differ among parties, scholars, courts, and even judges. Legal ML dataset annotation must combine legal correctness with technical usability despite this inherent variability.\"},{\"question\":\"How do existing legal machine learning datasets typically handle disagreement?\",\"answer\":\"Most analyzed datasets remove traces of disagreement instead of using information contained in conflicting annotations. Many accompanying publications also provide limited detail on how the “gold standard” is produced from initial labels.\"},{\"question\":\"What sources of disagreement does the article discuss in annotation?\",\"answer\":\"The article builds on a distinction between inattentive mistakes and “genuine subjectivity,” and argues that legal domains can also generate more objective disagreement from missing context or conflicting court decisions that are simultaneously valid.\"}]","I beg to differ - how disagreement is handled in the annotation of legal machine learning data sets | PDF",1785806070,60,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"i-beg-to-differ-how-disagreement-is-handled-in-the-annotation-of-legal-machine-learning-data-sets","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/i-beg-to-differ-how-disagreement-is-handled-in-the-annotation-of-legal-machine-learning-data-sets/121665/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-04",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why is disagreement central to legal proceedings and legal ML annotation?","Question",{"text":75,"@type":76},"Disagreement is the default state in legal proceedings because interpretations can differ among parties, scholars, courts, and even judges. Legal ML dataset annotation must combine legal correctness with technical usability despite this inherent variability.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How do existing legal machine learning datasets typically handle disagreement?",{"text":80,"@type":76},"Most analyzed datasets remove traces of disagreement instead of using information contained in conflicting annotations. Many accompanying publications also provide limited detail on how the “gold standard” is produced from initial labels.",{"name":82,"@type":73,"acceptedAnswer":83},"What sources of disagreement does the article discuss in annotation?",{"text":84,"@type":76},"The article builds on a distinction between inattentive mistakes and “genuine subjectivity,” and argues that legal domains can also generate more objective disagreement from missing context or conflicting court decisions that are simultaneously valid.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,109,114,119,122,127,130,134],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":29,"slug":108},5,"Comic","comic",{"id":110,"doc_module":4,"doc_module_name":46,"category_name":111,"show_sort_weight":112,"slug":113},6,"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]