[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82372-en":3,"doc-seo-82372-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},82372,687197207919,"Theodora","https://ap-avatar.wpscdn.com/avatar/a000253d6f5f7c60be?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779446848396160552",8,"Research & Report","Normalisation-Based Likelihood Ratio Estimation for Forensic Authorship Verification","Authorship verification (AV) determines whether two texts share the same author and, in forensics, evidence strength is quantified with likelihood ratios. Many AV pipelines rely on score calibration using a separate calibration model, which demands case-relevant data and adds time and complexity. This work introduces two calibration-free normalisation techniques—Square Root Correction and Hapax Correction—for LambdaG, targeting overestimation caused by long or highly repetitive texts. Across fifteen corpora and 100–9,500 token ranges, the methods match logistic-regression calibration; Hapax Correction is superior in about 45% of tests and often yields comparisons within 5% of the best calibrated result.","arXiv :2607 .0950 1v 1 [ cs .CL] 10 Jul 2026  \nNormalisation-Based Likelihood Ratio Estimation for Forensic Authorship Verification  \nSadie Barlowa,∗, Andrea Ninia , Edoardo Maninob  \na The University of Manchester, Department of Linguistics and English Language, Oxford  \nRoad, Manchester, M13 9PL  \nb The University of Manchester, Department of Computer Science, Oxford  \nRoad, Manchester, M13 9PL  \nAbstract  \nAuthorship verification (AV) is the task of determining whether two texts were written by the same author. In a forensic context, the strength of AV evidence can be quantified using likelihood ratios. Most AV methods are score-based and deriving well-calibrated likelihood ratios from these scores requires a separate calibration model. This, in turn, requires additional amounts of case-relevant data, which is often time-consuming to obtain and prepare. This study proposes two novel normalisation techniques, the Square Root Correction and the Hapax Correction, for deriving likelihood ratios from the AV method LambdaG without the need of a calibration model (Nini et al. , 2026) . These corrections are designed to mitigate the overestimation of evidential strength that may result from long or highly repetitive texts. Performance is evaluated against logistic regression calibration across fifteen corpora and a range of text lengths (100- 9,500 tokens), using the log-likelihood ratio cost (􀓸􀖇􀖇􀖍 ) . The proposed methods achieve performance comparable to logistic regression calibration, with the Hapax Correction outperforming it in approximately 45% of tests (weighted by corpora) . Furthermore, performance was more frequently close (within 5%) when the Hapax Correction was outperformed by logistic regression calibration, compared with the reverse comparison. Eliminating the need to train a calibration model reduces data-requirements, time and complexity, thereby increasing the accessibility and transparency of forensic text comparison. This combination of empirical performance and practical advantages supports the adoption of the proposed methods in forensic settings.  \n⋆ This work was supported by the North West Social Science Doctoral Training Partnership (NWSSDTP) [ES/P000665/1] .  \n∗ Corresponding author  \nEmail addresses: [sadie.barlow@manchester.ac.uk](sadie.barlow@manchester.ac.uk) (Sadie Barlow ),  \n[andrea.nini@manchester.ac.uk](andrea.nini@manchester.ac.uk) (Andrea Nini ), [edoardo.manino@manchester.ac.uk](edoardo.manino@manchester.ac.uk)  \n(Edoardo Manino )  \n1. Introduction  \nAcross forensic science, the Likelihood Ratio Framework has emerged as the’logically and legally’ endorsed standard for evaluation evidence (Ishihara et al. , 2022 , p.183; Forensic Science Regulator, 2021 , p.26) . The framework requires the comparison of the probability of evidence under two competing hypotheses. While the Likelihood Ratio Framework is firmly established in DNA analysis (Balding and Nichols, 1994 ; Taylor et al. , 2013) and well-validated in forensic voice comparison (Rose, 2006 ; Morrison, 2011), its application in forensic text comparison remains comparatively nascent (Ishihara, 2017 ; Grant, 2022 ; Niniet al. , 2026) .  \nFor score-based systems, Log-Likelihood Ratios (LLRs) can be derived through a process called calibration (Morrison, 2013 , p.174; van der Vloed, 2024) . This involves training a separate calibration model, which learns the necessary transformation to align predicted probabilities with empirical outcome frequencies (Guo et al. , 2017 ; Silva Filho et al. , 2023 , p.3215; Pull and Hurlin, 2025 , p.2) .  \nIn practice, calibration is most often achieved using logistic regression (Brümmer and du Preez, 2006 ; Gonzalez-Rodriguez et al. , 2007 ; Morrison, 2013) (detailed in Section 3), which requires comprehensive and case-specific training data. Compiling such datasets can often be complex and computationally demanding. Data should ideally reflect the case conditions (Morrison, 2024), however, what this means is ","cbCaitjtivggvJmx","https://ap.wps.com/l/cbCaitjtivggvJmx","pdf",327430,2,1,36,"English","en",105,"# Introduction\n## Likelihood Ratio Framework in Forensics\n## Calibration for Score-Based LLRs\n## Authorship Verification and LambdaG\n## Proposed Normalisation Techniques","[{\"question\":\"What problem does the paper address in forensic authorship verification?\",\"answer\":\"It addresses the difficulty of obtaining well-calibrated likelihood ratios for authorship verification without training a separate calibration model that requires additional case-relevant data.\"},{\"question\":\"How do the proposed normalisation techniques work?\",\"answer\":\"The Square Root Correction mitigates evidential strength overestimation driven by text length, while the Hapax Correction adjusts scores using a uniqueness measure relative to text size.\"},{\"question\":\"How is performance evaluated and how do the methods compare with logistic regression calibration?\",\"answer\":\"Performance is evaluated against logistic regression calibration across fifteen corpora and text lengths from 100 to 9,500 tokens using a log-likelihood ratio cost. The proposed methods are comparable overall, and Hapax Correction outperforms logistic regression calibration in roughly 45% of tests.\"}]",1784179984,91,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"normalisation-based-likelihood-ratio-estimation-for-forensic-authorship-verification","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/normalisation-based-likelihood-ratio-estimation-for-forensic-authorship-verification/82372/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-22","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the paper address in forensic authorship verification?","Question",{"text":75,"@type":76},"It addresses the difficulty of obtaining well-calibrated likelihood ratios for authorship verification without training a separate calibration model that requires additional case-relevant data.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How do the proposed normalisation techniques work?",{"text":80,"@type":76},"The Square Root Correction mitigates evidential strength overestimation driven by text length, while the Hapax Correction adjusts scores using a uniqueness measure relative to text size.",{"name":82,"@type":73,"acceptedAnswer":83},"How is performance evaluated and how do the methods compare with logistic regression calibration?",{"text":84,"@type":76},"Performance is evaluated against logistic regression calibration across fifteen corpora and text lengths from 100 to 9,500 tokens using a log-likelihood ratio cost. The proposed methods are comparable overall, and Hapax Correction outperforms logistic regression calibration in roughly 45% of tests.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]