[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83202-en":3,"doc-seo-83202-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},83202,1374391974468,"Eden","https://ap-avatar.wpscdn.com/davatar_29158cc5080c5b710cf443261637dec0",8,"Research & Report","Prior-Matched Evaluation of Operational Earth-Observation Classifiers: A Three-Number Reporting Method Demonstrated on Sentinel-1","Internal Waves Service screens the Sentinel-1 Wave-mode archive for internal solitary waves and routes detections to experts, where expert attention is the dominant cost of error and thus precision must lead. Classifiers trained and reported on balanced data can look good while failing operationally because balanced-test metrics hide the true positive rate (~1/20). A precision-first evaluation is required: a prior-matched three-number reporting method compares balanced-test, an operational-prior frozen test, and real post-deployment adjudication to measure the honest precision gap.","arXiv :2607 .07 146v 1 [ cs .LG] 8 Jul 2026  \nPrior-matched evaluation of operational Earth-observation classifiers: a three-number reporting method demonstrated on Sentinel-1  \ninternal-wave detection ∗  \nJoão Pinelo, João Gonçalves, Arun Shukla, Adriana Santos-Ferreira†  \n8 July 2026  \nAbstract  \nThe Internal Waves Service screens the Sentinel-1 Wave-mode archive for internal solitary waves, routing detections to experts whose adjudication time is the resource the effort exists to conserve. Because attention is the cost of error, precision leads. Its classifier was trained and reported at a one-to-one class balance, fixed before the operational rate could be known. That rate has since emerged at roughly one scene in twenty, anda balanced-test score badly overstates the precision a validator meets. A model that scores 0.794 balanced-test precision scores 0.192 in real operation: the gap is a systematic artefact of reporting at the wrong prior, invisible to the metric most work quotes. We show the mismatch to be an evaluation problem in the costume of a training one at a fixed recall, prior correction and calibration cannot move precision, and answer it with a prior-matched reporting method based on three numbers: balanced-test, operational-prior, and real post-deployment, whose contrast is the honest measure. A precision-first, leakage-controlled development cycle then improves the classifier lever by lever, each promoted only against a preregistered margin; added capacity not clearing it, calibration inert, feature aggregation the one real lift, so the honest negatives are as much a result asthe gain. Holding recall at a floor of 0.80 and certifying against a sealed, single-read lockbox, the promoted model reports 0.927 precision at the operational prior; an out-of-time check confirms discrimination transfers to unseen periods while a fixed operating point does not. Prior-matched reporting, begin balanced, then move to the prior as the stream reveals it, transfers to any operational Earth-observation service bootstrapping a rare-event detector under a prior it has yet to discover.  \n∗ This manuscript is being submitted to Geoscientific Model Development.  \n†Atlantic International Research Centre (AIR Centre), Azores, Portugal. ORCID—João Pinelo: 0000-0002-4890-0775; João Gonçalves: 0009-0001-4547-1696; Arun Shukla: 0009- 0001-5812-1660; Adriana Santos-Ferreira: 0000-0002-5704-6021 . Correspondence: João Pinelo ([joao.pinelo@aircentre.org](joao.pinelo@aircentre.org))  \n1. Introduction  \nInternal solitary waves are among the most energetic features of the ocean, and their surface signature — alternating bands of increased and reduced roughness as the wave-induced currents strain the short surface waves that syntheticaperture radar (SAR) responds to—is recorded in SAR imagery, which makes a satellite SAR archive a natural instrument for mapping where and how often they occur (e.g., Barintag et al. 2023) . The Internal Waves Service builds that map operationally from Sentinel-1 Wave-mode data, running a convolutional classifier over the incoming vignettes and passing its detections to domain experts who confirm or reject each one. The classifier’s purpose is not to be the final word but to spend the experts’ time well: the archive is far larger than any team could inspect exhaustively (Sect. 2), so the model exists to concentrate scarce expert attention on the scenes most likely to carry a wave.  \nThat framing sets the cost of error. A false positive lands a wave-free scene on a validator’s desk and consumes the attention the service is built to conserve; a missed wave is largely recoverable, because Wave-mode acquisitions recur over the same locations on the orbit repeat and the archive is reprocessed, so the same site returns for another look. Precision therefore leads and recall is held to a floor rather than maximised—the operating discipline the whole effort is organised around.  \nThe diﬀiculty is not the model so much as h","cbCais570HHBvYLZ","https://ap.wps.com/l/cbCais570HHBvYLZ","pdf",2396386,1,24,"English","en",105,"# Introduction\n## Operational setting and precision-first objective\n## Why balanced-test reporting fails under operational priors\n## Three-number prior-matched reporting method","[{\"question\":\"Why can a classifier with good balanced-test precision perform poorly in real Sentinel-1 operations?\",\"answer\":\"Because the classifier is evaluated on class-balanced data, while the operational stream has a much lower positive rate (~0.05). Balanced-test metrics do not expose the negatives-to-positives ratio seen in production, hiding the operational precision gap.\"},{\"question\":\"What is the core evaluation problem addressed by the paper?\",\"answer\":\"The mismatch between the prior assumed by reported metrics and the prior encountered in deployment. The paper argues this is an evaluation issue presented as a training issue, requiring honest reporting rather than simple retraining or calibration.\"},{\"question\":\"How does the proposed three-number prior-matched reporting method work?\",\"answer\":\"It characterizes an operational classifier with three figures: balanced-test performance, performance on a frozen test set drawn at the operational prior, and real post-deployment performance from expert adjudication. The contrast among these measures is treated as the true measure of operational precision.\"}]",1784185931,60,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"prior-matched-evaluation-of-operational-earth-observation-classifiers-a-three-number-reporting-method-demonstrated-on-sentinel-1","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/prior-matched-evaluation-of-operational-earth-observation-classifiers-a-three-number-reporting-method-demonstrated-on-sentinel-1/83202/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why can a classifier with good balanced-test precision perform poorly in real Sentinel-1 operations?","Question",{"text":75,"@type":76},"Because the classifier is evaluated on class-balanced data, while the operational stream has a much lower positive rate (~0.05). Balanced-test metrics do not expose the negatives-to-positives ratio seen in production, hiding the operational precision gap.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What is the core evaluation problem addressed by the paper?",{"text":80,"@type":76},"The mismatch between the prior assumed by reported metrics and the prior encountered in deployment. The paper argues this is an evaluation issue presented as a training issue, requiring honest reporting rather than simple retraining or calibration.",{"name":82,"@type":73,"acceptedAnswer":83},"How does the proposed three-number prior-matched reporting method work?",{"text":84,"@type":76},"It characterizes an operational classifier with three figures: balanced-test performance, performance on a frozen test set drawn at the operational prior, and real post-deployment performance from expert adjudication. The contrast among these measures is treated as the true measure of operational precision.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,109,114,119,122,127,130,134],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":28,"slug":108},5,"Comic","comic",{"id":110,"doc_module":4,"doc_module_name":45,"category_name":111,"show_sort_weight":112,"slug":113},6,"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":45,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":45,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":45,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":45,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":45,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]