[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85972-en":3,"doc-seo-85972-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},85972,13056703019404,"Miles","https://ap-avatar.wpscdn.com/davatar_29158cc5080c5b710cf443261637dec0",8,"Research & Report","Why Domain Matters: Domain-Aware Benchmarking of Underwater Object Detection and Annotation Quality","Underwater object detection is strongly affected by domain shift, where model performance varies across locations, habitats, and deployment conditions. Existing evaluations often rely on aggregate metrics that conceal failures in particular environments, while current benchmarks for domain generalization use synthetic changes that do not match real settings. A new framework assigns underwater domain labels based on appearance, scene composition, and acquisition geometry, enabling systematic study of both human annotation quality and detector performance. Results show substantial domain-dependent discrepancies and support interpretable measurement, benchmarking, and data/annotation planning for robust deployment.","Why Domain Matters: Domain-Aware Benchmarking of Underwater Object Detection and Annotation Quality  \nMelanie Wille Dimity Miller Tobias Fischer Scarlett Raine  \nQUT Centre for Robotics, Queensland University of Technology  \nBrisbane, Australia  \n{willemc, d24.miller, tobias.fischer, [sg.raine](sg.raine}@qut.edu.au)[}](sg.raine}@qut.edu.au)[@qut.edu.au](sg.raine}@qut.edu.au)  \narXiv :2607 . 10575v1 [ cs .CV] 12 Jul 2026  \nAbstract  \nUnderwater object detection is strongly affected by domain shift, where performance can vary significantly across different locations, habitats, and deployment conditions. However, detector performance is typically evaluated using aggregate metrics that hide failures in specific environments, while existing domain generalization benchmarks often rely on synthetic variations that do not reflect real-world conditions. We introduce a framework that characterizes underwater images by appearance, scene composition, and acquisition geometry to assign domain labels. Using this framework, we perform the first systematic study of how domain factors influence both human annotation quality in underwater object detection datasets and deep learning-based detector performance, revealing substantial domain-dependent discrepancies. By incorporating physically meaningful domain labels, domain shift becomes something we can characterize, measure, benchmark, and act on. We highlight how this can be used to guide data collection and annotation, design more informative benchmarks, and assess detector robustness across diverse underwater environments.  \n1. Introduction  \nUnderwater object detection is an important tool for marine science, enabling large-scale analysis of seafloor imagery to monitor benthic indicator species and assess ecosystem health under increasing human activity [6] . However, detector development often relies on collecting and annotating large amounts of training data, a process that is expensive and difficult to scale across new sites, habitats, and collection protocols [10, 31] . A major driver of data requirements is the large variation of underwater conditions, such as lighting, turbidity, depth, and environmental structure, which alter both image appearance and object characteristics, creating a wide range of distinct domains [24] . This often causes models trained on one set of conditions to perform  \nInitial Human Labels Revised Ground Truth YOLO Predictions  \nHolothurian – Echinus – Scallop – Starfish  \nFigure 1 . Example underwater images from visually distinct underwater domains. Left: original RUOD labels. Middle: revised RUOD-R labels (considered closest to reality) . Right: YOLO26n predictions. Some true objects are missed (orange crosses ), while others were hallucinated (red crosses ) . Domain conditions challenge both humans and detectors.  \npoorly when deployed in another, manifesting as domain shift, which challenges model generalization and limits data reusability [8, 10, 19] .  \nFig. 1 illustrates this challenge. Under changing underwater conditions, the same target classes may differ substantially in appearance and therefore detectability. In this paper, we investigate how such domain differences influence the performance of both object detectors and human annotators for consistently identifying and localizing targets inan image. Understanding the human side of the problem is particularly important because annotation quality directly determines the reliability of training and evaluation data, yet is typically assumed to be uniform across conditions.  \nExisting underwater object detection benchmarks provide valuable resources for training and evaluating detectors [9, 11, 15, 22], but they are not designed to explain domaindependent failures. In this work, we use the term domain to refer to a subset of images sharing common environmental,  \nscene, or acquisition conditions. Performance is typically summarized using a single dataset-level metric, such as mAP, which averages result","cbCaihWSAIfz9FUm","https://ap.wps.com/l/cbCaihWSAIfz9FUm","pdf",2000563,3,1,10,"English","en",105,"# Introduction\n## Problem: domain shift in underwater detection\n## Goal and contributions\n## Domain labeling framework","[{\"question\":\"What problem does the document address in underwater object detection?\",\"answer\":\"It addresses domain shift, where detection performance changes significantly across different underwater locations, habitats, and acquisition conditions.\"},{\"question\":\"Why are existing benchmark evaluations insufficient?\",\"answer\":\"They typically use aggregate dataset-level metrics (e.g., mAP) that hide failures specific to particular environments, and some domain generalization benchmarks rely on synthetic variations that do not reflect real deployment conditions.\"},{\"question\":\"How does the proposed framework characterize underwater domains?\",\"answer\":\"It assigns interpretable domain labels using physical image properties, including appearance, scene composition, and acquisition geometry, to enable domain-aware analysis of both detector performance and annotation quality.\"}]",1784207492,25,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"why-domain-matters-domain-aware-benchmarking-of-underwater-object-detection-and-annotation-quality","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/why-domain-matters-domain-aware-benchmarking-of-underwater-object-detection-and-annotation-quality/85972/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-25","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the document address in underwater object detection?","Question",{"text":75,"@type":76},"It addresses domain shift, where detection performance changes significantly across different underwater locations, habitats, and acquisition conditions.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"Why are existing benchmark evaluations insufficient?",{"text":80,"@type":76},"They typically use aggregate dataset-level metrics (e.g., mAP) that hide failures specific to particular environments, and some domain generalization benchmarks rely on synthetic variations that do not reflect real deployment conditions.",{"name":82,"@type":73,"acceptedAnswer":83},"How does the proposed framework characterize underwater domains?",{"text":84,"@type":76},"It assigns interpretable domain labels using physical image properties, including appearance, scene composition, and acquisition geometry, to enable domain-aware analysis of both detector performance and annotation quality.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":22,"slug":133},"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]