[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83340-en":3,"doc-seo-83340-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83340,687197207919,"Theodora","https://ap-avatar.wpscdn.com/avatar/a000253d6f5f7c60be?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779446848396160552",8,"Research & Report","Dual-Correlation Hypergraph Network for Unaligned RGBT Video Object Detection and a Large-Scale Benchmark","RGB-Thermal (RGBT) video object detection addresses the robustness limits of RGB-only methods under low-light, overexposure, and adverse weather such as rain and snow. Real RGBT pairs often suffer from spatial misalignment, harming cross-modal fusion. DHNet tackles this with a patch-based spatial alignment approach (PSAM) and a dual-correlation hypergraph fusion strategy (DHFM) that learns temporal and crossmodal complementary correlations. It also introduces DVT-VOD1000, a large-scale, scene-diverse benchmark with 1,000 sequences and 103,464 image pairs.","Dual-Correlation Hypergraph Network for Unaligned RGBT Video Object Detection and A Large-scale Benchmark  \nQishun Wang, Yapeng Li, Bin Luo, Zhengzheng Tu, and Chenglong Li, Senior Member, IEEE  \narXiv :2607 .08 19 1v 1 [ cs .CV] 9 Jul 2026  \nAbstract—RGB-Thermal (RGBT) Video Object Detection (VOD) has gained significant traction due to its ability to overcome the limitations of conventional RGB-based VOD under challenging conditions. However, spatial misalignment commonly exists between RGBT image pairs. To address this, we propose a Dual-Correlation Hypergraph Network (DHNet) that captures high-dimensional complementary information by explicitly modeling two types of correlations: temporal correlation across consecutive frames and spatial correlation from crossmodal features. Specifically, we first design a Patch-based Spatial Alignment Module (PSAM) to sequentially align the multimodal features at the local region level. Subsequently, we introduce a Dual Hypergraph Fusion Module (DHFM), which constructs separate temporal and multimodal hypergraphs to enhance object discriminability through dual-correlation learning. Furthermore, the field currently lacks a large-scale, scene-diverse benchmark dataset for comprehensive evaluation. To address this gap, we construct DVT-VOD1000, a large-scale RGBT VODdataset containing 1,000 video sequences with 103,464 RGBT image pairs. The dataset covers diverse scenarios, including campuses, parks, transportation, rural areas, night scenes, rain, and snow. Comprehensive experiments on VT-VOD50 and our DVT-VOD1000 demonstrate that DHNet achieves state-of-theart detection accuracy. The dataset and source code will be made publicly available on [https://github.com/tzz-ahu/ to support](https://github.com/tzz-ahu/ to support)[ ](https://github.com/tzz-ahu/ to support)academic research.  \nIndex Terms—RGB-Thermal, video object detection, multimodal fusion, hypergraph, benchmark dataset.  \nI. INTRODUCTION  \nAS a key visual perception task, Video Object Detection  \n(VOD) serves as a core component in fields like security surveillance and autonomous driving [1]–[3] . However, VOD that relies solely on RGB information remains susceptible to robustness issues in extreme scenarios, such as low-light conditions at night, overexposure, and adverse weather including rain, snow, and fog. To mitigate these limitations, the work by Wang et al. [4] introduces a method that combines RGB and thermal (RGBT) for VOD and creates a corresponding benchmark dataset, VT-VOD50 . The authors apply manual processing to this dataset to achieve consistency in both resolution and spatial distribution. However, the VTVOD50 dataset contains only 100 RGBT video sequencesand fewer than 10,000 image pairs. Furthermore, its data  \nCorresponding author: Zhengzheng Tu.  \nQishun Wang, Yapeng Li, Zhengzheng Tu, and Bin Luo are with Anhui Provincial Key Laboratory of Multimodal Cognitive Computation, School of Computer Science and Technology, Anhui University, Hefei 230601, China (email: [qishunahu@163.com](qishunahu@163.com); [zhengzhengahu@163.com](zhengzhengahu@163.com); [luobin@ahu.edu.cn](luobin@ahu.edu.cn)).  \nChenglong Li is with Anhui Provincial Key Laboratory of Multimodal Cognitive Computation, School of Artificial Intelligence, Anhui University, Hefei, 230601, China (e-mail: [lcl1314@foxmail.com](lcl1314@foxmail.com)) .  \n\n| \u003Cbr>VT-VOD50 (a) |  | \u003Cbr>\u003Cbr> |\n| --- | --- | --- |\n|  |  |  |\n\nFig. 1. Existing VT-VOD50 and our proposed DVT-VOD1000, commonly exhibit weak spatial alignment of objects. As shown in (b), this issue manifestsas significant regional variations in the degree of misalignment.  \noriginates exclusively from traffic scenarios, representing a single data domain. These limitations pose constraints on the comprehensive evaluation of RGBT VOD methods and adversely affect model generalization.  \nWang et al. [4] introduce EINet for RGBT VOD. The model employs a negative activation function to identify noise region","cbCaik7kGgCnlzRQ","https://ap.wps.com/l/cbCaik7kGgCnlzRQ","pdf",3769712,3,1,11,"English","en",105,"# Abstract\n# Index Terms\n# Introduction\n## Background: RGB-only limitations in VOD\n## Related Work and Dataset Gaps\n## Summary of Main Challenges and Proposed Approach","[{\"question\":\"What problem does the document target in RGBT video object detection?\",\"answer\":\"It targets spatial misalignment commonly existing between RGB and thermal frames, which weakens cross-modal fusion and reduces detection accuracy.\"},{\"question\":\"How does DHNet improve alignment and feature learning?\",\"answer\":\"DHNet uses a Patch-based Spatial Alignment Module (PSAM) for local region alignment and a Dual Hypergraph Fusion Module (DHFM) to learn dual correlations from temporal information and crossmodal features.\"},{\"question\":\"What benchmark dataset is introduced, and what does it include?\",\"answer\":\"The document introduces DVT-VOD1000, containing 1,000 video sequences and 103,464 RGBT image pairs, covering diverse scenes such as campuses, parks, transportation, rural areas, night conditions, rain, and snow.\"}]",1784186867,28,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"dual-correlation-hypergraph-network-for-unaligned-rgbt-video-object-detection-and-a-large-scale-benchmark","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/dual-correlation-hypergraph-network-for-unaligned-rgbt-video-object-detection-and-a-large-scale-benchmark/83340/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-25","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the document target in RGBT video object detection?","Question",{"text":75,"@type":76},"It targets spatial misalignment commonly existing between RGB and thermal frames, which weakens cross-modal fusion and reduces detection accuracy.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does DHNet improve alignment and feature learning?",{"text":80,"@type":76},"DHNet uses a Patch-based Spatial Alignment Module (PSAM) for local region alignment and a Dual Hypergraph Fusion Module (DHFM) to learn dual correlations from temporal information and crossmodal features.",{"name":82,"@type":73,"acceptedAnswer":83},"What benchmark dataset is introduced, and what does it include?",{"text":84,"@type":76},"The document introduces DVT-VOD1000, containing 1,000 video sequences and 103,464 RGBT image pairs, covering diverse scenes such as campuses, parks, transportation, rural areas, night conditions, rain, and snow.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]