[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-81763-en":3,"doc-seo-81763-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},81763,4398048949847,"Eliana","https://ap-avatar.wpscdn.com/avatar/400002536579ef2da7f?_k=1778318612642679267",8,"Research & Report","MVDGC Joint 3D and 2D Multi view Pedestrian Detection via Dual Geometric Constraints","Multi-view pedestrian detection (MVPD) depends on robust feature aggregation across viewpoints to reason under heavy occlusion. Prior work often projects view features to BEV for ground localization, but perspective transformation distorts spatial structure, blurring object features and weakening BEV point localization in crowded scenes. It also underutilizes the relationship between BEV ground points and image bounding boxes. MVDGC proposes a unified framework jointly estimating BEV pedestrian positions and 2D image boxes using sparse 3D cylindrical queries with dual geometric constraints.","arXiv :2607 .00273v1 [ cs .CV] 30 Jun 2026  \nMVDGC: Joint 3D and 2D Multi-view  \nPedestrian Detection via Dual Geometric Constraints  \nThinh Phan  \nAICV Lab, EECS Department, University of Arkansas  \nHao Vo  \nAICV Lab, EECS Department, University of Arkansas  \nKhoa Vo  \nAICV Lab, EECS Department, University of Arkansas  \nThanh Ngo  \nUniversity of Information Technology  \nCuong Pham  \nPosts and Telecommunications Institute of Technology (PTIT), Vietnam  \nNgan Le  \nAICV Lab, EECS Department, University of Arkansas  \n[thinhp@uark. edu](thinhp@uark. edu)  \n[haov@uark. edu](haov@uark. edu)  \n[khoavoho@uark. edu@uark. edu](khoavoho@uark. edu@uark. edu)  \n[ngoduc.thanh@uit. edu. vn](ngoduc.thanh@uit. edu. vn)  \n[cuongpv@ptit. edu. vn](cuongpv@ptit. edu. vn)  \n[thile@uark. edu](thile@uark. edu)  \nReviewed on OpenReview: [https: // openreview. net/ forum? id= 40cVQX5Mxc](https: // openreview. net/ forum? id= 40cVQX5Mxc)  \nAbstract  \nThe core challenge in multi-view pedestrian detection (MVPD) lies in effective aggregation of visual features from different viewpoints for robust occlusion reasoning. Recent approaches have addressed this by first projecting image-view features onto a Bird’s Eye View (BEV) map, where ground localization is then performed. Despite impressive performance, the perspective transformation induces severe distortion, causing spatial structure break and degrading the quality of object feature extraction. The blurred and ambiguous features hinder accurate BEV point localization, especially in densely populated regions. Moreover, the strong mutual relationship between the BEV ground point and image bounding boxes isnot capitalized on. Although multi-view consistency of 2D detections can serve as a powerful constraint in BEV space, these detections are commonly treated as auxiliary signals rather than being jointly optimized with the primary task.  \nIn this work, we propose MVDGC, a unified framework that jointly estimates pedestrian locations on the BEV plane and 2D bounding boxes in image views. MVDGC employs a sparse set of 3D cylindrical queries that embraces geometric context across both BEV and image views, enforcing dual spatial constraints for precise localization. Specifically, the geometric constraints is established by modeling each pedestrian as a vertical cylinder whose center lies on the BEV plane and whose projection casts a rectangular box in the image views. These queries function as shape anchors that directly extract 2D features from the intact image-view features using camera projection, eliminating projection-induced distortions. The 3D cylindrical query enables the unification of BEV and ImV localization into a single task: 3D cylinder position and shape refinement.  \nExtensive experiments and ablation studies demonstrate that MVDGC achieves state-ofthe-art performance across multiple evaluation metrics on MVPD benchmarks, including WildTrack and MultiViewX. On the generalized multi-view detection (GMVD) dataset, MVDGC achieves the highest MODP and precision, while maintaining competitive perfor-  \nmance on the remaining metrics, highlighting its robustness and generalization to unseen scene configurations. Code is available at: [https://github](https://github).com/UARK-AICV/MVDGC  \n1 Introduction  \nOcclusion remains a major challenge in monocular multi-object detection, particularly in crowded scenes. A common solution is to employ multi-camera systems with overlapping fields of view, where occluded objects in one view can be recovered through complementary perspectives from other cameras. This setting motivates multi-view object detection, which plays a key role in applications like surveillance Zhang et al. (2021b); Baqué et al. (2017); Zhang & Chan (2019), autonomous driving Liu et al. (2023); Li et al. (2024); Wang et al.(2023b), and 3D pose estimation Zhang et al. (2021a); Choudhury et al. (2023); Wang et al. (2023a) . The key challenge is how to effectively aggregate information across vi","cbCaisOB1debdHvR","https://ap.wps.com/l/cbCaisOB1debdHvR","pdf",18163557,3,1,23,"English","en",105,"# Abstract\n# Introduction","[{\"question\":\"What main problem does MVDGC address in multi-view pedestrian detection?\",\"answer\":\"MVDGC addresses the challenge of aggregating multi-view features effectively for reliable occlusion reasoning, especially when BEV projection distorts spatial structure and hinders accurate localization in dense crowds.\"},{\"question\":\"How does MVDGC unify BEV and image-view localization?\",\"answer\":\"MVDGC jointly estimates pedestrian locations on the BEV plane and 2D bounding boxes in image views using sparse 3D cylindrical queries that act as geometric anchors and extract 2D features via camera projection.\"},{\"question\":\"What are the “dual geometric constraints” used by the method?\",\"answer\":\"The method models each pedestrian as a vertical cylinder whose center lies on the BEV plane and whose projection forms a rectangular box in the image views, enforcing constraints that connect BEV positions with image bounding boxes.\"}]",1784175913,58,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"mvdgc-joint-3d-and-2d-multi-view-pedestrian-detection-via-dual-geometric-constraints","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/mvdgc-joint-3d-and-2d-multi-view-pedestrian-detection-via-dual-geometric-constraints/81763/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-25","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What main problem does MVDGC address in multi-view pedestrian detection?","Question",{"text":75,"@type":76},"MVDGC addresses the challenge of aggregating multi-view features effectively for reliable occlusion reasoning, especially when BEV projection distorts spatial structure and hinders accurate localization in dense crowds.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does MVDGC unify BEV and image-view localization?",{"text":80,"@type":76},"MVDGC jointly estimates pedestrian locations on the BEV plane and 2D bounding boxes in image views using sparse 3D cylindrical queries that act as geometric anchors and extract 2D features via camera projection.",{"name":82,"@type":73,"acceptedAnswer":83},"What are the “dual geometric constraints” used by the method?",{"text":84,"@type":76},"The method models each pedestrian as a vertical cylinder whose center lies on the BEV plane and whose projection forms a rectangular box in the image views, enforcing constraints that connect BEV positions with image bounding boxes.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]