[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83222-en":3,"doc-seo-83222-105":29,"detail-sidebar-cat-0-en-105":94},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},83222,962075114765,"Quinn","https://ap-avatar.wpscdn.com/davatar_a8503ba1806abce46bf441b54a3ca4cd",8,"Research & Report","HAJJv2-CrowdCount Zero-Shot Benchmark for Dense Crowd Counting","Automated crowd counting in Hajj video fails for reasons beyond model capacity: near-vertical camera viewpoints, heavy long-range occlusion, and frames containing well over a thousand people. The work revisits the HAJJv2 dataset and releases HAJJv2-CrowdCount with per-second, human-annotated totals for testing videos. Three recent zero-shot paradigms are benchmarked: YOLO-World, APGCC, and SAM3Count. SAM3Count achieves the lowest overall MAE (70.4), but in densest frames all degrade sharply, while APGCC degrades more gracefully, informing deployment decisions.","HAJJv2-CrowdCount Zero-Shot Benchmark for Dense Crowd Counting  \nReem AlYabis∗ , Fares AlTuwaim∗ , AlJawharh AlOtaibi∗ , Mohamed Eltahir∗  \n[Reem@daldata.ai](Reem@daldata.ai), [Fares@daldata.ai](Fares@daldata.ai), [AlJawharh@daldata.ai](AlJawharh@daldata.ai), [M.eltayeb@daldata.ai](M.eltayeb@daldata.ai)  \nRiyadh, Saudi Arabia  \narXiv :2607 .07322v 1 [ cs .CV] 8 Jul 2026  \nAbstract—Automated crowd counting in Hajj video is difficult not because current models lack capacity, but because the footage violates the assumptions those models were built on: cameras observe the crowd from steep, near-vertical angles, individuals occlude one another extensively, and a single frame can contain well over a thousand people. Benchmarks that test crowd counting in such environment are either private or not detailed per second. We revisit the HAJJv2 dataset and contribute HAJJv2-CrowdCount: per-second human-annotated crowd counts for its testing videos1. Using these annotations, we benchmark three recent zero-shot counting paradigms: an open-vocabulary detector (YOLO-World), a point-based counter (APGCC), and a promptable segmentation-based counter (SAM3Count). SAM3Count attains the lowest overall mean absolute error (MAE 70.4, 95% CI 56.0–86.1), ahead of YOLO-World (92.0) and APGCC (152.9). This ordering reverses, however, in the regime most relevant to deployment: on the densest frames, the detectionand segmentation-based counters both degrade sharply (MAE exceeding 300), while the point-based counter degrades far more gracefully (MAE 114.9). This inversion is decision-relevant for Hajj crowd management, where reliable counts are needed most precisely in the densest and most occluded scenes. The annotations are released to support reproduction and extension of these results.  \nIndex Terms—crowd counting, Hajj, zero-shot evaluation, HAJJv2, open-vocabulary detection, segmentation  \nI. INTRODUCTION  \nCrowd management during Hajj is fundamentally a safety problem. Accurate, near-real-time estimates of how many people occupy a corridor or courtyard directly inform operational decisions such as gate timing, flow direction, and crowd holds. Automated counting from existing camera infrastructure is therefore an attractive capability, but Hajj footage presents conditions under which most counting methods perform poorly: crowds are extremely dense, viewpoints are frequently near-vertical, and individuals are persistently occluded by one another.  \nA central practical question for a team seeking near-term deployment is whether a recent, general-purpose model can be applied directly, zero-shot without Hajj-specific training, and still achieve accuracy sufficient for an operations dashboard. Answering this question requires two resources that do not currently exist together: a Hajj test set with reliable per-frame counts, and a controlled, like-for-like comparison of current models evaluated on it.  \n1 Per-second crowd-count annotations available at: [https://github.com/](https://github.com/)[ ](https://github.com/)reem-8899/HAJJv2-CrowdCount  \nThis paper provides both. First, we annotate the HAJJv2 testing videos [4] with per-second crowd counts and release them publicly. Second, we benchmark three recent modelson these annotations under an identical zero-shot protocol, reporting not only aggregate accuracy but where each model fails, which we show is precisely where an aggregate ranking becomes misleading.  \nII. RELATED WORK  \nDensity-map regression has been the dominant paradigm in crowd counting since CSRNet [5] demonstrated that dilated convolutions on a VGG backbone can produce sharp density maps for congested scenes. Subsequent work has refined this approach considerably. More recently, point- and attentionbased formulations such as APGCC [1] have improved accuracy on standard ShanghaiTech-style benchmarks by predicting head locations directly rather than only a scalar total.  \nIn parallel, open-vocabulary detectors that accept a text prompt, e.g","cbCaia4fOtj4zpqG","https://ap.wps.com/l/cbCaia4fOtj4zpqG","pdf",3544690,1,5,"English","en",105,"# Introduction\n## Practical deployment problem\n## New resources and benchmark design\n# Related Work\n## Density-map regression\n## Point- and attention-based counting\n## Open-vocabulary detection and segmentation prompting\n# Per-second Annotations for HAJJV2\n## Dataset composition and sampling\n## Labeling protocols","[{\"question\":\"Why do common crowd-counting models struggle on Hajj footage?\",\"answer\":\"Because Hajj footage violates typical assumptions: cameras observe from steep, near-vertical angles, people occlude each other extensively, and many frames contain extremely large numbers of individuals.\"},{\"question\":\"What does HAJJv2-CrowdCount contribute to the benchmark?\",\"answer\":\"It provides per-second, human-annotated total crowd counts for the HAJJv2 testing videos, enabling a detailed zero-shot evaluation with reliable per-frame ground truth.\"},{\"question\":\"How do the zero-shot methods compare, especially on the densest frames?\",\"answer\":\"Overall, SAM3Count achieves the lowest mean absolute error, but on the densest and most occluded frames detection-and-segmentation methods degrade sharply (MAE above 300), while the point-based APGCC degrades more gracefully (MAE 114.9).\"},{\"question\":\"How were per-second labels created for the benchmark videos?\",\"answer\":\"Each testing video was sampled at one frame per second, and annotators recorded the number of visible people in each sampled frame using either independent counting or model-assisted counting based on detector outputs.\"}]",1784186034,13,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":89,"head_meta":91,"extra_data":93,"updated_unix":27},"hajjv2-crowdcount-zero-shot-benchmark-for-dense-crowd-counting","",{"@graph":35,"@context":88},[36,53,67],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/hajjv2-crowdcount-zero-shot-benchmark-for-dense-crowd-counting/83222/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":61,"encodingFormat":60,"isAccessibleForFree":62,"interactionStatistic":63},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-16",true,{"@type":64,"interactionType":65,"userInteractionCount":4},"InteractionCounter",{"@type":66},"ViewAction",{"@type":68,"mainEntity":69},"FAQPage",[70,76,80,84],{"name":71,"@type":72,"acceptedAnswer":73},"Why do common crowd-counting models struggle on Hajj footage?","Question",{"text":74,"@type":75},"Because Hajj footage violates typical assumptions: cameras observe from steep, near-vertical angles, people occlude each other extensively, and many frames contain extremely large numbers of individuals.","Answer",{"name":77,"@type":72,"acceptedAnswer":78},"What does HAJJv2-CrowdCount contribute to the benchmark?",{"text":79,"@type":75},"It provides per-second, human-annotated total crowd counts for the HAJJv2 testing videos, enabling a detailed zero-shot evaluation with reliable per-frame ground truth.",{"name":81,"@type":72,"acceptedAnswer":82},"How do the zero-shot methods compare, especially on the densest frames?",{"text":83,"@type":75},"Overall, SAM3Count achieves the lowest mean absolute error, but on the densest and most occluded frames detection-and-segmentation methods degrade sharply (MAE above 300), while the point-based APGCC degrades more gracefully (MAE 114.9).",{"name":85,"@type":72,"acceptedAnswer":86},"How were per-second labels created for the benchmark videos?",{"text":87,"@type":75},"Each testing video was sampled at one frame per second, and annotators recorded the number of visible people in each sampled frame using either independent counting or model-assisted counting based on detector outputs.","https://schema.org",{"og:url":51,"og:type":90,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":92,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":95},[96,100,104,108,112,117,122,125,130,133,137],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":97,"show_sort_weight":98,"slug":99},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":101,"show_sort_weight":102,"slug":103},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":105,"show_sort_weight":106,"slug":107},"Exam",70,"exam",{"id":21,"doc_module":4,"doc_module_name":45,"category_name":109,"show_sort_weight":110,"slug":111},"Comic",60,"comic",{"id":113,"doc_module":4,"doc_module_name":45,"category_name":114,"show_sort_weight":115,"slug":116},6,"Technology",50,"technology",{"id":118,"doc_module":4,"doc_module_name":45,"category_name":119,"show_sort_weight":120,"slug":121},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":123,"slug":124},30,"research-report",{"id":126,"doc_module":4,"doc_module_name":45,"category_name":127,"show_sort_weight":128,"slug":129},9,"Religion & Spirituality",20,"religion-spirituality",{"id":128,"doc_module":4,"doc_module_name":45,"category_name":131,"show_sort_weight":128,"slug":132},"World Cup","world-cup",{"id":134,"doc_module":4,"doc_module_name":45,"category_name":135,"show_sort_weight":134,"slug":136},10,"Lifestyle","lifestyle",{"id":138,"doc_module":4,"doc_module_name":45,"category_name":139,"show_sort_weight":21,"slug":140},19,"General","general"]