[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84803-en":3,"doc-seo-84803-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},84803,2336464648322,"Aria","https://ap-avatar.wpscdn.com/avatar/2200025388227c56fec?_k=1778556882303663488",8,"Research & Report","SteelBench Evaluating Vision Language Models in Real World Industrial Environments","Existing video benchmarks for action recognition often rely on consumer videos, egocentric recordings, or simulated industrial settings, where subjects are clearly visible and procedures are controlled. SteelBench targets real operational industrial CCTV conditions with distant workers, dust, steam, low light, glare, occlusion, and concurrent activities. The benchmark provides 1,345 densely annotated clips from 149 operational hours and 10,024 candidate videos, including per-worker actions, PPE attributes, spatial context, and safety-rule compliance. It also introduces a provenance-aware audit protocol to quantify label influence and reveal evaluation bias, showing limited model accuracy and frequent incorrect safety judgments. ","arXiv :2607 .05264v 1 [ cs .CV] 6 Jul 2026  \nSteelBench: Evaluating Vision-Language Models in Real-World Industrial Environments  \nSuryanarayana Reddy Yarrabothula 1 ,2   \nManisha Chawla 1 , Gagan Raj Gupta 1 , Kunal Sinha3 , Sashank Lekkala 1 , Ashirvadhan Dosapati 1 , Saikamal Nannuri 1 , Katragadda Ajay RamaSwamy Chowdary Gowtham 1  \n1Indian Institute of Technology Bhilai, 2 Steel Authority of India Limited, 3VIT Vellore  \n[yarrabothula@iitbhilai.ac.in](yarrabothula@iitbhilai.ac.in)  \nAbstract  \nExisting video benchmarks evaluate action recognition on consumer videos, egocentric recordings, or simulated industrial environments. They do not test visionlanguage models under the visual and procedural conditions of real industrial CCTV, where workers appear as distant figures amid dust, steam, low light, glare, occlusion, and overlapping activities. We introduce STEELBENCH, a diagnostic benchmark for industrial surveillance that jointly evaluates per-worker activity recognition, safety-rule reasoning, and annotation provenance. SteelBench contains 1,345 densely annotated clips, curated from 149 hours of operational plant footage and 10,024 candidate clips using temporal deduplication, class balancing, and visibility-aware stratified sampling. Each clip includes dense per-worker action labels, PPE attributes, spatial context, and safety-rule annotations.  \nBecause model-assisted annotation can shape the labels later used for model evaluation, SteelBench includes a provenance-aware audit protocol. The protocol measures label influence, evaluates sensitivity to ground-truth provenance, and reports a human reference from expert-reviewed labels. Applying this audit, we find that unaudited VLM-sourced ground truth can inflate same-family model accuracy by up to 17 percentage points. Across nine VLMs from four architectural families, the best model reaches only 42.6% action accuracy, compared with an 84.6% human benchmark. Performance also fragments across recognition, robustness, calibration, and safety reasoning. Even when models predict the correct action, 37-58% of cases still yield incorrect safety judgments, and no model passes more than 2 of 5 diagnostic checks. SteelBench shows that real industrial activity understanding requires provenance-aware and failure-modespecific evaluation rather than leaderboard accuracy alone. The dataset is publicly available on Hugging Face.  \n1 Introduction  \nVision-language models (VLMs) are routinely evaluated on action recognition, visual question answering, and scene understanding, but predominantly on data where subjects are clearly visible, centrally framed, and recorded under controlled or simulated conditions. Video benchmarks such as Kinetics [1], ActivityNet [2], and AVA [3] capture human activity at close range and high resolution, while egocentric datasets such as Ego4D [4] and Assembly101 [5] focus on detailed hand-object interactions. Many existing industrial video benchmarks use simulated or controlled environments [6], while related safety datasets often evaluate narrower tasks, such as PPE detection  \n∗ Corresponding author. Email: [yarrabothula@iitbhilai.ac.in](yarrabothula@iitbhilai.ac.in)  \nPreprint  \nfrom curated images [7–9] . None of these datasets, to our knowledge, combines dense per-worker action annotation, PPE assessment, spatial context, and safety-rule compliance from real operational industrial CCTV. Workers appear as 80–200 pixel figures at 7–10 meters. Visibility conditions arise from natural plant operations, including occlusion from worker activities and equipment placement, low light, dust from material handling, steam from cooling, and glare from molten metal. Multiple workers often perform different activities simultaneously. These conditions also make annotation difficult. When visual evidence is ambiguous, annotators may defer to VLM-generated pre-fills, creating a dependency between model-assisted label construction and model evaluation that most benchmark","cbCaius4ARZ4sZLJ","https://ap.wps.com/l/cbCaius4ARZ4sZLJ","pdf",6058208,1,40,"English","en",105,"# Abstract\n# Introduction\n## Industrial surveillance evaluation challenges\n## Dataset and annotation design\n## Provenance-aware audit protocol and findings","[{\"question\":\"What problem does SteelBench address in existing vision-language model evaluations?\",\"answer\":\"SteelBench addresses the gap between benchmark conditions and real industrial CCTV, where workers are distant and scenes include dust, steam, low light, glare, occlusion, and simultaneous activities that make reliable annotation difficult.\"},{\"question\":\"What does each SteelBench clip include for annotation?\",\"answer\":\"Each clip provides dense per-worker action labels, PPE attributes, spatial context, and safety-rule annotations, using a schema designed for multi-worker scenes.\"},{\"question\":\"How does SteelBench evaluate the reliability of model-assisted labels?\",\"answer\":\"SteelBench adds a provenance-aware audit protocol that measures label influence and sensitivity to ground-truth provenance, including a human reference from expert-reviewed labels to detect evaluation inflation.\"}]",1784198340,101,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"steelbench-evaluating-vision-language-models-in-real-world-industrial-environments","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/steelbench-evaluating-vision-language-models-in-real-world-industrial-environments/84803/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does SteelBench address in existing vision-language model evaluations?","Question",{"text":75,"@type":76},"SteelBench addresses the gap between benchmark conditions and real industrial CCTV, where workers are distant and scenes include dust, steam, low light, glare, occlusion, and simultaneous activities that make reliable annotation difficult.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What does each SteelBench clip include for annotation?",{"text":80,"@type":76},"Each clip provides dense per-worker action labels, PPE attributes, spatial context, and safety-rule annotations, using a schema designed for multi-worker scenes.",{"name":82,"@type":73,"acceptedAnswer":83},"How does SteelBench evaluate the reliability of model-assisted labels?",{"text":84,"@type":76},"SteelBench adds a provenance-aware audit protocol that measures label influence and sensitivity to ground-truth provenance, including a human reference from expert-reviewed labels to detect evaluation inflation.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,119,122,127,130,134],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":21,"slug":118},7,"Healthcare","healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":45,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":45,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":45,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":45,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]