[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82450-en":3,"doc-seo-82450-105":30,"detail-sidebar-cat-0-en-105":83},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},82450,7971461741311,"Ophelia","https://ap-avatar.wpscdn.com/avatar/74000253aff267980c6?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779345379180704826",8,"Research & Report","Evolution of Accuracy and Visual-Cognitive Errors in a Decade of Vision-Language AI Models","Vision–language models (VLMs) have advanced rapidly in visual reasoning, yet many evaluations rely on simple datasets such as MSCOCO, use limited non-curated human descriptions, and rarely analyze the model’s error types. This work introduces the Complex Social Behavior (CSB) dataset with 100 images of complex social interactions. Results across 2017–2025 show strong gains in scene-description accuracy, with MLLMs nearly eliminating most error categories except spatial dependence, and revealing which errors most affect caption quality.","arXiv :2607 .09654v1 [ cs .CV] 10 Jul 2026  \nEvolution of Accuracy and Visual-Cognitive Errors in a Decade of Vision-Language AI Models  \nShravan Murlidarana,1,∗, Miguel P. Ecksteina,b,c,1  \na Psychological & Brain Sciences, University of California, Santa Barbara, , Santa  \nBarbara, 93106, California, USA  \nb Department of Computer Science, University of California, Santa Barbara, , Santa  \nBarbara, 93106, California, USA  \nc Department of Electrical and Computer Engineering, University of California, Santa  \nBarbara, , Santa Barbara, 93106, California, USA  \nAbstract  \nVision–language models (VLMs) have made remarkable progress in visual reasoning during the last decade. Most evaluations have used simple scenes (MSCOCO) that do not showcase complex human interactions or behaviors, only a handful of non-curated human descriptions as a benchmark, and have not focused on understanding the model’s error types. Here, we introduce the Complex Social Behavior (CSB) dataset, containing 100 images depicting complex social interactions/behaviors. We analyze the progression of scene descriptions over a decade (2017-2025) of VLMs (four pre-Multimodal Large Language Models, MLLMs, and five MLLMs) . We evaluate the accuracy of the models and 20 human descriptions relative to a gold standard on the CSB dataset and on a sample from MS-COCO. We analyzed five visual-cognitive error types: object detection, recognition, hallucination, scene understanding, and spatial dependence. The CSB dataset showed a more pronounced improvement than MS-COCO in scene description accuracy, with pre-MLLMs achieving much lower accuracy than the bottom-ranked human descriptions and MLLMs attaining accuracies similar to the top-ranked human descriptions. We show that MLLMs have eliminated the gap in scene description accuracy between simpler MS-COCO scenes and scenes depicting complex behaviors (CSB) . MLLMs have almost eliminated all error types in our tested datasets, except for occasionally relying on different image regions for scene descriptions than humans do (spatial dependence error) . We also show that detection, recognition, and hallucination errors have the highest impact on scene description accuracy. Together, our findings provide a more thorough evaluation of how visual language models have advanced over the last decade.  \n∗ Corresponding author.  \nEmail addresses: [smurlidaran@ucsb.edu](smurlidaran@ucsb.edu) (Shravan Murlidaran), [eckstein@psych.ucsb.edu](eckstein@psych.ucsb.edu) (Miguel P. Eckstein)  \n1 These authors contributed equally to this work.  \nKeywords: Multi-Modal Large Language Model (MLLMs), Image Understanding, Captioning, Description Error Analysis,, Human and MLLM performance  \n1. Introduction  \nFigure 1: (a) The reported performance of DNNs over the years in object recognition compared to human performance. The performance of the then state-of-the-art models exceeded human performance. (b) Reported performance of Multi-Modal Large-Language Models (MLLMs) in the scene description task on the MS-COCO dataset. The model’s performance is comparable to human performance. (c) Model descriptions for scenes with and without social interaction. The image on the left depicts a simple scene in which the models’ descriptions are similar to those observed in humans. The image on the right depicts a scene of complex social interactions among humans. We can clearly see that the human description captures the social interaction in the scene. In contrast, the pre-MLLM’s description does not capture the social interactions, while MLLMs do.  \nArtificial Neural Networks have advanced significantly in visual reasoning since the development of deep Convolutional Neural Networks (CNNs) over a  \ndecade ago [1] . Figure 1a shows a glimpse of the early years of deep CNNs, with models that could outperform humans in tasks like object recognition. Deep Recurrent Neural Networks [2], which process and generate sequential data like text, combined with deep CNNs","cbCaiuDcDfw16WG3","https://ap.wps.com/l/cbCaiuDcDfw16WG3","pdf",25046559,3,1,25,"English","en",105,"# Abstract\n# Introduction\n## Motivation: limitations of standard benchmarks\n## Model progression across years and datasets","[{\"question\":\"Which visual-cognitive error types most impact scene-description accuracy?\",\"answer\":\"Detection, recognition, and hallucination errors have the highest impact on scene-description accuracy, while other categories improve substantially across tested models.\"}]",1784180445,63,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":78,"head_meta":80,"extra_data":82,"updated_unix":28},"evolution-of-accuracy-and-visual-cognitive-errors-in-a-decade-of-vision-language-ai-models","",{"@graph":36,"@context":77},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/evolution-of-accuracy-and-visual-cognitive-errors-in-a-decade-of-vision-language-ai-models/82450/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-21","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71],{"name":72,"@type":73,"acceptedAnswer":74},"Which visual-cognitive error types most impact scene-description accuracy?","Question",{"text":75,"@type":76},"Detection, recognition, and hallucination errors have the highest impact on scene-description accuracy, while other categories improve substantially across tested models.","Answer","https://schema.org",{"og:url":51,"og:type":79,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":81,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":84},[85,89,93,97,102,107,112,115,120,123,127],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":86,"show_sort_weight":87,"slug":88},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":90,"show_sort_weight":91,"slug":92},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Exam",70,"exam",{"id":98,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},5,"Comic",60,"comic",{"id":103,"doc_module":4,"doc_module_name":46,"category_name":104,"show_sort_weight":105,"slug":106},6,"Technology",50,"technology",{"id":108,"doc_module":4,"doc_module_name":46,"category_name":109,"show_sort_weight":110,"slug":111},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":113,"slug":114},30,"research-report",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},9,"Religion & Spirituality",20,"religion-spirituality",{"id":118,"doc_module":4,"doc_module_name":46,"category_name":121,"show_sort_weight":118,"slug":122},"World Cup","world-cup",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":124,"slug":126},10,"Lifestyle","lifestyle",{"id":128,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":98,"slug":130},19,"General","general"]