[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-117402-en":3,"doc-seo-117402-105":30,"detail-sidebar-cat-0-en-105":83},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},117402,7971461740909,"Levi","https://ap-avatar.wpscdn.com/davatar_155a257f0dc6eb9ab79c44ca47cae57d",8,"Research & Report","A comparison between humans and AI at recognizing objects in unusual poses","Deep learning continues narrowing the gap with human vision on object recognition benchmarks, yet robustness under challenging conditions remains uncertain. This study examines human versus AI performance on images where objects appear in unusual poses. Results show humans outperform state-of-the-art vision deep networks and most large vision-language models, with Gemini as the main exception. When viewing time is restricted, human accuracy declines to deep-network levels, indicating extra mental processing is required.","This is an electronic reprint of the original article.  \nThis reprint may differ from the original in pagination and typographic detail.  \nOllikka, Netta; Abbas, Amro; Perin, Andrea; Kilpeläinen, Markku; Deny, Stéphane  \nA comparison between humans and AI at recognizing objects in unusual poses  \nPublished in:  \nTransactions on Machine Learning Research  \nPublished: 01/01/2025  \nDocument Version  \nPublisher's PDF, also known as Version of record  \nPublished under the following license:  \nCC BY  \nPlease cite the original version:  \nOllikka, N. , Abbas, A. , Perin, A. , Kilpeläinen, M. , & Deny, S. (2025) . A comparison between humans and AI at recognizing objects in unusual poses. Transactions on Machine Learning Research, 2025(January), 1-32. [https://arxiv.org/abs/2402.03973](https://arxiv.org/abs/2402.03973)  \nThis material is protected by copyright and other intellectual property rights, and duplication or sale of all or part of any of the repository collections is not permitted, except that material may be duplicated by you foryour research use or educational purposes in electronic or print form. You must obtain permission for anyother use. Electronic or print copies may not be offered, whether for sale or otherwise to anyone who is not an authorised user.  \nA comparison between humans and AI at recognizing objects in unusual poses  \nNetta Ollikka  \nDepartment of Neuroscience and Biomedical Engineering Aalto University, Espoo, Finland  \nAmro Abbas  \nThe African Institute for Mathematical Sciences, Mbour-Thies, Senegal  \nAndrea Perin  \nDepartment of Computer Science Aalto University, Espoo, Finland  \nMarkku Kilpeläinen†  \nDepartment of Psychology and Logopedics University of Helsinki, Finland  \nStéphane Deny†  \nDepartment of Neuroscience and Biomedical Engineering Department of Computer Science  \nAalto University, Espoo, Finland  \n[netta. ollikka@aalto.fi](netta. ollikka@aalto.fi)  \n[afagiri@aimsammi. org](afagiri@aimsammi. org)  \n[xinandre@gmail. com](xinandre@gmail. com)  \n[markku.kilpelainen@helsinki.fi](markku.kilpelainen@helsinki.fi)  \n[stephane.deny.pro@gmail. com](stephane.deny.pro@gmail. com)  \nReviewed on OpenReview: [https: // openreview. net/ forum? id= XXXX](https: // openreview. net/ forum? id= XXXX)  \nAbstract  \nDeep learning is closing the gap with human vision on several object recognition benchmarks.  \nHere we investigate this gap in the context of challenging images where objects are seen in unusual poses. We find that humans excel at recognizing objects in such poses. In contrast, state-of-the-art deep networks for vision (EfficientNet, SWAG, ViT, SWIN, BEiT, ConvNext) and state-of-the-art large vision-language models (Claude 3.5, Gemini 1.5, GPT-4, SigLIP) are systematically brittle on unusual poses, with the exception of Gemini showing excellent robustness to that condition. As we limit image exposure time, human performance degrades to the level of deep networks, suggesting that additional mental processes (requiring additional time) are necessary to identify objects in unusual poses. An analysis of error patterns of humans vs. networks reveals that even time-limited humans are dissimilar to feed-forward deep networks. In conclusion, our comparison reveals that humans are overall more robust than deep networks and that they rely on different mechanisms for recognizing objects in unusual poses. Understanding the nature of the mental processes taking place during extra viewing time may be key to reproduce the robustness of human vision in silico.  \nAll code and data is available at [https://github.com/BRAIN-Aalto/unusual_poses](https://github.com/BRAIN-Aalto/unusual_poses).  \n1 Introduction  \nIn an era marked by the rapid advancement of deep learning for computer vision, a natural question arises: Can machines meet or even exceed the capabilities of the human visual system? Numerous recent studies have shown that deep networks outperform humans on well-known object recognition benchmarks (e.g., ImageNet:","cbCail0gtrcFMryZ","https://ap.wps.com/l/cbCail0gtrcFMryZ","pdf",2831292,1,33,"English","en",105,"# Abstract\n# Introduction\n## Motivation: deep learning vs human vision\n## Gap: robustness to global structural changes\n## Study approach and high-level findings","[{\"question\":\"Why does limiting viewing time affect human performance?\",\"answer\":\"As viewing time is reduced, human performance degrades toward the level of deep networks, suggesting additional mental processes are needed to analyze unusual poses.\"}]","A comparison between humans and AI at recognizing objects in unusual poses | PDF",1785675680,83,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":78,"head_meta":80,"extra_data":82,"updated_unix":28},"a-comparison-between-humans-and-ai-at-recognizing-objects-in-unusual-poses","",{"@graph":36,"@context":77},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/a-comparison-between-humans-and-ai-at-recognizing-objects-in-unusual-poses/117402/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-02",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71],{"name":72,"@type":73,"acceptedAnswer":74},"Why does limiting viewing time affect human performance?","Question",{"text":75,"@type":76},"As viewing time is reduced, human performance degrades toward the level of deep networks, suggesting additional mental processes are needed to analyze unusual poses.","Answer","https://schema.org",{"og:url":52,"og:type":79,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":81,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":84},[85,89,93,97,102,107,112,115,120,123,127],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":86,"show_sort_weight":87,"slug":88},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":90,"show_sort_weight":91,"slug":92},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Exam",70,"exam",{"id":98,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},5,"Comic",60,"comic",{"id":103,"doc_module":4,"doc_module_name":46,"category_name":104,"show_sort_weight":105,"slug":106},6,"Technology",50,"technology",{"id":108,"doc_module":4,"doc_module_name":46,"category_name":109,"show_sort_weight":110,"slug":111},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":113,"slug":114},30,"research-report",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},9,"Religion & Spirituality",20,"religion-spirituality",{"id":118,"doc_module":4,"doc_module_name":46,"category_name":121,"show_sort_weight":118,"slug":122},"World Cup","world-cup",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":124,"slug":126},10,"Lifestyle","lifestyle",{"id":128,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":98,"slug":130},19,"General","general"]