[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85072-en":3,"doc-seo-85072-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},85072,1099514067438,"River Wang","https://ap-avatar.wpscdn.com/avatar/100002539ee87300030?x-image-process=image/resize,m_fixed,w_180,h_180&k=1780474512215547542",8,"Research & Report","Texture Representations in Deep Vision Models: Comparing CNNs, Vision Transformers, and Human Perception","Texture perception research examines how visual texture representations relate to model architectures and to human psychophysics. This study generates textures of varying complexity using three algorithms applied to the same source images, then quantifies information in internal representations of CNNs and three Vision Transformers with a rank-based statistic. Representation alignment emerges across ViTs, not between ViTs and CNNs. For texture recognition, human performance is better predicted from ViT representations, suggesting ViTs capture human-like processing of texture patterns more faithfully than CNNs.","arXiv :2607 .0832 1v 1 [ cs .CV] 9 Jul 2026  \nTexture Representations in Deep Vision Models: Comparing CNNs, Vision Transformers, and Human  \nPerception  \nLudovica de Paolis 1,∗, Marco Baroni2,3 , Alessandro Laio4 , Eugenio Piasini 1  \n1Department of Neuroscience, International School for Advanced Studies (SISSA), Trieste, Italy  \n2Department of Language and Translation Sciences, Pompeu Fabra University, Barcelona, Spain  \n3ICREA, Barcelona, Spain  \n4Department of Data Science, International School for Advanced Studies (SISSA), Trieste, Italy  \nAbstract  \nIn computational vision science, Convolutional Neural Networks (CNNs) have emerged as a popular model of biological vision because of the alignment they can exhibit with neural and behavioral data in humans and animals. However, it remains unclear to what extent this alignment persists for visual tasks that extend beyond the canonical object-recognition paradigm based on well-defined semantic content. In this study, we diverge from the common object-centric view by focusing on another aspect of vision: texture perception. We consider textures of different complexity generated with three different algorithms from the same source images. Using a rank-based statistic, we quantify the information encoded in the internal representations of a CNN and three Vision Transformers (ViTs), and we compare the similarity of these representations to those inferred from human psychophysics data. We find that the representation of textures is aligned in different ViTs, but not between the ViTs and the CNN; that ViTs form similar representations for textures of different complexity; that human performance in recognizing textures can be better predicted from ViTs representations rather than CNN representations. Taken together, these results suggest that ViTs may capture more faithfully than CNNs how texture patterns are visually processed by humans, and that the representations of texture stimuli in computational models may be driven by the network architecture.  \n1 Introduction  \nConvolutional neural networks (CNNs) are ubiquitous in computer vision and in vision neuroscience, where they are commonly employed as models of the ventral visual stream and of the biological mechanisms underlying object recognition [Lindsay, 2021, Kriegeskorte, 2015]) . A pivotal moment in the history of CNNs was the release of AlexNet [Krizhevsky et al., 2012], which demonstrated significantly higher accuracy than other vision models in an object classification task connected to the ImageNet database [Deng et al., 2009] . This success reinforced the centrality of the object recognition paradigm in machine learning and computational vision neuroscience, which led to numerous advances in the following years. However, in real life scenarios image perception and representation is much more diversified than just recognizing objects. Consequently, parallel research directions in vision science have been focusing on aspects of vision that cannot be addressed within the framework of object recognition. One of these tasks is texture perception: visual textures are an acknowledged method to study complex real world vision phenomena [Victor et al., 2017] .  \n∗ Correspondence to Ludovica de Paolis: [ldepaoli@sissa.it](ldepaoli@sissa.it)  \nPreprint.  \nAn early formal definition of textures was provided by Julesz [Julesz, 1962, 1981], stating that textures can be seen as families of patterns that share certain local regularities. Julesz also conjectured that textures could be modeled via low-order statistics [Julesz et al., 1973] . This conjecture became very influential in the development of computational models of textures, which later resulted in the release of texture synthesis algorithms that operationalized the conjecture by means of low-level correlations. The most prominent models are those by Portilla & Simoncelli [Portilla and Simoncelli, 2000a], Victor & Conte [Victor and Conte, 2012], and Gatys et al.,[Gatys et al., ","cbCaiqgbIbuNvJLp","https://ap.wps.com/l/cbCaiqgbIbuNvJLp","pdf",5833803,3,1,20,"English","en",105,"# Abstract\n# 1 Introduction","[{\"question\":\"What question does the study address about deep vision models?\",\"answer\":\"To what extent representational alignment with biological vision persists when tasks move beyond object recognition, focusing specifically on texture perception.\"},{\"question\":\"How are the textures and model comparisons constructed?\",\"answer\":\"Textures of different complexity are generated from the same source images using three algorithms, and internal representations are compared across CNNs and three Vision Transformers using a rank-based statistic, then related to human psychophysics data.\"},{\"question\":\"What do the results show about CNNs versus Vision Transformers for textures?\",\"answer\":\"ViT representations are aligned across different ViTs, but not between ViTs and the CNN; human texture recognition performance is better predicted from ViT representations than from CNN representations.\"}]",1784200827,50,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"texture-representations-in-deep-vision-models-comparing-cnns-vision-transformers-and-human-perception","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/texture-representations-in-deep-vision-models-comparing-cnns-vision-transformers-and-human-perception/85072/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What question does the study address about deep vision models?","Question",{"text":75,"@type":76},"To what extent representational alignment with biological vision persists when tasks move beyond object recognition, focusing specifically on texture perception.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How are the textures and model comparisons constructed?",{"text":80,"@type":76},"Textures of different complexity are generated from the same source images using three algorithms, and internal representations are compared across CNNs and three Vision Transformers using a rank-based statistic, then related to human psychophysics data.",{"name":82,"@type":73,"acceptedAnswer":83},"What do the results show about CNNs versus Vision Transformers for textures?",{"text":84,"@type":76},"ViT representations are aligned across different ViTs, but not between ViTs and the CNN; human texture recognition performance is better predicted from ViT representations than from CNN representations.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,114,119,122,126,129,133],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":29,"slug":113},6,"Technology","technology",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":22,"slug":125},9,"Religion & Spirituality","religion-spirituality",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":127,"show_sort_weight":22,"slug":128},"World Cup","world-cup",{"id":130,"doc_module":4,"doc_module_name":46,"category_name":131,"show_sort_weight":130,"slug":132},10,"Lifestyle","lifestyle",{"id":134,"doc_module":4,"doc_module_name":46,"category_name":135,"show_sort_weight":106,"slug":136},19,"General","general"]