[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82171-en":3,"doc-seo-82171-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},82171,687197207057,"Sage","https://ap-avatar.wpscdn.com/davatar_29158cc5080c5b710cf443261637dec0",8,"Research & Report","On Locality and Length Generalization in Visual Reasoning","A key feature of human vision is sequential processing using local foveated glimpses rather than a single global computation. This work tests whether local sequential vision models provide fundamental computational advantages in visual state tracking and length generalization. Experiments on simple visual tasks requiring aggregation of local information show that, like language models, vision systems may rely on global shortcut strategies and fail to generalize to longer or more complex task lengths. Recurrent policies using strictly local perception reduce these failures and improve generalization. Results indicate local attention is crucial for robust compositional generalization.","arXiv :2607 .09061v1 [ cs .CV] 10 Jul 2026  \nON LOCALITY AND LENGTH GENERALIZATION IN VISUAL REASONING  \nPulkit Madan†  Sanjay Haresh†  Reza Ebrahimi  Sunny Panchal  Apratim Bhattacharyya  Roland Memisevic   \nQualcomm AI Research∗  \n{pmadan, sanjayh}@qti .qualcomm .com  \nABSTRACT  \nA striking feature of the human visual system is that it ingests visual information through a series of local foveated glimpses, rather than a single global computation. This makes human vision distinctly different from most popular computer vision models in use today, which input images globally and in a single shot. A natural question therefore is whether local, sequential vision models may provide any fundamental computational benefits in addition to being biologically more plausible than global models. In this work, we investigate this question from the perspective of visual state tracking and length generalization. Inspired by recent studies of length generalization in language models, we study the behavior of vision models trained on simple vision tasks that require the aggregation of local information across an image. Our experiments reveal that, similar to language models, vision models can learn to exploit global shortcuts and thereby fail to generalize over task length or complexity. We also show that recurrent vision policies based on strictly local perception can mitigate these failures, thereby allowing models to generalize on these tasks. Our results show that local attention may bean essential overlooked requirement for robust compositional generalization.  \n1 INTRODUCTION  \nCurrent state-of-the-art vision models have shown human level performance on tasks such as image captioning and visual question answering (Bai et al., 2025 ; OpenAI; Anthropic) . This success is built on models which ingest an image in a single forward pass to create an encoding of the global contents of the image. E.g., transformer based models encode images in a sequence of tokens and (self-)attend to these tokens at every step of a reasoning process. This mechanism is different from the way humans process images, which is based on local glimpses connected through saccades (e.g.,(Hayhoe & Ballard, 2005)) . This raises the question of whether this type of sequential processing is purely an evolutionary artifact or if it is beneficial, or even necessary, for human-level multimodal intelligence.  \nTo shed light onto this question, we introduce a set of simple visual reasoning tasks that humans would solve by following a trajectory of local glimpses over the image. The problems involve aggregating local 2D information to derive the state of a system represented diagrammatically in the image (Fig. 1) . The complexity of these visual reasoning problems can be defined by their length, which measures the minimum number of steps required to solve them. The main challenge we consider in this work is that of extrapolation—solving longer problems without explicit training—usually referred to as length generalization. Extrapolation requires models to go beyond memorization and towards true compositional understanding, requiring strategies that are independent of the length of problem. Following a trajectory of local glimpses is an example of a strategy that is independent of problem length.  \nOur visual reasoning problems are inspired by standard tasks widely used to study length generalization in language models, such as the task of determining the parity of a binary sequence. Recent  \n∗ Qualcomm AI Research is an initiative of Qualcomm Technologies, Inc.† Equal contribution.  \nFigure 1: Overview of our work’s central question: How should visual reasoning models process spatially distributed evidence when test-time task length exceeds the training range? Global models, that process the full image in a single pass, can learn shortcuts that fail out-of-distribution. Foveated recurrent processing, in contrast, decomposes the task into repeated local observations and state upda","cbCaii8vrzoZX84p","https://ap.wps.com/l/cbCaii8vrzoZX84p","pdf",3689112,2,1,25,"English","en",105,"# Introduction\n## Visual state tracking and length generalization\n## Global vs local processing\n## Shortcut failures and out-of-distribution behavior","[{\"question\":\"What problem does the paper investigate in visual reasoning models?\",\"answer\":\"The paper investigates how vision models perform visual state tracking and whether they can generalize when test-time task length exceeds the training range (length generalization).\"},{\"question\":\"Why do the authors expect vision models to fail on longer tasks?\",\"answer\":\"The authors find that vision models can learn global shortcut solutions that do not capture the compositional structure required for out-of-distribution generalization across task lengths or complexity.\"},{\"question\":\"How can models mitigate length generalization failures?\",\"answer\":\"The paper shows that recurrent vision policies based on strictly local perception—using sequential local glimpses with state updates—can mitigate shortcut reliance and allow better generalization to longer visual sequences.\"}]",1784178577,63,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"on-locality-and-length-generalization-in-visual-reasoning","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/on-locality-and-length-generalization-in-visual-reasoning/82171/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-21","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the paper investigate in visual reasoning models?","Question",{"text":75,"@type":76},"The paper investigates how vision models perform visual state tracking and whether they can generalize when test-time task length exceeds the training range (length generalization).","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"Why do the authors expect vision models to fail on longer tasks?",{"text":80,"@type":76},"The authors find that vision models can learn global shortcut solutions that do not capture the compositional structure required for out-of-distribution generalization across task lengths or complexity.",{"name":82,"@type":73,"acceptedAnswer":83},"How can models mitigate length generalization failures?",{"text":84,"@type":76},"The paper shows that recurrent vision policies based on strictly local perception—using sequential local glimpses with state updates—can mitigate shortcut reliance and allow better generalization to longer visual sequences.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]