[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86508-en":3,"doc-seo-86508-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86508,1099514067415,"Rowan","https://ap-avatar.wpscdn.com/avatar/100002539d78ffe74a7?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779092875211072502",8,"Research & Report","Artificial Foveated Perception for Mitigating Shortcut Learning in Robotic Foundation Models","Robotic foundation models advance multitask capability, cross-embodiment transfer, and language-conditioned control, yet robust real-world deployment remains challenging because policies can confuse causally relevant visual structure with spurious scene-level correlations. This failure is characterized as shortcut learning: exploiting predictive but non-causal correlations in training rather than task-relevant evidence. The paper introduces Artificial Foveated Perception (AFP), a lightweight, policy-agnostic module that predicts task-conditioned masks for grounding regions, uses them as auxiliary signals during fine-tuning, and leaves control unchanged afterward.","arXiv :2607 . 10655v1 [ cs .RO] 12 Jul 2026  \nArtificial Foveated Perception for Mitigating Shortcut Learning in Robotic Foundation Models  \nXiatao Sun 1†, Yuan Zhuang2 , Mateo Sanchez Lopez Negrete 1 , Matei-Victor Coldea 1 , Chen Liang 1 , Haoyang Zhang3 ,5 , Che Liu4 ,5 , Ziyao Zeng 1 , Shawn Li5 , Qian Wang 1 , Fei Miao2 , Daniel Rakita 1  \n1Yale University 2University of Connecticut  \n3Peking University 4Imperial College London 5Digients  \n†Corresponding author: [xiatao.sun@yale.edu](xiatao.sun@yale.edu)  \nWhy Direct Fine-tuning Fails when OOD The model relies on spurious correlations (shortcuts) and is misled by distractors.  \nHow AFP Helps  \nAFP focuses perception on task-relevant regions.  \nResult  \nBetter generalization and robustness to OOD distractors.  \nFigure 1: Artificial Foveated Perception (AFP) mitigates shortcut learning in robotic foundation models by forcing or guiding their perception towards task-relevant regions while suppressing distractors. Policies from direct fine-tuning often overfit to spurious correlations and fails under out-of-distribution (OOD) perturbations, whereas AFP improves visual grounding and enables more robust policy generalization across environmental variations.  \nAbstract: Robotic foundation models have recently made substantial progress in multi-task capability, cross-embodiment transfer, and language-conditioned control. Yet deploying these models robustly across diverse real-world settings remains difficult, in part because policies often fail to distinguish between causally relevant visual structure and spurious scene-level correlations. We identify this failure mode as shortcut learning: the tendency of a model to exploit predictive but non-causal correlations in the training distribution rather than the task-relevant visual evidence that determines successful action. Although shortcut learning has been extensively studied in computer vision and broader machine learning, its role in robotic foundation models remains comparatively underexplored. In this paper, we propose Artificial Foveated Perception (AFP), a lightweight, policy-agnostic module that takes the same vision and language inputs as existing Vision-LanguageAction and World Action Model pipelines and predicts task-conditioned masks over relevant objects, the robot, and other action-critical regions. We use these masks primarily as an auxiliary grounding signal during fine-tuning, aligning the policy’s visual attention with task-relevant regions while leaving the core policy architecture unchanged. Once fine-tuning is complete, the policy executes on the original observation stream without requiring AFP in the control loop. We evaluate AFP across state-of-the-art robotic foundation models and show that foveated perception reduces fine-tuning time, suppresses overfitting, and improves generalization under environmental perturbations. Through ablations over mask quality and grounding loss design, we further show that these gains arise from directing policy learning toward task-relevant visual evidence. These results suggest that task-conditioned  \nfoveated perception is a practical mechanism for making robotic foundation models more robust, data-efficient, and scalable.  \nKeywords: Imitation Learning, Shortcut Learning, Generalizability  \n1 Introduction  \nRobotic foundation models have made substantial progress in recent years. By leveraging representations from pretrained Vision-Language Models (VLMs) and World Models (WMs), state-of-the-art systems now demonstrate increasingly strong multi-task capability, cross-embodiment transfer, and language-conditioned control [1, 2, 3] . Despite this progress, robust deployment still often depends on task-specific adaptation. Unlike foundation models in other domains, such as Large Language Models (LLMs) with strong zero-shot inference [4], Vision-Language-Action (VLA) models and World Action Models (WAMs) typically require fine-tuning before reliable deployment in new robotic s","cbCaijcsHBn74DyJ","https://ap.wps.com/l/cbCaijcsHBn74DyJ","pdf",5156500,4,1,16,"English","en",105,"# Introduction\n## Motivation: Why Direct Fine-tuning Fails OOD\n## Proposed Method: Artificial Foveated Perception (AFP)\n## Results and Evaluation","[{\"question\":\"What is shortcut learning in robotic foundation models?\",\"answer\":\"Shortcut learning is the tendency of a model to exploit predictive but non-causal correlations from the training distribution instead of relying on task-relevant visual evidence that determines successful action.\"},{\"question\":\"How does Artificial Foveated Perception (AFP) mitigate shortcut learning?\",\"answer\":\"AFP predicts task-conditioned masks over objects, the robot, and other action-critical regions, then uses these masks as an auxiliary grounding signal during fine-tuning to align the policy’s visual attention with task-relevant evidence.\"},{\"question\":\"Does AFP need to run during the policy’s control loop after fine-tuning?\",\"answer\":\"No. After fine-tuning, the policy executes on the original observation stream without requiring AFP in the control loop.\"}]",1784212278,40,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"artificial-foveated-perception-for-mitigating-shortcut-learning-in-robotic-foundation-models","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/artificial-foveated-perception-for-mitigating-shortcut-learning-in-robotic-foundation-models/86508/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What is shortcut learning in robotic foundation models?","Question",{"text":75,"@type":76},"Shortcut learning is the tendency of a model to exploit predictive but non-causal correlations from the training distribution instead of relying on task-relevant visual evidence that determines successful action.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does Artificial Foveated Perception (AFP) mitigate shortcut learning?",{"text":80,"@type":76},"AFP predicts task-conditioned masks over objects, the robot, and other action-critical regions, then uses these masks as an auxiliary grounding signal during fine-tuning to align the policy’s visual attention with task-relevant evidence.",{"name":82,"@type":73,"acceptedAnswer":83},"Does AFP need to run during the policy’s control loop after fine-tuning?",{"text":84,"@type":76},"No. After fine-tuning, the policy executes on the original observation stream without requiring AFP in the control loop.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,119,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":29,"slug":118},7,"Healthcare","healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]