[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83487-en":3,"doc-seo-83487-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83487,687197100911,"Himbo","https://ap-avatar.wpscdn.com/avatar/a000239b6f1da00475?x-image-process=image/resize,m_fixed,w_180,h_180&k=1782698725881665579",8,"Research & Report","Information-Regularized Attention for Visual-Centric Reasoning","Vision–language models remain unreliable despite strong performance, suffering from object hallucination, weak visual grounding, and catastrophic forgetting after full-parameter instruction tuning. The work attributes these failures to missing explicit control of visual representation learning under the standard next-token prediction objective, leading to passive embedding optimization and spurious signal injection. It introduces Information-Regularized Attention (IRA), a stochastic attention method that regulates visual information injected into intermediate transformer layers. IRA converts representation uncertainty into local independent noise, yielding smoother embedding curvature trajectories and reduced attention-sink behavior, indicating more stable visual signal transformations. The results frame stochastic attention as a key driver of representation learning for more dependable VLMs.","arXiv :2607 .00434v 1 [ cs .CV] 1 Jul 2026  \nInformation-Regularized Attention for Visual-Centric Reasoning  \nGuohao Sun 1 ,2 ,∗ , Xiaofang Wang 1 , Yash Patel 1 , Mengchen Liu 1 , Zhiqiang Tao2 , Praveen Krishnan 1  \n1 FAIR at Meta, 2 Rochester Institute of Technology  \n∗Work done at Meta  \nVision–language models (VLMs) have become a paradigm for multimodal learning, yet remain unstable due to object hallucination, weak visual grounding, and catastrophic forgetting after fullparameter instruction tuning. We claim these failures result from a lack of explicit control over visual representation learning during the standard next-token prediction objective. As a result, visual embeddings thus become passively optimized and prone to injecting redundant or spurious signals. To counter this, we introduce Information-Regularized Attention (IRA), a stochastic attention mechanism that explicitly regulates the amount of visual information injected into the hidden states of intermediate transformer layers. This local reparameterization translates uncertainty about visual representations into local noise that is independent across data points. Beyond evaluating model performance, we also quantify embedding properties, where IRA produces smoother curvature trajectories and suppresses attention-sink across all layers, indicating a more stable transformation of the visual signal. Our results suggest that stochastic attention is not merely a regularizer but a key contributor to representation learning in a generative architecture, offering a new direction for building more reliable VLMs.  \nDate: July 2, 2026  \nCorrespondence: Praveen Krishnan at [pkrishnan@meta.com](pkrishnan@meta.com)   \n1 Introduction  \nVision-language models (VLMs) have emerged as a general-purpose framework for multimodal understanding, achieving strong performance across tasks such as visual question answering, image captioning, and multimodal dialogue Lu et al. (2019); Alayrac et al. (2022); Li et al. (2023a); Gan et al. (2022) . Despite this progress, modern VLMs remain limited by reliability issues, including object hallucination and unreliable grounding, where generated content is not supported by the visual input Li et al. (2023c); Rohrbach et al. (2018) . These failures suggest that standard next-text-token prediction objectives do not sufficiently align vision and language information in latent space.  \nCurrent VLM training paradigms are largely data-centric, relying on increasingly diverse forms of posttraining supervision, including visual instruction tuning Liu et al. (2023); Sun et al. (2024b), preference optimization Ouyang et al. (2022); Rafailov et al. (2023); Peng et al. (2025); Sun et al. (2024a, 2025b), and policy optimization Shao et al. (2024); Schulman et al. (2017); Sun et al. (2025a) . While effective, these approaches primarily improve model behavior by expanding the supervision signals, rather than directly regularizing the feature representations. Under the standard next-token prediction objective, all the embeddings are optimized only indirectly through language supervision. As a result, task-irrelevant or noisy visual signals can propagate through attention layers and interfere with cross-modal reasoning.  \nThis issue is reflected in recent observations of attention sinks Gu et al. (2025); Kang et al. (2025); de Llano et al. (2026); Barbero et al. (2025) and spike values in attention heads Sun et al. (2024c, 2026b); Xiao et al. (2023a), where attention collapses onto semantically uninformative tokens and produces noisy crossmodal interactions Rohrbach et al. (2018); Mahajan et al. (2025); Jiang et al. (2025) . Our study in Fig. 1 provides further evidence: the pretrained VLM fails to attend to the relevant visual regions, whereas standard supervised instructional fine-tuning improves alignment only partially and still exhibits biased, noisy attention. These findings suggest that improving VLM reliability requires moving beyond data-centric post","cbCaif9F6mco1aaT","https://ap.wps.com/l/cbCaif9F6mco1aaT","pdf",4530152,3,1,20,"English","en",105,"# Introduction\n## Reliability issues in modern VLMs\n## Data-centric training limitations\n## Attention sinks and noisy cross-modal interactions\n## Information-Regularized Attention (IRA) approach","[{\"question\":\"Why do vision-language models still show reliability problems like object hallucination and weak grounding?\",\"answer\":\"Because standard next-token prediction does not explicitly control visual representation learning in latent space, allowing passively optimized embeddings to inject redundant or spurious visual signals.\"},{\"question\":\"What is the core idea of Information-Regularized Attention (IRA)?\",\"answer\":\"IRA uses stochastic attention to regulate how much visual information is injected into hidden states of intermediate transformer layers, turning uncertainty into local noise independent across data points.\"},{\"question\":\"How does IRA change embedding and attention behavior compared with conventional training?\",\"answer\":\"IRA produces smoother curvature trajectories for embeddings and suppresses attention-sink behavior across transformer layers, indicating a more stable transformation of the visual signal.\"}]",1784188352,50,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"information-regularized-attention-for-visual-centric-reasoning","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/information-regularized-attention-for-visual-centric-reasoning/83487/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-26","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why do vision-language models still show reliability problems like object hallucination and weak grounding?","Question",{"text":75,"@type":76},"Because standard next-token prediction does not explicitly control visual representation learning in latent space, allowing passively optimized embeddings to inject redundant or spurious visual signals.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What is the core idea of Information-Regularized Attention (IRA)?",{"text":80,"@type":76},"IRA uses stochastic attention to regulate how much visual information is injected into hidden states of intermediate transformer layers, turning uncertainty into local noise independent across data points.",{"name":82,"@type":73,"acceptedAnswer":83},"How does IRA change embedding and attention behavior compared with conventional training?",{"text":84,"@type":76},"IRA produces smoother curvature trajectories for embeddings and suppresses attention-sink behavior across transformer layers, indicating a more stable transformation of the visual signal.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,114,119,122,126,129,133],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":29,"slug":113},6,"Technology","technology",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":22,"slug":125},9,"Religion & Spirituality","religion-spirituality",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":127,"show_sort_weight":22,"slug":128},"World Cup","world-cup",{"id":130,"doc_module":4,"doc_module_name":46,"category_name":131,"show_sort_weight":130,"slug":132},10,"Lifestyle","lifestyle",{"id":134,"doc_module":4,"doc_module_name":46,"category_name":135,"show_sort_weight":106,"slug":136},19,"General","general"]