[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84466-en":3,"doc-seo-84466-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},84466,687197100911,"Himbo","https://ap-avatar.wpscdn.com/avatar/a000239b6f1da00475?x-image-process=image/resize,m_fixed,w_180,h_180&k=1782698725881665579",8,"Research & Report","Context-Dependent Affordance Computation in Vision-Language Models","Vision-language models exhibit context-dependent affordance computation, evidenced by systematic context priming across agentic personas in Qwen3-VL-30B-A3B using 3,213 COCO-2017 scene-context pairs and a cross-model replication in LLaVA-1.5-13B. Results show massive affordance drift: lexical Jaccard mean similarity drops to 0.095 (95% CI 0.092–0.097), with >90% of lexical scene descriptions context-dependent, while semantic cosine similarity yields 58.5% context dependence. Stochastic baselines confirm the drift reflects genuine context effects, not generation noise. Stable latent factors emerge via Tucker decomposition, motivating robotics research toward dynamic, query-dependent ontological projection rather than static world modeling.","arXiv :2603 .044 19v2 [ cs .CL] 11 Jul 2026  \nDISSENSUS WORKING PAPER SERIES  \nDAI-2518  \nContext-Dependent Affordance Computation in Vision-Language Models  \nMurad Farzulla 1,2,*  \n1Dissensus, London, UK 2 King’s College London, London, UK  \n* [Correspondence: murad@dissensus.ai](Correspondence: murad@dissensus.ai) ORCID: 0009-0002-7164-8704  \nJanuary 2026  \nAbstract  \nWe characterize the phenomenon of context-dependent affordance computation in vision-language models (VLMs) . Our primary study uses Qwen3-VL-30B-A3B (n = 3 ,213 scene-context pairs from COCO-2017) subject to systematic context priming across 7 agentic personas, with a crossmodel replication on LLaVA-1.5-13B. We demonstrate massive affordance drift: in the Qwen3-VL data, mean Jaccard similarity between context conditions is 0.095 (95% CI [0.092, 0.097] across N = 479 images; 9 ,244 prime pairs; p \u003C 0.0001), indicating that > 90% of lexical scene description is context-dependent; the LLaVA replication reproduces the effect (mean J = 0. 160, 84% context-dependent) . Sentence-level cosine similarity confirms substantial drift at the semantic level (mean = 0.415, 58.5% context-dependent) . Stochastic baseline experiments (2,384 inference runs across 4 temperatures and 5 seeds) confirm this drift reflects genuine context effects rather than generation noise: within-prime variance is substantially lower than cross-prime variance across all conditions. Tucker decomposition with bootstrap stability analysis (n = 1 ,000 resamples) reveals stable orthogonal latent factors: a “Culinary Manifold” isolated to chef contexts and an “Access Axis”  \nspanning child-mobility contrasts. These findings establish that VLMs compute affordances in a substantially context-dependent manner—with the difference between lexical (90%) and semantic (58.5%) measures reflecting that surface vocabulary changes more than underlying meaning under context shifts—and suggest a direction for robotics research: dynamic, query-dependent ontological projection (JIT Ontology) rather than static world modeling. We do not claim to establish processing order or architectural primacy; such claims require internal representational analysis beyond output behavior.  \nKeywords: Vision-language models, Affordances, Context-dependent processing, Scene understanding, Functional semantics, Robotics  \nAcknowledgments  \nThe author thanks Claude (Anthropic) for assistance with analytical framework development, tensor decomposition analysis, and technical writing. The author also thanks the Visual Genome project for providing human affordance annotations for baseline comparison. This paper is part of the Adversarial Systems Research program at Dissensus and the Adversarial Systems & Complexity Research Initiative (ASCRI) . All errors, omissions, and interpretive limitations remain the author’s responsibility.  \nData & Code Availability. Analysis code and data are available at [https://github](https://github.com/studiofar)[.](https://github.com/studiofar)[com/studiofar](https://github.com/studiofar)[ ](https://github.com/studiofar)[zulla/semantic-vision](zulla/semantic-vision.)[.](zulla/semantic-vision.)  \n1 Introduction  \nContemporary computer vision operates on an implicit assumption: visual processing begins with geometric feature extraction from pixel-level data, proceeds through hierarchical abstraction to object recognition, and only subsequently—if at all—computes functional or semantic properties. This pipeline reflects a Cartesian conception of space as a neutral container:  \nPstd : I → Fpixel → Ffeature → O object →Ccontext → Aaffordance (1)  \nThis ordering is not theoretically neutral. It embeds assumptions about perception that have been challenged by ecological psychology (Gibson, 1979), phenomenology (Heidegger, 1927 ; Merleau-Ponty, 1945), and cognitive neuroscience (Goodale and Milner, 1992) . These traditions suggest an alternative architecture in which affordance computation precedes geometric decompos","cbCaimnt77idYakx","https://ap.wps.com/l/cbCaimnt77idYakx","pdf",291290,2,1,33,"English","en",105,"# Introduction\n## Standard vs semantic-first processing pipelines","[{\"question\":\"What phenomenon does the paper focus on in vision-language models?\",\"answer\":\"The paper characterizes context-dependent affordance computation, showing that the affordances inferred by VLMs change substantially when the surrounding scene context is primed.\"},{\"question\":\"How was the main study conducted and which models were used?\",\"answer\":\"The primary experiment uses Qwen3-VL-30B-A3B with 3,213 COCO-2017 scene-context pairs and systematic context priming across 7 agentic personas, with a cross-model replication on LLaVA-1.5-13B.\"},{\"question\":\"What evidence shows the effect is not just generation noise?\",\"answer\":\"Stochastic baseline experiments with many inference runs across temperatures and seeds report within-prime variance substantially lower than cross-prime variance, indicating genuine context effects.\"}]",1784195826,83,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"context-dependent-affordance-computation-in-vision-language-models","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/context-dependent-affordance-computation-in-vision-language-models/84466/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-21","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What phenomenon does the paper focus on in vision-language models?","Question",{"text":75,"@type":76},"The paper characterizes context-dependent affordance computation, showing that the affordances inferred by VLMs change substantially when the surrounding scene context is primed.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How was the main study conducted and which models were used?",{"text":80,"@type":76},"The primary experiment uses Qwen3-VL-30B-A3B with 3,213 COCO-2017 scene-context pairs and systematic context priming across 7 agentic personas, with a cross-model replication on LLaVA-1.5-13B.",{"name":82,"@type":73,"acceptedAnswer":83},"What evidence shows the effect is not just generation noise?",{"text":84,"@type":76},"Stochastic baseline experiments with many inference runs across temperatures and seeds report within-prime variance substantially lower than cross-prime variance, indicating genuine context effects.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]