[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84911-en":3,"doc-seo-84911-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},84911,1099514068035,"Ezra","https://ap-avatar.wpscdn.com/davatar_276721f389ce27ea32af1340a28f341c",8,"Research & Report","What Images Cannot Say: Language-Guided Olfactory Representation Learning","While existing datasets align visual scenes with electronic-nose signals, mapping smell to images remains difficult because many olfactory cues depend on contextual environmental factors that are not directly visible in pixels. SCENT introduces a multimodal framework that uses language guidance as a semantic bridge between vision and olfaction. Vision-Language Models generate scene descriptors, and a smell encoder aligns e-nose signals with shared visual and textual embeddings. A language-guided latent decomposition separates object-specific odors from contextual contributions.","arXiv :2607 .06402v 1 [ cs .CV] 7 Jul 2026  \nWhat Images Cannot Say: Language-Guided Olfactory Representation Learning  \nEleftherios Tsonis, Xi Wang, and Vicky Kalogeiton  \nLIX, École Polytechnique, IP Paris, CNRS {[firstname.lastname}@polytechnique.edu](firstname.lastname}@polytechnique.edu)[ ](firstname.lastname}@polytechnique.edu)[https://www.lix.polytechnique.fr/vista/projects/2026_scent_tsonis/](https://www.lix.polytechnique.fr/vista/projects/2026_scent_tsonis/)  \n[Abstract.](Abstract. Images tell us what a scene looks like)[ Images tell us what a scene looks like](Abstract. Images tell us what a scene looks like), [but rarely what it](but rarely what it)[ ](but rarely what it)[would feel like to be there. While recent datasets pair visual scenes](would feel like to be there. While recent datasets pair visual scenes)[ ](would feel like to be there. While recent datasets pair visual scenes)[with electronic-nose measurements](with electronic-nose measurements), [aligning smell signals with images](aligning smell signals with images)[ ](aligning smell signals with images)remains challenging because many olfactory cues arise from contextual environmental factors that are not directly visible in pixels. We introduce SCENT, a multimodal framework that uses language guidance asa semantic bridge between vision and olfaction. Our approach leverages Vision-Language Models (VLMs) to generate scene descriptors capturing objects, environmental context, and plausible ambient smell cues suggested by the visual scene. These descriptors provide semantic guidance for learning olfactory representations. We train a smell encoder that maps electronic-nose signals into a shared embedding space aligned with both visual and textual representations, and introduce a languageguided latent decomposition that separates object-specific odors from contextual environmental contributions. Experiments on the New York Smells dataset demonstrate that SCENT significantly improves crossmodal retrieval compared to vision-only baselines, achieving state-of-theart performance on smell-to-image and smell-to-text retrieval tasks. In addition, our framework produces interpretable olfactory representations that enable the disentanglement of complex smell mixtures. Our results reveal the importance of contextual semantic information for grounding olfactory perception in multimodal learning and pave the way for future research in this area.  \nKeywords: Cross-modal Retrieval · Olfactory Representation Learning  \n· Vision-Language Models  \n1 Introduction  \n“A picture is worth a thousand words” (att. Frederick R. Barnard)  \n. .. but it rarely tells us what it feels like to be there. .  \nIn real environments, perception extends beyond vision. A photograph of a busy street may evoke the smell of vehicle exhaust; a metro entrance may imply the metallic scent of ventilation air. Although such environmental cues are rarely  \n2 E. Tsonis et al.  \nFig. 1: Not everything we can smell is visible. Given only View 1 (in-sample) during training, the true olfactory context may lie outside the field of view, in View 2 (out-of-sample) . A VLM can bridge this gap by inferring plausible smells from semantic context alone. We use these language-derived signals as supervision to learn richer smell representations that go beyond what is directly seen.  \nvisible directly, humans routinely infer them from visual context. Despite years of progress in computer vision, AI systems today primarily capture what a scene looks like, while remaining largely unaware of what it might feel like to be there.  \nRecent advances in Vision-Language Models (VLMs) suggest that this longstanding idea may finally be realized: modern systems can generate rich descriptions of visual scenes [34, 35, 67], answer questions about images [19, 64, 71], and reason about complex environments through language [5, 30] . Their extensive world knowledge opens the possibility of inferring what exists in a scene, beyond what is ","cbCaiacyXXnxDQd7","https://ap.wps.com/l/cbCaiacyXXnxDQd7","pdf",11268878,1,32,"English","en",105,"# Abstract\n# Introduction\n## Motivation: limits of visual-only perception\n## Olfaction challenges with electronic noses\n## Language-guided vision-language models for scent context","[{\"question\":\"Why is aligning smell signals with images challenging?\",\"answer\":\"Many odor cues arise from contextual environmental factors that are not directly visible in pixels, and visual supervision often covers only a partial view of the scene while smells may come from outside the camera’s field of view.\"},{\"question\":\"What is SCENT and what does it do?\",\"answer\":\"SCENT is a multimodal framework that uses language guidance to connect vision and olfaction. It leverages vision-language models to generate scene descriptors and trains a smell encoder to map electronic-nose signals into a shared embedding space aligned with visual and textual representations.\"},{\"question\":\"How does SCENT improve cross-modal retrieval and interpretability?\",\"answer\":\"Experiments on the New York Smells dataset show improved cross-modal retrieval over vision-only baselines for smell-to-image and smell-to-text tasks. The framework also produces interpretable olfactory representations that disentangle complex smell mixtures into object-specific and contextual components.\"}]",1784199305,81,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"what-images-cannot-say-language-guided-olfactory-representation-learning","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/what-images-cannot-say-language-guided-olfactory-representation-learning/84911/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why is aligning smell signals with images challenging?","Question",{"text":75,"@type":76},"Many odor cues arise from contextual environmental factors that are not directly visible in pixels, and visual supervision often covers only a partial view of the scene while smells may come from outside the camera’s field of view.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What is SCENT and what does it do?",{"text":80,"@type":76},"SCENT is a multimodal framework that uses language guidance to connect vision and olfaction. It leverages vision-language models to generate scene descriptors and trains a smell encoder to map electronic-nose signals into a shared embedding space aligned with visual and textual representations.",{"name":82,"@type":73,"acceptedAnswer":83},"How does SCENT improve cross-modal retrieval and interpretability?",{"text":84,"@type":76},"Experiments on the New York Smells dataset show improved cross-modal retrieval over vision-only baselines for smell-to-image and smell-to-text tasks. The framework also produces interpretable olfactory representations that disentangle complex smell mixtures into object-specific and contextual components.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":45,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":45,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":45,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]