[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83743-en":3,"doc-seo-83743-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83743,549758252649,"Ivy","https://ap-avatar.wpscdn.com/avatar/8000253669c5317157?_k=1778319167496531819",8,"Research & Report","EmCom-Diffusion: Probing Visual Reflection in Emergent Languages via Image Generation","Measuring how emergent languages encode the visual information of their inputs remains an open challenge. Visual reflection is defined as the degree to which emergent messages preserve source-image details recoverable without using the original speaker–listener pair. Existing metrics rely on indirect proxies—concept inventories, captions, structural distance correlations, or referential accuracy—leading to omissions or misattributions. EmCom-Diffusion evaluates visual reflection directly by reconstructing input images from emergent messages using a fine-tuned text-to-image diffusion model and scoring perceptual similarity.","arXiv :2607 .03752v 1 [ cs .CV] 4 Jul 2026  \nEmCom-Diffusion: Probing Visual Reflection in Emergent Languages via Image Generation  \nHaruumi Omoto 1[0009−0004−3905−6412] and Tadahiro Taniguchi 1 ,2[0000−0002−5682−2076]  \n1 Graduate School of Informatics, Kyoto University, Kyoto, Japan  \n2 Research Organization of Science and Technology, Ritsumeikan University, Shiga, Japan  \n[omoto.haruumi.45u@st.kyoto-u.ac.jp](omoto.haruumi.45u@st.kyoto-u.ac.jp), [taniguchi@i.kyoto-u.ac.jp](taniguchi@i.kyoto-u.ac.jp)  \nAbstract. Measuring the extent to which emergent languages encode the visual content of their inputs is an open problem. We refer to this property as visual reflection: the extent to which emergent messages preserve information about their source images that can be recovered without appeal to the speaker–listener pair that produced them. Existing metrics measure it only indirectly, through proxies such as human-defined concept inventories, natural-language captions, structural distance correlations, or Referential Game accuracy, each of which can either miss visual content the message encodes or credit content it does not. We propose EmCom-Diffusion, an evaluation framework that measures visual reflection directly: it reconstructs each input image from its emergent message and compares the reconstruction with the original image itself, rather than with human-defined targets. Concretely, it fine-tunes apretrained text-to-image diffusion model on (image, emergent-message) pairs and scores visual reflection as the perceptual similarity between the reconstructed and original images, operating generatively rather than discriminatively. Instantiating it on MS-COCO with a Referential Game, we validate the metric against random and fixed-token baselines under three pretrained visual encoders, and compare it against four existing metrics (CBM, supervised translation, TopSim, and R@1) . EmComDiffusion captures visual content the other metrics miss. Our code is available at [https://github.com/Tanichu-Laboratory/EmCom-Diffusion](https://github.com/Tanichu-Laboratory/EmCom-Diffusion).  \nKeywords: Emergent Communication · Text-to-Image Generation  \n· Evaluation Metrics.  \n1 Introduction  \nEmergent Communication (EmCom) is a research framework in which agents develop an emergent language (EL) from scratch through interaction [15]; a central open problem is how to measure what this language encodes about its visual inputs. We call this property visual reflection, and understanding it is important for interpreting what an emergent language means. It also bears on a broader question: why language comes to encode the structure of the physical  \n2 H. Omoto and T. Taniguchi  \nworld [9, 29] . LLMs trained only on text appear to capture aspects of this physical structure [8, 20, 30], but this evidence comes from mature human language that has already converged on such encodings. Emergent communication instead lets us observe how this structure comes to be encoded as a language forms from scratch. Within EmCom, however, attention has largely gone to the formal properties of the resulting language, such as the compositionality, generalization, and syntactic structure of its token sequences [5, 22, 26], while visual reflection has received comparatively little attention. The few studies that do examine it report that emergent languages can drift from their visual input despite high task success [2, 13] . We therefore focus on visual reflection in emergent languages.  \nHowever, the lack of established evaluation metrics has limited systematic investigation of visual reflection. Metrics for evaluating semantic content in emergent languages fall into three families [24] . (i) Reference-based metrics compare emergent tokens against human-defined references [17, 22] . Concept-BestMatching (CBM) [4] matches tokens to predefined oracle concepts via bipartite matching; supervised translation methods [31] train translators on paired EL–NL data and report transl","cbCaioxgQ3c9t7AT","https://ap.wps.com/l/cbCaioxgQ3c9t7AT","pdf",9583728,3,1,15,"English","en",105,"# Introduction\n## Visual reflection and evaluation motivation\n## Existing metric families and limitations\n## EmCom-Diffusion overview","[{\"question\":\"What is visual reflection in emergent communication?\",\"answer\":\"Visual reflection is the extent to which emergent messages preserve information about their source images that can be recovered without relying on the speaker–listener pair that produced them.\"},{\"question\":\"Why do existing metrics for emergent languages have limitations?\",\"answer\":\"They use indirect proxies such as concept inventories, captions, distance correlations, or referential-game accuracy, which can miss visual content encoded by messages or credit content not reflected visually.\"},{\"question\":\"How does EmCom-Diffusion measure visual reflection?\",\"answer\":\"It fine-tunes a text-to-image diffusion model on (image, emergent-message) pairs, reconstructs each input image from its emergent message, and computes perceptual similarity between the reconstruction and the original image.\"}]",1784190158,38,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"emcom-diffusion-probing-visual-reflection-in-emergent-languages-via-image-generation","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/emcom-diffusion-probing-visual-reflection-in-emergent-languages-via-image-generation/83743/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What is visual reflection in emergent communication?","Question",{"text":75,"@type":76},"Visual reflection is the extent to which emergent messages preserve information about their source images that can be recovered without relying on the speaker–listener pair that produced them.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"Why do existing metrics for emergent languages have limitations?",{"text":80,"@type":76},"They use indirect proxies such as concept inventories, captions, distance correlations, or referential-game accuracy, which can miss visual content encoded by messages or credit content not reflected visually.",{"name":82,"@type":73,"acceptedAnswer":83},"How does EmCom-Diffusion measure visual reflection?",{"text":84,"@type":76},"It fine-tunes a text-to-image diffusion model on (image, emergent-message) pairs, reconstructs each input image from its emergent message, and computes perceptual similarity between the reconstruction and the original image.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]