[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82093-en":3,"doc-seo-82093-105":30,"detail-sidebar-cat-0-en-105":83},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},82093,1374391975076,"Riley","https://ap-avatar.wpscdn.com/avatar/14000253ca4ec9f6853?x-image-process=image/resize,m_fixed,w_180,h_180&k=1783305029341752051",8,"Research & Report","MemeBuddy Dialog-Style Audio Representations for Engaging Non-Visual Meme Experiences","Image memes are a widely used online form of humor and cultural reference, yet current accessibility support for blind users is limited. Prior caption-based approaches improve readability but often lose the timing, narrative structure, and contextual nuance that make memes engaging. MemeBuddy treats a meme as a dialog between role-based speakers, generating structured multi-turn audio using extracted meme text plus contextual knowledge from a multimodal LLM. A study with 14 blind participants shows higher engagement and satisfaction versus caption-style descriptions with comparable comprehension.","MemeBuddy: Dialog-Style Audio Representations for Engaging Non-Visual  \nMeme Experiences  \nChirag Bhansali1 , Vikas Ashok2 , Hae-Na Lee1  \n1Department of Computer Science and Engineering 2Department of Computer Science  \nMichigan State University Old Dominion University  \nEast Lansing, MI, USA Norfolk, VA, USA  \n{bhansal4, [leehaena}@msu.edu](leehaena}@msu.edu) [vganjigu@odu.edu](vganjigu@odu.edu)  \narXiv :2607 .089 12v 1 [ cs .HC] 9 Jul 2026  \nAbstract  \nImage memes are a pervasive form of online communication, widely used to convey humor, opinions, and cultural references. Prior work has explored making memes accessible to blind users, primarily through auto-generated descriptive captions. While these approaches improve comprehensibility and sometimes incorporate prosodic or emotional cues, they often fail to capture the humor, narrative structure, and contextual nuances that make memes engaging. We present MemeBuddy, a system that models memes as dialog, generating structured, multi-turn audio representations using role-based speakers. MemeBuddy reinterprets a meme as a conversation between two speakers, integrating extracted meme text with contextual knowledge implicitly inferred by a multimodal LLM (e.g., recognition of common meme templates and cultural references) to convey intent, timing, and implicit meaning through conversational interaction. We evaluate MemeBuddy in a user study with 14 blind participants. Results show that dialog-style meme representations consistently improve engagement and user satisfaction compared to caption-style descriptions, while maintaining comparable comprehension.  \n1 Introduction  \nImage memes are humorous and culturally referential visual artifacts that have become a dominant form of online communication across social media, forums, and messaging platforms (Davison, 2012 ; Molina, 2020 ; Bauckhage, 2011 ; KostadinovskaStojchevska and Shalevska, 2018 ; Morina and Bernstein, 2022) . Users rely on memes to express opinions, share experiences, and participate in collective cultural discourse (Chen, 2012 ; Blommaert and Varis, 2017) . Ensuring equitable access to memes is therefore critical, particularly for blind users who interact with content through synthesized speech via screen readers (e.g., JAWS, VoiceOver, NVDA) .  \nHowever, meme accessibility remains limited. Screen readers primarily depend on alt text, which is often missing, incomplete, or insufficient for conveying meaning of images to the blind users (Voykinska et al., 2016 ; Gleason et al., 2019a ; WebAIM, 2026) . To compensate, blind users increasingly rely on AI-based tools (e.g., ChatGPT, Seeing AI, Be My AI) that generate image descriptions on demand (Sharma et al., 2025 ; Adnin and Das, 2024) . While these systems improve comprehensibility, they typically produce static, captionstyle outputs that fail to capture the humor, timing, and contextual nuance central to memes (Dynel, 2016 ; Vásquez and Aslan, 2021) . As a result, the users often receive functionally correct but experientially diminished representations.  \nRecent work has explored enhancing engagement through emotional narration (Chen et al., 2025) . However, the approach largely retains a single-speaker, descriptive paradigm. We argue that the limitation lies not only in what is conveyed, but in how it is structured. Memes are inherently social and interpretative; their meaning often unfolds through reactions, timing, and shared understanding. This motivates modeling memes as dialogs rather than static descriptions.  \nTowards this, we present MemeBuddy, a system that generates dialog-style audio representations of memes using role-based speakers. Specifically, MemeBuddy reimagines a meme as a structured, multi-turn conversation between two speakers, enabling the gradual unfolding of humor, intent, and cultural context. As illustrated in Table 1, MemeBuddy supports two dialog variants: (i) commentator–commentator (CO), where two speakers collaboratively interpret","cbCailgUxtf8F4AR","https://ap.wps.com/l/cbCailgUxtf8F4AR","pdf",890869,2,1,19,"English","en",105,"# Abstract\n# Introduction\n## Problem: limited meme accessibility for blind users\n## Approach: MemeBuddy as dialog-style audio\n## Dialog variants and evaluation overview","[{\"question\":\"How is MemeBuddy evaluated, and what were the results?\",\"answer\":\"The system is evaluated through a user study with 14 blind participants. Dialog-style representations improve engagement and user satisfaction compared to caption-style descriptions while maintaining comparable comprehension.\"}]",1784178178,48,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":78,"head_meta":80,"extra_data":82,"updated_unix":28},"memebuddy-dialog-style-audio-representations-for-engaging-non-visual-meme-experiences","",{"@graph":36,"@context":77},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/memebuddy-dialog-style-audio-representations-for-engaging-non-visual-meme-experiences/82093/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-19","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71],{"name":72,"@type":73,"acceptedAnswer":74},"How is MemeBuddy evaluated, and what were the results?","Question",{"text":75,"@type":76},"The system is evaluated through a user study with 14 blind participants. Dialog-style representations improve engagement and user satisfaction compared to caption-style descriptions while maintaining comparable comprehension.","Answer","https://schema.org",{"og:url":51,"og:type":79,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":81,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":84},[85,89,93,97,102,107,112,115,120,123,127],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":86,"show_sort_weight":87,"slug":88},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":90,"show_sort_weight":91,"slug":92},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Exam",70,"exam",{"id":98,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},5,"Comic",60,"comic",{"id":103,"doc_module":4,"doc_module_name":46,"category_name":104,"show_sort_weight":105,"slug":106},6,"Technology",50,"technology",{"id":108,"doc_module":4,"doc_module_name":46,"category_name":109,"show_sort_weight":110,"slug":111},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":113,"slug":114},30,"research-report",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},9,"Religion & Spirituality",20,"religion-spirituality",{"id":118,"doc_module":4,"doc_module_name":46,"category_name":121,"show_sort_weight":118,"slug":122},"World Cup","world-cup",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":124,"slug":126},10,"Lifestyle","lifestyle",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":98,"slug":129},"General","general"]