[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83414-en":3,"doc-seo-83414-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83414,7971461741311,"Ophelia","https://ap-avatar.wpscdn.com/avatar/74000253aff267980c6?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779345379180704826",8,"Research & Report","How YouTube Frames ChatGPT Use in Education An Epistemic Network Analysis with Supporting Multimodal Metadata","Educational YouTube videos are analyzed through multimodal metadata—transcripts, titles, thumbnails, and viewer comments—to study how ChatGPT is framed across creator groups and how those framings shape audience response and platform reach. Following PRISMA, 52 videos are selected and organized into three structurally distinct discourse groups: conceptual scaffolding, retrieval-and-skill building, and output generation. Epistemic Network Analysis finds statistically significant group differences with large effect sizes, mirrored across modalities. Learning-oriented viewers describe ChatGPT as a thinking partner or tutor, while output-oriented viewers raise concerns about over-reliance, surface-level learning, and cognitive offloading. Output-oriented content reaches comparably to skill-oriented content yet shows weaker learning framing, revealing a scaling tension in self-directed AI literacy.","How YouTube Frames ChatGPT Use in Education: An Epistemic Network Analysis with Supporting Multimodal Metadata  \nShayla Sharmin  \n[shayla@udel.edu](shayla@udel.edu)[ ](shayla@udel.edu)University of Delaware Newark, Delaware, USA  \nMohammad Al-Ratrout  \n[mratrout@udel.edu](mratrout@udel.edu)[ ](mratrout@udel.edu)University of Delaware Newark, Delaware, USA  \nMohammad Fahim Abrar  \n[fahim@udel.edu](fahim@udel.edu)[ ](fahim@udel.edu)University of Delaware Newark, Delaware, USA  \nRoghayeh Leila Barmaki  \n[rlb@udel.edu](rlb@udel.edu)[ ](rlb@udel.edu)University of Delaware Newark, Delaware, USA  \narXiv :2607 .08698v 1 [ cs .HC] 9 Jul 2026  \nFigure 1: Transcripts from 52 YouTube videos were analyzed using Epistemic Network Analysis, showing three discourse groups: 􀀜1 (learning-oriented), 􀀜2 (skill-oriented), and 􀀜3 (productivity-oriented). For effective learning with LLM, the best practice is to combine the learning and skill-oriented prompts.  \nAbstract  \nWe examine educational YouTube videos through multimodal metadata, such as transcripts, titles, thumbnails, and viewer comments, to investigate how ChatGPT is framed across creator groups and how those framings relate to audience response and platform reach. Little is known about how large language models are presented to learners in informal, creator-driven public discourse. Following PRISMA, we selected 52 videos for analysis. We identified three structurally distinct discourse groups: (􀀜1) videos that positioned ChatGPT as a conceptual scaffold for thinking,(􀀜2) videos oriented toward retrieval practice and skill-building, and (􀀜3) videos that framed ChatGPT as a tool for output generation. Epistemic Network Analysis revealed statistically significant group differences with large effect sizes. Multimodal metadata consistently reflected these distinctions across transcript discourse, titles, and thumbnails. Viewers of learning-oriented content described ChatGPT asa thinking partner or tutor, whereas viewers of output-oriented content raised concerns about over-reliance, surface-level learning,  \nThis work is licensed under a Creative Commons Attribution 4 .0 International License. ICMI’26, Napoli, Italy  \n© 2025 Copyright held by the owner/author(s) .  \nACM ISBN XXXX [https://doi.org/XXXX](https://doi.org/XXXX)  \nand cognitive offloading. 􀀜 3 achieved comparable platform reach to 􀀜 2, yet with substantially weaker learning-oriented framing. This may suggest that output-oriented content competes for visibility despite lower pedagogical depth. These findings reveal a structural tension in self-directed AI learning: content that prioritizes quick outputs reaches far more learners than content that promotes deep engagement. This gap raises critical questions about whose vision of AI literacy scales and what learners are actually left with.  \nCCS Concepts  \n• Applied computing → E-learning; Interactive learning environments.  \nKeywords  \nYouTube; ChatGPT; Educational Videos; Epistemic Network; PRISMA  \nACM Reference Format:  \nShaylaSharmin, Mohammad Al-Ratrout, Mohammad Fahim Abrar, andRoghayeh Leila Barmaki. 2025. How YouTube Frames ChatGPT Use in Education: An Epistemic Network Analysis with Supporting Multimodal Metadata. In Proceedings of the 27th International Conference on Multimodal Interaction (ICMI’25), October 05–09, 2026, Napoli, Italy. ACM, New York, NY, USA, 9 pages. [https://doi.org/XXXX](https://doi.org/XXXX)  \n1 Introduction  \nThe rapid adoption of large language models (LLMs) such as ChatGPT has introduced new questions about how these tools are understood and used in educational contexts. Prior research on AI in education has predominantly examined how tools are introduced within designed instructional environments, such as classrooms and formal curricula, where pedagogical intent is explicit and learner behavior is observable. Educational use of LLMs, including ChatGPT, is increasing in formal classrooms, but it is still underexplored how ChatGPT has been framed in ","cbCaitUI4mS2YvXS","https://ap.wps.com/l/cbCaitUI4mS2YvXS","pdf",2137526,3,1,9,"English","en",105,"# Introduction\n## Research Questions\n## Contributions","[{\"question\":\"What multimodal data sources are used to analyze ChatGPT framing on YouTube?\",\"answer\":\"The study uses multimodal metadata including creator transcripts, titles, thumbnails, and viewer comments, along with engagement patterns to capture platform-level effects.\"},{\"question\":\"How many discourse groups are identified, and what distinguishes them?\",\"answer\":\"Three discourse groups are identified: learning-oriented framing as conceptual scaffolding, skill-oriented framing focused on retrieval practice and skill building, and productivity/output-oriented framing focused on generating outputs.\"},{\"question\":\"How does viewer response differ between learning-oriented and output-oriented content?\",\"answer\":\"Viewers of learning-oriented content describe ChatGPT as a thinking partner or tutor, while viewers of output-oriented content express concerns such as over-reliance, surface-level learning, and cognitive offloading.\"}]",1784187411,23,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"how-youtube-frames-chatgpt-use-in-education-an-epistemic-network-analysis-with-supporting-multimodal-metadata","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/how-youtube-frames-chatgpt-use-in-education-an-epistemic-network-analysis-with-supporting-multimodal-metadata/83414/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-25","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What multimodal data sources are used to analyze ChatGPT framing on YouTube?","Question",{"text":75,"@type":76},"The study uses multimodal metadata including creator transcripts, titles, thumbnails, and viewer comments, along with engagement patterns to capture platform-level effects.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How many discourse groups are identified, and what distinguishes them?",{"text":80,"@type":76},"Three discourse groups are identified: learning-oriented framing as conceptual scaffolding, skill-oriented framing focused on retrieval practice and skill building, and productivity/output-oriented framing focused on generating outputs.",{"name":82,"@type":73,"acceptedAnswer":83},"How does viewer response differ between learning-oriented and output-oriented content?",{"text":84,"@type":76},"Viewers of learning-oriented content describe ChatGPT as a thinking partner or tutor, while viewers of output-oriented content express concerns such as over-reliance, surface-level learning, and cognitive offloading.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]