[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86557-en":3,"doc-seo-86557-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86557,1374391974564,"Clementine","https://ap-avatar.wpscdn.com/avatar/14000253aa45c000a9e?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779874745381141002",8,"Research & Report","SLVMBench: Skill Learning from Video Memory","SLVMBench introduces the first benchmark that tests whether video large language models can learn procedural skills from long video memory and apply them to real-time tasks. Models receive up to 2-hour streams containing a tutorial video embedded among unrelated distractors, then answer time-locked questions about an ongoing target video. The benchmark evaluates the full pipeline of memorizing and extracting procedural knowledge, transferring it under temporal cutoff pressure.","arXiv :2607 . 11312v1 [ cs .CV] 13 Jul 2026  \nSLVMBench: Skill Learning from Video Memory  \nYudong Yang 1 , Guangzhi Sun2 , Yixuan Li 1 , Chao Zhang 1  \n1Tsinghua University 2University of Cambridge  \n[yang-yd21@mails.tsinghua.edu.cn](yang-yd21@mails.tsinghua.edu.cn) [cz277@tsinghua.edu.cn](cz277@tsinghua.edu.cn)  \nAbstract  \nWe introduce Skill Learning from Video Memory (SLVMBench ), the first benchmark that jointly evaluates whether video large language models (video-LLMs) can learn skills from long video memory and apply them to real-time tasks. SLVMBench presents models with up to 2-hour video streams that contain a tutorial video embedded in a stream of arbitrary irrelevant videos, resembling real-world human learning practices. Video-LLMs are asked to apply the acquired skill to answer real-time questions about an ongoing video. Unlike long-video understanding benchmarks that emphasize passive comprehension and skill-learning benchmarks that rely on short, immediate demonstrations, SLVMBench tests the full pipeline of memorizing and extracting procedural knowledge, as well as transferring it to real-time tasks. Moreover, rigorous human annotations feature sub-second-level temporal calibration, manually engineered questions eliminating common-sense guessing, and collated tutorials to ensure coverage of the required skills. Evaluations on state-of-the-art proprietary and open-source video LLMs show that video-LLMs struggle substantially with learning and applying skill knowledge from videos. Moreover, performance degrades markedly when the skill knowledge is placed within a long video memory. These results reveal a key limitation of existing video LLMs and position SLVMBench as the first benchmark for studying real-time skill acquisition and application from long-context video memory.  \n1 Introduction  \nRecent advances in audio-visual large language models (video-LLMs) have shown a remarkable ability to understand short videos spanning a couple of minutes, enabling progress in tasks such as captioning, question answering, and multimodal reasoning [1–9] . However, when video-LLMs are used to power AI agents in real-world scenarios, accurate real-time perception and effective long-term memory are both indispensable requirements for those models. For instance, when performing a certain operation on software, as humans do, the agent should be able to learn from the demonstration of this task that it has seen a couple of hours ago. The evaluation of such abilities is a critical indicator to expand the scope of application for embodied AI.  \nA stream of research focuses on improving long-term memory for video-LLMs under streaming settings to perform real-time question answering, using approaches such as KV cache compression [10–14] or the design of memory mechanisms to aggregate tokens [15–18] . Despite the progress, one of the key limitations lies in the evaluation of these approaches, which is often separated into two disjoint aspects: To evaluate real-time video perception, benchmarks such as StreamingBench [19] and OVOBench [20] are widely adopted [11, 14, 16], where the video memory required is often limited to 10-20 minutes without requiring long-term memory. On the other hand, to evaluate long video memory capabilities, a series of long-video benchmarks, such as VideoMME [21], LVBench [22], and LongVideoBench [23], are often used. However, these benchmarks do not require any real-time processing, and instead of assessing whether models can learn procedural knowledge (i.e., skills) from videos, they often primarily measure passive comprehension and information extraction.  \nPreprint.  \nFigure 1: Overview of the SLVMBench Task. The task evaluates multimodal episodic memory by requiring models to: (i) learn procedural knowledge from an initial Tutorial Video; (ii) maintain this knowledge throughout a long sequence of distractor videos; and (iii) perform predictive reasoning ina Target Video at a precise temporal cutoff.  \nTo","cbCaidzlbK0EGfnb","https://ap.wps.com/l/cbCaidzlbK0EGfnb","pdf",11885718,4,1,26,"English","en",105,"# Abstract\n# Introduction\n# Related Work","[{\"question\":\"What does SLVMBench benchmark for video-LLMs?\",\"answer\":\"SLVMBench jointly evaluates whether video-LLMs can learn procedural knowledge from long video memory and apply the learned skill to answer real-time questions about an ongoing video.\"},{\"question\":\"How is the SLVMBench input constructed for each sample?\",\"answer\":\"Each sample includes a tutorial video embedded inside a 2–3 hour stream of irrelevant distractor videos, followed by a target video where questions are asked at specific timestamps.\"},{\"question\":\"What do the evaluations show about current video-LLMs?\",\"answer\":\"Evaluations on state-of-the-art proprietary and open-source video LLMs indicate substantial difficulty in learning and transferring skill knowledge from videos, with performance degrading markedly as the temporal distance between tutorial and target increases.\"}]",1784212619,66,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"slvmbench-skill-learning-from-video-memory","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/slvmbench-skill-learning-from-video-memory/86557/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-28","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What does SLVMBench benchmark for video-LLMs?","Question",{"text":75,"@type":76},"SLVMBench jointly evaluates whether video-LLMs can learn procedural knowledge from long video memory and apply the learned skill to answer real-time questions about an ongoing video.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How is the SLVMBench input constructed for each sample?",{"text":80,"@type":76},"Each sample includes a tutorial video embedded inside a 2–3 hour stream of irrelevant distractor videos, followed by a target video where questions are asked at specific timestamps.",{"name":82,"@type":73,"acceptedAnswer":83},"What do the evaluations show about current video-LLMs?",{"text":84,"@type":76},"Evaluations on state-of-the-art proprietary and open-source video LLMs indicate substantial difficulty in learning and transferring skill knowledge from videos, with performance degrading markedly as the temporal distance between tutorial and target increases.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]