[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-134797-en":3,"doc-seo-134797-105":31,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":28,"seo_description":14,"update_tm":29,"read_time":30},134797,687207412472,"Angel","https://ap-avatar.wpscdn.com/davatar_155a257f0dc6eb9ab79c44ca47cae57d",8,"Research & Report","Generating High Quality Anime Videos with Diffusion - Multi-stage frame generation and refinement pipeline","Generating high-quality anime videos remains difficult because traditional animation is labor-intensive and preserving temporal coherence and smooth motion between scenes is complex. This project builds a multi-stage pipeline to generate and refine anime-style motion. Initial frames are produced from text or image using pretrained models such as SVD, ModelScopeT2V, and I2VGenXL, then extended with StreamingT2V via autoregressive conditioning. Final outputs use AnimeInterp for interpolation and AnimeSR for super-resolution, evaluated with FVD and LPIPS, showing that fine-tuning Stability AI’s diffusion model with SVD gives the best results.","Generating High Quality Anime Videos with Diffusion  \nJustin Lim Stanford University [jlim23@stanford.edu](jlim23@stanford.edu)  \nWinston Shum Stanford University [wshum@stanford.edu](wshum@stanford.edu)  \nJonathan Lee Stanford University [jezra@stanford.edu](jezra@stanford.edu)  \nAbstract  \nGenerating high-quality anime videos remains a significant challenge due to the labor-intensive nature of traditional animation processes, as well as complexities involved in maintaining temporal coherence and smooth motion dynamics between scenes. This project aims to improve anime generation through a multi-stage frame generation and refinement pipeline. We first propose leveraging pretrained models such as SVD, ModelScopeT2V, and I2VGenXL to generate initial video frames from text and image inputs. Subsequent video frames are then generated using the StreamingT2V framework, which employs autoregressive conditioning to ensure consistency and smooth transitions. The final refinement stage incorporates anime-specific interpolation (AnimeInterp) and super-resolution enhancement (AnimeSR) techniques to output videos specifically for anime studios. Evaluation metrics, including Frechet Video Distance (FVD) and Learned Perceptual Image Patch Similarity (LPIPS), are used to assess the perceptual similarity and quality of the generated videos. Our results show that fine-tuning Stability AI’s diffusion model, SVD, yields the best performance, demonstrating the potential of automating the production of high-quality anime videos that adhere closely to the original anime style.  \n1. Introduction  \nJapanese animation studios spend extensive hours animating and producing entire anime videos using a limited number of manga frames. Despite technological advancements, overworked animation studios often produce videos with low frame rates due to the labor-intensive process of drawing and designing each frame while meeting tight deadlines. However, recent developments in video understanding and generation within the deep learning domain have aimed to mitigate some of these challenges by focusing on frame generation or enhancing the quality of lowresolution frames. However, the task of generating highquality anime videos for extended durations remains largely unsolved due to several key challenges, including maintain-  \ning temporal coherence between frames, ensuring smooth optical flow, and effectively learning motion dynamics. As such, we aim to contribute to the field by experimenting with frameworks capable of converting single anime frames or context scripts to anime-style videos.  \nOur project involves the following steps: (1) Initial Frame Generation with models capable of image-to-video and/or text-to-video, (2) Autoregressive Fusion through feeding the initial frames from step 1 into StreamingT2V [2] to generate longer videos, and (3) Quality Refinement by refining the output with anime specific interpolation and resolution-enhancement methods from AnimeInterp [7] and AnimeSR [10] respectively. Using these approaches, we aim to experiment and search for the best pipeline that addresses the challenges of generating high resolution and high framerate anime videos that stay true the supplied context.  \n1.1. Problem Statement  \nOur problem statement is as follows: given either image or text, how can we generate an anime-style video that preserves optical flow, physical laws, and temporal coherence, incorporates context from the story, and maintains the animation style?  \n2. Related Works  \n2.1. A Survey on Long Video Generation: Challenges, Methods, Prospects  \nThis paper provides a comprehensive overview of techniques for generating long-form videos [3] . It highlightsa technique we plan to utilize in our project: the temporal autoregressive method. This method generates frames sequentially, with each frame conditioned on the preceding one, therefore creating a cohesive video that maintains temporal coherence across frames. The paper also discusses th","cbCaiuU80EJRo91F","https://ap.wps.com/l/cbCaiuU80EJRo91F","pdf",11598808,2,1,9,"English","en",105,"# Introduction\n## Problem Statement\n# Related Works\n## A Survey on Long Video Generation: Challenges, Methods, Prospects\n## Phenaki: Variable Length Video Generation from Open Domain Textual Descriptions\n## VideoDrafter: Content-Consistent MultiScene Video Generation with LLM","[{\"question\":\"为什么生成高质量的动漫视频仍然很难？\",\"answer\":\"主要原因在于传统动画制作耗时且逐帧绘制成本高，同时还要在跨场景生成中保持时间一致性与平滑运动。文中强调需要同时解决时间连贯性、光流平滑以及运动动态学习等挑战。\"},{\"question\":\"该项目采用了怎样的多阶段生成与优化流程？\",\"answer\":\"流程分为三步：先用支持图像到视频和/或文本到视频的模型生成初始帧；再将初始帧输入 StreamingT2V 进行自回归融合以生成更长视频；最后用 AnimeInterp 插值与 AnimeSR 超分辨率对质量进行专门细化。\"},{\"question\":\"论文如何评估生成视频的质量与感知相似度？\",\"answer\":\"使用 Frechet Video Distance（FVD）衡量感知层面的分布差异，并用 Learned Perceptual Image Patch Similarity（LPIPS）评估生成结果与参考之间的感知相似性。\"}]","Generating High Quality Anime Videos with Diffusion - Multi-stage frame generation and refinement pipeline | PDF",1787299653,23,{"code":4,"msg":32,"data":33},"ok",{"site_id":25,"language":24,"slug":34,"title":13,"keywords":35,"description":14,"schema_data":36,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":29},"generating-high-quality-anime-videos-with-diffusion-multi-stage-frame-generation-and-refinement-pipeline","",{"@graph":37,"@context":86},[38,54,69],{"@type":39,"itemListElement":40},"BreadcrumbList",[41,45,48,51],{"item":42,"name":43,"@type":44,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":46,"name":47,"@type":44,"position":20},"https://docshare.wps.com/document/","Document",{"item":49,"name":12,"@type":44,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":44,"position":53},"https://docshare.wps.com/document/generating-high-quality-anime-videos-with-diffusion-multi-stage-frame-generation-and-refinement-pipeline/134797/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":42,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-23","2026-08-21",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"为什么生成高质量的动漫视频仍然很难？","Question",{"text":76,"@type":77},"主要原因在于传统动画制作耗时且逐帧绘制成本高，同时还要在跨场景生成中保持时间一致性与平滑运动。文中强调需要同时解决时间连贯性、光流平滑以及运动动态学习等挑战。","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"该项目采用了怎样的多阶段生成与优化流程？",{"text":81,"@type":77},"流程分为三步：先用支持图像到视频和/或文本到视频的模型生成初始帧；再将初始帧输入 StreamingT2V 进行自回归融合以生成更长视频；最后用 AnimeInterp 插值与 AnimeSR 超分辨率对质量进行专门细化。",{"name":83,"@type":74,"acceptedAnswer":84},"论文如何评估生成视频的质量与感知相似度？",{"text":85,"@type":77},"使用 Frechet Video Distance（FVD）衡量感知层面的分布差异，并用 Learned Perceptual Image Patch Similarity（LPIPS）评估生成结果与参考之间的感知相似性。","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,121,124,128,131,135],{"id":21,"doc_module":4,"doc_module_name":47,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":47,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":47,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":47,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":47,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":47,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":47,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":22,"doc_module":4,"doc_module_name":47,"category_name":125,"show_sort_weight":126,"slug":127},"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":47,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":47,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":47,"category_name":137,"show_sort_weight":107,"slug":138},19,"General","general"]