[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"detail-sidebar-cat-1-en-105":3,"doc-seo-238597-105":53,"doc-detail-238597-en":126},{"code":4,"msg":5,"data":6},0,"success",[7,14,19,24,29,34,39,44,49],{"id":8,"doc_module":9,"doc_module_name":10,"category_name":11,"show_sort_weight":12,"slug":13},11,1,"Template","Presentations",90,"presentations",{"id":15,"doc_module":9,"doc_module_name":10,"category_name":16,"show_sort_weight":17,"slug":18},12,"Resumes",80,"resumes",{"id":20,"doc_module":9,"doc_module_name":10,"category_name":21,"show_sort_weight":22,"slug":23},14,"Invoices",70,"invoices",{"id":25,"doc_module":9,"doc_module_name":10,"category_name":26,"show_sort_weight":27,"slug":28},15,"Posters",60,"posters",{"id":30,"doc_module":9,"doc_module_name":10,"category_name":31,"show_sort_weight":32,"slug":33},16,"Social Media",50,"social-media",{"id":35,"doc_module":9,"doc_module_name":10,"category_name":36,"show_sort_weight":37,"slug":38},17,"Forms",40,"forms",{"id":40,"doc_module":9,"doc_module_name":10,"category_name":41,"show_sort_weight":42,"slug":43},18,"Letters",30,"letters",{"id":45,"doc_module":9,"doc_module_name":10,"category_name":46,"show_sort_weight":47,"slug":48},21,"Paper Templates",5,"papers-templates",{"id":50,"doc_module":9,"doc_module_name":10,"category_name":51,"show_sort_weight":4,"slug":52},158,"General","general-158",{"code":4,"msg":54,"data":55},"ok",{"site_id":56,"language":57,"slug":58,"title":59,"keywords":60,"description":61,"schema_data":62,"social_meta":119,"head_meta":121,"extra_data":123,"updated_unix":125},105,"en","autotransition-learning-to-recommend-video-transition-effects","AutoTransition - Learning to Recommend Video Transition Effects","","Video transition effects are widely used in video editing to connect shots and produce cohesive, visually appealing results, yet selecting high-quality transitions remains difficult for non-professionals due to limited cinematography and design knowledge. This paper presents an automatic video transition recommendation approach (VTR) that recommends transitions for each pair of neighboring shots using the raw video shots and companion audio. A large-scale dataset is built from publicly available editing templates, and VTR is formulated as a multi-modal retrieval problem with a matching framework that fuses vision and audio via a multi-modal transformer. Experiments and a user study show strong effectiveness, with efficiency gains up to 300× compared to manual editing.",{"@graph":63,"@context":118},[64,80,101],{"@type":65,"itemListElement":66},"BreadcrumbList",[67,71,74,77],{"item":68,"name":69,"@type":70,"position":9},"https://docshare.wps.com","Home","ListItem",{"item":72,"name":10,"@type":70,"position":73},"https://docshare.wps.com/template/",2,{"item":75,"name":51,"@type":70,"position":76},"https://docshare.wps.com/template/general/",3,{"item":78,"name":59,"@type":70,"position":79},"https://docshare.wps.com/template/autotransition-learning-to-recommend-video-transition-effects/238597/",4,{"url":78,"name":59,"@type":81,"image":82,"author":87,"headline":59,"publisher":90,"fileFormat":93,"inLanguage":57,"description":61,"dateModified":94,"datePublished":95,"encodingFormat":93,"isAccessibleForFree":96,"interactionStatistic":97},"DigitalDocument",{"url":83,"@type":84,"width":85,"height":86},"https://docshare.wps.com/thumbnails/autotransition-learning-to-recommend-video-transition-effects/238597.png","ImageObject",442,249,{"name":88,"@type":89},"Stanford","Person",{"url":68,"name":91,"@type":92},"DocShare","Organization","application/pdf","2026-09-26","2026-09-11",true,{"@type":98,"interactionType":99,"userInteractionCount":79},"InteractionCounter",{"@type":100},"ViewAction",{"@type":102,"mainEntity":103},"FAQPage",[104,110,114],{"name":105,"@type":106,"acceptedAnswer":107},"What is the video transition recommendation task (VTR) defined as?","Question",{"text":108,"@type":109},"Given a sequence of raw video shots and companion audio, VTR recommends a sequence of transitions for each neighboring shot pair, aiming to provide ranked candidate transitions rather than a single class.","Answer",{"name":111,"@type":106,"acceptedAnswer":112},"How is the training dataset for video transitions obtained?",{"text":113,"@type":109},"A large-scale dataset is collected from publicly available video templates on editing software, using comprehensive selection rules and preprocessing to refine high-quality samples.",{"name":115,"@type":106,"acceptedAnswer":116},"What model design is used to match vision/audio inputs to transitions?",{"text":117,"@type":109},"The method learns transition embeddings via a transition classification task, then uses a multi-modal transformer to fuse vision and audio information and capture contextual cues across sequential transition outputs.","https://schema.org",{"og:url":78,"og:type":120,"og:title":59,"og:site_name":91,"og:description":61},"article",{"robots":122,"canonical":78},"index,follow",{"doc_id":124,"site_id":56},238597,1789138336,{"code":4,"msg":5,"data":127},{"doc_id":124,"user_id":128,"nickname":88,"user_avatar":129,"doc_module":9,"category_id":50,"category_name":51,"doc_title":59,"doc_description":61,"doc_content":130,"file_id":131,"file_url":132,"file_type":133,"file_size":134,"view_count":79,"is_deleted":4,"is_public":9,"is_downloadable":9,"audit_status":9,"page_count":30,"language":135,"language_code":57,"site_id":56,"html_lang":57,"table_of_contents":136,"faqs":137,"seo_title":138,"seo_description":61,"update_tm":125,"read_time":139},2336477552062,"https://ap-avatar.wpscdn.com/davatar_994ba38a5ba835b3df7d355c54d3ed8d","AutoTransition: Learning to Recommend Video Transition Eﬀects  \nYaojie Shen1,2,3 , Libo Zhang1,2 , Kai Xu3 , and Xiaojie Jin3(B)  \n1 Institute of Software, Chinese Academy of Sciences, Beijing, China  \n2 University of Chinese Academy of Sciences, Beijing, China  \n3 ByteDance Inc. , Beijing, China  \n[jinxiaojie@bytedance.com](jinxiaojie@bytedance.com)  \nAbstract. Video transition eﬀects are widely used in video editing to connect shots for creating cohesive and visually appealing videos. However, it is challenging for non-professionals to choose best transitions due to the lack of cinematographic knowledge and design skills. In this paper, we present the premier work on performing automatic video transitions recommendation (VTR): given a sequence of raw video shots and companion audio, recommend video transitions for each pair of neighboring shots. To solve this task, we collect a large-scale video transition dataset using publicly available video templates on editing softwares. Then we formulate VTR as a multi-modal retrieval problem from vision/audio to video transitions and propose a novel multi-modal matching framework which consists of two parts. First we learn the embedding of video transitions through a video transition classiﬁcation task. Then we propose a model to learn the matching correspondence from vision/audio inputs to video transitions. Speciﬁcally, the proposed model employs a multi-modal transformer to fuse vision and audio information, as well as capture the context cues in sequential transition outputs. Through both quantitative and qualitative experiments, we clearly demonstrate the eﬀectiveness of our method. Notably, in the comprehensive user study, our method receives comparable scores compared with professional editors while improving the video editing eﬃciency by 300 × . We hope our work serves to inspire other researchers to work on this new task. The dataset and codes are public at [https://github.com/acherstyx/](https://github.com/acherstyx/)[ ](https://github.com/acherstyx/)AutoTransition.  \nKeywords: Video transition eﬀects recommendation · Multi-modal retrieval · Video editing  \nY. Shen, L. Zhang, K. Xu and X. Jin—Equal contribution.  \nSupplementary Information The online version contains supplementary material available at [https://doi.org/10.1007/978-3-031-19839-7](https://doi.org/10.1007/978-3-031-19839-7)   17.  \n􀀂c The Author(s), under exclusive license to Springer Nature Switzerland AG 2022  \nS. Avidan et al. (Eds.): ECCV 2022, LNCS 13698, pp. 285–300, 2022 .  \n[https://doi.org/10.1007/978-3-031-19839-7](https://doi.org/10.1007/978-3-031-19839-7_17)[_](https://doi.org/10.1007/978-3-031-19839-7_17)[17](https://doi.org/10.1007/978-3-031-19839-7_17)  \n286 Y. Shen et al.  \n1 Introduction  \nWith the advance of multimedia technology and network infrastructures, video is ubiquitous, occurring in numerous everyday activities such as education, entertainment, surveillance, etc. There is a massive amount of needs for people to edit videos and share with others. However, video editing is challenging for nonprofessionals since it is not only laborious but also needs a lot of cinematography and design knowledge. Some editing tools like Adobe Premier and Apple Final Cut Pro are developed to assist video editing, however their main target users are professionals while novices may ﬁnd it diﬃcult to learn. Moreover, they still lack the ability of automatic video editing, i.e., users have to manipulate videoson their own. Recently, popular video editing tools like InShot Video Editor and CapCut provide the function of creating videos in one-click. Nevertheless, since they only utilize simple strategies or ﬁxed video templates and ignore the content of input vision/audio, the quality of generated video is unsatisfactory.  \nThe video transition eﬀects play an important role in video editing to join shots for creating smooth and cohesive videos. In this paper, we introduce anew task of automatic video transitio","cbCaibURiaBHKm0c","https://ap.wps.com/l/cbCaibURiaBHKm0c","pdf",986610,"English","# Introduction\n## Problem Definition and Motivation\n## Challenges in Dataset Creation and Evaluation\n## Proposed Approach Overview\n## Paper Organization","[{\"question\":\"What is the video transition recommendation task (VTR) defined as?\",\"answer\":\"Given a sequence of raw video shots and companion audio, VTR recommends a sequence of transitions for each neighboring shot pair, aiming to provide ranked candidate transitions rather than a single class.\"},{\"question\":\"How is the training dataset for video transitions obtained?\",\"answer\":\"A large-scale dataset is collected from publicly available video templates on editing software, using comprehensive selection rules and preprocessing to refine high-quality samples.\"},{\"question\":\"What model design is used to match vision/audio inputs to transitions?\",\"answer\":\"The method learns transition embeddings via a transition classification task, then uses a multi-modal transformer to fuse vision and audio information and capture contextual cues across sequential transition outputs.\"}]","AutoTransition - Learning to Recommend Video Transition Effects | PDF",6]