[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-117456-en":3,"doc-seo-117456-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},117456,1374391975076,"Riley","https://ap-avatar.wpscdn.com/avatar/14000253ca4ec9f6853?x-image-process=image/resize,m_fixed,w_180,h_180&k=1783305029341752051",8,"Research & Report","Taylor Videos for Action Recognition - Proceedings of the 41st International Conference on Machine Learning","Effectively extracting motions from video is critical for action recognition but remains challenging due to the lack of explicit motion form, the presence of multiple motion concepts like displacement, velocity, and acceleration, and noise from unstable pixels. This work introduces the Taylor video format with Taylor frames that emphasize dominant motions and suppress static and small unstable changes via Taylor expansion on temporal blocks. Experiments show strong results with 2D CNNs, 3D CNNs, transformers, and fused improvements with RGB/optical flow, plus enhanced performance on skeleton-based action recognition using Taylor skeleton sequences.","Taylor Videos for Action Recognition  \nAuthor  \nWang , L , Yuan , X , Gedeon , T, Zheng , L  \nPublished 2024  \nConference Title  \nProceedings of the 41st International Conference on Machine Learning  \nVersion  \nVersion of Record (VoR)  \nRights statement  \nThis work is covered by copyright. You must assume that re-use is limited to personal use and that permission from the copyright owner must be obtained for all other uses. If the document is available under a specified licence , refer to the licence for details of permitted re-use. If you believe that this work infringes copyright please make a copyright takedown request using the form at [https://www.griffith.edu.au/copyright-matters](https://www.griffith.edu.au/copyright-matters).  \nDownloaded from  \n[https://hdl.handle.net/10072/436208](https://hdl.handle.net/10072/436208)  \nLink to published version  \n[https://proceedings.mlr.press/v235/wang24ck.html](https://proceedings.mlr.press/v235/wang24ck.html)  \nGriffith Research Online  \n[https://research-repository.griffith.edu.au](https://research-repository.griffith.edu.au)  \nTaylor Videos for Action Recognition  \nLei Wang * 1 Xiuyuan Yuan * 1 Tom Gedeon 2 Liang Zheng 1  \nAbstract  \nEffectively extracting motions from video is a critical and long-standing problem for action recognition. This problem is very challenging because motions (i) do not have an explicit form,(ii) have various concepts such as displacement, velocity, and acceleration, and (iii) often contain noise caused by unstable pixels. Addressing these challenges, we propose the Taylor video, a new video format that highlights the dominant motions (e.g., a waving hand) in each of its frames named the Taylor frame. Taylor video is named after Taylor series, which approximates a function at a given point using important terms. In the scenario of videos, we define an implicit motionextraction function which aims to extract motions from video temporal blocks. In these blocks, using the frames, the difference frames, and higherorder difference frames, we perform Taylor expansion to approximate this function at the starting frame. We show the summation of the higherorder terms in the Taylor series gives us dominant motion patterns, where static objects, small and unstable motions are removed. Experimentally, we show that Taylor videos are effective inputs to popular architectures including 2D CNNs, 3D CNNs, and transformers. When used individually, Taylor videos yield competitive action recognition accuracy compared to RGB videosand optical flow. When fused with RGB or optical flow videos, further accuracy improvement is achieved. Additionally, we apply Taylor video computation to human skeleton sequences, resulting in Taylor skeleton sequences that outperform the use of original skeletons for skeletonbased action recognition. Code is available at: [https://github.com/LeiWangR/video-ar](https://github.com/LeiWangR/video-ar).  \n*Equal contribution 1 School of Computing, Australian National University, Canberra, Australia 2 School of Electrical Engineering, Computing and Mathematical Sciences, Curtin University, Perth, Australia. Correspondence to: Lei Wang \u003C[lei.w@anu.edu.au](lei.w@anu.edu.au) >.  \nProceedings of the 41 st International Conference on Machine Learning, Vienna, Austria. PMLR 235, 2024 . Copyright 2024 by the author(s) .  \nFigure 1 . Visualizing different video formats. (Top left): RGB video and time-color reordering frames (Kim et al., 2022) . (Bottom left): U and V components of optical flow. (Right): proposed Taylor video frames. Taylor frames clearly (i) remove static objects and unstable motions and (ii) highlight motions.  \n1. Introduction  \nExtracting motions and thus recognizing actions from videos are important problems. Extensive efforts have been made in improving video inputs (Kim et al., 2022 ; Bilen et al., 2016 ; 2018 ; Wang & Koniusz, 2024), optimizing networks (Carreira & Zisserman, 2018 ; Wang et al., 2018 ; 2023b ; Lin et al., 2019), and t","cbCaibK7JzxbF9AP","https://ap.wps.com/l/cbCaibK7JzxbF9AP","pdf",3228390,1,18,"English","en",105,"# Introduction\n## Motivation and challenges in motion extraction\n## Proposed implicit motion modeling and Taylor expansion\n## Taylor video and Taylor frames overview\n# Technical approach\n## Computing Taylor frames from temporal blocks\n## Taylor skeleton sequences for action recognition","[{\"question\":\"What is the main challenge in extracting motion for action recognition?\",\"answer\":\"Motions have no explicit form, involve multiple related concepts (displacement, velocity, acceleration), and are often corrupted by noise from unstable pixels.\"},{\"question\":\"How does the proposed Taylor video represent motion?\",\"answer\":\"It introduces Taylor frames that encode dominant motion patterns and remove static objects and unstable small motions using Taylor expansion within temporal blocks.\"},{\"question\":\"How are Taylor videos evaluated against common input types?\",\"answer\":\"Taylor videos work effectively as inputs to 2D CNNs, 3D CNNs, and transformers, performing competitively alone, and improving further when fused with RGB or optical flow videos.\"}]","Taylor Videos for Action Recognition - Proceedings of the 41st International Conference on Machine Learning | PDF",1785675948,45,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"taylor-videos-for-action-recognition-proceedings-of-the-41st-international-conference-on-machine-learning","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/taylor-videos-for-action-recognition-proceedings-of-the-41st-international-conference-on-machine-learning/117456/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-02",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What is the main challenge in extracting motion for action recognition?","Question",{"text":75,"@type":76},"Motions have no explicit form, involve multiple related concepts (displacement, velocity, acceleration), and are often corrupted by noise from unstable pixels.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the proposed Taylor video represent motion?",{"text":80,"@type":76},"It introduces Taylor frames that encode dominant motion patterns and remove static objects and unstable small motions using Taylor expansion within temporal blocks.",{"name":82,"@type":73,"acceptedAnswer":83},"How are Taylor videos evaluated against common input types?",{"text":84,"@type":76},"Taylor videos work effectively as inputs to 2D CNNs, 3D CNNs, and transformers, performing competitively alone, and improving further when fused with RGB or optical flow videos.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]