[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-128110-en":3,"doc-seo-128110-105":31,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":28,"seo_description":14,"update_tm":29,"read_time":30},128110,687207022233,"Riley","https://ap-avatar.wpscdn.com/davatar_155a257f0dc6eb9ab79c44ca47cae57d",8,"Research & Report","ARTEMIS: animal recognition through enhanced multimodal integration system","ARTEMIS introduces an Animal Recognition Through Enhanced Multimodal Integration System: a transformer-based framework for multilabel animal action recognition that fuses video, image, and text. The system generates textual descriptions from video frames using BLIP2 and Llama 3, then feeds these descriptions into the model to boost accuracy. Ablation studies analyze model components and propose optimization of ensemble weights via genetic algorithms and reinforcement learning. Contrastive and cosine-similarity feature alignment further improve multimodal integration, reaching 79.82 mAP on Animal Kingdom.","International Journal of Machine Learning and Cybernetics [https://doi.org/10.1007/s13042-025-02602-3](https://doi.org/10.1007/s13042-025-02602-3)  \nARTEMIS: animal recognition through enhanced multimodal integration system  \nEdoardo Fazzari1,2,3 · Donato Romano1,2 · Fabrizio Falchi1,3 · Cesare Stefanini1,2  \nReceived: 28 October 2024 / Accepted: 3 March 2025 © The Author(s) 2025  \nAbstract  \nThis paper introduces Animal Recognition Through Enhanced Multimodal Integration System (ARTEMIS), a transformerbased framework designed for multilabel animal action recognition by fusing video, image, and textual modalities. ARTEMIS utilizes state-of-the-art captioning and language models, such as BLIP2 and Llama 3, to generate textual descriptions from video frames, which are input to the model, significantly enhancing its performance unlikely previous results that do not consider this modality. Through comprehensive ablation studies, we explore the contribution of various model components and propose optimization strategies, including genetic algorithms and reinforcement learning, to dynamically adjust ensemble weights. Our feature alignment techniques-using contrastive and cosine similarity losses-further improve multimodal integration. Evaluations on the Animal Kingdom dataset, which includes 30,100 clips across 140 action classes, demonstrate that ARTEMIS achieves a new state-of-the-art mAP of 79.82, outperforming existing methods. The combination of multimodal fusion and ensemble strategies makes ARTEMIS a robust solution for complex animal action recognition tasks. The code of our fusion method is available at [https://github.com/edofazza/ARTEMIS](https://github.com/edofazza/ARTEMIS).  \nKeywords Animal action recognition · Multi-modal deep learning · Ensemble · Genetic algorithm · Reinforcement learning · Contrastive learning  \n1 Introduction  \nAnimals exhibit a wide range of behaviors, from affectionate and defensive actions to feeding, movement, and aggression. These behaviors are of great interest in ethological studies [1], as they can be indicative of an animal’s health and  \n* Edoardo Fazzari [edoardo.fazzari@santannapisa.it](edoardo.fazzari@santannapisa.it)  \nDonato Romano  \n[donato.romano@santannapisa.it](donato.romano@santannapisa.it)  \nFabrizio Falchi  \n[fabrizio.falchi@cnr.it](fabrizio.falchi@cnr.it)  \nCesare Stefanini  \n[cesare.stefanini@santannapisa.it](cesare.stefanini@santannapisa.it)  \n1 The BioRobotics Institute, Sant’Anna School of Advanced Studies, Viale Rinaldo Piaggio, 56025 Pontedera, Italy  \n2 Department of Excellence in Robotics and AI, Sant’Anna School of Advanced Studies, Piazza Martiri della Libertà, 56127 Pisa, Italy  \n3 Institute of Information Science and Technologies, National Research Council of Italy, via G. Moruzzi, 56124 Pisa, Italy  \nwell-being [2] . By associating specific actions with health conditions, researchers can potentially identify pathological behaviors or detect anomalies in wildlife and livestock [3, 4] . This is particularly valuable for environmental monitoring and improving animal welfare [5] .  \nIn pursuit of understanding these behaviors, researchers have published datasets [6] and developed methodologies for extracting meaningful information from sensor data, images, and videos [7, 8] . However, much of this research has been limited to individual species and a narrow set of actions [9, 10]. The release of the Animal Kingdom [11] dataset marked a significant shift in this field. Animal Kingdom introduced a dataset encompassing a wide variety of animal species and reorganized actions in a way that allows for cross-species application. The hope is that models trained on this benchmark dataset will generalize across species, paving the way for the development of a foundational model for animal action recognition. Such a model could have widespread applications in agriculture [12],(neuro-)ethology [6], and other research fields, significantly reducing the need for retraining or fi","cbCaifqgNGw3PzEr","https://ap.wps.com/l/cbCaifqgNGw3PzEr","pdf",2683364,2,1,16,"English","en",105,"# Abstract\n# Introduction\n## Problem background and motivation\n## Datasets and limitations of existing methods\n## Proposed approach: ARTEMIS and fusion strategy\n## Ensemble optimization with genetic algorithms and reinforcement learning","[{\"question\":\"What does ARTEMIS aim to recognize and how is it structured?\",\"answer\":\"ARTEMIS targets multilabel animal action recognition and is built as a transformer-based framework with two fusion stages across video, image, and text modalities.\"},{\"question\":\"How does ARTEMIS use textual information from video?\",\"answer\":\"It employs captioning and language models (e.g., BLIP2 and Llama 3) to generate textual descriptions from video frames, which are then used as inputs to enhance recognition performance.\"},{\"question\":\"How are ensemble weights optimized in ARTEMIS?\",\"answer\":\"ARTEMIS addresses weighted ensemble selection using genetic algorithms and reinforcement learning to dynamically adjust ensemble weights based on performance.\"}]","ARTEMIS: animal recognition through enhanced multimodal integration system | PDF",1785944884,40,{"code":4,"msg":32,"data":33},"ok",{"site_id":25,"language":24,"slug":34,"title":13,"keywords":35,"description":14,"schema_data":36,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":29},"artemis-animal-recognition-through-enhanced-multimodal-integration-system","",{"@graph":37,"@context":86},[38,54,69],{"@type":39,"itemListElement":40},"BreadcrumbList",[41,45,48,51],{"item":42,"name":43,"@type":44,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":46,"name":47,"@type":44,"position":20},"https://docshare.wps.com/document/","Document",{"item":49,"name":12,"@type":44,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":44,"position":53},"https://docshare.wps.com/document/artemis-animal-recognition-through-enhanced-multimodal-integration-system/128110/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":42,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-28","2026-08-05",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What does ARTEMIS aim to recognize and how is it structured?","Question",{"text":76,"@type":77},"ARTEMIS targets multilabel animal action recognition and is built as a transformer-based framework with two fusion stages across video, image, and text modalities.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How does ARTEMIS use textual information from video?",{"text":81,"@type":77},"It employs captioning and language models (e.g., BLIP2 and Llama 3) to generate textual descriptions from video frames, which are then used as inputs to enhance recognition performance.",{"name":83,"@type":74,"acceptedAnswer":84},"How are ensemble weights optimized in ARTEMIS?",{"text":85,"@type":77},"ARTEMIS addresses weighted ensemble selection using genetic algorithms and reinforcement learning to dynamically adjust ensemble weights based on performance.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":47,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":47,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":47,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":47,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":47,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":47,"category_name":118,"show_sort_weight":30,"slug":119},7,"Healthcare","healthcare",{"id":11,"doc_module":4,"doc_module_name":47,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":47,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":47,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":47,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":47,"category_name":137,"show_sort_weight":107,"slug":138},19,"General","general"]