[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-81620-en":3,"doc-seo-81620-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},81620,34359740700684,"Finn","https://ap-avatar.wpscdn.com/avatar/1f400023980c374ae676?_k=1777273430885731487",8,"Research & Report","Test-Time Adaptation via Cache Personalization for Facial Expression Recognition in Videos","Facial expression recognition (FER) in videos demands model personalization to handle large inter-subject variations and distribution shifts that can degrade vision-language model transfer. While test-time adaptation (TTA) improves robustness, many methods depend on unsupervised parameter optimization, creating compute costs unsuitable for real deployment. The paper proposes TTA through Cache Personalization (TTA-CaP), a gradient-free VLM personalization approach using three complementary caches and a tri-gate update policy. Experiments on BioVid, StressID, and BAH show consistent gains under subject and environmental shifts with low memory and compute overhead.","Test-Time Adaptation via Cache Personalization for Facial Expression  \nRecognition in Videos  \nMasoumeh Sharafi 1 , Muhammad Osama Zeeshan 1 , Soufiane Belharbi 1 , Alessandro Lameiras Koerich2 , Marco Pedersoli 1 , Eric Granger 1  \n1LIVIA, Department of Systems Engineering, ´ETS Montreal, Canada  \n2LIVIA, Department of Software and IT Engineering, ´ETS Montreal, Canada  \n{masoumeh.sharafi.1,[muhammad-osama.zeeshan.1](muhammad-osama.zeeshan.1}@ens.etsmtl.ca)[}](muhammad-osama.zeeshan.1}@ens.etsmtl.ca)[@ens.etsmtl.ca](muhammad-osama.zeeshan.1}@ens.etsmtl.ca)[ ](muhammad-osama.zeeshan.1}@ens.etsmtl.ca){soufiane.belharbi,marco.pedersoli,alessandro.koerich,[eric.granger](eric.granger}@etsmtl.ca)[}](eric.granger}@etsmtl.ca)[@etsmtl.ca](eric.granger}@etsmtl.ca)  \narXiv :2603 .2 1309v 3 [ cs .CV] 9 Jul 2026  \nAbstract  \nFacial expression recognition (FER) in videos requires model personalization to capture the considerable variations across subjects. Vision-language models (VLMs) offer strong transfer to downstream tasks through image-text alignment, but their performance can still degrade under inter-subject distribution shifts. Personalizing models using test-time adaptation (TTA) methods can mitigate this challenge. However, most state-of-the-art TTA methods rely on unsupervised parameter optimization, introducing computational overhead that is impractical in many real-world applications. This paper introduces TTA through Cache Personalization (TTA-CaP), a cache-based TTA method that enables cost-effective (gradient-free) personalization of VLMs for video FER. Prior cache-based TTA methods rely solely on dynamic memories that store test samples, which can accumulate errors and drift due to noisy pseudolabels. To address this limitation, TTA-CaP introduces three complementary caches – a personalized static cache constructed through feature-statistics matching, a positive target cache that accumulates reliable subject-specific samples, and a negative target cache that stores low-confidence cases as negative samples. To prevent target-cache corruption, a tri-gate mechanism controls cache updates based on temporal stability, confidence, and consistency with the personalized static cache. Together, the caches provide complementary, subject-matched positive and negative evidence for robust online personalization. Finally, TTA-CaP refines predictions by fusing embeddings, yielding representations that support temporally stable video-level predictions. Our experiments 1 on three challenging video FER datasets—BioVid, StressID, and BAH—indicate that  \n[1](1 Our code is publicly available at: github.com/MasoumehSharafi/TTA)[ Our code is publicly available at: github.com/MasoumehSharafi/TTA](1 Our code is publicly available at: github.com/MasoumehSharafi/TTA)CaP.  \nTTA-CaP can outperform state-of-the-art TTA methods under subject-specific and environmental shifts, while maintaining low computational and memory overhead for realworld deployment.  \n1. Introduction  \nVideo-based FER is a key component of various affective computing applications, such as human–computer interaction [29], health monitoring [11], and clinical assessment of pain, depression, and stress [3] . Unlike image-based FER, video FER must account for both facial appearance and the temporal evolution of expressions. Expressions may develop gradually, remain subtle across several frames, or involve brief transitions between neutral and expressive states. Consequently, individual frames are not equally informative, and predictions may fluctuate within the same video. These challenges are further amplified in subjectindependent settings due to variations in facial morphology, expression intensity, behavioral patterns, and acquisition conditions. As a result, models trained on a fixed group of subjects may fail to capture the expression characteristics of previously unseen individuals. This motivates the use of large-scale pre-trained models that can transfer effectively to FER w","cbCaijdiRn74rEKw","https://ap.wps.com/l/cbCaijdiRn74rEKw","pdf",5233049,5,1,17,"English","en",105,"# Abstract\n# Introduction\n## Video-based FER challenges\n## Vision-language transfer and domain shift\n## Motivation for personalization and TTA","[{\"question\":\"What problem does TTA-CaP address in video facial expression recognition?\",\"answer\":\"TTA-CaP targets performance degradation under inter-subject and acquisition-condition shifts by personalizing the vision-language model at test time for each subject video.\"},{\"question\":\"Why are many existing TTA methods impractical for real-world use?\",\"answer\":\"Many state-of-the-art TTA approaches rely on unsupervised parameter optimization, which introduces computational overhead.\"},{\"question\":\"How does TTA-CaP perform cache-based personalization without gradients?\",\"answer\":\"TTA-CaP builds a personalized static cache from feature-statistics matching and maintains positive and negative target caches using reliable and low-confidence samples, while a tri-gate mechanism controls updates to prevent cache corruption.\"}]",1784174868,43,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"test-time-adaptation-via-cache-personalization-for-facial-expression-recognition-in-videos","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/test-time-adaptation-via-cache-personalization-for-facial-expression-recognition-in-videos/81620/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-26","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What problem does TTA-CaP address in video facial expression recognition?","Question",{"text":76,"@type":77},"TTA-CaP targets performance degradation under inter-subject and acquisition-condition shifts by personalizing the vision-language model at test time for each subject video.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"Why are many existing TTA methods impractical for real-world use?",{"text":81,"@type":77},"Many state-of-the-art TTA approaches rely on unsupervised parameter optimization, which introduces computational overhead.",{"name":83,"@type":74,"acceptedAnswer":84},"How does TTA-CaP perform cache-based personalization without gradients?",{"text":85,"@type":77},"TTA-CaP builds a personalized static cache from feature-statistics matching and maintains positive and negative target caches using reliable and low-confidence samples, while a tri-gate mechanism controls updates to prevent cache corruption.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":20,"slug":138},19,"General","general"]