[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82821-en":3,"doc-seo-82821-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},82821,5909877438554,"Maeve","https://ap-avatar.wpscdn.com/avatar/5600025385ad2bf12a7?_k=1778553567797529272",8,"Research & Report","UniSkip Mamba Frequency-Aware State Space Model for Audio-Visual Temporal Forgery Localization","Audio-Visual Temporal Forgery Localization (AV-TFL) is increasingly urgent as AI-generated multimedia enables evidence fabrication and opinion manipulation. Existing approaches that use transformer temporal modeling or channel-wise multimodal fusion treat all frequency components equally, overfitting to high-frequency noise and losing robustness under real-world degradation. Frequency-domain ablation shows forgery-discriminative patterns concentrate in low/mid bands (0–0.15), while high-frequency parts mainly add noise; removing them improves detection. UniSkip-Mamba integrates unified multimodal sequence fusion and skip-scanning state space blocks with group-scan-merge regularization, achieving SOTA results on LAV-DF and AV-Deepfake1M with faster inference.","UniSkip-Mamba: A Frequency-Aware State Space Model for Audio-Visual Temporal Forgery Localization  \nCangjin Qiu∗ , Quan Zhang∗ , Dan Jiang, and Ke Zhang†  \narXiv :2607 .04498v 1 [ cs .CV] 5 Jul 2026  \nAbstract—With the proliferation of AI-generated content, sophisticated multimedia manipulation has raised critical concerns about malicious applications such as opinion manipulation and evidence fabrication, making Audio-Visual Temporal Forgery Localization (AV-TFL) an urgent research frontier. Existing TFL methods have progressed along two main paradigms: Transformer-based temporal modeling and channel-wise multimodal fusion. While these approaches capture temporal dependencies and cross-modal correlations, they process all frequency components indiscriminately, leading to overfitting on highfrequency noise and limited robustness under real-world data degradation. Through systematic frequency domain analysis, we find that forgery-discriminative patterns concentrate in the low/mid-frequency range (normalized frequency 0–0.15), while high-frequency components primarily introduce noise—removing them even improves detection performance by +1.4% . Based on this phenomenon, we propose UniSkip-Mamba, a frequencyaware State Space Model framework that incorporates Unified Multimodal Sequence Fusion to preserve cross-modal phase relationships, and Skip-Scanning Mamba Blocks that implement frequency-aware regularization through a novel Group-ScanMerge mechanism, naturally biasing learning toward discriminative low/mid-frequency patterns (0–0.15) while maintaining representational completeness. We achieve state-of-the-art (SOTA) performance: 63.4% AP@0.95 on LAV-DF (+9.8% improvement) and 63.58% mAP on AV-Deepfake1M (+14.32% improvement), with 6× faster inference. Our frequency-domain analysis provides theoretical justification from a signal processing perspective for why skip-scanning inherently improves both accuracy and robustness.  \nIndex Terms—Temporal forgery localization, Mamba, Unified sequence.  \nI. INTRODUCTION  \nThe proliferation of Artificial Intelligence Generated Content (AIGC) [1]–[3] has enabled sophisticated multimedia manipulation, raising critical concerns about malicious applications such as opinion manipulation and evidence fabrication. While the multimedia forensics community has made significant progress in Deepfake detection [4]–[8], existing methods predominantly rely on binary classification, failing to identify where manipulation occurs temporally. This limitation severely constrains their utility in practical scenarios such as judicial forensics and content moderation, where precise temporal boundaries of forged segments are indispensable. Consequently, Audio-Visual Temporal Forgery Localization (AV-TFL) has emerged as a critical research frontier.  \nBefore addressing the technical challenges of AV-TFL,  \nwe pose a fundamental question: In which frequency bands ∗ Cangjin Qiu and Quan Zhang contributed equally to this work.  \n†Corresponding author: Ke Zhang.  \nCangjin Qiu and Ke Zhang are with Soochow University, Suzhou, China. Quan Zhang and Dan Jiang are with Tsinghua University, Beijing, China.  \nFig. 1. Motivating Observation: Frequency Band Contribution Analysis. We systematically remove each frequency band from input features and measure performance impact (mAP@0.95) . Key findings: (1) Low/midfrequency bands (0–0.15) are critical for detection, with removal causing up to 92.8% performance drop; (2) High-frequency removal (> 0. 15) has minimal impact (−0 .7% to +1 .4%), with stride-4 models actually improving—confirming that high frequencies introduce noise rather than discriminative information. This motivates our Skip-Scanning design as a frequency-aware regularization strategy. Comprehensive experimental setup and detailed analysis in Section IV-G.  \ndoes discriminative forgery information reside? Through spectral analysis and systematic ablation experiments, we reveal that forgery-discriminative patterns","cbCaivLY6wtfYYKi","https://ap.wps.com/l/cbCaivLY6wtfYYKi","pdf",2868180,1,11,"English","en",105,"# Introduction\n## Frequency-band contribution analysis\n## Motivation and challenges in existing methods\n## Proposed framework and contributions","[{\"question\":\"What problem does AV-TFL address, and why is it important?\",\"answer\":\"AV-TFL localizes the temporal segments where audio-visual content is forged. It is crucial for practical use in judicial forensics and content moderation where precise temporal boundaries are required.\"},{\"question\":\"Which frequency bands contain the most forgery-discriminative information?\",\"answer\":\"Forgery-discriminative patterns concentrate in the low/mid-frequency range with normalized frequency 0–0.15. Removing these bands causes large performance drops, while removing high frequencies (\\u003e0.15) has minimal impact and can even improve results.\"},{\"question\":\"How does UniSkip-Mamba improve accuracy and robustness for AV-TFL?\",\"answer\":\"UniSkip-Mamba uses unified multimodal sequence fusion to preserve cross-modal phase relationships and skip-scanning Mamba blocks with a group-scan-merge mechanism to apply frequency-aware regularization. This biases learning toward discriminative 0–0.15 patterns while filtering high-frequency noise.\"}]",1784183191,28,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"uniskip-mamba-frequency-aware-state-space-model-for-audio-visual-temporal-forgery-localization","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/uniskip-mamba-frequency-aware-state-space-model-for-audio-visual-temporal-forgery-localization/82821/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does AV-TFL address, and why is it important?","Question",{"text":75,"@type":76},"AV-TFL localizes the temporal segments where audio-visual content is forged. It is crucial for practical use in judicial forensics and content moderation where precise temporal boundaries are required.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"Which frequency bands contain the most forgery-discriminative information?",{"text":80,"@type":76},"Forgery-discriminative patterns concentrate in the low/mid-frequency range with normalized frequency 0–0.15. Removing these bands causes large performance drops, while removing high frequencies (>0.15) has minimal impact and can even improve results.",{"name":82,"@type":73,"acceptedAnswer":83},"How does UniSkip-Mamba improve accuracy and robustness for AV-TFL?",{"text":84,"@type":76},"UniSkip-Mamba uses unified multimodal sequence fusion to preserve cross-modal phase relationships and skip-scanning Mamba blocks with a group-scan-merge mechanism to apply frequency-aware regularization. This biases learning toward discriminative 0–0.15 patterns while filtering high-frequency noise.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":45,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":45,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":45,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]