[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84753-en":3,"doc-seo-84753-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},84753,4398048950312,"Violet","https://ap-avatar.wpscdn.com/avatar/400002538284de19e3c?_k=1778320343897328908",8,"Research & Report","CompressedVQA-AEV Full-Reference and No-Reference Quality Assessment Models for Asymmetric Encoded Videos","This report presents solutions to the QoMEX 2026 Grand Challenge on Video Quality Assessment for Asymmetric Encoded Videos, proposing a full-reference model CompressedVQA-AEV-FR and a no-reference model CompressedVQA-AEV-NR. The FR system uses a SwinB backbone to compute multi-stage texture and structure similarity statistics between reference and distorted videos, then aggregates frame estimates by temporal mean pooling. The NR system combines SigLIP2 and Swin-B frame encoders with cross-fold ensembling to predict perceptual quality without reference data. Results rank first on the FR track and fourth on the NR track.","CompressedVQA-AEV: Full-Reference and No-Reference Quality Assessment Models for Asymmetric Encoded Videos  \nWei Sun 1 , Xingwei Liu 1 , Dandan Zhu 1 , Xiangyang Zhu3 , Weixia Zhang2 , Guangtao Zhai2  \n1East China Normal University, 2 Shanghai Jiao Tong University  \n3 Shanghai Artificial Intelligence Laboratory  \narXiv :2607 .04606v1 [ ee ss .IV] 6 Jul 2026  \nAbstract—This report presents our solutions to the QoMEX 2026 Grand Challenge on Video Quality Assessment for Asymmetric Encoded Videos, comprising a full-reference (FR) model, CompressedVQA-AEV-FR, and a no-reference (NR) model, CompressedVQA-AEV-NR. The FR approach leverages a SwinB backbone to extract multi-stage similarity statistics between reference and distorted videos for quality prediction. For the NR setting, our model employs complementary frame-level encoders based on SigLIP2 and Swin-B, followed by temporal mean pooling and cross-fold ensembling to estimate perceptual quality without reference data. Our CompressedVQA-AEV-FR achieves first place in the FR track of QoMEX 2026 Grand Challenge, while CompressedVQA-AEV-NR secures fourth place in the NR track, demonstrating the effectiveness of our proposed models. The code is available at [https://github.com/sunwei925/](https://github.com/sunwei925/)[ ](https://github.com/sunwei925/)CompressedVQA-AEV.  \nI. INTRODUCTION  \nVideo quality assessment (VQA) [1]–[6] is essential for optimizing modern video encoding and streaming systems. Traditional video quality metrics such as VMAF [7] and ITUT P.1204.3 [8] have been primarily developed and validatedon conventionally (symmetrically) encoded content, where all spatial regions are treated with uniform encoding parameters.  \nWith the advancement of semantic segmentation and visualsaliency prediction techniques, Region-of-Interest (ROI) encoding [9] has emerged as a promising strategy that allocates higher bitrates to perceptually important regions while compressing background areas more aggressively. Such asymmetric encoding can achieve significant bitrate savings or improve perceived quality at the same bitrate. However, the reliability of existing VQMs on asymmetrically encoded videos—where spatial quality varies substantially across regions—remains largely underexplored.  \nTo promote research in this direction, the QoMEX 2026 Grand Challenge on Video Quality Assessment for Asymmetric Encoded Videos [10] was organized, providing the SportROI dataset containing both symmetrically and asymmetrically encoded sports videos with subjective quality annotations. Participants were invited to develop full-reference (FR) and no-reference (NR) models that generalize across both encoding paradigms.  \nIn this report, we present our solutions to both tracks. For the FR track, we propose CompressedVQA-AEV-FR,  \nwhich extracts multi-stage hierarchical features from a Swin Transformer [11] backbone and computes per-stage texture and structure similarity descriptors between reference and distorted frames. For the NR track, we propose CompressedVQA-AEVNR, which employs complementary SigLIP2 [12] and SwinB [11] encoders with cross-fold ensembling to predict perceptual quality without reference information. Our FR model achieves first place in the FR track, and our NR model secures fourth place in the NR track, demonstrating the effectiveness of the proposed approaches.  \nII. MODEL ARCHITECTURE  \n1) CompressedVQA-AEV-FR: CompressedVQA-AEV-FR is a full-reference video quality assessment framework designed to predict the perceptual quality of compressed videos by leveraging both reference and distorted content. The framework follows a multi-stage feature similarity paradigm [13]–  \n[15]: for each pair of reference and distorted frames, deep hierarchical features are independently extracted through apretrained visual backbone, and per-stage similarity descriptors capturing texture and structure fidelity are computed and concatenated into a unified quality-aware representation. A lightweight re","cbCaiu2XNUDRpgpQ","https://ap.wps.com/l/cbCaiu2XNUDRpgpQ","pdf",191096,1,5,"English","en",105,"# Introduction\n# Model Architecture\n## CompressedVQA-AEV-FR\n## Multi-Stage Feature Extraction\n## Per-Stage Similarity Computation","[{\"question\":\"What are the two proposed video quality assessment models in the report?\",\"answer\":\"The report proposes CompressedVQA-AEV-FR (full-reference) and CompressedVQA-AEV-NR (no-reference) for assessing quality of asymmetric encoded videos.\"},{\"question\":\"How does the full-reference model CompressedVQA-AEV-FR estimate quality?\",\"answer\":\"It extracts multi-stage hierarchical features with a SwinB backbone, computes per-stage texture and structure similarity descriptors between reference and distorted frames, uses a regression head for frame quality, and applies temporal mean pooling across sampled frames.\"},{\"question\":\"How does the no-reference model CompressedVQA-AEV-NR work without reference videos?\",\"answer\":\"It employs complementary frame-level encoders based on SigLIP2 and Swin-B, then performs temporal mean pooling and cross-fold ensembling to estimate perceptual quality without using reference information.\"}]",1784198045,13,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"compressedvqa-aev-full-reference-and-no-reference-quality-assessment-models-for-asymmetric-encoded-videos","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/compressedvqa-aev-full-reference-and-no-reference-quality-assessment-models-for-asymmetric-encoded-videos/84753/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What are the two proposed video quality assessment models in the report?","Question",{"text":75,"@type":76},"The report proposes CompressedVQA-AEV-FR (full-reference) and CompressedVQA-AEV-NR (no-reference) for assessing quality of asymmetric encoded videos.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the full-reference model CompressedVQA-AEV-FR estimate quality?",{"text":80,"@type":76},"It extracts multi-stage hierarchical features with a SwinB backbone, computes per-stage texture and structure similarity descriptors between reference and distorted frames, uses a regression head for frame quality, and applies temporal mean pooling across sampled frames.",{"name":82,"@type":73,"acceptedAnswer":83},"How does the no-reference model CompressedVQA-AEV-NR work without reference videos?",{"text":84,"@type":76},"It employs complementary frame-level encoders based on SigLIP2 and Swin-B, then performs temporal mean pooling and cross-fold ensembling to estimate perceptual quality without using reference information.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,109,114,119,122,127,130,134],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":21,"doc_module":4,"doc_module_name":45,"category_name":106,"show_sort_weight":107,"slug":108},"Comic",60,"comic",{"id":110,"doc_module":4,"doc_module_name":45,"category_name":111,"show_sort_weight":112,"slug":113},6,"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":45,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":45,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":45,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":45,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":45,"category_name":136,"show_sort_weight":21,"slug":137},19,"General","general"]