[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82188-en":3,"doc-seo-82188-105":29,"detail-sidebar-cat-0-en-105":90},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},82188,2336464648746,"Skyler","https://ap-avatar.wpscdn.com/davatar_276721f389ce27ea32af1340a28f341c",8,"Research & Report","Beyond Time Shifts Adapting Omni LLM as a Reference-Free Evaluator for Generative Audio-Visual Models","Audio-visual generative models are evolving toward world simulators, where cross-modal synchronization becomes a key proxy for evaluating consistency of generated world dynamics and causal relations. Existing metrics largely assume structural correctness and reduce synchronization to temporal alignment, failing under structural hallucinations and asymmetric audio-video relations that currently require expert human annotation. The work resolves the relative-vs-absolute evaluation paradox by converting ranked human judgments into a continuous, globally consistent reference-free metric using SynthSync data, an Omni-LLM-based latent projection, and RGRPO.","arXiv :2607 .09091v1 [ cs .CV] 10 Jul 2026  \nBeyond Time Shifts: Adapting Omni-LLM asa Reference-Free Evaluator for Generative Audio-Visual Models  \nYijie Qian 1 ,3⋆‡, Juncheng Wang2⋆, Chao Xu3 ,4 , Huihan Wang2 , Yuxiang Feng 1 , Yang Liu3 ,4 , Baigui Sun3 ,4 , Yong Liu 1†, and Shujun Wang2†  \n1 Zhejiang University, Hangzhou, China  \n2 The Hong Kong Polytechnic University, Hong Kong, China  \n3 IROOTECH TECHNOLOGY, China  \n4 Wolf 1069 b Lab, Sany Group, China ⋆ Equal contribution. † Corresponding authors.  \n{[yijieqian@zju.edu.cn](yijieqian@zju.edu.cn), [wjc2830@gmail.com}](wjc2830@gmail.com})⋆ , {[yongliu@iipc.zju.edu.cn](yongliu@iipc.zju.edu.cn),  \n[shu-jun.wang@polyu.edu.hk}](shu-jun.wang@polyu.edu.hk})†  \nAbstract. As audio-visual generative models evolve into world simulators, cross-modal synchronization stands as a critical proxy for assessing the consistency of world dynamics and causality in generated content.  \nHowever, existing evaluation metrics presume structural correctness, reducing synchronization to mere temporal alignment. Consequently, they fail on generative outputs, especially when exhibiting structural hallucinations and asymmetric cross-modal relations, which currently mandate expert human annotation to assess synchronization. This dependency introduces a critical paradox: human evaluators rely on relative, reference-dependent comparisons, whereas automated metrics require reference-free, absolute scalars. We resolve this paradox by proposing a framework that distills relative human perception into a continuous, globally consistent metric. First, we introduce SynthSync, a dataset of generative failures ranked via pairwise human annotations. Second, we adapt the Omni-LLM equipped with a continuous latent projection to translate relative human rankings into continuous absolute values.  \nThird, we propose Real-Valued Group Relative Policy Optimization (RGRPO) to internalize the global causal structure of synchronization via listwise score distributions. Empirically, our metric achieves state-ofthe-art human preference alignment. We leverage this estimator to establish a standardized benchmark, advancing AV-Gen assessment from low-level signal correlation to visually grounded causality. Project page:  \n[https://chenhaoqcdyq.github.io/BeyondTimeShifts](https://chenhaoqcdyq.github.io/BeyondTimeShifts)  \nKeywords: Joint Audio-Video Generation · Audio-Visual Synchronization · Omni-LLM  \n1 Introduction  \nDriven by systems like Sora-2 [3,47], Veo-3 [18,46], and Seedance-2 [4,53], audiovisual generative models [32, 35, 39, 56, 61] are evolving from visual synthesizers  \n‡ This work was conducted in collaboration with IROOTECH TECHNOLOGY.  \n2 Y. Qian et al.  \nTwo Cases generated by Veo-3.1   \nContributions of this paper  \nFig. 1: Upper row: Two Veo-3.1 [18]-generated audio-video samples with evident mismatches between visual events and audio, showcasing structural hallucinations where the synthesized impact sounds (e.g., the golf strike) exhibit semantic and physical textures inconsistent with the visual cues, and a breakdown in asymmetric relationships where the visual (e.g., dog barking) lacks corresponding audio signal. Lower row: Our synchronization evaluation pipeline. We first introduce SynthSync (left), a dataset of authentic generative synchronization failures with pairwise preference annotations. We then fine-tune an omni-LLM as a reference-free continuous synchronization scorer using a new training paradigm (middle), followed by reinforcement post-training to better capture the global causal structure of synchronization judgments (right) .  \ninto nascent “world simulators” [3] . While these models achieve single-modality fidelity, the synchronization of cross-modal events remains fragile (shown in upper row of Fig. 1) . Such failures matter not only for perceptual realism, but also as a key observable proxy for whether a model has learned how actions and sounds relate in the real world.  \nHowever, measuring","cbCaipaWwTwQ6ZkL","https://ap.wps.com/l/cbCaipaWwTwQ6ZkL","pdf",6212886,1,28,"English","en",105,"# Abstract\n# Introduction\n# Contributions of this paper","[{\"question\":\"What problem does the paper address in evaluating generative audio-visual synchronization?\",\"answer\":\"Current evaluation metrics assume structural correctness and treat synchronization mainly as temporal alignment. This breaks down for modern generative outputs that show structural hallucinations, asymmetric cross-modal relations, and temporally diffuse events, making reliable evaluation hard without expert judgment.\"},{\"question\":\"How does the proposed method avoid reference-based evaluation?\",\"answer\":\"The approach distills relative human preference comparisons into continuous absolute scores. It uses SynthSync pairwise human annotations, then adapts Omni-LLM with a continuous latent projection to produce a globally consistent reference-free synchronization metric.\"},{\"question\":\"What is RGRPO and what role does it play?\",\"answer\":\"Real-Valued Group Relative Policy Optimization internalizes the global causal structure of synchronization judgments. It does so via listwise score distributions during reinforcement post-training to better reflect how humans assess causal alignment.\"}]",1784178695,71,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":85,"head_meta":87,"extra_data":89,"updated_unix":27},"beyond-time-shifts-adapting-omni-llm-as-a-reference-free-evaluator-for-generative-audio-visual-models","",{"@graph":35,"@context":84},[36,53,67],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/beyond-time-shifts-adapting-omni-llm-as-a-reference-free-evaluator-for-generative-audio-visual-models/82188/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":61,"encodingFormat":60,"isAccessibleForFree":62,"interactionStatistic":63},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-16",true,{"@type":64,"interactionType":65,"userInteractionCount":4},"InteractionCounter",{"@type":66},"ViewAction",{"@type":68,"mainEntity":69},"FAQPage",[70,76,80],{"name":71,"@type":72,"acceptedAnswer":73},"What problem does the paper address in evaluating generative audio-visual synchronization?","Question",{"text":74,"@type":75},"Current evaluation metrics assume structural correctness and treat synchronization mainly as temporal alignment. This breaks down for modern generative outputs that show structural hallucinations, asymmetric cross-modal relations, and temporally diffuse events, making reliable evaluation hard without expert judgment.","Answer",{"name":77,"@type":72,"acceptedAnswer":78},"How does the proposed method avoid reference-based evaluation?",{"text":79,"@type":75},"The approach distills relative human preference comparisons into continuous absolute scores. It uses SynthSync pairwise human annotations, then adapts Omni-LLM with a continuous latent projection to produce a globally consistent reference-free synchronization metric.",{"name":81,"@type":72,"acceptedAnswer":82},"What is RGRPO and what role does it play?",{"text":83,"@type":75},"Real-Valued Group Relative Policy Optimization internalizes the global causal structure of synchronization judgments. It does so via listwise score distributions during reinforcement post-training to better reflect how humans assess causal alignment.","https://schema.org",{"og:url":51,"og:type":86,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":88,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":91},[92,96,100,104,109,114,119,122,127,130,134],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":93,"show_sort_weight":94,"slug":95},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":97,"show_sort_weight":98,"slug":99},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":101,"show_sort_weight":102,"slug":103},"Exam",70,"exam",{"id":105,"doc_module":4,"doc_module_name":45,"category_name":106,"show_sort_weight":107,"slug":108},5,"Comic",60,"comic",{"id":110,"doc_module":4,"doc_module_name":45,"category_name":111,"show_sort_weight":112,"slug":113},6,"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":45,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":45,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":45,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":45,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":45,"category_name":136,"show_sort_weight":105,"slug":137},19,"General","general"]