[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86235-en":3,"doc-seo-86235-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":11,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},86235,1374391974585,"Genevieve","https://ap-avatar.wpscdn.com/davatar_276721f389ce27ea32af1340a28f341c",8,"Research & Report","Omni-Decision: A Progressive Evidence-State Agent System for Omni-Modal QA","Omni-modal evidence-seeking question answering requires an agent to answer queries whose supporting evidence is sparsely distributed across videos, audio, images, web pages, and computation outputs. Existing agentic multimodal systems often lose traceability of what is grounded, what is missing, and when evidence is sufficient. Omni-Decision introduces a training-free evidence-state framework that converts each query into an evidence-closure process with structured state tracking confirmed facts, conflicts, dependencies, and open needs, enabling inspectable acquisition, validation, repair, and stopping.","arXiv :2607 . 1 1433v 1 [ cs .AI] 13 Jul 2026  \nOmni-Decision: A Progressive Evidence-State Agent System for Omni-Modal QA  \nMing Ma 1 ,2 Yi Zhu3 ,∗ Yiran Zhong3 ,∗ Feida Zhu3 Weigao Sun3  \nJunhan Shi4 Lingrui Mei5 Tianming Yang 1 Steven Hoi3  \n1 Institute of Neuroscience, Chinese Academy of Sciences  \n2 University of Chinese Academy of Sciences  \n3 Tongyi Lab, Alibaba Group  \n4 Tsinghua University  \n5 Institute of Computing Technology, Chinese Academy of Sciences  \n[mam2022@ion.ac.cn](mam2022@ion.ac.cn), [zhu.yee@outlook.com](zhu.yee@outlook.com), [zhongyiran@gmail.com](zhongyiran@gmail.com)  \n∗ Corresponding authors.  \nAbstract  \nOmni-modal evidence-seeking QA requires agents to answer questions whose evidence is sparsely distributed across videos, audio, images, web pages, and computation results. Existing agentic multimodal systems often leave evidence in scratchpads, tool trajectories, or free-form histories, making it difficult to track what has been grounded, what remains missing, and when the evidence is sufficient to answer. We propose Omni-Decision, a training-free evidence-state system that turns omni-modal QA into a query-scoped evidence-closure process. For each query, Omni-Decision maintains a structured evidence state containing confirmed evidence, unresolved conflicts, fact and computation dependencies, and open evidence needs. A shared state view conditions planning, evidence acquisition, validation, repair, and finalization. Heterogeneous observations from media, web, computation, and verification modules are normalized, judged, and committed through deterministic state updates. This design enables targeted evidence acquisition, preserves sparse cross-modal cues, and provides inspectable control over repair and stopping. Omni-Decision achieves 45.6% accuracy on OmniGAIA and 58.3% on WorldSense, improving over the baselines by +27.3 and +30.2 percentage points, respectively. No-state ablations and trajectory audits further support the role of explicit evidence-state control in multi-step omni-modal evidence seeking.  \n1 Introduction  \nOmni-modal question answering is moving beyond closed-form perceptual understanding toward evidence-seeking QA [Fu et al., 2024, Wu et al., 2024, Hong et al., 2025, Li et al., 2026] . This setting is closer to practical agentic problem solving: a user asks a question grounded in heterogeneous media, and the answer may require the system to identify relevant visual or acoustic evidence, complete missing attributes from external sources, check consistency across modalities, and sometimes compute a derived value before responding [Nakano et al., 2022, Schick et al., 2023, Yu et al., 2026] .  \nThis setting is difficult for three reasons. First, evidence is sparse and distributed [Zhong et al., 2022, Ranasinghe et al., 2025, Ren et al., 2025, Tang et al., 2025, Zhang et al., 2025, Yu et al., 2026] . A relevant clue may appear in a short video segment, a subtitle span, an audio event, an image region, a web page, or a computation result. Second, the evidence chain is partially observable. The system usually does not know in advance which entity attribute, relation, or intermediate value  \nPreprint.  \nwill be needed until earlier evidence has been grounded. Third, answer readiness is itself a control problem. A system must decide not only what to acquire next, but also whether a candidate answer is supported, whether a conflict should trigger repair, and whether the remaining gap is unfillable under the available actions [Shinn et al., 2023, Yao et al., 2023, Han et al., 2025, Wang et al., 2026a, Zhang et al., 2026] .  \nThese properties make omni-modal evidence-seeking QA a state-maintenance problem rather than merely a context-scaling problem [Wu et al., 2024, Chen et al., 2025, He et al., 2025, Li et al., 2025, Yang et al., 2025, Wang et al., 2026b] . A longer context window or a larger number of frames can expose more raw observations, but it does not by itself specify which entity h","cbCairavP98k3Lht","https://ap.wps.com/l/cbCairavP98k3Lht","pdf",784758,1,25,"English","en",105,"# Introduction\n## Motivation: evidence sparsity and partial observability\n## Control of answer readiness and stopping\n## State-maintenance view and cognitive-control analogy\n## Benchmarks and task variants\n## Omni-Decision overview: query-scoped evidence state","[{\"question\":\"What problem does Omni-Decision address in omni-modal question answering?\",\"answer\":\"Omni-Decision targets evidence-seeking QA where supporting information is sparsely distributed across multiple modalities and sources, making it hard to know what has been grounded, what is missing, and when to stop.\"},{\"question\":\"How does Omni-Decision represent and manage evidence during a query?\",\"answer\":\"For each query, it maintains a structured evidence state that tracks confirmed evidence, unresolved conflicts, fact/computation dependencies, and open evidence needs.\"},{\"question\":\"What makes Omni-Decision effective compared with existing approaches?\",\"answer\":\"By using deterministic state updates that normalize and commit heterogeneous observations into the same evidence state, it enables targeted evidence acquisition, explicit repair control, and improved stopping behavior, leading to higher accuracy on OmniGAIA and WorldSense.\"}]",1784209701,63,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"omni-decision-a-progressive-evidence-state-agent-system-for-omni-modal-qa","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/omni-decision-a-progressive-evidence-state-agent-system-for-omni-modal-qa/86235/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-27","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":11},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does Omni-Decision address in omni-modal question answering?","Question",{"text":75,"@type":76},"Omni-Decision targets evidence-seeking QA where supporting information is sparsely distributed across multiple modalities and sources, making it hard to know what has been grounded, what is missing, and when to stop.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does Omni-Decision represent and manage evidence during a query?",{"text":80,"@type":76},"For each query, it maintains a structured evidence state that tracks confirmed evidence, unresolved conflicts, fact/computation dependencies, and open evidence needs.",{"name":82,"@type":73,"acceptedAnswer":83},"What makes Omni-Decision effective compared with existing approaches?",{"text":84,"@type":76},"By using deterministic state updates that normalize and commit heterogeneous observations into the same evidence state, it enables targeted evidence acquisition, explicit repair control, and improved stopping behavior, leading to higher accuracy on OmniGAIA and WorldSense.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":45,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":45,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":45,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]