[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-160256-en":3,"doc-seo-160256-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},160256,137451211410,"\tCallum ","https://ap-avatar.wpscdn.com/avatar/2000bb0a9246f588df?x-image-process=image/resize,m_fixed,w_180,h_180&k=1786362646172706240",8,"Research & Report","PCA-Bench: Evaluating Multimodal Large Language Models in Perception-Cognition-Action Chain - ACL 2024","PCA-Bench is introduced as a multimodal decisionmaking benchmark to evaluate the integrated capabilities of Multimodal Large Language Models (MLLMs). It goes beyond simplistic or single-skill evaluations by defining three complex embodied/open-world scenarios: autonomous driving, domestic robotics, and open-world games. Models must jointly perform Perception, Cognition, and Action in a reasoning chain, while PCA-Bench also localizes errors across perception, knowledge, and reasoning. An automatic PCA-Eval protocol and Embodied-InstructionEvolution (EIE) improve evaluation efficiency and open-source performance, with public benchmark data and code released.","PCA-Bench: Evaluating Multimodal Large Language Models in Perception-Cognition-Action Chain  \nLiang Chen 1 , Yichi Zhang 1 , Shuhuai Ren 1 , Haozhe Zhao 1 , Zefan Cai 1 , Yuchi Wang 1 , Peiyi Wang 1 , Xiangdi Meng 1 , Tianyu Liu2 , Baobao Chang 1†  \n1 National Key Laboratory for Multimedia Information Processing, Peking University  \n2 Alibaba Group  \n{leo.liang.chen, yczhang, [shuhuai_ren}@stu.pku.edu.cn](shuhuai_ren}@stu.pku.edu.cn)  \n[tianyu0421@alibaba-inc.com](tianyu0421@alibaba-inc.com) , [chbb@pku.edu.cn](chbb@pku.edu.cn)  \n􀂇 PCA-EVAL  PCA-Bench-V1  PCA-Bench-Action-V1  \nAbstract  \nWe present PCA-Bench, a multimodal decisionmaking benchmark for evaluating the integrated capabilities of Multimodal Large Language Models (MLLMs) . Departing from previous benchmarks focusing on simplistic tasks and individual model capability, PCA-Bench introduces three complex scenarios: autonomous driving, domestic robotics, and open-world games. Given task instructions and diverse contexts, the model is required to seamlessly integrate multiple capabilities of Perception, Cognition, and Action in a reasoning chain to make accurate decisions. Moreover, PCA-Bench features error localization capabilities, scrutinizing model inaccuracies in areas such as perception, knowledge, or reasoning. This enhances the reliability of deploying MLLMs. To balance accuracy and efficiency in evaluation, we propose PCA-Eval, an automatic evaluation protocol, and assess 10 prevalent MLLMs. The results reveal significant performance disparities between open-source models and powerful proprietary models like GPT-4 Vision. To address this, we introduce Embodied-InstructionEvolution (EIE), an automatic framework for synthesizing instruction tuning examples in multimodal embodied environments. EIE generates 7,510 training examples in PCA-Bench and enhances the performance of open-source MLLMs, occasionally surpassing GPT-4 Vision (+3% in decision accuracy), thereby validating the effectiveness of EIE. Our findings suggest that robust MLLMs like GPT4-Vision show promise for decision-making in embodied agents, opening new avenues for MLLM research. All benchmark data and evaluation code are made public.  \n1 Introduction  \nMultimodal Large Language Models (MLLMs) have demonstrated remarkable capabilities in tackling complex tasks that necessitate a chain of in-  \n†  \nCorresponding author.  \nMultimodal LLM  \nFigure 1: Example of decision making with MLLMs in the Perception-Cognition-Action Chain.  \ntegrated skills, including visual perception, world knowledge, reasoning, action, and more (OpenAI, 2023 ; Dai et al., 2023a ; Liu et al., 2023b ; Li et al., 2023c ; Zhao et al., 2023 ; Cheng et al., 2024b ; Zhang et al., 2024) .  \nHowever, current MLLM benchmarks often evaluate these capabilities individually (Fu et al., 2023 ; Liu et al., 2023e), overlooking the significant integrated potential that Large Language Models (LLMs) contribute to multimodal models. While some benchmarks like MMMU (Yue et al., 2023) and MathVista (Lu et al., 2023a) require abilities from both the vision and language part, they lack error localization techniques beyond accuracy assessments. This complicates identifying which part of the MLLM malfunctioned when making mistakes—whether it was the visual or the language component—and determines which aspect requires enhancement to enhance overall performance.  \n1086  \nFindings of the Association for Computational Linguistics: ACL 2024 , pages 1086–1104 August 11-16, 2024 ©2024 Association for Computational Linguistics  \nTo address the challenges of insufficient integrated benchmarking and error localization problems, we introduce PCA-Bench. It arises with MLLM’s applications in embodied AI and decision making, where models called agents need to first process multimodal observation from different environments, reason with the current situation and goal, and finally make an action from a given action space. The abilities in the complex decision makin","cbCaiaVDN4otiiMQ","https://ap.wps.com/l/cbCaiaVDN4otiiMQ","pdf",9976146,1,19,"English","en",105,"# Abstract\n# Introduction","[{\"question\":\"What is PCA-Bench designed to evaluate for MLLMs?\",\"answer\":\"PCA-Bench evaluates the integrated Perception-Cognition-Action capabilities of Multimodal Large Language Models in multimodal decisionmaking tasks.\"},{\"question\":\"What scenarios are included in PCA-Bench?\",\"answer\":\"PCA-Bench includes autonomous driving, domestic robotics, and open-world games, each requiring multi-capability reasoning to select accurate actions.\"},{\"question\":\"How does PCA-Bench support error localization?\",\"answer\":\"PCA-Bench uses a 6-element annotation tuple with anchors that score and localize mistakes in Action, Cognition, and Perception, enabling more precise diagnosis than accuracy-only benchmarks.\"}]","PCA-Bench: Evaluating Multimodal Large Language Models in Perception-Cognition-Action Chain - ACL 2024 | PDF",1788052881,48,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"pca-bench-evaluating-multimodal-large-language-models-in-perception-cognition-action-chain-acl-2024","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/pca-bench-evaluating-multimodal-large-language-models-in-perception-cognition-action-chain-acl-2024/160256/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-30",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What is PCA-Bench designed to evaluate for MLLMs?","Question",{"text":75,"@type":76},"PCA-Bench evaluates the integrated Perception-Cognition-Action capabilities of Multimodal Large Language Models in multimodal decisionmaking tasks.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What scenarios are included in PCA-Bench?",{"text":80,"@type":76},"PCA-Bench includes autonomous driving, domestic robotics, and open-world games, each requiring multi-capability reasoning to select accurate actions.",{"name":82,"@type":73,"acceptedAnswer":83},"How does PCA-Bench support error localization?",{"text":84,"@type":76},"PCA-Bench uses a 6-element annotation tuple with anchors that score and localize mistakes in Action, Cognition, and Perception, enabling more precise diagnosis than accuracy-only benchmarks.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":21,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},"General","general"]