[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-seo-148787-105":3,"detail-sidebar-cat-0-en-105":81,"doc-detail-148787-en":130},{"code":4,"msg":5,"data":6},0,"ok",{"site_id":7,"language":8,"slug":9,"title":10,"keywords":11,"description":12,"schema_data":13,"social_meta":74,"head_meta":76,"extra_data":78,"updated_unix":80},105,"en","plug-and-play-vqa-pnp-vqa-zero-shot-vqa-by-conjoining-large-pretrained-models-with-zero-training","Plug-and-Play VQA - PNP-VQA - Zero-shot VQA by Conjoining Large Pretrained Models with Zero Training","","Visual question answering (VQA) under the zero-shot setting is a difficult vision-language reasoning task. This paper proposes Plug-and-Play VQA (PNP-VQA), a modular framework that enables zero-shot VQA without additional training of large pretrained language models. By using question-guided informative image captions as an intermediate representation, the method connects pretrained vision-language and language models for question answering. It achieves state-of-the-art results on zero-shot VQAv2 and GQA, outperforming strong end-to-end baselines.",{"@graph":14,"@context":73},[15,34,56],{"@type":16,"itemListElement":17},"BreadcrumbList",[18,23,27,31],{"item":19,"name":20,"@type":21,"position":22},"https://docshare.wps.com","Home","ListItem",1,{"item":24,"name":25,"@type":21,"position":26},"https://docshare.wps.com/document/","Document",2,{"item":28,"name":29,"@type":21,"position":30},"https://docshare.wps.com/document/research-report/","Research & Report",3,{"item":32,"name":10,"@type":21,"position":33},"https://docshare.wps.com/document/plug-and-play-vqa-pnp-vqa-zero-shot-vqa-by-conjoining-large-pretrained-models-with-zero-training/148787/",4,{"url":32,"name":10,"@type":35,"image":36,"author":41,"headline":10,"publisher":44,"fileFormat":47,"inLanguage":8,"description":12,"dateModified":48,"datePublished":49,"encodingFormat":47,"isAccessibleForFree":50,"interactionStatistic":51},"DigitalDocument",{"url":37,"@type":38,"width":39,"height":40},"https://docshare.wps.com/thumbnails/plug-and-play-vqa-pnp-vqa-zero-shot-vqa-by-conjoining-large-pretrained-models-with-zero-training/148787.png","ImageObject",300,407,{"name":42,"@type":43},"Liam","Person",{"url":19,"name":45,"@type":46},"DocShare","Organization","application/pdf","2026-09-16","2026-08-26",true,{"@type":52,"interactionType":53,"userInteractionCount":55},"InteractionCounter",{"@type":54},"ViewAction",6,{"@type":57,"mainEntity":58},"FAQPage",[59,65,69],{"name":60,"@type":61,"acceptedAnswer":62},"What is Plug-and-Play VQA (PNP-VQA)?","Question",{"text":63,"@type":64},"PNP-VQA is a modular framework for zero-shot visual question answering that combines pretrained vision-language and language models without additional training.","Answer",{"name":66,"@type":61,"acceptedAnswer":67},"How does PNP-VQA bridge vision and language modalities?",{"text":68,"@type":64},"It generates question-guided informative image captions using a network interpretability technique, then feeds the captions to a pretrained language model to answer the question.",{"name":70,"@type":61,"acceptedAnswer":71},"What performance improvements does PNP-VQA achieve?",{"text":72,"@type":64},"PNP-VQA reaches state-of-the-art zero-shot results on VQAv2 and GQA, including an improvement over Flamingo and FewVLM at comparable parameter scales.","https://schema.org",{"og:url":32,"og:type":75,"og:title":10,"og:site_name":45,"og:description":12},"article",{"robots":77,"canonical":32},"index,follow",{"doc_id":79,"site_id":7},148787,1787786118,{"code":4,"msg":82,"data":83},"success",[84,88,92,96,101,105,110,114,119,122,126],{"id":22,"doc_module":4,"doc_module_name":25,"category_name":85,"show_sort_weight":86,"slug":87},"Story & Novel",90,"story-novel",{"id":26,"doc_module":4,"doc_module_name":25,"category_name":89,"show_sort_weight":90,"slug":91},"Literature",80,"literature",{"id":33,"doc_module":4,"doc_module_name":25,"category_name":93,"show_sort_weight":94,"slug":95},"Exam",70,"exam",{"id":97,"doc_module":4,"doc_module_name":25,"category_name":98,"show_sort_weight":99,"slug":100},5,"Comic",60,"comic",{"id":55,"doc_module":4,"doc_module_name":25,"category_name":102,"show_sort_weight":103,"slug":104},"Technology",50,"technology",{"id":106,"doc_module":4,"doc_module_name":25,"category_name":107,"show_sort_weight":108,"slug":109},7,"Healthcare",40,"healthcare",{"id":111,"doc_module":4,"doc_module_name":25,"category_name":29,"show_sort_weight":112,"slug":113},8,30,"research-report",{"id":115,"doc_module":4,"doc_module_name":25,"category_name":116,"show_sort_weight":117,"slug":118},9,"Religion & Spirituality",20,"religion-spirituality",{"id":117,"doc_module":4,"doc_module_name":25,"category_name":120,"show_sort_weight":117,"slug":121},"World Cup","world-cup",{"id":123,"doc_module":4,"doc_module_name":25,"category_name":124,"show_sort_weight":123,"slug":125},10,"Lifestyle","lifestyle",{"id":127,"doc_module":4,"doc_module_name":25,"category_name":128,"show_sort_weight":97,"slug":129},19,"General","general",{"code":4,"msg":82,"data":131},{"doc_id":79,"user_id":132,"nickname":42,"user_avatar":133,"doc_module":4,"category_id":111,"category_name":29,"doc_title":10,"doc_description":12,"doc_content":134,"file_id":135,"file_url":136,"file_type":137,"file_size":138,"view_count":55,"is_deleted":4,"is_public":22,"is_downloadable":22,"audit_status":22,"page_count":139,"language":140,"language_code":8,"site_id":7,"html_lang":8,"table_of_contents":141,"faqs":142,"seo_title":143,"seo_description":12,"update_tm":80,"read_time":144},8796095461564,"https://ap-avatar.wpscdn.com/davatar_155a257f0dc6eb9ab79c44ca47cae57d","Plug-and-Play VQA: Zero-shot VQA by Conjoining Large Pretrained  \nModels with Zero Training  \nAnthony Meng Huat Tiong 1 ,2 , Junnan Li 1 , Boyang Li2 ,  \nSilvio Savarese 1 , and Steven C.H. Hoi 1  \n1 Salesforce Research 2Nanyang Technological University, Singapore {anthony.tiong, [junnan.li](junnan.li) , ssavarese, [shoi}@salesforce.com](shoi}@salesforce.com)  \n[boyang.li@ntu.edu.sg](boyang.li@ntu.edu.sg)  \n[https://github.com/salesforce/LAVIS/tree/main/projects/pnp-vqa](https://github.com/salesforce/LAVIS/tree/main/projects/pnp-vqa)  \nAbstract  \nVisual question answering (VQA) is a hallmark of vision and language reasoning and a challenging task under the zero-shot setting. We propose Plug-and-Play VQA (PNP-VQA), a modular framework for zero-shot VQA. In contrast to most existing works, which require substantial adaptation of pretrained language models (PLMs) for the vision modality, PNP-VQA requires no additional training of the PLMs.  \nInstead, we propose to use natural language and network interpretation as an intermediate representation that glues pretrained models together. We first generate question-guided informative image captions, and pass the captions to a PLM as context for question answering. Surpassing end-to-end trained baselines, PNP-VQA achieves state-of-the-art results on zero-shot VQAv2 (Goyal et al., 2017) and GQA (Hudson and Manning, 2019) . With 11B parameters, it outperforms the 80B-parameter Flamingo model (Alayrac et al., 2022) by 8.5% on VQAv2 . With 738M PLM parameters, PNPVQA achieves an improvement of 9.1% on GQA over FewVLM (Jin et al., 2022) with 740M PLM parameters.  \n1 Introduction  \nRecent years have witnessed unprecedented performance gains on many natural language reasoning tasks, especially in zero-shot and few-shot settings, being derived from scaling up pretrained language models (PLMs) and their training data (Devlin et al., 2019 ; Liu et al., 2019 ; Brown et al., 2020 ; Raffel et al., 2020 ; Black et al., 2022 ; Sanhet al., 2022 ; Wei et al., 2021) . Inspired by their success, a natural thought is that utilizing PLMs should also boost zero-shot performance in visionlanguage reasoning tasks.  \nHowever, to leverage PLMs for vision-language tasks, most existing methods require non-trivial adaptation of the PLMs for the vision modality, which necessitates the design of new network components and training objectives. For example, Sung  \net al. (2022) and Alayrac et al. (2022) insert into the PLMs new layers that are trained from scratch. Tsimpoukelli et al. (2021) train vision encoders that output soft prompts to frozen PLMs. Chen et al.(2022) and Eichenberg et al. (2021) train both the vision encoders and new layers inserted into PLMs. In the zero-shot setting, various vision-language pretraining objectives are employed, such as image captioning (Alayrac et al., 2022) and imageconditioned masked language modeling (Jin et al., 2022) .  \nFrom the perspective of general-purpose AI, the ability to perform new tasks by simply recombining large-scale pretrained models, or foundation models (Bommasani et al., 2021), without architectural changes or extra training would be highly desirable. Such a system would be able to dynamically adjust to previously unknown tasks by simply rewiring a small number of foundation models. However, to obtain high performance without some form of end-to-end training would seem difficult, if not impossible.  \nWe present Plug-and-Play VQA (PNP-VQA), a framework for zero-shot visual question answering which conjoins large pretrained models with zero additional training and achieves state-of-the-art performance on zero-shot VQAv2 (Goyal et al., 2017) and GQA (Hudson and Manning, 2019) . For the purpose of bridging the vision and language modalities, we employ a pretrained vision-language model (PVLM) (Li et al., 2022b) that describes visual information with textual captions. In order to obtain relevant and informative captions, we apply a network interpretability techniq","cbCaii0NrTNUEH3h","https://ap.wps.com/l/cbCaii0NrTNUEH3h","pdf",8786376,17,"English","# Abstract\n# Introduction\n# Related Work","[{\"question\":\"What is Plug-and-Play VQA (PNP-VQA)?\",\"answer\":\"PNP-VQA is a modular framework for zero-shot visual question answering that combines pretrained vision-language and language models without additional training.\"},{\"question\":\"How does PNP-VQA bridge vision and language modalities?\",\"answer\":\"It generates question-guided informative image captions using a network interpretability technique, then feeds the captions to a pretrained language model to answer the question.\"},{\"question\":\"What performance improvements does PNP-VQA achieve?\",\"answer\":\"PNP-VQA reaches state-of-the-art zero-shot results on VQAv2 and GQA, including an improvement over Flamingo and FewVLM at comparable parameter scales.\"}]","Plug-and-Play VQA - PNP-VQA - Zero-shot VQA by Conjoining Large Pretrained Models with Zero Training | PDF",43]