[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84193-en":3,"doc-seo-84193-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},84193,962075114765,"Quinn","https://ap-avatar.wpscdn.com/davatar_a8503ba1806abce46bf441b54a3ca4cd",8,"Research & Report","Comparative Study of Domain-adapted VLMs for General Document Visual Question Answering","Document Visual Question Answering (DocVQA) is a demanding multimodal task that requires models to use visual, textual, and layout signals from documents. Despite strong results of Vision-Language Models (VLMs) on related text-vision problems, robustness and transfer across document domains remains limited. This study evaluates eight open-source pretrained VLMs on DocVQA across industrial documents, infographics, and presentation slides using zero-shot, fully supervised finetuning, and few-shot domain transfer. Large pretrained VLMs excel on structured layouts but degrade on visually complex layouts; supervised finetuning helps smaller models. Cross-domain and few-shot findings indicate visual understanding—not missing knowledge—forms the main bottleneck, and adaptation with 50 target samples can surpass some fully supervised results.","arXiv :2607 .07 179v 1 [ cs .CV] 8 Jul 2026  \nComparative Study of Domain-adapted VLMs for General Document Visual Question Answering  \nMiguel Lopez-Duran, Elena Marrero, Julian Fierrez, Marta Robledo-Moreno, Ruben Vera-Rodriguez, Daniel DeAlcala, Aythami Morales, Ruben Tolosana, Oscar Delgado, Alvaro Ortigosa, and Javier Ortega-Garcia  \nBiometricsAI, Universidad Autónoma de Madrid (UAM), Spain [miguel.lopezd@uam.es](miguel.lopezd@uam.es) , [julian.fierrez@uam.es](julian.fierrez@uam.es)  \nAbstract. Document Visual Question Answering (DocVQA) presentsa complex multimodal challenge, requiring models to exploit visual, textual, and layout information from documents. Although Vision-Language Models (VLMs) have shown remarkable performance in text-vision tasks, their robustness and transferability to different document domains remains underexplored. In this study, we present a comprehensive evaluation of 8 open-source pretrained VLMs on DocVQA in three different document domains: industrial documents of varying type, infographics, and presentation slides. We systematically assess model performance under zero-shot evaluations, fully supervised finetuning with inter-and intra-dataset evaluations, and few-shot learning evaluations of knowledge transfer between domains. Our findings demonstrate that while large pretrained VLMs possess strong zero-shot baselines for structured layouts, their performance strongly decreases on visually complex layouts of infographics and slides. Although parameter scaling is a dominant factor on performance, supervised finetuning yields higher relative gains in smaller architectures. Furthermore, our cross-domain and few-shot experiments show that visual understanding is the main bottleneck for DocVQA, not a lack of knowledge from the VLMs. Using 50 target domain samples, the models finetuned in DocVQA with datasets of different domains rapidly adapt to the target domain documents, even surpassing their fully supervised counterparts in some cases.  \nKeywords: Document Visual Question Answering, Vision Language Models, Domain Adaptation.  \n1 Introduction  \nVisual Question Answering (VQA) is a core multi-modal task in machine learning consisting of answering text-based questions about an image. Generally speaking, VQA tasks can be classified as extractive and abstractive Question Answering (QA) . Extractive QA consists of answering the question by extracting a subset of tokens within the image. On the other hand, abstractive QA aims to answer the question based on the image content, but the answer does not need to be extracted from it.  \n2 M. Lopez-Duran, E. Marrero, J. Fierrez, et al.  \nThere are different VQA tasks depending on the application scenario, such as maps [3], daily photos [15], or scientific papers [37] . Among all these tasks, Document VQA (DocVQA), which aims to answer questions based on document images, is much more challenging. It requires detecting the layout objects [35,36,23] and extracting the relationships between them to extract relevant information to answer the question correctly.  \nThis complexity is even greater when different document domains are taken into account. Different document domains differ greatly from each other. For example, scientific papers usually present a structured layout with one or two columns, same fonts for all the objects, and the relations between objects are simple and much more direct. However, other document types, like presentation slides, do not have a structured layout and the relations between objects are less structured in order to be visually appealing.  \nFor all these reasons, Document Understanding (DU) models need to be able to adapt to different domains, not only based on the document image but also on the internal relationships between objects in different document domains. This adaptation should be present during training, but also during inference, when models may be given documents from scarce domains that have never or barely been seen d","cbCaibgQJOYTa16s","https://ap.wps.com/l/cbCaibgQJOYTa16s","pdf",896055,4,1,17,"English","en",105,"# Introduction\n## Visual Question Answering and DocVQA\n## Document domains and the need for adaptation\n## Vision-Language Models and current limitations\n## Motivation and study goal","[{\"question\":\"What problem does DocVQA address in the document domain?\",\"answer\":\"DocVQA answers questions based on document images, requiring layout understanding and relation extraction among document elements to retrieve relevant information.\"},{\"question\":\"How do pretrained VLMs generally perform across different document domains?\",\"answer\":\"They provide strong zero-shot baselines on structured layouts, but performance drops substantially on visually complex layouts such as infographics and presentation slides.\"},{\"question\":\"What does the study find about the main bottleneck for DocVQA?\",\"answer\":\"Cross-domain and few-shot experiments suggest the key limitation is visual understanding of layouts, not a lack of task knowledge in the VLMs themselves.\"}]",1784193848,43,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"comparative-study-of-domain-adapted-vlms-for-general-document-visual-question-answering","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/comparative-study-of-domain-adapted-vlms-for-general-document-visual-question-answering/84193/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does DocVQA address in the document domain?","Question",{"text":75,"@type":76},"DocVQA answers questions based on document images, requiring layout understanding and relation extraction among document elements to retrieve relevant information.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How do pretrained VLMs generally perform across different document domains?",{"text":80,"@type":76},"They provide strong zero-shot baselines on structured layouts, but performance drops substantially on visually complex layouts such as infographics and presentation slides.",{"name":82,"@type":73,"acceptedAnswer":83},"What does the study find about the main bottleneck for DocVQA?",{"text":84,"@type":76},"Cross-domain and few-shot experiments suggest the key limitation is visual understanding of layouts, not a lack of task knowledge in the VLMs themselves.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]