[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-81660-en":3,"doc-seo-81660-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},81660,16904993612988,"Olivia Brown","https://ap-avatar.wpscdn.com/davatar_a8503ba1806abce46bf441b54a3ca4cd",8,"Research & Report","Explaining is Harder Than Predicting Alone Evaluating Concept-based Explanations of MLLMs as ICL Visual Classifiers","In-context learning (ICL) enables multimodal large language models (MLLMs) to classify images from a few labeled examples, yet the use of provided context is unclear. This paper evaluates concept-based explainability for frozen MLLMs under few-shot ICL using five increasingly formal conditions, from baseline classification to Description Logics (DL) axiom generation, across four state-of-the-art models via an LLM-as-a-judge pipeline. Results show explaining is harder than predicting: forcing structured explanations monotonically reduces accuracy, while explanation quality correlates with correct predictions when discriminative visual features are articulated. The study concludes current MLLMs lack instruction-tuning for machine-verifiable formal explanations.","Explaining is Harder Than Predicting Alone: Evaluating Concept-based Explanations of MLLMs as ICL Visual Classifiers  \narXiv :2605 .282 15v2 [ cs .AI] 10 Jul 2026  \nCarmen Quiles-Ramrez 1 Leticia L. Rodrguez 2  \nAbstract  \nIn-context learning (ICL) enables multimodal large language models (MLLMs) to classify images from a few labelled examples. Yet how these models use the provided context remains opaque. While Chain-of-Thought prompting is widely used, recent work argues it may not reflect true internal computation (Barez et al., 2025) . In this paper, we systematically evaluate the conceptbased explainability of frozen MLLMs under fewshot ICL using five conditions of increasing formal rigour: from baseline classification to Description Logics (DL) axiom generation. Evaluating four state-of-the-art MLLMs via an independent LLM-as-a-judge pipeline, we demonstrate that explaining is genuinely harder than predicting alone. Surprisingly, forcing models to generate formally structured, concept-based explanations degrades predictive accuracy monotonically (from 93.8% to 90.1%), contradicting the assumption that explicit reasoning universally aids performance. However, when models successfully articulate class-discriminative visual features, explanation quality strongly correlates with correct predictions. Our findings suggest that while MLLMs excel at visual classification, they lack the specific instruction-tuning required for formal, machineverifiable explainability.  \n1. Introduction  \nIn-context learning (ICL) (Brown et al., 2020) has emerged as a powerful paradigm for adapting large language models to new tasks without parameter updates. Extended to the multimodal setting, MLLMs can now classify images from a few labelled examples provided directly in the context window (Dong et al., 2024) . Despite impressive few-shot  \n1Universidad de Granada, Spain 2Universidad de Buenos Aires, Argentina. Correspondence to: Carmen Quiles-Ram´ırez \u003Ccar[menqr@ugr.es](menqr@ugr.es) >.  \nAccepted to the 2nd Workshop on Compositional Learning atICML 2026, Seoul, South Korea. Copyright 2026 by the author(s) .  \nNicols Martorell 2 Natalia Daz-Rodrguez 1  \naccuracy, the reasoning process underlying these predictions remains a black box: models produce answers without any principled account of why a particular image belongs to a given class.  \nThe dominant approach to eliciting explanations from large models is chain-of-thought (CoT) prompting, which asks the model to verbalise intermediate reasoning steps. However, Barez et al. (2025) argue that CoT does not constitute genuine explainability: the generated text may diverge from the model’s actual internal computation, offering only a plausible-sounding post-hoc narrative. Similarly, Huang et al. (2025) show that VLMs tend to mimic rather than reason from ICL context, further questioning whether CoTstyle outputs constitute genuine understanding. This gap motivates the search for concept-based explanations (Poeta et al., 2025)—formal statements that are both humanreadable and machine-verifiable.  \nWe approach this problem through the lens of eXplainable AI (XAI) (Ali et al., 2023) and Description Logics (Baader et al., 2009), a decidable fragment of first-order logic widely used for knowledge representation. By designing a hierarchy of five explanation conditions of increasing formal complexity, we systematically investigate whether state-ofthe-art MLLMs can extract the discriminative visual features of an image class, formalise them as IF–THEN rules, and ultimately express them as DL. We evaluate explanation quality using an independent LLM-as-a-judge pipeline anda set of contributed XAI metrics.  \nStudying explanations within few-shot image classification provides a highly controlled setting. It encourages the model to ground its responses in newly introduced classes, helping to mitigate the influence of broad pre-training biases. Our primary goal is to assess how MLLMs explain the knowledge derive","cbCaitSxb1s5QkWL","https://ap.wps.com/l/cbCaitSxb1s5QkWL","pdf",2638643,4,1,25,"English","en",105,"# Introduction\n# Related Work\n# Methodology and Explanation Conditions\n# Evaluation Framework and Judge Model\n# Experimental Design\n# Experimental Results\n# Discussion\n# Conclusion","[{\"question\":\"What problem does the paper address about multimodal ICL image classification?\",\"answer\":\"It studies why the reasoning behind image-classification answers is opaque, even though models can produce labels from a few examples in the context.\"},{\"question\":\"How does the paper evaluate concept-based explanations of MLLMs?\",\"answer\":\"It systematically tests five explanation conditions with increasing formal rigor, from free-text explanations to generating Description Logics (DL) axioms, using an LLM-as-a-judge evaluation pipeline and multiple quality metrics.\"},{\"question\":\"What are the main findings about the relationship between explanations and prediction accuracy?\",\"answer\":\"Generating formally structured, concept-based explanations degrades predictive accuracy monotonically, but when models successfully articulate class-discriminative visual features, explanation quality strongly correlates with correct predictions.\"}]",1784175239,63,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"explaining-is-harder-than-predicting-alone-evaluating-concept-based-explanations-of-mllms-as-icl-visual-classifiers","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/explaining-is-harder-than-predicting-alone-evaluating-concept-based-explanations-of-mllms-as-icl-visual-classifiers/81660/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-25","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the paper address about multimodal ICL image classification?","Question",{"text":75,"@type":76},"It studies why the reasoning behind image-classification answers is opaque, even though models can produce labels from a few examples in the context.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the paper evaluate concept-based explanations of MLLMs?",{"text":80,"@type":76},"It systematically tests five explanation conditions with increasing formal rigor, from free-text explanations to generating Description Logics (DL) axioms, using an LLM-as-a-judge evaluation pipeline and multiple quality metrics.",{"name":82,"@type":73,"acceptedAnswer":83},"What are the main findings about the relationship between explanations and prediction accuracy?",{"text":84,"@type":76},"Generating formally structured, concept-based explanations degrades predictive accuracy monotonically, but when models successfully articulate class-discriminative visual features, explanation quality strongly correlates with correct predictions.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]