[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"detail-sidebar-cat-0-en-105":3,"doc-seo-147237-105":59,"doc-detail-147237-en":130},{"code":4,"msg":5,"data":6},0,"success",[7,13,18,23,28,33,38,43,48,51,55],{"id":8,"doc_module":4,"doc_module_name":9,"category_name":10,"show_sort_weight":11,"slug":12},1,"Document","Story & Novel",90,"story-novel",{"id":14,"doc_module":4,"doc_module_name":9,"category_name":15,"show_sort_weight":16,"slug":17},2,"Literature",80,"literature",{"id":19,"doc_module":4,"doc_module_name":9,"category_name":20,"show_sort_weight":21,"slug":22},4,"Exam",70,"exam",{"id":24,"doc_module":4,"doc_module_name":9,"category_name":25,"show_sort_weight":26,"slug":27},5,"Comic",60,"comic",{"id":29,"doc_module":4,"doc_module_name":9,"category_name":30,"show_sort_weight":31,"slug":32},6,"Technology",50,"technology",{"id":34,"doc_module":4,"doc_module_name":9,"category_name":35,"show_sort_weight":36,"slug":37},7,"Healthcare",40,"healthcare",{"id":39,"doc_module":4,"doc_module_name":9,"category_name":40,"show_sort_weight":41,"slug":42},8,"Research & Report",30,"research-report",{"id":44,"doc_module":4,"doc_module_name":9,"category_name":45,"show_sort_weight":46,"slug":47},9,"Religion & Spirituality",20,"religion-spirituality",{"id":46,"doc_module":4,"doc_module_name":9,"category_name":49,"show_sort_weight":46,"slug":50},"World Cup","world-cup",{"id":52,"doc_module":4,"doc_module_name":9,"category_name":53,"show_sort_weight":52,"slug":54},10,"Lifestyle","lifestyle",{"id":56,"doc_module":4,"doc_module_name":9,"category_name":57,"show_sort_weight":24,"slug":58},19,"General","general",{"code":4,"msg":60,"data":61},"ok",{"site_id":62,"language":63,"slug":64,"title":65,"keywords":66,"description":67,"schema_data":68,"social_meta":123,"head_meta":125,"extra_data":127,"updated_unix":129},105,"en","llmdet-learning-strong-open-vocabulary-object-detectors-under-the-supervision-of-large-language-models-cvpr-2025-supplemental-materials","LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models - CVPR 2025 Supplemental Materials","","Supplementary materials for LLMDet present ablation evidence on how different large language model backbones and vision-language components affect detector performance, including rare-class gains when using LLaVA-OneVision as a multimodal foundation. The document details a multi-step training pipeline that builds a stronger large vision-language model by combining vision experts (LLMDet and SigLIP), concatenating their features, and learning a projector into the LLM input space. It also reports benchmark comparisons across VQA, hallucination, and comprehensive understanding settings, and summarizes observed limitations in long-form image description generation.",{"@graph":69,"@context":122},[70,84,105],{"@type":71,"itemListElement":72},"BreadcrumbList",[73,77,79,82],{"item":74,"name":75,"@type":76,"position":8},"https://docshare.wps.com","Home","ListItem",{"item":78,"name":9,"@type":76,"position":14},"https://docshare.wps.com/document/",{"item":80,"name":40,"@type":76,"position":81},"https://docshare.wps.com/document/research-report/",3,{"item":83,"name":65,"@type":76,"position":19},"https://docshare.wps.com/document/llmdet-learning-strong-open-vocabulary-object-detectors-under-the-supervision-of-large-language-models-cvpr-2025-supplemental-materials/147237/",{"url":83,"name":65,"@type":85,"image":86,"author":91,"headline":65,"publisher":94,"fileFormat":97,"inLanguage":63,"description":67,"dateModified":98,"datePublished":99,"encodingFormat":97,"isAccessibleForFree":100,"interactionStatistic":101},"DigitalDocument",{"url":87,"@type":88,"width":89,"height":90},"https://docshare.wps.com/thumbnails/llmdet-learning-strong-open-vocabulary-object-detectors-under-the-supervision-of-large-language-models-cvpr-2025-supplemental-materials/147237.png","ImageObject",300,407,{"name":92,"@type":93},"Liam","Person",{"url":74,"name":95,"@type":96},"DocShare","Organization","application/pdf","2026-09-17","2026-08-26",true,{"@type":102,"interactionType":103,"userInteractionCount":29},"InteractionCounter",{"@type":104},"ViewAction",{"@type":106,"mainEntity":107},"FAQPage",[108,114,118],{"name":109,"@type":110,"acceptedAnswer":111},"How do different large language models affect LLMDet’s detector performance?","Question",{"text":112,"@type":113},"The ablations compare LLM backbones and show that using the LLaVA-OneVision vision-language setup yields improvements, with noticeable gains on rare classes. Increasing LLM size only slightly improves performance, suggesting benefits concentrate on reasoning rather than visual representations.","Answer",{"name":115,"@type":110,"acceptedAnswer":116},"What training pipeline does LLMDet use to build a stronger large multimodal model?",{"text":117,"@type":113},"The pipeline uses multi-step training: first pretraining a projector, then finetuning the multimodal model with visual instruction tuning. It combines visual features from two encoders (LLMDet and SigLIP) and maps them into the LLM’s input space.",{"name":119,"@type":110,"acceptedAnswer":120},"What limitations are observed in the co-trained LLM-detector output?",{"text":121,"@type":113},"Even with prompts for detailed descriptions, the co-trained system tends to produce relatively short descriptions for whole images. Region-level outputs are also described as relatively simple grounding phrases, motivating collection of more informative region data.","https://schema.org",{"og:url":83,"og:type":124,"og:title":65,"og:site_name":95,"og:description":67},"article",{"robots":126,"canonical":83},"index,follow",{"doc_id":128,"site_id":62},147237,1787760626,{"code":4,"msg":5,"data":131},{"doc_id":128,"user_id":132,"nickname":92,"user_avatar":133,"doc_module":4,"category_id":39,"category_name":40,"doc_title":65,"doc_description":67,"doc_content":134,"file_id":135,"file_url":136,"file_type":137,"file_size":138,"view_count":29,"is_deleted":4,"is_public":8,"is_downloadable":8,"audit_status":8,"page_count":34,"language":139,"language_code":63,"site_id":62,"html_lang":63,"table_of_contents":140,"faqs":141,"seo_title":142,"seo_description":67,"update_tm":129,"read_time":143},8796095461564,"https://ap-avatar.wpscdn.com/davatar_155a257f0dc6eb9ab79c44ca47cae57d","Supplementary Materials for  \nLLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models  \nShenghao Fu 1 ,3 ,4 , Qize Yang3 , Qijie Mo 1 ,4 , Junkai Yan 1 ,4 , Xihan Wei3 ,  \nJingke Meng 1 ,4 , Xiaohua Xie 1 ,4 ,5 *, Wei-Shi Zheng 1 ,2 ,4 ,6∗  \n1 School of Computer Science and Engineering, Sun Yat-sen University, China;  \n2Peng Cheng Laboratory, China; 3Tongyi Lab, Alibaba Group;  \n4 Key Laboratory of Machine Intelligence and Advanced Computing, Ministry of Education, China;  \n5 Guangdong Province Key Laboratory of Information Security Technology, China;  \n6Pazhou Laboratory (Huangpu), China  \n[fushh7@mail2.sysu.edu.cn](fushh7@mail2.sysu.edu.cn), [xiexiaoh6@mail.sysu.edu.cn](xiexiaoh6@mail.sysu.edu.cn), [wszheng@ieee.org](wszheng@ieee.org)  \n\n| LLM | AP | APr | APc | APf |\n| --- | --- | --- | --- | --- |\n| Qwen2-0.5b-instruct [12] | 44.4 | 36.4 | 39.2 | 50.5 |\n| LLaVA-OneVision-0.5b-ov [4] | 44.5 | 38.6 | 39.3 | 50.3 |\n| Qwen2-1.5b-instruct [12] | 44.6 | 35.3 | 39.5 | 50.8 |\n\nTable 1-1 . Ablations on large language models.  \n1. More Ablation Studies  \nEffect of different large language models. By default, we use the LLM in LLaVA-OneVision-0.5b-ov [4], which is finetuned from Qwen2-0 .5b-instruct [12] . Since the LLMin LLaVA-OneVision-0.5b-ov is pretrained with abundant multi-modal data but with a different vision encoder, thepretraining can still improve the performance, especially for rare classes (+2 .2% APr ), as shown in Table 1-1. But we find that increasing the size of the LLM only slightly improves the performance, perhaps larger language models mainly improve in reasoning ability which does not benefit the detector’s visual representations.  \n2. LLMDet Builds a Stronger Large VisionLanguage Model  \nIn this subsection, we show that LLMDet can serve as a general vision foundation model and in turn gets a strong large multi-modal model. Recent large multi-modal models (LMM) are based on pretrained large language models and pretrained vision foundation models. Different vision foundation models will significantly affect the performance ofLMMs [16] . Since LLMDetis enhanced under the  \n* : Corresponding authors are Xiaohua Xie and Wei-Shi Zheng.  \nPart of the work was done when Shenghao Fu was an intern at Alibaba.  \n| goal | Alignment | Training LMM |\n| --- | --- | --- |\n| task | Image Captioning | Instruction Following |\n| loss | llm loss | llm loss |\n| dataset | LSC 558k recap | LLaVA 1.5 instruction tuning dataset (665K) |\n\nFigure 2-1 . The multi-step training pipeline of using LLMDet to build a strong large multi-modal model. The large multi-modal model uses a mixture of vision encoders, including LLMDet and SigLIP. In each step, modules in orange color are tunable while modules in blue color are frozen. We first pretrain a new projector and then finetune the large multi-modal model with visual instruct tuning.  \nsupervision of long detailed image-level captions and prealigned with LLM, LLMDet inherits great potential to build a stronger LMM. Following recent advances [10, 11, 16], we build the LMM using a mixture of vision experts, i.e. a SigLIP [14] vision encoder and our LLMDet. As shown in Figure 2-1, the visual features from two vision encoders are concatenated along the channel dimension, and then a projector is utilized to map the features to the LLM’s input space. We start from LLaVA-OneVision-0.5b-ov [4] and  \n\n| Method | GQA\u003Cbr>[3] | POPE [6] |  |  | MME [2] |  |\n| --- | --- | --- | --- | --- | --- | --- |\n|  |  | rand | pop | adv | perception | cognition |\n| OneVision-0.5b | 56.9 | 87.5 | 86.3 | 85.0 | 1238 | 240 |\n| OneVision-0.5b\u003Cbr>+MM-GDINO | 61.2 | 88.9 | 88.1 | 86.6 | 1207 | 256 |\n| OneVision-0.5b +LLMDet | 61.2 | 88.8 | 88.0 | 86.0 | 1297 | 264 |\n\nTable 2-2 . Multi-modal performance using different vision encoders. OneVision-0 .5b is short for LLaVA-OneVision-0 .5bov [4] .  \ninsert our LLMDet to it as shown in Figure 2-1. We first pretrain a new projector a","cbCaie8Cwzd8yJZi","https://ap.wps.com/l/cbCaie8Cwzd8yJZi","pdf",954662,"English","# Ablations\n## More Ablation Studies\n## LLMDet Builds a Stronger Large VisionLanguage Model\n## Limitations\n# Implement Details of Zero-Shot Test on Referring Expression Comprehension","[{\"question\":\"How do different large language models affect LLMDet’s detector performance?\",\"answer\":\"The ablations compare LLM backbones and show that using the LLaVA-OneVision vision-language setup yields improvements, with noticeable gains on rare classes. Increasing LLM size only slightly improves performance, suggesting benefits concentrate on reasoning rather than visual representations.\"},{\"question\":\"What training pipeline does LLMDet use to build a stronger large multimodal model?\",\"answer\":\"The pipeline uses multi-step training: first pretraining a projector, then finetuning the multimodal model with visual instruction tuning. It combines visual features from two encoders (LLMDet and SigLIP) and maps them into the LLM’s input space.\"},{\"question\":\"What limitations are observed in the co-trained LLM-detector output?\",\"answer\":\"Even with prompts for detailed descriptions, the co-trained system tends to produce relatively short descriptions for whole images. Region-level outputs are also described as relatively simple grounding phrases, motivating collection of more informative region data.\"}]","LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models - CVPR 2025 Supplemental Materials | PDF",18]