[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83453-en":3,"doc-seo-83453-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83453,1099513958607,"Jiven","https://ap-avatar.wpscdn.com/avatar/100002390cf8733938c?x-image-process=image/resize,m_fixed,w_180,h_180&k=1778829742770036399",8,"Research & Report","Steal the Patch Size: Adversarially Manipulate Vision-Language Models","A black-box model-stealing attack recovers private vision-tokenizer configurations of deployed vision-language models, specifically the vision patch size and the input preprocessing pipeline. The method exploits a task-level side channel created by ViT-style patchification: synthetic grid images aligned with the hidden patch grid erase boundary cues at tokenization, producing periodic accuracy drops. Sweeping grid cell sizes infers patch size, while padding and a consistency-check identify whether preprocessing is dynamic or fixed-resolution and recover resize targets. Results across Qwen-VL variants and models such as GPT and Claude show reliable tokenizer recovery and enable preprocessing-aware transfer attacks and model-targeted adversarial manipulation.","Steal the Patch Size: Adversarially Manipulate Vision-Language Models  \nKai Hu 1 Akash Bharadwaj 1 Weichen Yu 1 Matt Fredrikson 1  \narXiv :2607 .00174v1 [ cs .CV] 30 Jun 2026  \nAbstract  \nWe present a black-box model-stealing attack that recovers private vision-tokenizer configurations of deployed vision-language models (VLMs), including the visual patch size and input preprocessing pipeline. The key idea is a task-level side channel induced by ViT-style patchification:  \nwhen a synthetic grid image is aligned with the hidden patch grid, boundary cues are erased at tokenization, causing periodic accuracy drop. By sweeping the grid cell size and measuring these collapses, we infer the patch size; by introducing padding and a consistency-check test, we further identify whether preprocessing is dynamicor fixed-resolution and recover the target resize resolution. Across open-source Qwen-VL variants and proprietary models including GPT and Claude, we reliably recover tokenizer-related parameters. Finally, we show that such leakage enables preprocessing-aware transfer attacks and model-targeted adversarial manipulation.  \n1. Introduction  \nVision-Language Models (VLMs) deployed via public APIs rely on complex and often undocumented visual preprocessing pipelines. Beyond model weights, these pipelines include architectural and system-level design choices such as vision patch size, input resizing strategies, padding/cropping rules, and target resolutions. These parameters materially affect efficiency, accuracy, and robustness, yet are typically treated as private deployment-time information and are not disclosed in APIs or model documentation.  \nA “blind spot” in black-box VLMs. Despite impressive capabilities, black-box VLMs exhibit striking failures on tasks that are trivial for humans. Consider a simple grid-size counting query: given an image of a colored N × N grid  \n1Carnegie Mellon University, Pittsburgh, USA. Correspondence to: Kai Hu \u003C[kaihu@cmu.edu](kaihu@cmu.edu) > .  \nProceedings of the 43 rd International Conference on Machine Learning, Seoul, South Korea. PMLR 306, 2026 . Copyright 2026 by the author(s) .  \n(e.g., Figure 8), ask the model to report N. Humans can answer this at a glance. However, we observe that state-ofthe-art models (e.g., GPT and Claude) can abruptly collapse on this task for specific cell sizes, producing incorrect countseven when the image remains crisp and unambiguous.  \nThe “Visual Strawberry” analogy Just as LLMs struggle to count letters in “strawberry” because BPE tokenization hides individual characters, we show that VLMs fail to count grid cells because patch tokenization (patchification) hides visual edges. In both cases, the structural granularity of the tokenizer mismatches the granularity of the information, creating a blind spot.  \nMechanism: Patch-Size Matching (PSM). Concretely, most VLMs tokenize images via a vision transformer (ViT) that partitions inputs into non-overlapping square patches, followed by a patch projection. When salient boundaries (e.g., grid lines) consistently fall between patch interiors after preprocessing, boundary cues can be suppressed at the tokenization stage, yielding nearly uniform patch tokens that lack edge information. This effect is not accidental: as we sweep the grid cell size D , the alignment condition recurs, causing periodic accuracy collapses. We refer to this phenomenon as Patch-Size Matching (PSM): failures occur when the grid frequency becomes commensurate with the patch sampling frequency (e.g., D = kP after preprocessing), creating a repeatable task-level side channel.  \nFrom phenomenon to stealing: recovering private hyperparameters. This paper studies the following question:  \nCan private architectural and preprocessing parameters of black-box VLMs be systematically recovered through API access alone?  \nWe answer this in the affirmative. We present a black-box model stealing attack that recovers the visual patch size and the input preproces","cbCaidO07aof11Gq","https://ap.wps.com/l/cbCaidO07aof11Gq","pdf",5059142,5,1,15,"English","en",105,"# Introduction\n## Blind spot in black-box VLMs\n## Visual Strawberry analogy\n## Mechanism: Patch-Size Matching (PSM)\n## From phenomenon to stealing\n## Stealing patch size and preprocessing under unknown resizing/padding\n## Implications for model security","[{\"question\":\"What private information does the attack recover from deployed vision-language models?\",\"answer\":\"It recovers vision-tokenizer-related hyperparameters, including the visual patch size and the input preprocessing pipeline.\"},{\"question\":\"How does Patch-Size Matching (PSM) cause periodic accuracy drops?\",\"answer\":\"When a synthetic grid aligns with the hidden patch grid, boundary cues can be suppressed at tokenization, yielding patch tokens that lack edge information. Sweeping grid cell sizes makes this alignment recur and accuracy collapses become periodic.\"},{\"question\":\"How does the method handle unknown resizing, padding, or cropping during preprocessing?\",\"answer\":\"It uses a three-stage approach: detect dynamic versus fixed-resolution preprocessing, recover patch size up to a scaling factor under unknown resizing, and identify the target input resolution using a consistency-check based on hypothesis testing.\"}]",1784188064,38,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"steal-the-patch-size-adversarially-manipulate-vision-language-models","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/steal-the-patch-size-adversarially-manipulate-vision-language-models/83453/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-25","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What private information does the attack recover from deployed vision-language models?","Question",{"text":76,"@type":77},"It recovers vision-tokenizer-related hyperparameters, including the visual patch size and the input preprocessing pipeline.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How does Patch-Size Matching (PSM) cause periodic accuracy drops?",{"text":81,"@type":77},"When a synthetic grid aligns with the hidden patch grid, boundary cues can be suppressed at tokenization, yielding patch tokens that lack edge information. Sweeping grid cell sizes makes this alignment recur and accuracy collapses become periodic.",{"name":83,"@type":74,"acceptedAnswer":84},"How does the method handle unknown resizing, padding, or cropping during preprocessing?",{"text":85,"@type":77},"It uses a three-stage approach: detect dynamic versus fixed-resolution preprocessing, recover patch size up to a scaling factor under unknown resizing, and identify the target input resolution using a consistency-check based on hypothesis testing.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":20,"slug":138},19,"General","general"]