[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85117-en":3,"doc-seo-85117-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},85117,2336464648746,"Skyler","https://ap-avatar.wpscdn.com/davatar_276721f389ce27ea32af1340a28f341c",8,"Research & Report","Format Sensitivity Index Token-Controlled Prompt Wrapper Robustness and Schema Compliance in LLM Benchmarking","Prompt wrappers often differ only in formatting, yet they can materially change model scores and reverse leaderboard conclusions. The study introduces the Format Sensitivity Index (FSI) to quantify wrapper-induced accuracy variation and the Parseability Sensitivity Index (PSI) to measure wrapper-driven changes in answer parseability. Using a token-controlled protocol across 140,000 OpenRouter generations for 7 QA tasks, 5 wrapper families, and 4 instruct models (7B–72B), mean FSI varies by over 30×, driven largely by compliance failures. Regression analysis shows parseability strongly mediates accuracy, motivating statistically robust reporting and practical recommendations for benchmarking and structured-output deployments.","arXiv :2607 .09665v 1 [ cs .AI] 2 May 2026  \nFormat Sensitivity Index: Token-Controlled Prompt Wrapper Robustness and Schema Compliance in LLM Benchmarking  \nDeep Pankajbhai Mehta  \nAdobe Inc.  \nAbstract  \nPrompt wrappers often differ only in formatting, yet they can change model scores enough to flip leaderboard conclusions. We study this variance under a token-controlled protocol and introduce two complementary metrics: the Format Sensitivity Index (FSI), the accuracy range induced by wrapper choice, and the Parseability Sensitivity Index (PSI), the corresponding range in answer parseability. Across 140,000 OpenRouter generations spanning 7 question-answering tasks, 5 wrapper families, and 4 instruct models from 7B to 72B parameters, we find that mean FSI varies by more than 30 times across models and is largely explained by compliance failures. A fixed-effects regression shows that parseability remains a strong predictor of accuracy even after controlling for task, model, and wrapper. We argue that reporting accuracy without wrapper variance and compliance is statistically fragile, and we provide practical recommendations for both benchmarking and structured-output deployments.  \n1 Introduction  \nBenchmarks increasingly evaluate large language models (LLMs) through prompts that include a wrapper: delimiters, JSON instructions, step-by-step templates, and similar formatting that is meant to be semantically irrelevant. In practice, wrappers are rarely standardized across papers, toolkits, or vendors. This creates a methodological failure mode: a model can look strong or weak depending on the wrapper rather than the underlying capability.  \nOur experiments expose an extreme example. On the same set of 7 tasks, the same model can score near random accuracy under a strict JSON wrapper while exceeding 0.75 accuracy under a delimiter-based structured wrapper. The difference is not subtle prompting artistry. It is frequently a compliance problem: if the output is not parseable by the benchmark’s answer extractor, the run is scored as incorrect.  \nPrior work has documented sensitivity to prompt formatting in few-shot settings (Sclar et al. , 2024; Chatterjee et al. , 2024) . However, most benchmarking pipelines still report a single number per model per task and often ignore structured-output failure modes that dominate real deployments, including tool calls, schema adherence, and database queries. We contribute a simple, reproducible way to quantify wrapper-driven variance in a setting that resembles common evaluation practice.  \nContributions.  \n• Two sensitivity metrics. We define FSI and PSI as wrapper-induced ranges of accuracy and parseability, and we report bootstrap confidence intervals plus a normalized sensitivity variant to reduce dependence on mean accuracy.  \n• Token-controlled evaluation. We implement a protocol that pads prompts to a fixed character budget, logs realized prompt token counts, and quantifies residual token spread.  \n• Large empirical study. We evaluate 4 instruct models across 7 tasks and 5 wrapper families, yielding 140,000 generations, and show that format sensitivity differs sharply across models.  \n• Mechanism analysis. We show that parseability strongly mediates accuracy differences, supported by correlation and fixed-effects regression.  \n2 Related Work  \nPrompt-format sensitivity. LLMs can be brittle to meaning-preserving prompt changes, including formatting (Sclar et al. , 2024; Chatterjee et al. , 2024) . PromptBench studies adversarial prompt robustness (Zhu et al. , 2023), and surveys catalog broader evaluation risks (Chang et al. , 2023) . Our setting is narrower but common: instruction-following, single-turn QA, and wrappers that aim to change only output format.  \nBenchmarking methodology. Evaluation frameworks such as HELM emphasize standardization and transparency (Liang et al. , 2022) . Toolkits such as lm-evaluation-harness provide a shared interface for benchmark execution (EleutherAI","cbCaihTmEnr5cRSe","https://ap.wps.com/l/cbCaihTmEnr5cRSe","pdf",1087784,2,1,12,"English","en",105,"# Introduction\n# Related Work\n# Experimental Setup\n## Tasks and Models\n## Prompt Wrappers\n## Results and Metrics","[{\"question\":\"What problem does the paper identify in LLM benchmarking?\",\"answer\":\"Prompt wrappers that differ only in formatting can change reported scores enough to flip leaderboard rankings, creating a methodological failure mode unrelated to underlying capability.\"},{\"question\":\"How do FSI and PSI quantify wrapper-induced effects?\",\"answer\":\"FSI measures the accuracy range induced by wrapper choice, while PSI measures the corresponding range in answer parseability under the same wrapper variations.\"},{\"question\":\"What main factor explains differences in accuracy across wrappers and models?\",\"answer\":\"Compliance failures are identified as a dominant driver, and parseability is shown to strongly mediate accuracy differences even after controlling for task, model, and wrapper.\"}]",1784201209,30,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"format-sensitivity-index-token-controlled-prompt-wrapper-robustness-and-schema-compliance-in-llm-benchmarking","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/format-sensitivity-index-token-controlled-prompt-wrapper-robustness-and-schema-compliance-in-llm-benchmarking/85117/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the paper identify in LLM benchmarking?","Question",{"text":75,"@type":76},"Prompt wrappers that differ only in formatting can change reported scores enough to flip leaderboard rankings, creating a methodological failure mode unrelated to underlying capability.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How do FSI and PSI quantify wrapper-induced effects?",{"text":80,"@type":76},"FSI measures the accuracy range induced by wrapper choice, while PSI measures the corresponding range in answer parseability under the same wrapper variations.",{"name":82,"@type":73,"acceptedAnswer":83},"What main factor explains differences in accuracy across wrappers and models?",{"text":84,"@type":76},"Compliance failures are identified as a dominant driver, and parseability is shown to strongly mediate accuracy differences even after controlling for task, model, and wrapper.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":29,"slug":121},"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]