[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85867-en":3,"doc-seo-85867-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},85867,1099514068035,"Ezra","https://ap-avatar.wpscdn.com/davatar_276721f389ce27ea32af1340a28f341c",8,"Research & Report","One Token Is Enough: Fingerprinting and Verifying Large Language Models from Single-Token Output Distributions","Large language models are increasingly accessed through opaque inference chains where clients cannot verify that the returned output comes from the advertised model. This work proposes a behavioral fingerprint defined as the empirical distribution of answers to trivial one-word prompts, collected at a cost of one output token per query. Experiments on 165 models via a commercial aggregator show non-uniform, model-specific distributions, successful lineage recovery via Jensen–Shannon divergence, and a biometric-style verification protocol with low error rates.","One Token Is Enough: Fingerprinting and Verifying Large Language Models from Single-Token Output  \nDistributions  \nTomˇs Bruckner  \narXiv :2607 . 10252v 1 [ cs .CR] 11 Jul 2026  \nAbstract—Large language models (LLMs) are increasingly consumed through opaque serving chains – API aggregators, resellers, and inference providers – in which the client has no technical means to confirm that the model answering is the model advertised, and recent audits show that a substantial fraction of commercial endpoints deviate from the vendor’s reference weights. Existing identification techniques require long generated texts, token-level log-probabilities, adversarially crafted prompts, or the model owner’s cooperation. We show that far weaker evidence suffices. We define a behavioral fingerprint of an LLMas the empirical distribution of its answers to trivial one-word prompts – “name a random number between 1 and 100” – collected across four languages at a cost of one output token per query. Measuring 165 models served via a large commercial aggregator (OpenRouter), we find that (i) these distributions are highly non-uniform (median cell entropy 1.0 bit) and modelspecific: split halves of the same model’s samples lie an order of magnitude closer than samples of different models; (ii) Jensen– Shannon divergence between fingerprints recovers model lineage, assigning a model to its documented family with 59.5% leaveone-out accuracy against an 18.4% chance rate; and (iii) a biometric-style verification protocol achieves a 7.3% equal error rate with the full 40-cell battery, and below 11% with eight probe cells – roughly a hundred single-token queries per audit. We further report ecosystem anomalies, including a proprietarybranded flagship endpoint distributionally indistinguishable from an open-weight Qwen model. The protocol, prompts, raw data, and analysis code are released for reproduction and operational use.  \nIndex Terms—Large language models, model fingerprinting, model attribution, API auditing, black-box verification, information forensics.  \nI. INTRODUCTION  \nTHE market for large language model inference has  \nrapidly stratified. Between the model creator and the end application now sit inference providers, resellers, and aggregators that route requests among dozens of upstream deployments. The client addresses a model by name – a string – and receives text. Nothing in this exchange proves that the advertised model produced the answer: the provider may substitute a cheaper model, an aggressively quantized variant, or an older version, and pocket the margin. This is not a hypothetical concern. Gao et al. [1] found that 11 of 31 commercial endpoints serving Llama models produced output distributions statistically incompatible with the vendor’s reference weights; Cai et al. [2] formalize this model substitution threat and show  \nT. Bruckner is with the Faculty of Informatics and Statistics, Prague University of Economics and Business, Prague, Czech Republic (e-mail: [bruckner@vse.cz](bruckner@vse.cz)) .  \nthat nave output-based checks are brittle under production nondeterminism; Zhu et al. [3] document the same concern for quantized variants. Aggregators themselves acknowledge heterogeneous quantization among their upstream providers.1 The economic incentive is structural: inference cost falls steeply with quantization and model size, while detection risk has so far been low.  \nVerifying which model is behind an API is therefore a forensic attribution problem, and it is harder than it looks. The client typically has (i) no access to weights or logits – many production APIs return text only; (ii) no ability to fine-tune or watermark the model, ruling out cooperative techniques [4], [5], [6]; and (iii) a budget: continuous auditing of many endpoints must cost cents, not dollars. Existing noncooperative identification methods fall short of at least oneof these constraints. Classifier-based attribution of generated text [7], [8] needs long ou","cbCaidr7ogRaoFsq","https://ap.wps.com/l/cbCaidr7ogRaoFsq","pdf",336219,5,1,9,"English","en",105,"# Abstract\n# Introduction\n## Threat model and verification challenge\n## Research questions (RQ1, RQ2)","[{\"question\":\"What problem does the paper address about large language model APIs?\",\"answer\":\"Clients cannot technically confirm that an API’s answers come from the advertised model when providers, resellers, and aggregators may substitute models, versions, or quantized variants.\"},{\"question\":\"How does the proposed fingerprinting method work?\",\"answer\":\"It defines a behavioral fingerprint as the empirical distribution of single-token answers to trivial one-word prompts, such as requesting a random number.\"},{\"question\":\"What evidence does the paper provide that fingerprints identify and verify models?\",\"answer\":\"Across 165 models, the distributions are highly non-uniform and model-specific, fingerprint distances recover model lineage with 59.5% leave-one-out accuracy, and a verification protocol achieves a 7.3% equal error rate with the full probe set.\"}]",1784206785,23,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"one-token-is-enough-fingerprinting-and-verifying-large-language-models-from-single-token-output-distributions","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/one-token-is-enough-fingerprinting-and-verifying-large-language-models-from-single-token-output-distributions/85867/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-26","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What problem does the paper address about large language model APIs?","Question",{"text":76,"@type":77},"Clients cannot technically confirm that an API’s answers come from the advertised model when providers, resellers, and aggregators may substitute models, versions, or quantized variants.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How does the proposed fingerprinting method work?",{"text":81,"@type":77},"It defines a behavioral fingerprint as the empirical distribution of single-token answers to trivial one-word prompts, such as requesting a random number.",{"name":83,"@type":74,"acceptedAnswer":84},"What evidence does the paper provide that fingerprints identify and verify models?",{"text":85,"@type":77},"Across 165 models, the distributions are highly non-uniform and model-specific, fingerprint distances recover model lineage with 59.5% leave-one-out accuracy, and a verification protocol achieves a 7.3% equal error rate with the full probe set.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,110,115,120,123,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":20,"slug":137},19,"General","general"]