[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83280-en":3,"doc-seo-83280-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83280,1374391974564,"Clementine","https://ap-avatar.wpscdn.com/avatar/14000253aa45c000a9e?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779874745381141002",8,"Research & Report","Does Bielik Know What It Doesn’t Know? Activation Dispersion Separates Entity Familiarity from Factual Reliability Across Model Scale","Large language models (LLMs) can hallucinate most readily about entities they have never seen. This work tests whether post-MLP activation patterns reveal entity familiarity before any answer is produced, and whether that signal predicts factual reliability. Using four Polish Bielik models (1.5B–11B) across athletes, cities, writers, and musicians, inverse participation ratio and spectral entropy separate known from fabricated entities with AUROC 0.95–1.00, transferring across domains. Yet behavioral factual reliability scales differently and is difficult to separate within known entities, indicating distinct phenomena.","arXiv :2607 .07670v 1 [ cs .CL] 8 Jul 2026  \nDoes Bielik Know What It Doesn’t Know? Activation Dispersion Separates Entity Familiarity  \nfrom Factual Reliability Across Model Scale ∗  \nGrzegorz Brzezinka  \nProsit AS  \n[greg@prosit. no](greg@prosit. no)  \nJuly 2026  \nAbstract  \nLarge language models (LLMs) hallucinate most readily about entities they have never seen. We ask whether a model’s activations betray entity familiarity before a single answer token is generated-and whether that signal says anything about the factual reliability of the answers. On four Polish Bielik models (1.5B–11B parameters), we probe four entity domains-athletes, cities, writers, and musicians-each with 42 well-known, 42 obscure-but-real, and 42 fabricated entities addressed by a one-sentence question (504 prompts per model) . Two unsupervised, single-forward-pass dispersion measures over post-SwiGLUMLP activations-inverse participation ratio and spectral entropy-separate known from fabricated entities with AUROC 0.95–1.00 across all four domains and every scale; a supervised linear probe reaches 0.99–1.00. Both clear a selection-aware permutation floor of about 0.70–0.74 (95th percentile of a null that re-selects the best layer per permutation; empirical p ≤ 10 −3), survive held-out layer selection (0.93–0.99), and persist on real names: known versus real-but-obscure entities separate at 0.96–1.00, so fabricated-string artifacts are not the driver. The signal transfers across entity types: a probe trained on one domain and evaluated on another retains a mean off-diagonal AUROC of 0.92–0.99, with the only large drops occurring where the prompt template also changes (cities use a different question stem) . A matched-template counterfactual-re-extracting cities and writers under one shared neutral stem-shows these drops are template-caused: transfer into cities recovers to 0.999–1.000 at every scale, leaving only a residual cities-as-source asymmetry. Per-head attention-entropy analysis shows the signal is diffuse rather than carried by a small fixed set of heads. This representational signal, taken as the better of the two metrics, is at ceiling on this contrast already at 1.5B-whether familiarity itself still improves with scale is not distinguishable at this ceiling. Behavioral factual reliability instead scales sharply: 0, 2, 10, and 19 of 42 known athletes are answered fully correctly by the 1.5B, 4.5B, 7B, and 11B models under a strict judge (6, 16, 24, 33 under a soft key-facts rubric) . Within known entities, separating correct answers from hallucinations is much harder: the probe reaches 0.93, and dispersion metrics do no better than a first-token-entropy baseline. A five-sample semantic-entropy baseline attains 0.71–0.83 on the known-vs-fabricated contrast at five times the inference cost, though it wins the correctness contrast (up to 0.87) where dispersion fails. Despite this representational awareness, the models almost never hold back: an LLM audit of all 2,520 sampled athlete answers finds 2 refusals and 1 hedged answer (99.88% direct assertions), both refusals from the largest model. Entity familiarity and factual reliability are distinct phenomena moving on different scaling curves. (Behavioral labels, semantic entropy, and the refusal audit are reported for athletes only; the other three domains carry condition-based detection labels.)  \n1 Introduction  \nWhen an LLM is asked about something it does not know, does its internal state look different from when it is on familiar ground? A physically motivated intuition says yes: knowledge retrieval should behave like a  \n∗ Preprint. Code and data are available at the project repository: [https://github.com/agentGreg/](https://github.com/agentGreg/)[ ](https://github.com/agentGreg/)bielik-hallucination-detection.  \nlocalized excitation: a compact, selective activation pattern (memory retrieval), whereas confabulation should look delocalized, with activation smeared broadly across the netwo","cbCaigk9HkDcbnhm","https://ap.wps.com/l/cbCaigk9HkDcbnhm","pdf",843058,2,1,23,"English","en",105,"# Introduction\n## Activation dispersion and the familiarity vs reliability hypothesis\n## Motivation from sparse context and mechanistic knowledge circuits\n## Prior work on unsupervised internal-state detection","[{\"question\":\"What question does the paper investigate about Bielik models?\",\"answer\":\"It asks whether a model’s internal activations reveal entity familiarity before generating any answer token, and whether this signal relates to the factual reliability of the responses.\"},{\"question\":\"How do the authors measure entity familiarity without labels?\",\"answer\":\"They compute unsupervised dispersion metrics from post-SwiGLU MLP activations in a single forward pass, specifically inverse participation ratio and spectral entropy, to distinguish known from fabricated entities.\"},{\"question\":\"Do activation-based familiarity signals also predict behavioral factual reliability?\",\"answer\":\"Not fully. The representational familiarity signal separates known from fabricated entities strongly, but behavioral factual reliability scales on a different curve and within-known distinctions are much harder than the dispersion metrics capture.\"}]",1784186470,58,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"does-bielik-know-what-it-doesnt-know-activation-dispersion-separates-entity-familiarity-from-factual-reliability-across-model-scale","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/does-bielik-know-what-it-doesnt-know-activation-dispersion-separates-entity-familiarity-from-factual-reliability-across-model-scale/83280/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-25","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What question does the paper investigate about Bielik models?","Question",{"text":75,"@type":76},"It asks whether a model’s internal activations reveal entity familiarity before generating any answer token, and whether this signal relates to the factual reliability of the responses.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How do the authors measure entity familiarity without labels?",{"text":80,"@type":76},"They compute unsupervised dispersion metrics from post-SwiGLU MLP activations in a single forward pass, specifically inverse participation ratio and spectral entropy, to distinguish known from fabricated entities.",{"name":82,"@type":73,"acceptedAnswer":83},"Do activation-based familiarity signals also predict behavioral factual reliability?",{"text":84,"@type":76},"Not fully. The representational familiarity signal separates known from fabricated entities strongly, but behavioral factual reliability scales on a different curve and within-known distinctions are much harder than the dispersion metrics capture.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]