[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83420-en":3,"doc-seo-83420-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83420,7971461741311,"Ophelia","https://ap-avatar.wpscdn.com/avatar/74000253aff267980c6?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779345379180704826",8,"Research & Report","Validity of LLMs as Data Annotators: AMALIA on Authority","A national language model offers a linguistic community an instrument for quantifying what citizens say and value. Portugal’s publicly funded AMALIA, a 9B-parameter European Portuguese model, matches trained human coders on agreement, within six points, when asked to annotate the moral foundation of authority. The work distinguishes agreement from construct validity and tests decomposition-based calibration to recover validity. Calibration transfer to smaller models and Portuguese yields negative results: decomposition restores at most about half of holistic performance, suggesting surface-pattern miscoding and limiting measurement-only reliability.","arXiv :2607 .0873 1v 1 [ cs .CL] 9 Jul 2026  \nValidity of LLMs as data annotators: AMALIA on authority  \nA Preprint  \nManuel Pita  \nArtificial Intelligence, Social Interaction and Complexity Laboratory  \nCICANT, Universidade Lusófona  \n[manuel.pita@ulusofona.pt](manuel.pita@ulusofona.pt)  \nJuly 2026  \nAbstract  \nA national language model offers a linguistic community its own instrument for measuring what its citizens say and value. Portugal’s publicly funded AMALIA, a 9B-parameter model for European Portuguese, appears competitive on agreement alone. Asked to code the moral foundation of authority—a construct that must be inferred rather than read off the text—it agrees with trained human coders to within six points of open models eight to thirteen times its size. Yet with LLMs, the distinction between reliability, as agreement, and validity, as faithfulness to the construct’s theory, becomes critical. The instrument is a black box: we see the codes but not the inferential path that produced them, so agreement cannot distinguish measurement from pattern-matching. Decomposing the prompt into the codebook’s atomic clauses lets us probe the box systematically; the recovery gap is the extent to which decomposition fails to reproduce the original prompt’s performance. Calibration aims to close that gap. If calibration closes the gap, some portability should survive across models and languages, since the target theoretical construct does not change; where it does not, the construct–model instrument is the most likely locus of failure. We test whether a calibrated English instrument transfers to a smaller national model, AMALIA-9B, and to European Portuguese. For a single construct, we find evidence that it does not. At best, decomposition recovers only about half of AMALIA’s holistic performance. Analysis of miscoding suggests that AMALIA often relies on surface correlates, such as moral outrage near authority figures, for the specific annotation task. Crucially, an open multilingual LLM closes the gap on the same Portuguese corpus, under the same instructions, which points away from the corpus as the main explanation. The implications extend to the sovereign-LLM programmes that a growing number of countries are prioritizing. AMALIA can still screen and pre-code at scale; it cannot yet measure theoretical constructs—even one as explicitly specified as authority—well enough to stand on its own. The evidence is one construct and one corpus, a single counterexample. It leaves open the question of whether native-language instruction narrows the gap. Still, it specifies what a sovereignty programme must settle before entrusting a model with measurement: not only that the model agrees with human coders, but what evidential route warrants that agreement.  \nKeywords large language models · text as data · LLM sovereignty · construct validity · data annotation  \n1 Introduction  \nOn 1 July 2026, Portugal released AMALIA, a nine-billion-parameter language model 1 , built with public funds to serve European Portuguese [Simplício et al., 2026] . The release joins a growing family of national and sovereign models. Thailand’s Typhoon project, for instance, treats sovereign LLMs as a class of their own, models that a national institution can hold, inspect and operate without depending on the small set of private organizations that currently own these systems [Pipatanakul and Taveekitworachai, 2026] . The motivation is both empirical and sociopolitical. Empirically, the dominant LLMs were trained mostly for use in English [Rao et al., 2025] . Sociopolitically, their operation remains concentrated in a restricted set of private organizations, which makes a nation’s capacity to hold and audit these systems part of a wider agenda of digital sovereignty. This second dimension has reached the institutional level as well. Greece has trained a national LLM on its own parliamentary corpora as part of its approach to digital sovereignty [Fitsilis et al. , 202","cbCaiaviJeFlww10","https://ap.wps.com/l/cbCaiaviJeFlww10","pdf",466405,4,1,15,"English","en",105,"# Abstract\n# 1 Introduction\n## Local and sovereign LLM motivation\n## Text-as-data and construct annotation\n## Why AMALIA’s release record matters","[{\"question\":\"What question does the paper address about AMALIA’s use in text-as-data annotation?\",\"answer\":\"It evaluates whether AMALIA’s annotations are valid as measurements of an inferred theoretical construct (authority), not merely whether they agree with human coders.\"},{\"question\":\"How do agreement and validity differ in the authors’ framework?\",\"answer\":\"Agreement measures coder concordance, while validity requires faithfulness to the construct’s underlying theory, which may not be identifiable from the text alone.\"},{\"question\":\"What do the results show about using a calibrated instrument across models and languages?\",\"answer\":\"The calibrated English instrument does not reliably transfer to AMALIA-9B or to European Portuguese for the single construct tested, and decomposition recovers only about half of holistic performance at best.\"}]",1784187476,38,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"validity-of-llms-as-data-annotators-amalia-on-authority","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/validity-of-llms-as-data-annotators-amalia-on-authority/83420/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What question does the paper address about AMALIA’s use in text-as-data annotation?","Question",{"text":75,"@type":76},"It evaluates whether AMALIA’s annotations are valid as measurements of an inferred theoretical construct (authority), not merely whether they agree with human coders.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How do agreement and validity differ in the authors’ framework?",{"text":80,"@type":76},"Agreement measures coder concordance, while validity requires faithfulness to the construct’s underlying theory, which may not be identifiable from the text alone.",{"name":82,"@type":73,"acceptedAnswer":83},"What do the results show about using a calibrated instrument across models and languages?",{"text":84,"@type":76},"The calibrated English instrument does not reliably transfer to AMALIA-9B or to European Portuguese for the single construct tested, and decomposition recovers only about half of holistic performance at best.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]