[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83446-en":3,"doc-seo-83446-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83446,1099513958607,"Jiven","https://ap-avatar.wpscdn.com/avatar/100002390cf8733938c?x-image-process=image/resize,m_fixed,w_180,h_180&k=1778829742770036399",8,"Research & Report","Benchmarking Frontier LLMs on Arabic Cultural and Sociolinguistic Knowledge: A Cross-Evaluation Framework with Human SME Ground Truth","Human expert evaluation is a key bottleneck for deploying language models in specialized, high-stakes settings, especially for Arabic sociolinguistic knowledge where credible grading requires both linguistic competence and deep cultural familiarity. A cross-evaluation framework is built for Egyptian and Iraqi Arabic, producing 103 SME-validated prompt–rubric pairs (Cultural vs Linguistic) with penalty-weighted rubrics. Multiple frontier LLMs are evaluated as targets and automated judges across 302 prompt–response pairs using judge reliability metrics (MAD and signed error).","arXiv :2607 .00139v1 [ cs .CL] 30 Jun 2026  \nJuly 2, 2026  \nBenchmarking Frontier LLMs on Arabic Cultural and Sociolinguistic Knowledge: A Cross-Evaluation Framework with Human SME Ground Truth  \nSajjad Abdoli1,*,†1 , Ghassan Al-Sumaidaee1,*,†2 , Ahmad ElShiekh1,*3 , Clayton W. Taylor1,*4 , Ahmed Rashad1  \n1 Perle AI * Equal contribution; names sorted alphabetically. † Corresponding authors [sajjad@perle.ai](sajjad@perle.ai) [ghassan.al-sumaidaee@perle.ai](ghassan.al-sumaidaee@perle.ai) [clayton@perle.ai](clayton@perle.ai)[mad.elshiekh@perle.ai](mad.elshiekh@perle.ai) [ahmed@perle.ai](ahmed@perle.ai)  \nThe cost of human expert evaluation is a principal bottleneck to deploying language models in specialized, often high-stakes domains. This bottleneck is particularly acute regarding Arabic sociolinguistic knowledge: credible grading requires not only linguistic fluency but deep cultural familiarity that cannot be approximated by surface-level, black-and-white metrics and training data. We address this with across-evaluation framework instantiated on two underrepresented Arabic dialect communities: Egyptian and Iraqi Arabic. We contribute a dataset of 103 validated prompt–rubric pairs (70 Egyptian and 33 Iraqi; 53 of which are Cultural, the other 50 Linguistic), authored and graded by native-speaker SMEs using penalty-weighted rubrics that distinguish positive content requirements from answer-specific negative error criteria. Three frontier LLMs serve as target models (whose responses are graded by human SMEs across 302 unique prompt–response pairs), while five frontier LLMs serve as automated judges grading target model responses against the rubric, enforcing a provider-level self-evaluation guard. A dual-metric scheme combining Mean Absolute Deviation (MAD) with a Signed Mean Error separates directional grading bias from symmetric noise. Across 1,307 judge evaluations: GPT-5.4 is the most reliable judge (MADj = 10 .21 pp, Signed Error = −1 . 12%); four of five judges show systematic leniency (+2 .01% to +6 .56%); Cultural tasks are harder to grade than Linguistics tasks for all judges (MAD gap 1.83–4.78 pp); and models substantially outperform on Egyptian prompts when compared to outputs for Iraqi prompts . However, given the difference in leniency between the Iraqi and Egyptian SMEs scoring model outputs, we cannot solely attribute the Egyptian-Iraqi performance gap to model knowledge alone. We have therefore chosen to emphasize findings that do not assume identical leniency across human graders. Across all samples, regardless of subdomain or judge leniency, implicit cultural reasoning, which requires models simulate native-speaker judgment rather than rely on lexical verification, emerges as the primary failure mode for automated grading across all judge models.  \n1. Introduction  \nObjective, verifiable domains like math and science, in which experts generally agree on what distinguishes“true” from “false,” are well-suited for reinforcement learning (“RL”) with binary reward signals. Language, however, and the experiences we use it to describe, are intrinsically subjective. Subjective domains (e.g.  \n1 Corresponding author: [sajjad@perle.ai](sajjad@perle.ai)  \n2 Corresponding author: [ghassan.al-sumaidaee@perle.ai](ghassan.al-sumaidaee@perle.ai); ORCID: 0000-0002-5536-0252  \n3 ORCID: 0009-0001-6837-6202  \n4 ORCID: 0009-0006-6478-8994  \n© [2026 Perle.ai. All](2026 Perle.ai. All) rights reserved. 1  \nJuly 2, 2026  \nbelief systems, cultures, linguistic dialects and subdialects) therefore elude binary benchmarking techniques employed for more “objective,” verifiable domains. This necessitates the use of a limited pool of human subject matter experts (“SMEs”) for model evaluation and feedback, which in turn embed certain bottlenecks in RL environments themselves. The adaptation of techniques that’ve helped scale feedback and reward signaling in more “objective,” verifiable domains, such as LLMs-as-judges, is therefore essential to","cbCainjKWMJv2c65","https://ap.wps.com/l/cbCainjKWMJv2c65","pdf",671437,4,1,19,"English","en",105,"# Introduction\n# Cross-Evaluation Framework\n## Datasets and Rubrics\n## Judge and Target Models\n# Evaluation Metrics and Findings\n## Reliability and Bias Analysis","[{\"question\":\"Why is human SME evaluation particularly necessary for Arabic sociolinguistic knowledge?\",\"answer\":\"Arabic sociolinguistic grading depends on cultural familiarity beyond surface linguistic fluency, and simple black-and-white metrics or training-data statistics cannot reliably capture credible judgment.\"},{\"question\":\"What does the proposed cross-evaluation framework do with frontier LLMs?\",\"answer\":\"Each frontier model is evaluated twice: as a target whose responses are graded by human SMEs (gold standard) and as a judge that grades other models’ responses against the rubric.\"},{\"question\":\"Which metric framework is used to analyze judge reliability and bias?\",\"answer\":\"Judge reliability uses Mean Absolute Deviation (MAD) to quantify noise and a Signed Mean Error to separate directional grading bias from symmetric noise; results indicate GPT-5.4 is the most reliable judge while others show systematic leniency.\"}]",1784187866,48,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"benchmarking-frontier-llms-on-arabic-cultural-and-sociolinguistic-knowledge-a-cross-evaluation-framework-with-human-sme-ground-truth","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/benchmarking-frontier-llms-on-arabic-cultural-and-sociolinguistic-knowledge-a-cross-evaluation-framework-with-human-sme-ground-truth/83446/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-26","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why is human SME evaluation particularly necessary for Arabic sociolinguistic knowledge?","Question",{"text":75,"@type":76},"Arabic sociolinguistic grading depends on cultural familiarity beyond surface linguistic fluency, and simple black-and-white metrics or training-data statistics cannot reliably capture credible judgment.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What does the proposed cross-evaluation framework do with frontier LLMs?",{"text":80,"@type":76},"Each frontier model is evaluated twice: as a target whose responses are graded by human SMEs (gold standard) and as a judge that grades other models’ responses against the rubric.",{"name":82,"@type":73,"acceptedAnswer":83},"Which metric framework is used to analyze judge reliability and bias?",{"text":84,"@type":76},"Judge reliability uses Mean Absolute Deviation (MAD) to quantify noise and a Signed Mean Error to separate directional grading bias from symmetric noise; results indicate GPT-5.4 is the most reliable judge while others show systematic leniency.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},"General","general"]