[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83466-en":3,"doc-seo-83466-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83466,1099513958762,"Logic","https://ap-avatar.wpscdn.com/avatar/1000023916a998db790?x-image-process=image/resize,m_fixed,w_180,h_180&k=1784791008015729253",8,"Research & Report","EPC A Standardized Protocol for Measuring Evaluator Preference Dynamics in LLM Agent Systems","EPC (Evaluator Preference Coupling) introduces a standardized, RFC-style protocol for measuring how evaluator biases propagate through LLM agent closed-loop adaptation via evaluator feedback. The protocol specifies a four-phase isolation paradigm, including executor/evaluator configuration, text and visual strategy/task design, the TTRL update rule, and metric computation (γ, JSD, ECE, Brier) plus an output schema. It is paired with a timebound aversioned reference snapshot that version-tags coupling measurements across evaluator conditions and documents expected measurement decay as evaluators update.","arXiv :2607 .00297v 1 [ cs .LG] 1 Jul 2026  \nEPC: A Standardized Protocol for Measuring Evaluator Preference Dynamics in LLM Agent Systems  \nAnonymous authors  \nPaper under double-blind review  \nAbstract  \nWhen LLM agents use evaluator feedback to adapt their behavior in closed loops, evaluator biases propagate through the agent’s strategy distribution—a phenomenon known as evaluator preference coupling. Prior work has documented coupling across multiple evaluator families and model versions, but the field lacks a standardized protocol that enables third-party researchers to (i) reproduce coupling measurements,(ii) compare results across evaluators and time points, and (iii) detect measurement decay as proprietary evaluators silently update. This paper provides the protocol.  \nWe specify EPC (Evaluator Preference Coupling)—a detailed, RFC-style protocol specification for the four-phase isolation paradigm, covering executor and evaluator configuration, strategy and task design, the TTRL update rule, metric computation (γ, JSD, ECE, Brier), and output schema. We accompany the protocol with aversioned Reference Snapshot v1.0: coupling measurements for eight evaluator conditions (N = 122 unique experimental repetitions across GPT-4o, Qwen, DeepSeek, and others) derived from five independent studies, annotated with evaluator version identifiers, API endpoints, and measurement dates. The snapshot is explicitly timebound: all values are conditional on specific model versions and are expected to decay as proprietary evaluators update. We define a versioning convention (vX .Y-Z , encoding protocol version, snapshot version, and evaluator generation) and provide a usage guide covering adoption, interpretation, and known pitfalls. The protocol, reference snapshot, and implementation code are released as open infrastructure.  \n1 Introduction  \nEvaluator-driven preference dynamics have been documented across multiple LLM agent configurations Liu (2026a;b;c) . In the standard setup, an agent maintains a strategy weight distribution, receives pairwise feedback from an evaluator, and adapts via test-time reinforcement learning (TTRL) . The coupling coefficient γ and Jensen-Shannon divergence (JSD) quantify how strongly the evaluator’s preferences transfer across task domains and how concentrated the agent’s strategy distribution becomes.  \nHowever, the field currently operates without a standardized protocol. Each study uses slightly different protocol variants, task sets, strategy definitions, and metric implementations. Crossstudy comparison is impossible. More critically, proprietary evaluators silently update, causing measurements to decay within weeks Liu (2026a) . Without versioned baselines and explicit expiration dates, the literature accumulates measurements that are no longer valid for current model versions. This problem is not unique to evaluator coupling: a recent audit of 26 AI benchmarks found that the median benchmark has a longevity score of just 5 out of 100 BenchRisk (2026), and the ML community  \nis shifting toward continuous, versioned, community-governed evaluation infrastructure MLCommons (2026); SWE-rebench (2025); HF Community (2026) .  \nThis paper provides the protocol, the reference snapshot, and the versioning convention.  \nThis paper is not a claim of new empirical findings. It is a protocol specification paper—analogous to an RFC in the networking community or a measurement standard in the physical sciences. The coupling measurements in the reference snapshot have been previously reported in domain-specific studies Liu (2026a) . Our contribution is the standardization, versioning, and community infrastructure that transforms these measurements from one-off observations into a reproducible, comparable, and auditable measurement system. We explicitly do not introduce new metrics, new experimental conditions, or new scientific claims. We introduce a discipline—a protocol that enables the community to collectively ma","cbCaicZkV9N30zfM","https://ap.wps.com/l/cbCaicZkV9N30zfM","pdf",387078,4,1,10,"English","en",105,"# Introduction\n# Protocol Specification\n## Overview\n## Agent Configuration\n## TTRL Algorithm","[{\"question\":\"What problem does EPC address in evaluator-driven LLM agent experiments?\",\"answer\":\"Evaluator preference coupling causes evaluator biases to propagate through the agent’s strategy distribution during closed-loop adaptation. EPC standardizes how to measure this coupling so results can be reproduced and compared across evaluators and time points.\"},{\"question\":\"How does the EPC protocol measure evaluator preference coupling?\",\"answer\":\"It uses a four-phase isolation paradigm: pure-text, pure-visual, T→V coupling, and V→T coupling. The coupling coefficient γA→B quantifies how the agent’s weight distribution shifts relative to the pure-domain reference.\"},{\"question\":\"Why does the reference snapshot emphasize versioning and measurement decay?\",\"answer\":\"Proprietary evaluators can silently update, making earlier coupling measurements no longer valid for current evaluator and model versions. EPC therefore requires explicit evaluator identity recording and timebound, version-tagged baselines to keep results auditable and current.\"}]",1784188169,25,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"epc-a-standardized-protocol-for-measuring-evaluator-preference-dynamics-in-llm-agent-systems","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/epc-a-standardized-protocol-for-measuring-evaluator-preference-dynamics-in-llm-agent-systems/83466/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-26","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does EPC address in evaluator-driven LLM agent experiments?","Question",{"text":75,"@type":76},"Evaluator preference coupling causes evaluator biases to propagate through the agent’s strategy distribution during closed-loop adaptation. EPC standardizes how to measure this coupling so results can be reproduced and compared across evaluators and time points.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the EPC protocol measure evaluator preference coupling?",{"text":80,"@type":76},"It uses a four-phase isolation paradigm: pure-text, pure-visual, T→V coupling, and V→T coupling. The coupling coefficient γA→B quantifies how the agent’s weight distribution shifts relative to the pure-domain reference.",{"name":82,"@type":73,"acceptedAnswer":83},"Why does the reference snapshot emphasize versioning and measurement decay?",{"text":84,"@type":76},"Proprietary evaluators can silently update, making earlier coupling measurements no longer valid for current evaluator and model versions. EPC therefore requires explicit evaluator identity recording and timebound, version-tagged baselines to keep results auditable and current.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":22,"slug":133},"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]