[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84338-en":3,"doc-seo-84338-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":11,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},84338,1099514068365,"Aurelia","https://ap-avatar.wpscdn.com/avatar/10000253d8d9f28188e?_k=1776742907772140068",8,"Research & Report","Compete Then Collaborate: Frontier AI Teachers Build a Verifiable Curriculum to Improve a Coding Student Beyond Imitation","Large language models increasingly serve as teachers by generating training data for smaller student coders. Prior multi-teacher distillation merges outputs but often lacks a principled way to identify the best frontier teacher and may rely on an LLM judge biased toward its own generations. This work proposes compete-then-collaborate: four lab-spanning frontier teachers are execution-ranked with fairness controls, then jointly construct a verifiable curriculum for Qwen2.5-Coder, producing measurable gains via verifiable rewards rather than imitation.","arXiv :2607 .08255v 1 [ cs .AI] 9 Jul 2026  \nCompete then Collaborate: Frontier AI Teachers Build a Verifiable Curriculum to Improve a Coding Student Beyond  \nImitation  \nMiseong (Shawn) Kim  \nGenesis Cortex AI Inc.  \nORCID: 0009-0003-7566-3212  \nJuly 2026  \nCode, data, tests, and verification harness:  \n[https://github.com/shawnkim678/compete-then-collaborate](https://github.com/shawnkim678/compete-then-collaborate)[ ](https://github.com/shawnkim678/compete-then-collaborate)Every number is recomputable from the released files or by regenerating teacher solutions under your own  \nprovider terms.  \nAbstract  \nLarge language models are increasingly used as teachers that generate training data for smaller student models. Prior multi-teacher knowledge distillation merges several teachers’ outputs, but does not ask which frontier model teaches best, and typically relies on an LLM judge that is known to favor its own outputs. We introduce a compete-then-collaborate framework in which four frontier AI teachers spanning the major labs (Claude/Anthropic, CodexGPT/OpenAI, Grok/xAI, Gemini/Google) are first ranked head-to-head by an executionbased judge (unit tests / stdin–stdout checks) with fairness controls (a shared task bank, teacher self-correction, and an intersection-controlled training set), and then collaborate to build a verifiable curriculum for a single coding student (Qwen2.5-Coder) . We report three findings.  \n(1) Under execution verification, all four teachers solve standard problems near-perfectly after self-correction (≈99–100%)—a saturation effect, not a skill difference—while harder competition problems separate them (Gemini 77% > Claude 69% ≈ Codex 69% > Grok 50%); still, the most robust results are on the student side and do not depend on the teacher ranking.  \n(2) Imitation (SFT) on the teachers’ verified solutions does not improve—and can degrade—an already-competent coder student at both 7B and 32B (e.g., 76.7%→72.7% on MBPP-test, 5.9%→2.9% on competition problems for the union of all teachers) . (3) The same collaborative curriculum used as a reinforcement-learning-with-verifiable-rewards (RLVR) environment instead improves the student (5.9%→8.8% peak on held-out competition problems over a 1000-step run, +49% relative), reversing the direction of SFT. Our central claim is that the value of AI-teacher collaboration is not pooling answers to imitate, but jointly constructing a verifiable environment in which the student learns by doing. We release a fully reproducible on-prem pipeline (NVIDIA GB10) including framework patches required to run GRPO on a bleeding-edge stack.  \n1 Introduction  \nDistilling capable “teacher” LLMs into smaller “student”models is now standard practice. Two questions are under-explored: (i) which commercial frontier model is the better teacher, measured by the student’s real downstream ability rather than by an LLM judge; and (ii) whether combining teachers is best done by merging their answers or by some other mechanism.  \nWe study these on Python coding. Our judge is code execution—objective and free of the self-preference bias documented for LLM-as-judge setups. We first run a competition: teachers solve a shared task bank; a teacher’s output enters the student’s data only if it passes hidden tests. We then run a collaboration: all teachers’ verified work forms one curriculum, used two ways—as imitation targets (SFT) and as a verifiable reward environment (RLVR) .  \nContributions. (1) An execution-verified, bias-free ranking of frontier AI teachers with three fairness controls (shared tasks, self-correction, intersection) . (2) A controlled comparison of collaboration modes showing imitation-SFT fails/degrades competent coder students while verifiable-reward RL improves them. (3) A reproducible on-prem (GB10) pipeline, includingthe framework patches needed to run GRPO on transformers 5.5 / torch 2.11 / cu130 .  \n2 Related Work  \nMulti-teacher KD. Prior work merges multiple teachers’ r","cbCaijjBdOP5xD0O","https://ap.wps.com/l/cbCaijjBdOP5xD0O","pdf",363762,9,1,"English","en",105,"# Abstract\n# Introduction\n# Related Work\n# Method\n## Execution-verified generation","[{\"question\":\"How does the compete-then-collaborate framework rank and select among frontier AI teachers?\",\"answer\":\"Teachers are first head-to-head ranked using execution-based judgment, such as unit tests and stdin–stdout checks, together with fairness controls like a shared task bank, teacher self-correction, and an intersection-controlled training set.\"},{\"question\":\"What is the impact of imitation training (SFT) using teachers’ verified solutions on a competent coding student?\",\"answer\":\"Imitation (SFT) on the teachers’ verified solutions does not improve the student and can degrade performance for both 7B and 32B, with reported drops on MBPP-test and competition problems.\"},{\"question\":\"How does the verifiable reward approach (RLVR) differ from SFT, and why does it improve the student?\",\"answer\":\"The same collaborative, execution-verified curriculum is used as a reinforcement-learning-with-verifiable-rewards environment; this reverses the direction of SFT and increases the student’s peak performance on held-out competition problems.\"}]",1784194919,20,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"compete-then-collaborate-frontier-ai-teachers-build-a-verifiable-curriculum-to-improve-a-coding-student-beyond-imitation","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/compete-then-collaborate-frontier-ai-teachers-build-a-verifiable-curriculum-to-improve-a-coding-student-beyond-imitation/84338/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-28","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"How does the compete-then-collaborate framework rank and select among frontier AI teachers?","Question",{"text":75,"@type":76},"Teachers are first head-to-head ranked using execution-based judgment, such as unit tests and stdin–stdout checks, together with fairness controls like a shared task bank, teacher self-correction, and an intersection-controlled training set.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What is the impact of imitation training (SFT) using teachers’ verified solutions on a competent coding student?",{"text":80,"@type":76},"Imitation (SFT) on the teachers’ verified solutions does not improve the student and can degrade performance for both 7B and 32B, with reported drops on MBPP-test and competition problems.",{"name":82,"@type":73,"acceptedAnswer":83},"How does the verifiable reward approach (RLVR) differ from SFT, and why does it improve the student?",{"text":84,"@type":76},"The same collaborative, execution-verified curriculum is used as a reinforcement-learning-with-verifiable-rewards environment; this reverses the direction of SFT and increases the student’s peak performance on held-out competition problems.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,126,129,133],{"id":21,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":20,"doc_module":4,"doc_module_name":45,"category_name":124,"show_sort_weight":28,"slug":125},"Religion & Spirituality","religion-spirituality",{"id":28,"doc_module":4,"doc_module_name":45,"category_name":127,"show_sort_weight":28,"slug":128},"World Cup","world-cup",{"id":130,"doc_module":4,"doc_module_name":45,"category_name":131,"show_sort_weight":130,"slug":132},10,"Lifestyle","lifestyle",{"id":134,"doc_module":4,"doc_module_name":45,"category_name":135,"show_sort_weight":106,"slug":136},19,"General","general"]