[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83093-en":3,"doc-seo-83093-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83093,1099514067415,"Rowan","https://ap-avatar.wpscdn.com/avatar/100002539d78ffe74a7?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779092875211072502",8,"Research & Report","RuBench A Repository Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications","RuBench 1.0 evaluates coding agents in a setting where developers express maintenance requests in native Russian, using real customer-style task statements rather than English-curated issues. The benchmark provides 25 repository-level fix tasks mined from five live open-source projects across multiple languages, with scoring performed by upstream maintainer regression tests withheld from public release. Deployed product configurations are measured via repeated runs, pass@1 with task-level confidence intervals, and detailed cost/token accounting to ensure auditable uncertainty and freshness against training cutoffs.","arXiv :2607 .064 1 1v 1 [ cs . SE] 7 Jul 2026  \nRuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications  \nEvgeny Shilov  \nIndependent Researcher  \n[https://vibecoding.ru/rubench](https://vibecoding.ru/rubench)  \nCode and data: [https://github.com/eugeneshilow/rubench](https://github.com/eugeneshilow/rubench)  \nJuly 7, 2026  \nAbstract  \nDevelopers increasingly delegate real maintenance work to product-grade coding agents, and many of them state tasks in their native language, in the style of a customer request rather thana curated English issue. Existing repository-level agentic benchmarks do not measure this setting:  \ntheir task statements are English by design. We introduce RuBench 1.0, a benchmark of 25 tasks mined from recent fix commits in five live open-source repositories (aiohttp, aiogram, Laravel, NestJS, Fastify; Python, PHP, TypeScript, JavaScript), where each task is specified natively in Russian—written from scratch in the style of an actual customer request, not translated from an English issue—and judged by the upstream maintainer’s regression tests, which we withhold from public release. All 25 fix commits postdate the training-data cutoffs of every evaluated model, giving a contamination argument that holds task-by-task rather than on average. We evaluate deployed product configurations (CLI agent + model + reasoning effort)—Claude Code with Opus 4 .8, Sonnet 5, and Haiku 4 .5, and Codex CLI with GPT-5 .5—with three independent runs each, reporting pass@1 with task-level confidence intervals, paired comparisons, dollar cost, and token usage. The best configuration resolves 78 .7% of tasks; at N =25 only the gaps to the weakest model are statistically resolvable, which we state explicitly. Auditing full trajectories of a fifth, hors-concours configuration (Claude Code + Fable 5, July 2, 2026 release), we caught the product silently substituting the model: on 5 of 25 tasks (20%) an official safeguard fallback re-routed routine [HTTP-protocol fixes to Opus 4.8](HTTP-protocol fixes to Opus 4.8)—direct, reproducible evidence that the deployed product, not the model, is the unit actually being measured. We release task statements, metadata, full agent trajectories, and diffs; grading oracles are withheld, with a SHA-256 manifest committed at publication time.  \n1 Introduction  \nCoding agents have become consumer products. Tools such as Claude Code and Codex CLI take a natural-language task, operate on a full repository checkout with shell access, and return a patch. A large share of their users are not native English speakers, and the natural way for such a user to state a task is in their own language and their own words: “the bot started crashing on every channel post after we made it an admin —fix the framework, and don’t break existing chats”—not a curated English GitHub issue. Working from a native-language, customer-style specification is a distinct capability: the agent must recover intent, technical vocabulary, and acceptance criteria from text whose surface form never appears in English-centric training pipelines for software tasks.  \nExisting benchmarks do not measure this. Repository-level suites in the SWE-bench line [1] draw task statements from English GitHub issues; the “multilingual” variants—SWE-bench Multilingual  \nand Multi-SWE-bench [11]—are multilingual in programming languages while explicitly selecting repositories “where the primary language is English,” 1 so English statements are a design criterion, not an accident of sampling. Benchmarks that do contain non-English natural language (mCoNaLa [27], ODEX [28], HumanEval-XL [29], MERA Code [8], EDIT-Bench [9]) operate at snippet or single-file level, not as agents in a live repository. A separate line of evidence shows that static repository benchmarks decay: solution leakage and weak tests in SWE-bench [2], memorization effects on SWE-bench-Verified [3 , 4], and the eventual deprecation of Verif","cbCaijuXbWRgKLSQ","https://ap.wps.com/l/cbCaijuXbWRgKLSQ","pdf",540285,6,1,16,"English","en",105,"# Abstract\n# Introduction","[{\"question\":\"What makes RuBench different from existing repository-level agentic coding benchmarks?\",\"answer\":\"RuBench tasks are specified natively in Russian as customer-style requests, not translated from English GitHub issues. Tasks are evaluated in live repositories using the maintainers’ regression tests, which are withheld from public release.\"},{\"question\":\"How are RuBench tasks selected and freshness handled?\",\"answer\":\"RuBench 1.0 contains 25 fix tasks mined from recent fix commits merged between February and June 2026. All tasks postdate the training-data cutoffs of every evaluated model, enforced at import time.\"},{\"question\":\"How is agent performance measured in RuBench?\",\"answer\":\"The benchmark evaluates deployed product configurations (CLI agent + model + reasoning effort) with three independent runs each, reporting pass@1 with task-level 95% confidence intervals. It also includes paired comparisons plus per-task cost, token, and time accounting.\"}]",1784185170,40,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"rubench-a-repository-level-agentic-coding-benchmark-with-natively-authored-russian-task-specifications","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/rubench-a-repository-level-agentic-coding-benchmark-with-natively-authored-russian-task-specifications/83093/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What makes RuBench different from existing repository-level agentic coding benchmarks?","Question",{"text":76,"@type":77},"RuBench tasks are specified natively in Russian as customer-style requests, not translated from English GitHub issues. Tasks are evaluated in live repositories using the maintainers’ regression tests, which are withheld from public release.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How are RuBench tasks selected and freshness handled?",{"text":81,"@type":77},"RuBench 1.0 contains 25 fix tasks mined from recent fix commits merged between February and June 2026. All tasks postdate the training-data cutoffs of every evaluated model, enforced at import time.",{"name":83,"@type":74,"acceptedAnswer":84},"How is agent performance measured in RuBench?",{"text":85,"@type":77},"The benchmark evaluates deployed product configurations (CLI agent + model + reasoning effort) with three independent runs each, reporting pass@1 with task-level 95% confidence intervals. It also includes paired comparisons plus per-task cost, token, and time accounting.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,111,115,119,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":29,"slug":118},7,"Healthcare","healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":107,"slug":137},19,"General","general"]