[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83137-en":3,"doc-seo-83137-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83137,687197207057,"Sage","https://ap-avatar.wpscdn.com/davatar_29158cc5080c5b710cf443261637dec0",8,"Research & Report","Cost-Effective Agent Harnesses for Abstract Reasoning and Generalization on ARC-AGI-1","Recent progress on ARC-AGI-1 from disclosed architectures comes from two regimes: heavy test-time compute over frontier models or benchmark-specific training using small fine-tuned models and task-specialized architectures. This work studies a third regime: an open-weight model in non-thinking mode (DeepSeek V3.2) under a strict budget, without ARC-specific fine-tuning. Agentic harnesses explicitly decompose pattern discovery and program-synthesis. An Explorer-Definer Pipeline attains 57.50% pass@2 at $0.25/task; a Reflective Orchestrator reaches 67.25% pass@2 at $0.62/task.","Cost-Effective Agent Harnesses for Abstract Reasoning and Generalization on ARC-AGI-1  \nKabir Moghe  \nDepartment of Computer Science Dartmouth College [kabir.moghe.26@dartmouth.edu](kabir.moghe.26@dartmouth.edu)  \nPeter Chin  \nDepartment of Computer Science Dartmouth College [peter.chin@dartmouth.edu](peter.chin@dartmouth.edu)  \narXiv :2607 .06764v 1 [ cs .AI ] 7 Jul 2026  \nAbstract  \nRecent progress on ARC-AGI-1 from disclosed architectures has come broadly from two regimes: heavy test-time compute over frontier models (evolutionary search, exhaustive sampling, extended chain-of-thought), or benchmark-specific training in which small models are fine-tuned on ARC data, often with taskspecialized architectures. We study a third regime: an open-weight model in non-thinking mode (DeepSeek V3.2) under a strict budget, with no ARC-specific fine-tuning. We study what is recoverable through architecture alone, building agentic harnesses that decompose pattern-discovery and program-synthesis stages explicitly. First, we introduce an Explorer-Definer Pipeline that separates pattern discovery from executable transformation synthesis, implemented as a two-stage agent pipeline. Next, we present the Reflective Orchestrator, which augments the pipeline with autonomous exploration of new transformations when previous hypotheses fail on training pairs. On the ARC-AGI-1 public 400-task evaluation set, the pipeline reaches 57.50% pass@2 at $0.25 per task, and the orchestrator reaches 67.25% pass@2 at $0.62 per task. Together these architectures lift a 15.50% one-shot baseline by ∼52 points without benchmark-specific training or heavy test-time compute. Furthermore, the orchestrator-driven lift tests a falsifiable diagnostic the pipeline produces; unbiased pass@k analysis suggests the pipeline is generation-bound, not selection-bound (selection via training-pair accuracy captures ∼95% of the candidate ceiling) and predicts that significant improvement requires broader generation, not better ranking. The orchestrator implements this prediction via adaptive re-exploration and confirms it (unbiased pass@1 lift +9.8 pp, matching selection-mediated pass@2 lift) . An additional pipeline ablation identifies its think tool as a significant component, with removal reducing pass@2 by 5.75 pp.  \n1 Introduction  \nIn the years since its inception, ARC has become approachable, but at a cost. Strong reported performance from public work has generally come from one of two regimes. The first relies on heavy test-time compute over frontier large language models (LLMs): evolutionary search over candidate solutions, exhaustive program enumeration, or extended chain-of-thought reasoning, with reported costs ranging from a few dollars to several hundred dollars per task [1–4] . The second relies on benchmark-specific training: small models trained from scratch or fine-tuned directly on ARC-distribution data, typically combined with test-time architectures employing fine-tuning on the target task and aggressive data augmentation [5–7] . Both regimes have successfully advanced leaderboard numbers, but each trades strong performance for substantial compute spent either on significant inference per task or on benchmark-specific training.  \nPreprint.  \nContextualizing the “no RL” or fine-tuning condition. The scope distinction relevant to this paper is between benchmark-specific training, whether supervised, test-time-finetuned, or RL-based, and general-purpose post-training. The above second-regime systems all fine-tune or train from scratch on ARC training tasks. The model we use—DeepSeek V3.2 in non-thinking mode—has undergone general post-training on broad task mixtures, including significant reinforcement learning components (e.g., RLHF, agentic-capability RL, reasoning-style RL), but has not been specialized to ARC, with the exception of any public data leakage. We hold that general post-training constant across every condition in this paper, so the architectural de","cbCaifVh2M5VzzeP","https://ap.wps.com/l/cbCaifVh2M5VzzeP","pdf",1837330,2,1,24,"English","en",105,"# Introduction\n## Motivation and two existing regimes\n## Scope and model setting\n## Central architectural findings","[{\"question\":\"What are the two main existing regimes for high performance on ARC-AGI-1?\",\"answer\":\"Prior approaches typically use heavy test-time compute over frontier LLMs or benchmark-specific training/fine-tuning on ARC data with task-specialized architectures.\"},{\"question\":\"What is the third regime proposed in this paper?\",\"answer\":\"The paper studies an open-weight model (DeepSeek V3.2) in non-thinking mode under a strict budget, without ARC-specific fine-tuning, focusing on what can be recovered through architecture alone.\"},{\"question\":\"How do the Explorer-Definer Pipeline and Reflective Orchestrator improve results?\",\"answer\":\"The Explorer-Definer Pipeline reaches 57.50% pass@2 at $0.25/task, and the Reflective Orchestrator reaches 67.25% pass@2 at $0.62/task by enabling autonomous exploration when hypotheses fail.\"}]",1784185540,60,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"cost-effective-agent-harnesses-for-abstract-reasoning-and-generalization-on-arc-agi-1","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/cost-effective-agent-harnesses-for-abstract-reasoning-and-generalization-on-arc-agi-1/83137/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What are the two main existing regimes for high performance on ARC-AGI-1?","Question",{"text":75,"@type":76},"Prior approaches typically use heavy test-time compute over frontier LLMs or benchmark-specific training/fine-tuning on ARC data with task-specialized architectures.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What is the third regime proposed in this paper?",{"text":80,"@type":76},"The paper studies an open-weight model (DeepSeek V3.2) in non-thinking mode under a strict budget, without ARC-specific fine-tuning, focusing on what can be recovered through architecture alone.",{"name":82,"@type":73,"acceptedAnswer":83},"How do the Explorer-Definer Pipeline and Reflective Orchestrator improve results?",{"text":84,"@type":76},"The Explorer-Definer Pipeline reaches 57.50% pass@2 at $0.25/task, and the Reflective Orchestrator reaches 67.25% pass@2 at $0.62/task by enabling autonomous exploration when hypotheses fail.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,109,114,119,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":29,"slug":108},5,"Comic","comic",{"id":110,"doc_module":4,"doc_module_name":46,"category_name":111,"show_sort_weight":112,"slug":113},6,"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]