[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-81876-en":3,"doc-seo-81876-105":31,"detail-sidebar-cat-0-en-105":93},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":28,"seo_description":14,"update_tm":29,"read_time":30},81876,2336464648322,"Aria","https://ap-avatar.wpscdn.com/avatar/2200025388227c56fec?_k=1778556882303663488",8,"Research & Report","ArchEval Measuring AI Agents as Computer Architects","Computer architecture progress has relied on benchmarks, but LLM agents introduce a new measurement challenge: when tasked to act as computer architects, success depends on interpreting workloads, selecting mechanisms, driving simulators, predicting performance before feedback, satisfying hard constraints, and judging which feasible designs are worth evaluating. The paper presents ArchEval, a benchmark platform with 20 challenges spanning CPU cores, systems, memory, accelerators, and compute-in-memory, using eight simulators and three evaluation settings. Results reveal a sharp capability boundary, with limited performance modeling and pre-feedback design judgment.","ArchEval: Measuring AI Agents as Computer Architects  \nChenyu Wang∗1, Zishen Wan∗1, Jeffrey Ma 1 , Shvetank Prakash 1 , Zhenting Qi 1 , Haebin Do 1 , Andy Cheng 1 , Arya Tschand 1 , Jiahe Shi2 , Yilun Du 1 , Vijay Janapa Reddi 1  \n1Harvard University 2Massachusetts Institute of Technology  \nUSA  \narXiv :2607 .0360 1v 1 [ cs .AR] 3 Jul 2026  \nAbstract  \nComputer architecture has long used benchmarksto make progress measurable. LLM agents create a different measurement problem: when asked to act as computer architects, success is not merely writing code or tuning parameters. The agent must interpret workloads, choose mechanisms, use simulators, predict performance before feedback, satisfy hard constraints, and decide which feasible design is worth evaluating.  \nThis paper introduces ArchEval, a benchmark and platform for evaluating LLM agents on computer architecture design and optimization. It contains 20 challenges across CPU core mechanisms, system architecture, memory systems, accelerators, and computein-memory, backed by eight simulators. Each challenge is posed under three evaluation settings: L1 full harness, with a prepared harness and repeated simulator feedback; L2 simulator-code container, where simulator source is available but the agent must assemble its own experimental workflow; and L3 agent-only, where the agent receives static workload evidence and constraints but norunnable simulator feedback before final submission. Each run reports baseline-normalized verifier performance and records the full trajectory, connecting final results to workload analysis, simulatortool use, prediction, constraint handling, and artifact integrity.  \nInitial results show a sharp boundary in current agents. With L1 support, all four evaluated agents reach or exceed the challenge baseline and improve real architecture designs across diverse simulators. Removing support exposes diagnosable weaknesses: many agents do not turn simulator source into useful local experiments, and their L3 performance predictions often disagree with verifier results. In L3, only GPT-5.5 + Codex remains above baseline, reaching 1. 21× geomean baseline-normalized performance and a 65% win rate; the other three agents fall below baseline. Even GPT-5.5 + Codex has only a 15% performance-modeling pass rate. ArchEval therefore frames today’s agents as useful optimization assistants rather than autonomous architects, and identifies the capabilities needed next: simulator-tool use, calibrated performance prediction, pre-feedback design judgment, and useful mechanism discovery.  \n1 Introduction  \nFor decades, computer architecture has marked progress by building benchmarks for the systems it learned to design: SPEC for single processors, PARSEC and CloudSuite for workloads, and MLPerf for machine-learning hardware [2, 6, 10, 40] . These suites evaluate artifacts and workloads. They do not evaluate the designer who chose the objective, mechanism, simulator, and constraints. For human designers, that separation was natural: judging the designer was morea matter of education and hiring than an architecture-benchmark problem. That boundary is beginning to move. LLM agents now fix  \n∗ Equal contribution.  \nFigure 1: ArchEval evaluates LLM agents as computer architects. (a) Existing architecture benchmarks primarily measure completed designs, not the design process. (b) ArchEval asks whether LLM agents can perform the architecture reasoning process behind them: analyzing workloads, proposing mechanisms, modeling performance, and iterating across simulator-backed design domains.  \nreal software bugs [49], reason through hard problems [28], solve open problems in mathematics [32], and have begun to take on computer architecture design itself [4, 12, 45, 47] . Recent systems work frames this software-to-silicon shift as a recurring challenge around tool interfaces, verification, and feedback loops [44] . Computer architecture now faces a new question: can an LLM agent do the","cbCaivjYFb7lUPcJ","https://ap.wps.com/l/cbCaivjYFb7lUPcJ","pdf",2688690,6,1,26,"English","en",105,"# Abstract\n# Introduction\n## Measurement challenges for agent-based architecture design\n## Architecture design as an iterative loop\n## Why building the benchmark is hard","[{\"question\":\"What makes measuring LLM agents as computer architects different from traditional architecture benchmarks?\",\"answer\":\"Traditional benchmarks measure completed artifacts, while LLM-agent performance must reflect the design process—workload interpretation, mechanism choice, simulator use, performance prediction before feedback, and constraint handling.\"},{\"question\":\"How does ArchEval evaluate agents across different levels of information and feedback?\",\"answer\":\"ArchEval defines three settings: L1 full harness with repeated simulator feedback, L2 simulator-code container where agents build their own experimental workflow, and L3 agent-only with static workload evidence and no runnable simulator feedback before final submission.\"},{\"question\":\"What do the initial ArchEval results indicate about current agent capabilities?\",\"answer\":\"Results show a sharp boundary: with L1 support agents meet or exceed baselines and improve designs, while removing support exposes weaknesses—many agents fail to turn simulator source into useful experiments and L3 predictions often disagree with verifier outcomes.\"}]","ArchEval Measuring AI Agents as Computer Architects | PDF",1784176803,66,{"code":4,"msg":32,"data":33},"ok",{"site_id":25,"language":24,"slug":34,"title":13,"keywords":35,"description":14,"schema_data":36,"social_meta":88,"head_meta":90,"extra_data":92,"updated_unix":29},"archeval-measuring-ai-agents-as-computer-architects","",{"@graph":37,"@context":87},[38,55,70],{"@type":39,"itemListElement":40},"BreadcrumbList",[41,45,49,52],{"item":42,"name":43,"@type":44,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":46,"name":47,"@type":44,"position":48},"https://docshare.wps.com/document/","Document",2,{"item":50,"name":12,"@type":44,"position":51},"https://docshare.wps.com/document/research-report/",3,{"item":53,"name":13,"@type":44,"position":54},"https://docshare.wps.com/document/archeval-measuring-ai-agents-as-computer-architects/81876/",4,{"url":53,"name":13,"@type":56,"author":57,"headline":13,"publisher":59,"fileFormat":62,"inLanguage":24,"description":14,"dateModified":63,"datePublished":64,"encodingFormat":62,"isAccessibleForFree":65,"interactionStatistic":66},"DigitalDocument",{"name":9,"@type":58},"Person",{"url":42,"name":60,"@type":61},"DocShare","Organization","application/pdf","2026-07-29","2026-07-16",true,{"@type":67,"interactionType":68,"userInteractionCount":20},"InteractionCounter",{"@type":69},"ViewAction",{"@type":71,"mainEntity":72},"FAQPage",[73,79,83],{"name":74,"@type":75,"acceptedAnswer":76},"What makes measuring LLM agents as computer architects different from traditional architecture benchmarks?","Question",{"text":77,"@type":78},"Traditional benchmarks measure completed artifacts, while LLM-agent performance must reflect the design process—workload interpretation, mechanism choice, simulator use, performance prediction before feedback, and constraint handling.","Answer",{"name":80,"@type":75,"acceptedAnswer":81},"How does ArchEval evaluate agents across different levels of information and feedback?",{"text":82,"@type":78},"ArchEval defines three settings: L1 full harness with repeated simulator feedback, L2 simulator-code container where agents build their own experimental workflow, and L3 agent-only with static workload evidence and no runnable simulator feedback before final submission.",{"name":84,"@type":75,"acceptedAnswer":85},"What do the initial ArchEval results indicate about current agent capabilities?",{"text":86,"@type":78},"Results show a sharp boundary: with L1 support agents meet or exceed baselines and improve designs, while removing support exposes weaknesses—many agents fail to turn simulator source into useful experiments and L3 predictions often disagree with verifier outcomes.","https://schema.org",{"og:url":53,"og:type":89,"og:title":13,"og:site_name":60,"og:description":14},"article",{"robots":91,"canonical":53},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":94},[95,99,103,107,112,116,121,124,129,132,136],{"id":21,"doc_module":4,"doc_module_name":47,"category_name":96,"show_sort_weight":97,"slug":98},"Story & Novel",90,"story-novel",{"id":48,"doc_module":4,"doc_module_name":47,"category_name":100,"show_sort_weight":101,"slug":102},"Literature",80,"literature",{"id":54,"doc_module":4,"doc_module_name":47,"category_name":104,"show_sort_weight":105,"slug":106},"Exam",70,"exam",{"id":108,"doc_module":4,"doc_module_name":47,"category_name":109,"show_sort_weight":110,"slug":111},5,"Comic",60,"comic",{"id":20,"doc_module":4,"doc_module_name":47,"category_name":113,"show_sort_weight":114,"slug":115},"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":47,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":47,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":125,"doc_module":4,"doc_module_name":47,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":47,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":133,"doc_module":4,"doc_module_name":47,"category_name":134,"show_sort_weight":133,"slug":135},10,"Lifestyle","lifestyle",{"id":137,"doc_module":4,"doc_module_name":47,"category_name":138,"show_sort_weight":108,"slug":139},19,"General","general"]