[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85355-en":3,"doc-seo-85355-105":30,"detail-sidebar-cat-0-en-105":95},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},85355,7971461741311,"Ophelia","https://ap-avatar.wpscdn.com/avatar/74000253aff267980c6?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779345379180704826",8,"Research & Report","Which Optimizer, At What Budget? A Tournament of Optimizers for Search-Based SE","Software configuration and tuning are expensive and error-prone because modern systems expose hundreds of interacting options, making single evaluations costly. Prior practice often recommends specific optimizers (e.g., NSGA-II), yet comparable results may require far fewer evaluations with alternatives. This study clusters 20 optimizers by six data assumptions, then tournaments them on 106 SBSE tasks under four labeling budgets (14,000+ CPU hours). No optimizer wins universally: the best choice migrates with budget, and a cheap two-attribute lookup plus budget predicts the winner.","Which Optimizer, At What Budget? A Tournament of Optimizers for Search-Based SE  \nKishan Kumar Ganguly and Tim Menzies  \nNorth Carolina State University  \n[kgangul@ncsu.edu](kgangul@ncsu.edu), [timm@ieee.org](timm@ieee.org)  \narXiv :2607 . 1 1705v 1 [ cs . SE] 13 Jul 2026  \nAbstract—Configuring and tuning modern software is unavoidable, expensive, and error-prone: a single system can expose hundreds of interacting options, and scoring one setting can mean a full build or test run. The standard response is automated optimization, but the number of available optimizers is large and growing. And some of the guidance for selecting among them is misleading: NSGA-II, for example, is widely recommended, yet other algorithms reach the same results using only 1/20th as many evaluations.  \nTo help practitioners make better choices about tools to configure their systems, we cluster 20 optimizers, based on six assumptions about the data. Next, we run a tournament across those optimizers, using 106 SE optimization tasks at four labeling budgets (taking 14,000+ CPU hours). We find that no optimizer wins outright. The best one migrates with the budget (from a geometric active learner when labels are scarce to differential evolution when labels are plentiful) so a winner “crowned” atone budget is wrong at another on up to half our tasks.  \nRunning such a tournament for every new domain is impractical due to its CPU cost. Fortunately, we find that those 14,000 hours can be replaced by a table lookup over two cheap-to-obtain task attributes (plus the labeling budget). Predictions from this table tie or beat a hindsight oracle on ≈ 75% of held-out tasks.  \nTo support open science, our tournament and replication package are open-sourced for SBSE researchers and practitioners at [https://github.com/KKGanguly/OptimizerTournament](https://github.com/KKGanguly/OptimizerTournament).  \nIndex Terms—Search-based SE, algorithm selection, empirical study  \nI. INTRODUCTION  \nConfiguring and tuning modern software is unavoidable and expensive. A modern system stacks languages, libraries, compilers, and deployment pipelines, each exposing its own flags and defaults: one open-source database used in this study carries 460 binary options, creating a configuration space of 2460 , larger than the number of stars in the sky [1] . Exhaustive search is impractical at this scale, so the field turns to searchbased software engineering (SBSE) to formulate this as an optimization problem. Under this paradigm,a fitness function scores candidate configurations and a search algorithm walks the space looking for good ones [10], [13] .  \nWe study black-box optimization: reasoning about a system from inputs and outputs alone. It is the workhorse of SE configuration, used to tune compilers and databases [46],[47], configure product lines and cloud systems [48], [49], and generate adversarial tests [50] . Over 100 such tasks are documented in the MOOT repository (Table I), representing more than a decade of SBSE work [2] .  \nA practitioner who reaches for these tools confronts an overwhelming set of choices. Table II lists 20 optimizers drawn from the recent SE literature, and the No-Free-Lunch theorem guarantees that none of them wins everywhere [51], [52] . Which optimizer suits which task is, today, largely guesswork.  \nThe labeling budget compounds the problem. A single SE fitness evaluation can be costly: it may run an entire test suite for program repair [24], execute thousands of unit tests to score LLM-generated code [53], or profile system-level time and energy. Agentic and LLM-driven workflows push this cost higher. Still, An engineer may afford only dozens of evaluations, and the best optimizer at thirty evaluations need not be the best at two hundred. This exposes a gap:  \nThe gap. SBSE has no systematic way to organize optimizers by the assumptions they encode, and no cheap, budget-aware rule for choosing one on a given task.  \nWe close this gap with an assumption-","cbCaibD5BgGqbz8H","https://ap.wps.com/l/cbCaibD5BgGqbz8H","pdf",553861,2,1,12,"English","en",105,"# Introduction\n## Research gap and tournament design\n## Key research questions and contributions","[{\"question\":\"Why does optimizer selection become difficult in search-based software engineering?\",\"answer\":\"Modern software exposes many configurable options, and each optimizer’s effectiveness depends on assumptions about the data. The labeling budget further changes what is feasible, so the best optimizer under one budget may not remain best under another.\"},{\"question\":\"How is the tournament organized in the study?\",\"answer\":\"The study clusters 20 optimizers using six assumptions, then runs a tournament across four labeling budgets on 106 SBSE tasks from the MOOT repository, repeated 20 times (about 14,000+ CPU hours).\"},{\"question\":\"What is the main finding about whether a single optimizer is best?\",\"answer\":\"No optimizer wins outright across all tasks. The best optimizer changes with labeling budget, moving from a geometric active learner when labels are scarce to differential evolution when labels are plentiful.\"},{\"question\":\"Can the winner be predicted without running the full tournament?\",\"answer\":\"Yes. A lookup model using two cheap-to-obtain task attributes plus the labeling budget ties or beats a hindsight oracle on about 75% of held-out tasks.\"}]",1784202744,30,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":90,"head_meta":92,"extra_data":94,"updated_unix":28},"which-optimizer-at-what-budget-a-tournament-of-optimizers-for-search-based-se","",{"@graph":36,"@context":89},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/which-optimizer-at-what-budget-a-tournament-of-optimizers-for-search-based-se/85355/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-21","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81,85],{"name":72,"@type":73,"acceptedAnswer":74},"Why does optimizer selection become difficult in search-based software engineering?","Question",{"text":75,"@type":76},"Modern software exposes many configurable options, and each optimizer’s effectiveness depends on assumptions about the data. The labeling budget further changes what is feasible, so the best optimizer under one budget may not remain best under another.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How is the tournament organized in the study?",{"text":80,"@type":76},"The study clusters 20 optimizers using six assumptions, then runs a tournament across four labeling budgets on 106 SBSE tasks from the MOOT repository, repeated 20 times (about 14,000+ CPU hours).",{"name":82,"@type":73,"acceptedAnswer":83},"What is the main finding about whether a single optimizer is best?",{"text":84,"@type":76},"No optimizer wins outright across all tasks. The best optimizer changes with labeling budget, moving from a geometric active learner when labels are scarce to differential evolution when labels are plentiful.",{"name":86,"@type":73,"acceptedAnswer":87},"Can the winner be predicted without running the full tournament?",{"text":88,"@type":76},"Yes. A lookup model using two cheap-to-obtain task attributes plus the labeling budget ties or beats a hindsight oracle on about 75% of held-out tasks.","https://schema.org",{"og:url":51,"og:type":91,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":93,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":96},[97,101,105,109,114,119,124,126,131,134,138],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":106,"show_sort_weight":107,"slug":108},"Exam",70,"exam",{"id":110,"doc_module":4,"doc_module_name":46,"category_name":111,"show_sort_weight":112,"slug":113},5,"Comic",60,"comic",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},6,"Technology",50,"technology",{"id":120,"doc_module":4,"doc_module_name":46,"category_name":121,"show_sort_weight":122,"slug":123},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":29,"slug":125},"research-report",{"id":127,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":129,"slug":130},9,"Religion & Spirituality",20,"religion-spirituality",{"id":129,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":129,"slug":133},"World Cup","world-cup",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":135,"slug":137},10,"Lifestyle","lifestyle",{"id":139,"doc_module":4,"doc_module_name":46,"category_name":140,"show_sort_weight":110,"slug":141},19,"General","general"]