[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85215-en":3,"doc-seo-85215-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},85215,13056703019662,"Evangeline","https://ap-avatar.wpscdn.com/avatar/be000253a8e92610077?_k=1778726343310543188",8,"Research & Report","Is Model Instability Just Noise to be Tolerated or a Property that can be Managed","Software analytics rerunning the same analysis can produce different models and conclusions, reducing trust and limiting adoption in practice. A large study across 127 multi-objective software engineering optimization problems (12,700 test cases) shows repeated runs of a state-of-the-art optimizer agree on only 13.7% of test cases. The work argues instability is measurable and manageable, not merely noise: refining label spending, model complexity, and split scoring increases agreement 4.8× and reduces average optimization error standard deviation by 22% while improving recommendation quality. Causal and data-locality interventions help only partially, indicating data-imposed stability limits. Instability should be reported as a standard evaluation axis.","Is Model Instability just Noise to be Tolerated or a Property that can be Managed?  \nAmirali Rayegan  \nNorth Carolina State University Raleigh, North Carolina, USA [arayega@ncsu.edu](arayega@ncsu.edu)  \nLunxiao Li  \nNorth Carolina State University Raleigh, North Carolina, USA [lli66@ncsu.edu](lli66@ncsu.edu)  \nTim Menzies  \nNorth Carolina State University Raleigh, North Carolina, USA [tjmenzie@ncsu.edu](tjmenzie@ncsu.edu)  \narXiv :2607 . 10420v 1 [ cs . SE] 11 Jul 2026  \nAbstract—In software analytics, rerunning the same analysis twice often yields different models and conclusions. This reduces trust in the model and limits its use. We find that model instability is a major problem. Across 127 multi-objective SE optimization problems (12,700 test cases), repeated runs of a state-of-theart optimizer agree on only 13.7% of test cases, even under improved settings. We argue that this instability is not merely noise to tolerate, but a property that can be measured and managed. By adjusting how labels are spent, how complex the models become, and how splits are scored, we obtain models that agree 4.8 times as often as the default configuration. The standard deviation of optimization error falls by 22% on average (mean std 17.4 to 13.6), while recommendation quality improves rather than degrades. In terms of quality, the refined settings are statistically top-ranked on 119 of 127 datasets, compared to 74 for the defaults. We then test causal and data-locality interventionsand find that they help only partially, suggesting a residual stability floor. Our evidence suggests there are fundamental limits to stability set by the data itself (noise, scarce labels, proxy objectives, and the many near-equivalent models a dataset admits). We conclude that instability should be treated as a standard evaluation axis in SE optimization, which should be routinely measured, reported alongside performance, and used to calibrate trust in any single run. The methods in this paper provide a baseline against which future efforts to reduce SBSE instability can be judged.  \nTo support open science, we offer the following reproduction package: [https://tinyurl.com/Model-Instability](https://tinyurl.com/Model-Instability)  \nIndex Terms—Software Analytics, Rashomon Effect, SearchBased Software Engineering, Optimization, Causal Reasoning  \nI. INTRODUCTION  \nSoftware analytics is an influential research area with much practical value [1]–[6] . By mining historical project data, teams allocate testing effort, prioritize high-risk modules, and cut maintenance costs. Eventually, though, stakeholders aska deceptively simple question: “What did you learn from all that data?” The answer is usually a symbolic explanation, a decision tree, a rule set, or a causal graph. Unfortunately, these explanations are often unstable. Rerunning the same analysis yields different models and, therefore, different conclusions. This reduces trust in the model and, therefore, limits its use. For example, Figure 2 shows model variability seen after running the same optimizer several times on the same data using different random number seeds. As seen in Figure 2, the same learner can produce models with different structures and use different attributes. The problem of model instability is  \nnot just a quirk of the learner used in Figure 2 . Rather, it is endemic. As SE problems grow larger and larger, state-of-the-art methods inject randomness to scale, and conclusion instability has been reported across regression [7], text mining [8], defect prediction [9], causal discovery [10], and LLMs [11] .  \nA natural reaction is to treat instability as a bug to fix with more data, tuning, or ensembling. But the literature suggests a harder truth. Often, there is no single “best” explanation to recover. Breiman’s “two cultures” essay notes that structurally different models can fit the same data nearly as well [13], the Rashomon effect, which Xin et al. formalize as the (often enormous) Rashomo","cbCaiq05HoEuUDFi","https://ap.wps.com/l/cbCaiq05HoEuUDFi","pdf",2216329,1,12,"English","en",105,"# Introduction\n## Research questions and contributions","[{\"question\":\"What problem does the paper identify in software analytics and SE optimization?\",\"answer\":\"Repeated runs of the same analysis often produce different models and conclusions. This creates performance and structural instability that reduces trust and limits practical use.\"},{\"question\":\"How prevalent is performance instability according to the study?\",\"answer\":\"Across 127 multi-objective SE optimization problems and 12,700 test cases, repeated runs agree on only 13.7% of test cases even under improved settings.\"},{\"question\":\"How can instability be managed, and what evidence is provided about its limits?\",\"answer\":\"Adjusting label spending, model complexity, and split scoring yields models that agree 4.8 times more often than the default configuration, and reduces optimization error standard deviation by 22% on average. Causal and data-locality interventions help only partially, suggesting a residual stability floor driven by the data itself.\"}]",1784201794,30,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"is-model-instability-just-noise-to-be-tolerated-or-a-property-that-can-be-managed","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/is-model-instability-just-noise-to-be-tolerated-or-a-property-that-can-be-managed/85215/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the paper identify in software analytics and SE optimization?","Question",{"text":75,"@type":76},"Repeated runs of the same analysis often produce different models and conclusions. This creates performance and structural instability that reduces trust and limits practical use.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How prevalent is performance instability according to the study?",{"text":80,"@type":76},"Across 127 multi-objective SE optimization problems and 12,700 test cases, repeated runs agree on only 13.7% of test cases even under improved settings.",{"name":82,"@type":73,"acceptedAnswer":83},"How can instability be managed, and what evidence is provided about its limits?",{"text":84,"@type":76},"Adjusting label spending, model complexity, and split scoring yields models that agree 4.8 times more often than the default configuration, and reduces optimization error standard deviation by 22% on average. Causal and data-locality interventions help only partially, suggesting a residual stability floor driven by the data itself.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,122,127,130,134],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":28,"slug":121},"research-report",{"id":123,"doc_module":4,"doc_module_name":45,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":45,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":45,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":45,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]