[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-117777-en":3,"doc-seo-117777-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},117777,1099513958762,"Logic","https://ap-avatar.wpscdn.com/avatar/1000023916a998db790?x-image-process=image/resize,m_fixed,w_180,h_180&k=1784791008015729253",8,"Research & Report","Approaches to Improving the Accuracy of Machine Learning Models in Requirements Elicitation Techniques Selection - Preprint","Selecting techniques is central to business analysis approach planning in IT projects, especially when choosing requirements elicitation techniques. The effectiveness of machine learning models used for this selection depends strongly on training data balance, which becomes problematic when popular techniques dominate and the dataset is imbalanced. This paper analyzes using the Synthetic Minority Over-sampling Technique in ML models for elicitation-technique selection under imbalanced training data and explores ways to improve positive feature-importance selection. Computational experiments confirm improved accuracy and propose reusable ML approaches for planning business analysis activities.","Preprint  \nApproaches to Improving the Accuracy of Machine Learning Models in Requirements Elicitation Techniques Selection  \nDenys Gobov1, Olga Solovei2  \n1 National Technical University of Ukraine \"Igor Sikorsky Kyiv Polytechnic Institute\", Kyiv, Ukraine  \n2 Kyiv National University of Construction and Architecture, Ukraine [d.gobov@kpi.ua](d.gobov@kpi.ua), [solovey.ol@knuba.edu.ua](solovey.ol@knuba.edu.ua)  \nAbstract: Selecting techniques is a crucial element of the business analysis approach planning in IT projects. Particular attention is paid to the choice of techniques for requirements elicitation. One of the promising methods for selecting techniques is using machine learning algorithms trained on the practitioners' experience considering different projects' contexts. The effectiveness of ML models is significantly affected by the balance of the training dataset, which is violated in the case of popular techniques. The paper aims to analyze the efficiency of the Synthetic Minority Over-sampling Technique usage in Machine Learning models for elicitation technique selection in case of the imbalanced training dataset and possible ways for positive feature importance selection. The computational experiment results confirmed the effectiveness of using the proposed approaches to improve the accuracy of machine learning models for selecting requirements elicitation techniques. Proposed approaches can be used to build Machine Learning models for business analysis activities planning in IT projects.  \nKeywords: requirement elicitation technique, machine learning, decision tree, over-sampling technique, binary classification problem.  \n1. Introduction  \nThe choice of techniques for effectively identifying requirements in developing IT solutions is essential in planning business analysis work. A thorough understanding of the variety of techniques available, their advantages and disadvantages assists the business analyst in adapting to a particular project context [1] . The business analyst must create a combination of techniques to guarantee the success of software  \nrequirements identification activities, as it is impossible to fulfill all project stakeholders' needs using just one technique [2] . One approach to solving this problem is to use machine learning models, where training samples are formed based on practitioners' experience or recommendations. For example, in studies [3,4], a machine learning model was built that recommends the usage of elicitation techniques depending on the combinations of factors. However, the use of machine learning models for commonly used techniques and the Accuracy of their work is associated with several difficulties. If we select the most frequently used techniques as a target class, the gathered dataset got imbalanced, i.e., most observations belong to a target class with a positive value equal to \" 1,\" which indicates that the technique was used in the project.  \nThe mentioned model from [3] was constructed with Decision Jungle Tree (DJT) algorithm that was empirically selected as the most efficient. DJT learner, like a binary decision tree learner, creates a tree that is biased to the majority class when a dataset is imbalanced [5]. The reason causes are the following: while recursively partitioning the dataset so that the observations with similar target values are grouped together, it qualifies the candidate split of the node m using the parameters that minimize the impurity Q = argmint G (Qm , t )(1), where  \nG (Qm , t ) = r~~m~~rlefmt~~ ~~ H left ( Qleftm (t )) + ~~ ~~rr~~m~~rigmht~~ ~~ H right ( Qrightm (t )) (2)  \nH left ( Qleftm (t )) (3), H right ( Qrightm (t )) (4)– are gini or entropy measures of split's impurity for the classification task. The lower the value of H left ( Qleftm (t )) (5), H right ( Qrightm (t )) (6), the better the split. When the dataset is imbalanced, there is a significant probability that the majority of the class samples are included in the same nodes, wh","cbCairVhJBsFrELS","https://ap.wps.com/l/cbCairVhJBsFrELS","pdf",594749,1,11,"English","en",105,"# Introduction\n## Requirements elicitation technique selection in IT projects\n## Machine learning models and dataset imbalance\n## Decision tree bias under imbalanced data\n## Accuracy vs AUC under imbalance\n## Proposed improvements: data balancing and feature importance","[{\"question\":\"Why is selecting requirements elicitation techniques important in IT project planning?\",\"answer\":\"Selecting appropriate techniques is essential for effectively identifying requirements. Business analysts combine multiple techniques because no single technique can satisfy all stakeholders' needs in a project context.\"},{\"question\":\"How does imbalanced training data affect machine learning accuracy for technique selection?\",\"answer\":\"When frequently used techniques dominate the target class, the dataset becomes imbalanced. This can bias decision-tree-based learners toward the majority class, degrading model performance and producing misleading metrics.\"},{\"question\":\"What approaches does the paper propose to improve ML model accuracy under imbalance?\",\"answer\":\"The paper proposes adding a data preprocessing step to balance the training dataset using the Synthetic Minority Over-sampling Technique. It also proposes machine-learner-independent methods for computing feature importance to avoid errors caused by decision tree bias.\"}]","Approaches to Improving the Accuracy of Machine Learning Models in Requirements Elicitation Techniques Selection - Preprint | PDF",1785679497,28,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"approaches-to-improving-the-accuracy-of-machine-learning-models-in-requirements-elicitation-techniques-selection-preprint","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/approaches-to-improving-the-accuracy-of-machine-learning-models-in-requirements-elicitation-techniques-selection-preprint/117777/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-05","2026-08-02",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"Why is selecting requirements elicitation techniques important in IT project planning?","Question",{"text":76,"@type":77},"Selecting appropriate techniques is essential for effectively identifying requirements. Business analysts combine multiple techniques because no single technique can satisfy all stakeholders' needs in a project context.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How does imbalanced training data affect machine learning accuracy for technique selection?",{"text":81,"@type":77},"When frequently used techniques dominate the target class, the dataset becomes imbalanced. This can bias decision-tree-based learners toward the majority class, degrading model performance and producing misleading metrics.",{"name":83,"@type":74,"acceptedAnswer":84},"What approaches does the paper propose to improve ML model accuracy under imbalance?",{"text":85,"@type":77},"The paper proposes adding a data preprocessing step to balance the training dataset using the Synthetic Minority Over-sampling Technique. It also proposes machine-learner-independent methods for computing feature importance to avoid errors caused by decision tree bias.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,121,124,129,132,136],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":46,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":46,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":46,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":133,"doc_module":4,"doc_module_name":46,"category_name":134,"show_sort_weight":133,"slug":135},10,"Lifestyle","lifestyle",{"id":137,"doc_module":4,"doc_module_name":46,"category_name":138,"show_sort_weight":107,"slug":139},19,"General","general"]