[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83484-en":3,"doc-seo-83484-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},83484,687197100911,"Himbo","https://ap-avatar.wpscdn.com/avatar/a000239b6f1da00475?x-image-process=image/resize,m_fixed,w_180,h_180&k=1782698725881665579",8,"Research & Report","BT-APE: A Computationally Light Backtracking Approach to Automatic Prompt Engineering for Requirements Classification","Large language models increasingly support requirements engineering (RE) tasks, yet prompt design for requirements classification is often manually tuned, yielding inconsistent results and leaving automated prompt construction underexplored. BT-APE frames prompt creation as an optimization problem, iteratively refining prompts using LLM-generated candidates, backtracking search, and dynamic example selection. Evaluations on three benchmark datasets with five instruction-tuned LLMs show BT-APE matches a stronger but resource-intensive baseline, while using about 72% fewer input tokens and 66% less wall-clock time at equivalent accuracy.","arXiv :2607 .00427v 1 [ cs . SE] 1 Jul 2026  \nBT-APE: A Computationally Light Backtracking Approach to Automatic Prompt Engineering for Requirements Classification  \nMOHAMMAD AMIN ZADENOORI, Department of Statistics, University of Padova, Italy WAAD ALHOSHAN, Imam Mohammad Ibn Saud Islamic University (IMSIU), Saudi Arabia JACEK DĄBROWSKI, Lero, the Research Ireland Centre for Software, University of Limerick, Ireland LIPING ZHAO, University of Manchester, United Kingdom  \nALESSIO FERRARI, University College Dublin (UCD), Ireland and Istituto di Scienza e Tecnologie dell’Informazione “A. Faedo”(ISTI), Consiglio Nazionale delle Ricerche (CNR), Italy  \nLarge language models (LLMs) are increasingly applied to requirements engineering (RE) tasks, including requirements classification, model generation, trace-link detection and others. Prompts, which guide LLM behavior, are typically designed manually through trial and error, often leading to inconsistent and suboptimal performance on RE tasks. Despite the importance of prompt design, prior RE research largely relies on manually constructed prompts and does not systematically optimize them; moreover, automated methods for prompt construction remain largely unexplored, leaving their effectiveness unclear. To address this gap, we propose a lightweight Automatic Prompt Engineering (APE) approach named Backtracking APE (BTAPE) and apply it to requirements classification as a representative RE task. We frame prompt design as an optimization problem and iteratively refine prompts using LLM-generated candidates, backtracking search, and dynamic example selection. We evaluate BT-APE on three benchmark datasets with five instruction-tuned LLMs against four classical prompting baselines (zero-shot, few-shot, chain-of-thought, and CoT+few-shot) and a state-of-the-art, yet more resource intensive, APE baseline (PE2) . Our results show that BT-APE and PE2 achieve nearly identical performance, both substantially outperforming the four classical prompting baselines across datasets and models, with large effect sizes. However, compared with PE2, BT-APE imposes a substantially lighter computational footprint, consuming approximately 72% fewer cumulative input tokensand 66% less wall-clock time at equivalent accuracy (see Appendix C), making it better suited to deployment on small or resource-constrained servers. We also find that domain-informed prompt definitions enhance early performance, while iterative optimization partly compensates for weaker initial prompts. The contribution of this work is threefold: (i) a lightweight APE framework, together with an open interactive tool and replication package that operationalize the full pipeline; (ii) a comprehensive empirical evaluation across datasets and instruction-tuned LLMs that provides the first systematic comparison of APE against classical prompting for requirements classification; and (iii) insights into the impact of class definitions and prompt evolution on classification performance.  \nCCS Concepts: • Computing methodologies → Natural language generation; • Software and its engineering → Software organization and properties.  \nAuthors’ Contact Information: Mohammad Amin Zadenoori, [amin.zadenoori@unipd.it](amin.zadenoori@unipd.it), Department of Statistics, University of Padova, Padova,, Italy; WaadAlhoshan, [wmaboud@imamu.edu.sa](wmaboud@imamu.edu.sa), Imam Mohammad Ibn Saud Islamic University (IMSIU), Riyadh,, Saudi Arabia; Jacek Dąbrowski, [jacek.dabrowski@lero.ie](jacek.dabrowski@lero.ie), Lero, the Research Ireland Centre for Software, University of Limerick, Limerick, Ireland; Liping Zhao, [liping.zhao@manchester.ac.uk](liping.zhao@manchester.ac.uk), University of Manchester, Manchester,, United Kingdom; Alessio Ferrari, [alessio.ferrari@ucd.ie](alessio.ferrari@ucd.ie), University College Dublin (UCD), Dublin,, Ireland and Istituto di Scienza e Tecnologie dell’Informazione “A. Faedo”(ISTI), Consiglio Nazionale delle Ricerche (CNR), Pis","cbCailBRFxz1Awik","https://ap.wps.com/l/cbCailBRFxz1Awik","pdf",2354842,1,58,"English","en",105,"# Introduction\n## Prompt Engineering for Requirements Classification\n## BT-APE Method and Optimization Framing\n## Experimental Evaluation and Baselines\n## Efficiency, Findings, and Contributions","[{\"question\":\"What problem does BT-APE address in requirements engineering with LLMs?\",\"answer\":\"BT-APE targets the fact that prompt design for requirements classification is usually manual and trial-and-error, which can produce inconsistent and suboptimal performance, while automated prompt construction has not been systematically studied in this RE setting.\"},{\"question\":\"How does BT-APE optimize prompts?\",\"answer\":\"BT-APE treats prompt design as an optimization problem, iteratively refining prompts with LLM-generated candidates, a backtracking search strategy, and dynamic example selection.\"},{\"question\":\"What are BT-APE’s main results compared with the classical prompting baselines and PE2?\",\"answer\":\"BT-APE substantially outperforms classical baselines (zero-shot, few-shot, chain-of-thought, and CoT+few-shot) with large effect sizes, and it achieves nearly identical performance to PE2 while requiring about 72% fewer cumulative input tokens and 66% less wall-clock time at equivalent accuracy.\"}]",1784188343,146,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"bt-ape-a-computationally-light-backtracking-approach-to-automatic-prompt-engineering-for-requirements-classification","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/bt-ape-a-computationally-light-backtracking-approach-to-automatic-prompt-engineering-for-requirements-classification/83484/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does BT-APE address in requirements engineering with LLMs?","Question",{"text":75,"@type":76},"BT-APE targets the fact that prompt design for requirements classification is usually manual and trial-and-error, which can produce inconsistent and suboptimal performance, while automated prompt construction has not been systematically studied in this RE setting.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does BT-APE optimize prompts?",{"text":80,"@type":76},"BT-APE treats prompt design as an optimization problem, iteratively refining prompts with LLM-generated candidates, a backtracking search strategy, and dynamic example selection.",{"name":82,"@type":73,"acceptedAnswer":83},"What are BT-APE’s main results compared with the classical prompting baselines and PE2?",{"text":84,"@type":76},"BT-APE substantially outperforms classical baselines (zero-shot, few-shot, chain-of-thought, and CoT+few-shot) with large effect sizes, and it achieves nearly identical performance to PE2 while requiring about 72% fewer cumulative input tokens and 66% less wall-clock time at equivalent accuracy.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":45,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":45,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":45,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]