[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82971-en":3,"doc-seo-82971-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},82971,687197207639,"Asher","https://ap-avatar.wpscdn.com/davatar_a8503ba1806abce46bf441b54a3ca4cd",8,"Research & Report","Do It Right! A Methodology for Successful NLP System Development","Natural language processing (NLP) supports clinical research and decision-making by extracting information from electronic medical records. While textbooks and tutorials cover text-processing algorithms, successful NLP projects depend on more than technical knowledge. The paper proposes a stepwise approach that applies the Systems Development Life Cycle (SDLC) to language-processing projects focused on information extraction. It emphasizes structured process discipline to reduce failure risk in clinical settings, especially when using modern large language models.","arXiv :2607 .05644v 1 [ cs .CL] 6 Jul 2026  \nDo It Right! A Methodology for Successful NLP System Development  \nPatterson OVa,∗, South Ba , Workman TLa , DuVall SLa  \na [Affiliation placeholder], , , , ,  \nAbstract  \nNatural language processing (NLP) is a common method for supplying data to clinical research and decision making by extracting information from electronic medical records [1, 2, 3] . Numerous textbooks and tutorials describe specific algorithms and applications for text processing [4, 5, 6, 7, 8, 9, 10, 11], yet algorithmic knowledge is only one ingredient of a successful NLP project. Drawing on the available literature, this paper presents a stepwise approach that applies the Systems Development Life Cycle (SDLC) to projects that rely on data extraction through language processing.  \nKeywords: clinical natural language processing, systems development lifecycle, information extraction, large language models, clinical text  \n1. Introduction  \nNLP encompasses a broad range of algorithms for the computerized processing of unstructured text. For roughly its first 40 years after emerging asa discipline in the 1950s, NLP was largely a promise of the future: linguists studied language structure, computer scientists developed algorithms, and hardware engineers expanded computational capacity. Many systems were built, but few crossed from academic research into practical use [2, 4, 5, 12] . Since the mid-1980s, the growth of electronic data, greater computing power, and an expanding open-source community have lowered barriers to entry and broadened NLP’s use, including clinical text processing for health outcomes research [13] .  \n∗ Corresponding author  \nEmail address: [ovpatterson@gmail.com](ovpatterson@gmail.com) (Patterson OV)  \nNLP has long since moved beyond computer science to become central to clinical informatics and biomedical research [7] . As interest has grown, experts have produced a wealth of learning materials covering individual tasks, including parsing, part-of-speech tagging, and semantic role labeling, as well as machine learning and system architecture [4, 5, 6, 7, 8, 9, 10, 11] . This abundance of resources and freely available implementations can suggest that building an information extraction system for a given use case requires little specialized effort.  \nThe recent proliferation of large language models (LLMs) has renewed this impression. Models capable of answering clinical questions, summarizing notes, and extracting structured data are now accessible to any researcher with an application programming interface (API) key. The apparent ease suggests that reliable extraction is now within routine reach. That picture is an illusion. These models fabricate plausible content not present in the source, omit information that is there, and produce inconsistent results from prompts that say the same thing differently. These models are impressive but do not remove the need for the underlying process. Instead, they make that need less visible.  \nNLP development carries the same project risks as any other software undertaking. The clinical literature rarely reports failures, but the Management Information Systems literature documents them extensively [14], and clinical NLP has no particular immunity. Business practitioners address these risks through the Systems Development Life Cycle (SDLC), a formal sequence of steps for developing computerized solutions. Applying the SDLC deliberately gives clinical researchers a structured means of reducing the likelihood of failure.  \nApplying clinical NLP also resembles retrospective manual chart abstraction. The primary difference is the use of computerized algorithms in place of human reviewers. The extensive chart-abstraction literature, therefore, offers relevant lessons in project success and failure [15, 16, 17] . The recent arrival of large language models draws this parallel even closer. Where earlier ruleand feature-based systems matched patterns that bore li","cbCaipVkKh03806Z","https://ap.wps.com/l/cbCaipVkKh03806Z","pdf",579508,3,1,31,"English","en",105,"# Introduction\n# Systems Development Life Cycle\n## Planning\n## Analysis\n## Design\n## Implementation\n## Testing\n## Deployment\n## Maintenance","[{\"question\":\"为什么仅掌握NLP算法不足以保证成功的信息抽取项目？\",\"answer\":\"算法知识只是成功NLP项目的一部分。论文指出，真正的成功还需要用系统化流程管理项目，从而降低失败概率。\"},{\"question\":\"论文如何将Systems Development Life Cycle（SDLC）用于临床NLP的信息抽取？\",\"answer\":\"论文提出把SDLC的规划、分析、设计、实现、测试、部署和维护等步骤应用到依赖语言处理的数据抽取项目中，并逐步指导从启动到完成的执行。\"},{\"question\":\"为什么大语言模型（LLMs）没有消除临床NLP开发的风险？\",\"answer\":\"论文认为LLMs可能会编造源文本中不存在的内容、遗漏已存在的信息，并且对提示措辞敏感导致结果不一致。因此仍需要底层流程与过程纪律来保证可靠抽取。\"}]",1784184388,78,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"do-it-right-a-methodology-for-successful-nlp-system-development","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/do-it-right-a-methodology-for-successful-nlp-system-development/82971/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"为什么仅掌握NLP算法不足以保证成功的信息抽取项目？","Question",{"text":75,"@type":76},"算法知识只是成功NLP项目的一部分。论文指出，真正的成功还需要用系统化流程管理项目，从而降低失败概率。","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"论文如何将Systems Development Life Cycle（SDLC）用于临床NLP的信息抽取？",{"text":80,"@type":76},"论文提出把SDLC的规划、分析、设计、实现、测试、部署和维护等步骤应用到依赖语言处理的数据抽取项目中，并逐步指导从启动到完成的执行。",{"name":82,"@type":73,"acceptedAnswer":83},"为什么大语言模型（LLMs）没有消除临床NLP开发的风险？",{"text":84,"@type":76},"论文认为LLMs可能会编造源文本中不存在的内容、遗漏已存在的信息，并且对提示措辞敏感导致结果不一致。因此仍需要底层流程与过程纪律来保证可靠抽取。","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]