[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-119692-en":3,"doc-seo-119692-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},119692,1099513958607,"Jiven","https://ap-avatar.wpscdn.com/avatar/100002390cf8733938c?x-image-process=image/resize,m_fixed,w_180,h_180&k=1778829742770036399",8,"Research & Report","Transforming Unstructured Text into Data with Context Rule Assisted Machine Learning (CRAML)","Transforming Unstructured Text into Data with Context Rule Assisted Machine Learning (CRAML) presents a method and no-code tools that let domain experts build custom structured, labeled datasets from unstructured document text. CRAML uses context rule–assisted workflows to create accurate, reproducible labeling at massive scale while making the resulting text classification models traceable to expert-written rules. The paper demonstrates three use cases, producing document-level coded tabular datasets for quantitative research and scalable niche schemes for qualitative studies. It also releases open-source software, classifiers, and a replicable training example, aimed at low-resource, flexible supervised ML.","arXiv :2301 .08549v1 [ cs .CL] 20 Jan 2023  \nTransforming Unstructured Text into Data with Context Rule Assisted  \nMachine Learning (CRAML)  \nStephen Meisenbacher∗†  \nPeter Norlander∗‡  \nJanuary 23, 2023  \nAbstract  \nAbstract: We describe a method and new no-code software tools enabling domain experts to build custom structured, labeled datasets from the unstructured text of documents and build niche machine learning text classi􀀜cation models traceable to expert-written rules. The Context Rule Assisted Machine Learning (CRAML) method allows accurate and reproducible labeling of massive volumes of unstructured text. CRAML enables domain experts to access uncommon constructs buried within a document corpus, and avoids limitations of current computational approaches that often lack context, transparency, and interpetability. In this research methods paper, we present three use cases for CRAML: we analyze recent management literature that draws from text data, describe and release new machine learning models from an analysis of proprietary job advertisement text, and present 􀀜ndings of social and economic interest from a public corpus of franchise documents. CRAML produces document-level coded tabular datasets that can be used for quantitative academic research, and allows qualitative researchers to scale niche classi􀀜cation schemes over massive text data. CRAML is a low-resource, 􀀝exible, and scalable methodology for building training data for supervised ML. We make available as open-source resources: the software, job advertisement text classi􀀜ers, a novel corpus of franchise documents, and a fully replicable start-to-􀀜nish trained example in the context of no poach clauses.  \nKeywords: labor markets, data creation, text classi􀀜cation, hybrid system, big data  \n∗ Equally contributing authors. We are grateful for support from the Economic Security Project Anti-Monopoly Fund; Loyola Rule of Law Institute; Loyola Quinlan School of Business; and Loyola University Chicago. We also thank Patricia Tabarani, Eric George, Steve Sauerwald, and Chris Erickson.  \n†Technical University of Munich, School of Computation, Information and Technology ‡Loyola University Chicago, Quinlan School of Business  \n1 Introduction  \nAdvances in computational methods o􀀛er new ways to gain insight from large volumes of unstructured text, and yet these methods each have signi􀀜cant limitations, tradeo􀀛s, and a lack clear guidelines [99] . Despite the advances in Natural Language Processing (NLP), Machine Learning (ML), Arti􀀜cial Intelligence (AI) and more recently Deep Learning (DL), there is still 􀀐the herculean task of 􀀜nding constructs using full-text search􀀑 [79, p. 547] .  \nUnstructured text lack a schema, are not standardized, have multiple formats, and come from diverse sources [9] . Greater attention and systematic research is needed to develop processes to create structured data from unstructured text [45] . A lack of structured information is a barrier to understanding for researchers and organizations: 􀀐as much as 80% of an organization’s data is ‘dark􀀑’  \n[78] . For knowledge professionals, just-in-time access to domain speci􀀜c document repositories is vital, but knowledge must 􀀜rst be codi􀀜ed [114] . Zettabytes (ZB) of new data are produced daily [20], and roughly 80% is unstructured [67] . While the manual expert labor required for qualitatively coding novel classi􀀜cation schemes is hard to scale, statistical packages, many information systems, and quantitative social scientists expect data to be codi􀀜ed and in a regular format: 􀀐tabular data 􀀕 variables in columns, cases in rows􀀑 [80] .  \nML and AI are potential solutions, but require structured training data. Hand-coding by experts is required in many domains, but does not easily scale. To address the disjunction between these, we describe a method that bridges expert-built set of context rules scaled to classify unstructured text and build training data for machine learning. Context Rule Assis","cbCaicUGMk5PAXuW","https://ap.wps.com/l/cbCaicUGMk5PAXuW","pdf",1149380,1,53,"English","en",105,"# Introduction\n## Problem: Unstructured text and lack of structured schemas\n## Need: Training data for ML/AI\n## Solution: CRAML method and workflow\n## Contributions and transparency goals","[{\"question\":\"What problem does CRAML address in machine learning from text?\",\"answer\":\"CRAML targets the difficulty of converting unstructured documents into structured, labeled training data, especially when context-specific constructs are hard to find with standard full-text search or generic NLP pipelines.\"},{\"question\":\"How does Context Rule Assisted Machine Learning work?\",\"answer\":\"CRAML uses expert-written context rules inside a workflow that supports building structured labeled datasets from unstructured text, then trains ML models whose outputs remain traceable to those rules.\"},{\"question\":\"What kinds of outputs does CRAML produce for researchers?\",\"answer\":\"CRAML produces document-level coded tabular datasets suitable for quantitative research and enables qualitative researchers to scale niche text classification schemes over large corpora.\"}]","Transforming Unstructured Text into Data with Context Rule Assisted Machine Learning (CRAML) | PDF",1785725801,134,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"transforming-unstructured-text-into-data-with-context-rule-assisted-machine-learning-craml","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/transforming-unstructured-text-into-data-with-context-rule-assisted-machine-learning-craml/119692/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-03",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does CRAML address in machine learning from text?","Question",{"text":75,"@type":76},"CRAML targets the difficulty of converting unstructured documents into structured, labeled training data, especially when context-specific constructs are hard to find with standard full-text search or generic NLP pipelines.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does Context Rule Assisted Machine Learning work?",{"text":80,"@type":76},"CRAML uses expert-written context rules inside a workflow that supports building structured labeled datasets from unstructured text, then trains ML models whose outputs remain traceable to those rules.",{"name":82,"@type":73,"acceptedAnswer":83},"What kinds of outputs does CRAML produce for researchers?",{"text":84,"@type":76},"CRAML produces document-level coded tabular datasets suitable for quantitative research and enables qualitative researchers to scale niche text classification schemes over large corpora.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]