[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-81704-en":3,"doc-seo-81704-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},81704,3848291630094,"Emma Wilson","https://eur-avatar.wpscdn.com/davatar_085a072bc5b1113ac321206ff7593b45",8,"Research & Report","SchemaRAG: Dynamic Large Schema Reduction for LLM-driven Structured Information Extraction","Extracting structured data from unstructured text with large language models (LLMs) becomes difficult when target schemas are large, complex, and hierarchical. Including the full schema in prompts can raise cost and latency, worsen lost-in-the-middle effects, and hit context length limits. SchemaRAG introduces a retrieval-augmented generation framework that dynamically prunes the output schema space using schema metadata and optional few-shot examples. Evaluations on healthcare and e-commerce datasets show up to an 8.8% micro-F1 gain, 47% lower latency, and 48% reduced token costs.","SchemaRAG: Dynamic Large Schema Reduction for LLM-driven Structured Information Extraction  \nSin Yu Bonnie Ho* Arlie Coles* Erik Larsson  \nEric Marshall Nathan Bodenstab Paul Vozila  \nMicrosoft  \narXiv :2607 .00008v 1 [ cs .IR] 4 May 2026  \nAbstract  \nExtracting structured data from unstructured text using large language models (LLMs) becomes challenging when the target schemas are large and complex. In such cases, including the full schema in the prompt increases cost and latency, risks lost-in-the-middle performance degradation, and can exceed context length limits. We propose SchemaRAG, a retrievalaugmented generation (RAG) framework that dynamically prunes the output schema space for schema-conditioned information extraction tasks by leveraging schema metadata and fewshot examples (when available) . We evaluate SchemaRAG on real-world healthcare ande-commerce datasets. Our results show that SchemaRAG can achieve up to an 8.8% increase in micro-F 1 , a 47% reduction in latency, and a 48% reduction in token costs, demonstrating its practicality for large-schema extraction.  \n1 Introduction  \nStructured information extraction (IE) pairs values from unstructured text with schema-defined keys. In healthcare, this includes medical attribute extraction from unstructured clinical notes (Jiang et al., 2011 ; Agrawal et al., 2022) and accelerated human annotation of medical narrative data (Goel et al., 2023) . Similarly, e-commerce applications focus on product attribute value extraction from unstructured text descriptions with and without supplementary multimodal data (Brinkmann et al., 2024 ; Zhu et al., 2020) .  \nFurther applications of structured IE exist for problems that can be cast as form-filling, where models need to populate complex, hierarchical schemas from unstructured text. These forms could have numerous rows, and each row could have complex constraints on potential values (e.g., required  \n*Equal contribution. Corresponding authors:  \n[siho@microsoft.com](siho@microsoft.com), [arliecoles@microsoft.com](arliecoles@microsoft.com)  \ntypes including strings, numbers, single-select radio buttons, multi-select checkboxes, or other custom types) . In this work, we consider two such use cases: 1) populating electronic health records (EHRs) from medical narratives and 2) filling a product metadata form from e-commerce product description text.  \nInstruction-tuned large language models (LLMs) are well-suited for these tasks due to their broad semantic knowledge (Singhal et al., 2023) and ability to follow structured prompt instructions (Liu et al., 2024b) . However, naively injecting large schemas into the prompt can exceed context limits or introduce many irrelevant rows. This leads to lost-in-the-middle effects (Liu et al., 2024a) and performance degradation as the model struggles to isolate pertinent rows. The resulting latency and inference costs can preclude real-time applications. To mitigate these issues, we propose SchemaRAG, a retrieval-augmented generation approach that retrieves and prunes the output schema space.  \nBy leveraging schema metadata and, when available, limited labeled data, SchemaRAG dynamically selects a relevant schema subset for just-intime in-context prompted LLM use (Figure 1) . A well-selected reduced schema shrinks prompt size, decreases the chance of the LLM filling in irrelevant rows, and improves overall extraction performance.  \nThe framework is schema-agnostic and supports arbitrary levels of hierarchy. Because the embeddings driving the retrieval can be computed offline, the system is efficient and easily integrated into existing pipelines. Crucially, SchemaRAG operates as a training-free solution, requiring no pretraining or fine-tuning. We explicitly focus on schemaconditioned extraction using hosted LLMs under real-world deployment scenarios including large, complex schemas and strict resource budgets. By utilizing prompting rather than retraining, our approach allows for rapid iteratio","cbCaileJQkHx6wN5","https://ap.wps.com/l/cbCaileJQkHx6wN5","pdf",862867,2,1,14,"English","en",105,"# Introduction\n# Related Work","[{\"question\":\"Why does injecting a large schema into an LLM prompt hurt structured information extraction?\",\"answer\":\"Large schemas increase prompt length, cost, and latency, can trigger context length limits, and can cause lost-in-the-middle performance degradation by distracting the model from relevant rows or fields.\"},{\"question\":\"What is SchemaRAG and how does it address large-schema extraction?\",\"answer\":\"SchemaRAG is a retrieval-augmented generation framework that dynamically selects and prunes a reduced subset of the output schema, guided by schema metadata and optional few-shot examples, before running a just-in-time prompted LLM extraction.\"},{\"question\":\"What results does SchemaRAG achieve on real-world datasets?\",\"answer\":\"On healthcare and e-commerce datasets, SchemaRAG reports up to an 8.8% increase in micro-F1, along with up to 47% latency reduction and 48% token-cost reduction, though outcomes vary by dataset.\"}]",1784175509,35,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"schemarag-dynamic-large-schema-reduction-for-llm-driven-structured-information-extraction","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/schemarag-dynamic-large-schema-reduction-for-llm-driven-structured-information-extraction/81704/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why does injecting a large schema into an LLM prompt hurt structured information extraction?","Question",{"text":75,"@type":76},"Large schemas increase prompt length, cost, and latency, can trigger context length limits, and can cause lost-in-the-middle performance degradation by distracting the model from relevant rows or fields.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What is SchemaRAG and how does it address large-schema extraction?",{"text":80,"@type":76},"SchemaRAG is a retrieval-augmented generation framework that dynamically selects and prunes a reduced subset of the output schema, guided by schema metadata and optional few-shot examples, before running a just-in-time prompted LLM extraction.",{"name":82,"@type":73,"acceptedAnswer":83},"What results does SchemaRAG achieve on real-world datasets?",{"text":84,"@type":76},"On healthcare and e-commerce datasets, SchemaRAG reports up to an 8.8% increase in micro-F1, along with up to 47% latency reduction and 48% token-cost reduction, though outcomes vary by dataset.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]