[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83365-en":3,"doc-seo-83365-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83365,1099514068365,"Aurelia","https://ap-avatar.wpscdn.com/avatar/10000253d8d9f28188e?_k=1776742907772140068",8,"Research & Report","Grounded Event Extraction from SEC 8-K Filings with a Fine-Grained Taxonomy","Form 8-K filings are the primary mechanism for U.S. public companies to disclose material events, yet SEC item codes are coarse and often combine multiple economically distinct actions into a single label. Fine-grained event labeling is feasible with large language models, but reliability requires traceability to source text and constraint to a fixed label set. The proposed two-stage system tags disclosures using a three-tier taxonomy of 119 event types, grounding each tag in verbatim quotes via fuzzy n-gram validation and assigning quality scores by re-grading against category definitions. Applied to 292,984 filings (2022–2026), it produces 601,088 grounded tags and releases stratified results. Precision is calibrated by score and validated with an event study.","Grounded Event Extraction from SEC 8-K Filings with a  \nFine-Grained Taxonomy  \nRian Dolphin∗ [Massive.com](Massive.com)[ ](Massive.com)Dublin, Ireland  \nJoe Dursun  \n[Massive.com](Massive.com)[ ](Massive.com)Atlanta, Georgia, USA  \nJarrett Blankenship  \n[Massive.com](Massive.com)[ ](Massive.com)Atlanta, Georgia, USA  \nKatie Adams  \n[Massive.com](Massive.com)[ ](Massive.com)Atlanta, Georgia, USA  \nQuinton Pike  \n[Massive.com](Massive.com)[ ](Massive.com)Atlanta, Georgia, USA  \narXiv :2607 .08346v 1 [ cs .CL] 9 Jul 2026  \nAbstract  \n[Form 8-K filings are the primary channel through which U.S. pub](Form 8-K filings are the primary channel through which U.S. pub)lic companies disclose material events, but the SEC item codes attached to them are coarse: a single item spans routine administrative changes and chief executive departures, and many of the most market-moving disclosures fall into a catch-all item. Large language models make fine-grained labelling feasible at corpus scale, but only if the labels can be traced to the source text and shown to be reliable. We present a two-stage system that tags 8-K disclosures against a three-tier taxonomy of 119 event types. The first stage constrains output to valid taxonomy entries and anchors every tag to a verbatim quote via fuzzy n-gram validation; the second re-grades each cited quote against the category definition to produce a quality score. Applying the system to 292,984 filings from 2022 to 2026 yields 601,088 grounded event tags, which we release. Over 5,125 stratified tags, an LLM judge finds precision rises monotonically with the quality score, from 12% to 96%, while unsupported tags fall from 8% to near zero. Ablation shows the score is calibrated only when assigned in a dedicated second pass. An event study on unsigned abnormal returns confirms, without any language model, that the taxonomy separates economically distinct events sharing an item code.  \nKeywords  \nSEC filings, event extraction, large language models, event study, LLM-as-judge  \n1 Introduction  \nU.S. public companies must report material corporate events on Form 8-K within four business days. The 32 SEC item codes attached to each filing, such as Item 5.02 for officer and director changes or Item 1.01 for material agreements, are the standard machinereadable event labels: they determine legal filing obligations, organize disclosure databases, and serve as event-type controls in a large empirical literature [8, 10] . The codes were designed as legal filing categories rather than as economic event types, and they are coarse along two dimensions. Firstly, a single item conflates economically distinct events: Item 5.02 covers both the departure of a chief executive and the routine retirement of a single outside director. Secondly, a large share of substantive disclosure carries  \n∗ [rian@massive.com](rian@massive.com)  \nThe processed data and taxonomy are available at: [massive.com/docs/rest/stocks/filings/8-k-disclosures](massive.com/docs/rest/stocks/filings/8-k-disclosures)  \nno informative item at all: Item 8 .01 (“Other Events”) is a voluntary catch-all, and we find that it is the modal location of many of the most price-relevant event types, including clinical trial results, merger completions, and regulatory decisions.  \nLarge language models make it feasible to assign fine-grained, economically motivated event labels at corpus scale. Doing so for research and applied use raises two requirements that generic prompting does not meet. The labels must be constrained, so that outputs map onto a fixed vocabulary that downstream users can rely on, and they must be auditable, so that every label can be traced to specific language in the source document and assessed for reliability without rerunning the model.  \nThis paper presents and evaluates an LLM-based extraction system built around these requirements. The system places a three-tier taxonomy of 119 corporate event types in the prompt of a compact instruction-","cbCaitD6VAPL9S4q","https://ap.wps.com/l/cbCaitD6VAPL9S4q","pdf",725669,3,1,9,"English","en",105,"# Introduction\n## Form 8-K and coarse SEC item codes\n## Need for fine-grained, auditable labels\n# Two-stage LLM extraction system\n## Three-tier taxonomy of 119 event types\n## Schema validation and quote grounding\n## Second-pass quality scoring\n# Evaluation\n## Intrinsic evaluation with LLM-as-judge\n## Economic evaluation via event study","[{\"question\":\"Why are SEC item codes in Form 8-K insufficient for event research?\",\"answer\":\"SEC item codes are designed as legal filing categories and are coarse: a single code can conflate distinct economic events, and many market-moving disclosures appear in catch-all items with little informative structure.\"},{\"question\":\"How does the proposed system ensure event labels are auditable and reliable?\",\"answer\":\"Each predicted event tag is anchored to a verbatim quote from the filing using fuzzy n-gram validation, and the first stage enforces schema validation so only taxonomy-valid labels are produced.\"},{\"question\":\"What is the role of the second stage and quality scores?\",\"answer\":\"The second stage re-reads each cited quote against the category definition to assign a calibrated quality score from 1 to 5, turning the score into a reliability dial that predicts precision.\"}]",1784187015,23,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"grounded-event-extraction-from-sec-8-k-filings-with-a-fine-grained-taxonomy","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/grounded-event-extraction-from-sec-8-k-filings-with-a-fine-grained-taxonomy/83365/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-25","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why are SEC item codes in Form 8-K insufficient for event research?","Question",{"text":75,"@type":76},"SEC item codes are designed as legal filing categories and are coarse: a single code can conflate distinct economic events, and many market-moving disclosures appear in catch-all items with little informative structure.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the proposed system ensure event labels are auditable and reliable?",{"text":80,"@type":76},"Each predicted event tag is anchored to a verbatim quote from the filing using fuzzy n-gram validation, and the first stage enforces schema validation so only taxonomy-valid labels are produced.",{"name":82,"@type":73,"acceptedAnswer":83},"What is the role of the second stage and quality scores?",{"text":84,"@type":76},"The second stage re-reads each cited quote against the category definition to assign a calibrated quality score from 1 to 5, turning the score into a reliability dial that predicts precision.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]