[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-187683-en":3,"doc-seo-187683-105":30,"detail-sidebar-cat-1-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":11,"category_id":12,"category_name":13,"doc_title":14,"doc_description":15,"doc_content":16,"file_id":17,"file_url":18,"file_type":19,"file_size":20,"view_count":11,"is_deleted":4,"is_public":11,"is_downloadable":11,"audit_status":11,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":15,"update_tm":28,"read_time":29},187683,2336477405376,"Stanley","https://ap-avatar.wpscdn.com/davatar_29158cc5080c5b710cf443261637dec0",1,158,"General","Document Metadata Extraction","This document appears to be a technical report or research paper detailing a process for document metadata extraction, likely within the context of machine learning or Natural Language Processing. It outlines a workflow that includes 'Vertical Classifier Training,' 'Field Classifier Training,' 'Validation,' and 'Production.' The progress and performance are tracked through metrics such as 'Template Coverage,' 'Vertical Coverage,' and 'Rule Coverage,' with specific percentages (100%, 50%, 30%) indicating the achieved levels for each. A graphical representation shows the 'Objective' score over numerous 'Trial id's, plotting the performance of different experimental runs. The document also presents tables comparing the performance of various models (Bills, Offers, Hotels) using statistical measures like AUC ROC and AUC Precision Recall, and detailing template effectiveness (Existing vs. New) and template discovery rates. The overall focus is on evaluating and improving the accuracy and coverage of automated document analysis systems.","| Feature Name | Description |\n| --- | --- |\n| subject-text | Words in the subject line |\n| sender-text | Tokens in the sender field |\n| top-text | Top 150 words in the body |\n| strong-text | Text marked header, title, bold etc. |\n| alt-text | Alt-text supplied for image content |\n| footer-text | Last 100 words in the body |\n| html-tag-count | Number of HTML tags in the body |\n| text-token-count | Number of text tokens in the body |\n| link-tag-count | Number of link tags |\n| image-tag-count | Number of image tags |\n| script-tag-count | Number of script tags |\n| table-tag-count | Number of table tags |\n| datetime-count | Number of candidate date-time spans |\n| salient-entities | Top entity IDs |\n\n| Feature Name | Description |\n| --- | --- |\n| {5/10/20}-w-before | 5/10/20 words before the candidate span |\n| {5/10/20}-w-after | 5/10/20 words after the candidate span |\n| field-text | Contents ofthe candidate span |\n| doc-index | Position of the field in the document (0-1) |\n| candidate-index | Positional rank relative to all candidates |\n\n| Statistic | Bills | Offers | Hotels |\n| --- | --- | --- | --- |\n| AUC ROC\u003Cbr>AUC Precision Recall | 0.9727\u003Cbr>0.8010 | 0.9995\u003Cbr>0.9998 | 0.9999\u003Cbr>0.8837 |\n\n\n| Model | Structural Templates |  | Sender-Subject Templates |  |\n| --- | --- | --- | --- | --- |\n|  | Existing | New | Existing | New |\n| Bills | 84.6 | 66.8 | 94.1 | 76.2 |\n| Offers | 100.0 | 87.8 | 100.0 | 87.4 |\n| Hotels | 97.7 | 98.7 | 100.0 | 78.0 |\n\n\n| Description | Bills | Offers | Hotels |\n| --- | --- | --- | --- |\n| Emails matching a positive template\u003Cbr>Emails matching templates with rules | 73.3\u003Cbr>58.7 | 70.0\u003Cbr>40.0 | 75.9\u003Cbr>38.2 |\n| New templates discovered as a percentage of pre-existing templates | 79.0 | 8.0 | 57.6 |","cbCaiiaVcktULy6a","https://ap.wps.com/l/cbCaiiaVcktULy6a","pdf",695617,10,"English","en",105,"# Document Metadata Extraction (Multimodal)\n## Analysis and Workflow\n## Performance Metrics\n### Model Comparison\n#### Template Effectiveness","[{\"question\":\"What are the key stages in the document metadata extraction workflow presented?\",\"answer\":\"The workflow includes Vertical Classifier Training, Field Classifier Training, Validation, and Production stages. These are followed by Bottleneck Analysis and review of coverage metrics.\"},{\"question\":\"What performance metrics are used to evaluate the extraction process?\",\"answer\":\"The evaluation involves metrics such as Template Coverage (100%), Vertical Coverage (50%), and Rule Coverage (30%). A graph also plots the 'Objective' score against 'Trial id' to track experimental performance.\"},{\"question\":\"How does the document compare the performance of different models and templates?\",\"answer\":\"Tables compare models for Bills, Offers, and Hotels using AUC ROC and AUC Precision Recall. They also detail the effectiveness of existing and new structural and sender-subject templates, and the discovery rate of new templates.\"}]","Document Metadata Extraction | PDF",1788384785,4,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":14,"keywords":34,"description":15,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"document-metadata-extraction-187683","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":11},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/template/","Template",2,{"item":49,"name":13,"@type":43,"position":50},"https://docshare.wps.com/template/general/",3,{"item":52,"name":14,"@type":43,"position":29},"https://docshare.wps.com/template/document-metadata-extraction-187683/187683/",{"url":52,"name":14,"@type":54,"author":55,"headline":14,"publisher":57,"fileFormat":60,"inLanguage":23,"description":15,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-09-04","2026-09-02",true,{"@type":65,"interactionType":66,"userInteractionCount":47},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What are the key stages in the document metadata extraction workflow presented?","Question",{"text":75,"@type":76},"The workflow includes Vertical Classifier Training, Field Classifier Training, Validation, and Production stages. These are followed by Bottleneck Analysis and review of coverage metrics.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What performance metrics are used to evaluate the extraction process?",{"text":80,"@type":76},"The evaluation involves metrics such as Template Coverage (100%), Vertical Coverage (50%), and Rule Coverage (30%). A graph also plots the 'Objective' score against 'Trial id' to track experimental performance.",{"name":82,"@type":73,"acceptedAnswer":83},"How does the document compare the performance of different models and templates?",{"text":84,"@type":76},"Tables compare models for Bills, Offers, and Hotels using AUC ROC and AUC Precision Recall. They also detail the effectiveness of existing and new structural and sender-subject templates, and the discovery rate of new templates.","https://schema.org",{"og:url":52,"og:type":87,"og:title":14,"og:site_name":58,"og:description":15},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,98,103,108,113,118,123,128,133],{"id":94,"doc_module":11,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},11,"Presentations",90,"presentations",{"id":99,"doc_module":11,"doc_module_name":46,"category_name":100,"show_sort_weight":101,"slug":102},12,"Resumes",80,"resumes",{"id":104,"doc_module":11,"doc_module_name":46,"category_name":105,"show_sort_weight":106,"slug":107},14,"Invoices",70,"invoices",{"id":109,"doc_module":11,"doc_module_name":46,"category_name":110,"show_sort_weight":111,"slug":112},15,"Posters",60,"posters",{"id":114,"doc_module":11,"doc_module_name":46,"category_name":115,"show_sort_weight":116,"slug":117},16,"Social Media",50,"social-media",{"id":119,"doc_module":11,"doc_module_name":46,"category_name":120,"show_sort_weight":121,"slug":122},17,"Forms",40,"forms",{"id":124,"doc_module":11,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},18,"Letters",30,"letters",{"id":129,"doc_module":11,"doc_module_name":46,"category_name":130,"show_sort_weight":131,"slug":132},21,"Paper Templates",5,"papers-templates",{"id":12,"doc_module":11,"doc_module_name":46,"category_name":13,"show_sort_weight":4,"slug":134},"general-158"]