[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-186251-en":3,"doc-seo-186251-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},186251,1374404997633,"Dipper","https://ap-avatar.wpscdn.com/davatar_a8503ba1806abce46bf441b54a3ca4cd",8,"Research & Report","Document Metadata Extraction","This document outlines the critical importance of meticulously extracting metadata from various document formats, emphasizing a multimodal approach that integrates both textual and visual content. It details a structured JSON output requirement mandating specific fields such as title, title_en, language, abstract, keywords, category, table of contents (toc), and frequently asked questions (faqs). The guidelines rigorously define language detection rules, prioritizing dominant language and ensuring strict consistency across all text-based fields as per the detected language code. Special attention is given to title generation, enforcing rules for handling orphan serials, file names, and document titles, with strict formatting requirements including hyphenation for multi-component titles and a 100-character limit. The abstract generation mandates a direct, SEO-friendly description within a defined character range, while keywords should highlight the most relevant terms. The toc should provide a concise, at most two-level outline, and faqs must present diverse, distinct, and relevant question-answer pairs. Category classification relies on a predefined JSON mapping, mandating a single best-match or a fallback to 'General'. The overall objective is to standardize metadata extraction for enhanced document organization, retrieval, and analysis across diverse multimodal inputs.","| Struktural | Komunikatif | Literasi |\n| --- | --- | --- |\n| Mengetahui (knowing) | Mengerjakan (doing) | Mengerjakan dan merefleksiberdasarkan pengetahuan (doing and reflecting on doing in terms of knowing) |\n| Pengetahuan tentang bahasa (usage) | Penggunaan bahasa (use) | Keduanya (usage/use relation) |\n| Bentuk bahasa (language forms) | Fungsi bahasa (language function) | Keduanya (form-function relation) |\n| Pencapaian pengetahuan tentang bahasa (achievement i.e. display of knowledge) | Kemampuan fungsionaluntuk berkomunikasi (functional ability to communicate) | Kemampuan berkomunikasidengan wajar yang didorong oleh kesadaran metakomunikatif (communicative appropriateness informed by metacommunicative awareness) |","cbCaidy48tRq3cHu","https://ap.wps.com/l/cbCaidy48tRq3cHu","pdf",397216,1,24,"English","en",105,"# Document Metadata Extraction (Multimodal)\n## Instructions\n## Language Rules\n## Field Specifications\n### Title Rules\n### `title_en`\n### `language`\n### abstract\n### `keywords`\n### `toc`\n### `faqs`\n### `category`\n## Output Requirements (STRICT)\n## Inputs","[{\"question\":\"What is the primary objective of this metadata extraction guideline?\",\"answer\":\"The primary objective is to standardize metadata extraction from multimodal documents, ensuring a structured JSON output that integrates textual and visual content for enhanced organization and analysis.\"},{\"question\":\"What are the key requirements for language handling in the metadata?\",\"answer\":\"The guideline mandates strict language detection, prioritizing the dominant language and ensuring all text fields (title, abstract, keywords, toc, faqs) consistently adhere to the detected language code for output.\"},{\"question\":\"How are titles to be formatted and what are the constraints?\",\"answer\":\"Titles must be in the document's detected language, using hyphens to separate multi-component titles and optimizing internal punctuation. They must also be 100 characters or fewer, with a specific format for titles containing feature identifiers.\"}]","Document Metadata Extraction | PDF",1788372159,60,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"document-metadata-extraction-186251","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/document-metadata-extraction-186251/186251/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-09-04","2026-09-02",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What is the primary objective of this metadata extraction guideline?","Question",{"text":76,"@type":77},"The primary objective is to standardize metadata extraction from multimodal documents, ensuring a structured JSON output that integrates textual and visual content for enhanced organization and analysis.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"What are the key requirements for language handling in the metadata?",{"text":81,"@type":77},"The guideline mandates strict language detection, prioritizing the dominant language and ensuring all text fields (title, abstract, keywords, toc, faqs) consistently adhere to the detected language code for output.",{"name":83,"@type":74,"acceptedAnswer":84},"How are titles to be formatted and what are the constraints?",{"text":85,"@type":77},"Titles must be in the document's detected language, using hyphens to separate multi-component titles and optimizing internal punctuation. They must also be 100 characters or fewer, with a specific format for titles containing feature identifiers.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":93},[94,98,102,106,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":29,"slug":109},5,"Comic","comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":107,"slug":138},19,"General","general"]