[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85382-en":3,"doc-seo-85382-105":29,"detail-sidebar-cat-0-en-105":89},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},85382,1099513958607,"Jiven","https://ap-avatar.wpscdn.com/avatar/100002390cf8733938c?x-image-process=image/resize,m_fixed,w_180,h_180&k=1778829742770036399",8,"Research & Report","Profiling and Evolution of Intellectual Property","Internet growth accelerates the expansion of scientific and technological resources, but the larger quantity and variety of mixed information raise the cost of acquiring useful knowledge. For technology-driven enterprises and users, science and technology policy documents add an important resource type beyond papers and patents. Efficient extraction and accurate, fast retrieval from noisy policy corpora reduce information barriers and acquisition costs, carrying significant social value. The work discusses challenges in science and technology policy information and presents related technologies and developments.","Profiling and Evolution of Intellectual Property  \nBowen Yu  \nSchool of Computer Science (National Pilot School of Software Engineering), Beijing University of Posts and Telecommunications; Beijing Key Laboratory of Intelligent Telecommunication Software and Multimedia Beijing, China  \nYingxia Shao∗ School of Computer Science (National Pilot School of Software Engineering), Beijing University of Posts and Telecommunications; Beijing Key Laboratory of Intelligent Telecommunication Software and Multimedia Beijing, China  \nAng Li  \nSchool of Computer Science (National Pilot School of Software Engineering), Beijing University of Posts and Telecommunications; Beijing Key Laboratory of Intelligent Telecommunication Software and Multimedia Beijing, China  \narXiv :2204 .09333v 3 [ cs .IR] 11 Jul 2026  \nAbstract  \nIn recent years, with the rapid growth of Internet data, the number and types of scientific and technological resources are also rapidly expanding. However, the increase in the number and category of information data will also increase the cost of information acquisition. For technology-based enterprises or users, in addition to general papers, patents, and other resources, policies related to technology or the development of their industries should also belong to a type of scientific and technological resource. Extracting valuable science and technology policy resources from a huge amount of mixed-content data and providing accurate and fast retrieval will help break down information barriers and reduce informationacquisition costs, which has profound social significance and utility. This article focuses on the difficulties and problems in the field of science and technology policy and introduces related technologies and developments.  \nKeywords  \npolicy data, content extraction, text classification, text matching, language model  \n1 Introduction  \nIn recent years, with the rapid growth of information data on the Internet, the number and types of scientific and technological resources have rapidly expanded [1] . The increase in the number and category of information data sometimes increases the cost of information acquisition. For any individual’s value standard, disorganized data means that a large amount of information is not of interest and that more time is required to identify valid information. Multi-view clustering illustrates how heterogeneous scientific information can be organized through complementary subspaces [2] . For scholar-oriented resources, dynamic interest tracking can further connect multi-view clustering with the evolution of researchers’interests [3]. Taking technology-based enterprises as an example, in addition to papers and patents, policies related to science and technology or supporting the development of their industries are also scientific and technological resources. Such resources are mixed with a large amount of irrelevant policy data, increasing acquisition costs and difficulty. Interpretable machine-learning models are valuable in such decision-support settings because they make automated judgments easier for users to understand [4]. Extracting valuable resources from mixed-content data and providing accurate  \n∗Corresponding author: [shaoyx@bupt.edu.cn](shaoyx@bupt.edu.cn).  \nand fast retrieval can break down information barriers and reduce acquisition costs.  \nPolicy resources usually come from multiple fields and disciplines. The characteristics of multiple data sources lead to inherent difficulties in collecting and obtaining policy resources. Webcontent mining surveys summarize the diversity of extraction strategies required for these sources [5, 6] . Sequential market-state modeling is another example of extracting structured signals from large, noisy online data collections [7] . Sentiment-variation-aware analysis can similarly explain abrupt sentiment spikes in temporally evolving public-event data [8]. Community-detection methods based on deep modularity optimization offer a comple","cbCaihOQkw4fHDZb","https://ap.wps.com/l/cbCaihOQkw4fHDZb","pdf",372811,1,4,"English","en",105,"# Introduction\n# Web Content Extraction of Science and Technology Policy Resources","[{\"question\":\"Why is retrieving science and technology policy resources becoming costly?\",\"answer\":\"The number and variety of Internet information increase the difficulty and expense of identifying relevant items. Policy resources are mixed with substantial irrelevant content, raising acquisition costs and reducing efficiency.\"},{\"question\":\"What makes policy resources challenging to collect and extract?\",\"answer\":\"Policy resources come from multiple fields and disciplines, and different data-source structures require different collection rules. The main body text must be accurately extracted to support downstream algorithm training.\"},{\"question\":\"What are the main approaches for web content extraction mentioned in the document?\",\"answer\":\"The document covers vision-based page segmentation (VIPS) that combines visual representation with DOM structure, and template-based extraction that relies on repeated template parts. It also notes that some methods need complete page rendering, which can be resource-intensive.\"}]",1784203044,10,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":84,"head_meta":86,"extra_data":88,"updated_unix":27},"profiling-and-evolution-of-intellectual-property","",{"@graph":35,"@context":83},[36,52,66],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":21},"https://docshare.wps.com/document/profiling-and-evolution-of-intellectual-property/85382/",{"url":51,"name":13,"@type":53,"author":54,"headline":13,"publisher":56,"fileFormat":59,"inLanguage":23,"description":14,"dateModified":60,"datePublished":60,"encodingFormat":59,"isAccessibleForFree":61,"interactionStatistic":62},"DigitalDocument",{"name":9,"@type":55},"Person",{"url":40,"name":57,"@type":58},"DocShare","Organization","application/pdf","2026-07-16",true,{"@type":63,"interactionType":64,"userInteractionCount":4},"InteractionCounter",{"@type":65},"ViewAction",{"@type":67,"mainEntity":68},"FAQPage",[69,75,79],{"name":70,"@type":71,"acceptedAnswer":72},"Why is retrieving science and technology policy resources becoming costly?","Question",{"text":73,"@type":74},"The number and variety of Internet information increase the difficulty and expense of identifying relevant items. Policy resources are mixed with substantial irrelevant content, raising acquisition costs and reducing efficiency.","Answer",{"name":76,"@type":71,"acceptedAnswer":77},"What makes policy resources challenging to collect and extract?",{"text":78,"@type":74},"Policy resources come from multiple fields and disciplines, and different data-source structures require different collection rules. The main body text must be accurately extracted to support downstream algorithm training.",{"name":80,"@type":71,"acceptedAnswer":81},"What are the main approaches for web content extraction mentioned in the document?",{"text":82,"@type":74},"The document covers vision-based page segmentation (VIPS) that combines visual representation with DOM structure, and template-based extraction that relies on repeated template parts. It also notes that some methods need complete page rendering, which can be resource-intensive.","https://schema.org",{"og:url":51,"og:type":85,"og:title":13,"og:site_name":57,"og:description":14},"article",{"robots":87,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":90},[91,95,99,103,108,113,118,121,126,129,132],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":92,"show_sort_weight":93,"slug":94},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":96,"show_sort_weight":97,"slug":98},"Literature",80,"literature",{"id":21,"doc_module":4,"doc_module_name":45,"category_name":100,"show_sort_weight":101,"slug":102},"Exam",70,"exam",{"id":104,"doc_module":4,"doc_module_name":45,"category_name":105,"show_sort_weight":106,"slug":107},5,"Comic",60,"comic",{"id":109,"doc_module":4,"doc_module_name":45,"category_name":110,"show_sort_weight":111,"slug":112},6,"Technology",50,"technology",{"id":114,"doc_module":4,"doc_module_name":45,"category_name":115,"show_sort_weight":116,"slug":117},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":119,"slug":120},30,"research-report",{"id":122,"doc_module":4,"doc_module_name":45,"category_name":123,"show_sort_weight":124,"slug":125},9,"Religion & Spirituality",20,"religion-spirituality",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":127,"show_sort_weight":124,"slug":128},"World Cup","world-cup",{"id":28,"doc_module":4,"doc_module_name":45,"category_name":130,"show_sort_weight":28,"slug":131},"Lifestyle","lifestyle",{"id":133,"doc_module":4,"doc_module_name":45,"category_name":134,"show_sort_weight":104,"slug":135},19,"General","general"]