[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-81730-en":3,"doc-seo-81730-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},81730,4810365810221,"Aurora","https://ap-avatar.wpscdn.com/davatar_155a257f0dc6eb9ab79c44ca47cae57d",8,"Research & Report","Making Failure Safe: A Constrained, Verifiable Agent Framework for Open-Web Data Collection","LLMs and agents can generate web scrapers from natural-language requirements, but free-form generation often breaks due to dependency errors, broken selectors, schema mismatches, and heterogeneous page structures. A constrained, verifiable agent framework shifts LLM output to typed JSON collector configurations, using a six-type taxonomy, template and utility-function constraints, static Airflow DAG execution, rule-based quality checks, and structured feedback correction. Experiments on 138 tasks show strong requirement-to-typing support and deterministic instantiation needs completion of source, field, and execution constraints.","arXiv :2607 .00035v 1 [ cs .AI] 25 Jun 2026  \nMaking Failure Safe: A Constrained, Verifiable Agent Framework for Open-Web Data Collection  \nBo Chen  \nInstitute of Computing Technology, Chinese Academy of Sciences, Beijing, China  \n[chenbo01@ict.ac.cn](chenbo01@ict.ac.cn)  \nAbstract. LLMs and agents can generate web scrapers from naturallanguage requirements, but direct generation remains unreliable because of dependency errors, broken selectors, schema mismatches, and heterogeneous page structures. We propose a constrained, verifiable agent framework that shifts LLM output from free-form code to typed JSON collector configurations, combining a six-type collector taxonomy, template and utility-function constraints, static Airflow DAG execution, rule-based quality checking, and structured feedback correction. Experiments on 138 tasks show that the taxonomy supports description-based requirement typing, while confirming that stable instantiation requires completing source, field, and execution constraints beyond the initial description. On 80 independently source-verified tasks, the framework runs with zero execution-stage LLM tokens and the lowest average wall-clock time, trading moderate one-shot quality for a reusable, deterministic, and verifiable execution path suited to repeated scheduled collection. These results position the framework as a reusable, low-cost, and verifiable execution path for repeated open-web data collection.  \nKeywords: Data collection framework · Agent · Constrained generation · Data quality validation  \n1 Introduction  \nOpen-web data—publicly accessible information from news, government notices, e-commerce, and academic publications—drives growing demand for automated collection at scale. Traditional pipelines depend on manual effort: requirement analysis, site-structure inspection, scraper coding, and validation, suffering from long development cycles and low reusability.  \nRecent LLM and agent advances offer a new path: understanding naturallanguage requirements, generating scraper code, and executing validation tasks. However, direct application proves unreliable. The core tension is between the structural brittleness of heterogeneous web pages and the stochastic inconsistency of LLM code generation. Without collector-type constraints, agents deviate in task decomposition; LLM-generated scrapers may contain dependency errors, broken selectors, and inconsistent output schemas.  \n2 B. Chen  \nTo address this, we argue that the key is not to pursue perfect agent reasoning, but to make the collection pipeline verifiable—so that every stage’s inputsand outputs can be checked, measured, and audited. We achieve this through three pillars: (1) typed modeling of collection tasks with explicit functional boundaries; (2) template and utility-function constraints that shift the agent from free-form code generation to configuration generation; and (3) small-sample validation and quality feedback before scale execution, forming a “generate– execute–check–fix” closed loop.  \nOur contributions are:  \n1. A six-type collector taxonomy (search, list, detail, API, interactive, file) with explicit functional boundaries and composition rules.  \n2. A constrained agent framework organizing requirement understanding, instantiation, validation, quality checking, and feedback correction into a verifiable closed loop, generating JSON configurations under template, slot, and Schema constraints—not free-form code.  \n3. Multi-level experimental validation covering requirement-type classification, collector instantiation, quality–cost comparison against runtime LLM baselines on 80 verified tasks, feedback correction, and Airflow compatibility testing.  \nThe research is organized around three questions: RQ1: How to model collection requirements as typed tasks; RQ2: How to instantiate collectors under constraints for validatable, reusable, schedulable configurations; RQ3: How to intercept low-quality configurations before scale exec","cbCairVjMcjSBPmb","https://ap.wps.com/l/cbCairVjMcjSBPmb","pdf",392337,4,1,15,"English","en",105,"# Introduction\n# Related Work\n## Open-Web Data Collection\n## LLMs, Code Generation, and Constraints","[{\"question\":\"为什么直接用 LLM/Agent 生成网页爬虫在实践中不可靠？\",\"answer\":\"直接生成容易出现依赖错误、选择器失效、输出模式不匹配以及网页结构差异带来的不一致，从而导致整体流程脆弱且难以稳定运行。\"},{\"question\":\"框架如何实现“可验证”的数据采集执行路径？\",\"answer\":\"通过将自由代码生成替换为带约束的类型化 JSON collector 配置，并结合模板/工具函数约束、静态 Airflow DAG 执行、规则化质量检查以及结构化反馈闭环。\"},{\"question\":\"实验结果表明该框架在什么方面更有优势？\",\"answer\":\"在 138 个任务上，六类 taxonomy 支持基于描述的需求类型化；在 80 个源已独立验证的任务上，框架可在执行阶段实现零 LLM token 使用，并获得最低平均 wall-clock 时间，同时保持可复用与可审计的确定性执行路径。\"}]",1784175687,38,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"making-failure-safe-a-constrained-verifiable-agent-framework-for-open-web-data-collection","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/making-failure-safe-a-constrained-verifiable-agent-framework-for-open-web-data-collection/81730/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"为什么直接用 LLM/Agent 生成网页爬虫在实践中不可靠？","Question",{"text":75,"@type":76},"直接生成容易出现依赖错误、选择器失效、输出模式不匹配以及网页结构差异带来的不一致，从而导致整体流程脆弱且难以稳定运行。","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"框架如何实现“可验证”的数据采集执行路径？",{"text":80,"@type":76},"通过将自由代码生成替换为带约束的类型化 JSON collector 配置，并结合模板/工具函数约束、静态 Airflow DAG 执行、规则化质量检查以及结构化反馈闭环。",{"name":82,"@type":73,"acceptedAnswer":83},"实验结果表明该框架在什么方面更有优势？",{"text":84,"@type":76},"在 138 个任务上，六类 taxonomy 支持基于描述的需求类型化；在 80 个源已独立验证的任务上，框架可在执行阶段实现零 LLM token 使用，并获得最低平均 wall-clock 时间，同时保持可复用与可审计的确定性执行路径。","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]