[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83946-en":3,"doc-seo-83946-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83946,687197207639,"Asher","https://ap-avatar.wpscdn.com/davatar_a8503ba1806abce46bf441b54a3ca4cd",8,"Research & Report","Search Beyond What Can Be Taught: Evolving the Knowledge Boundary in Agentic Visual Generation","Visual generators excel at rendering, yet they fabricate unknown details because user requests are open-ended, evolving, and deeply long-tailed, creating a structural world-knowledge bottleneck. SEARCHGEN-20K and SEARCHGEN-BENCH provide 20,839 prompts across twelve failure categories and twenty-two domains, supported by SEARCHGEN-CORPUS-1M for offline research. On SEARCHGEN-BENCH, frontier open generators score only 21–28/100, revealing a major evaluation gap. Naive search harms prompts, and teach-then-search co-training discovers an evolving generator-specific knowledge boundary enabling recursive, world-knowledge-grounded generation.","arXiv :2607 .05382v 3 [ cs .CV] 9 Jul 2026  \nSearch Beyond What Can Be Taught: Evolving the Knowledge Boundary in Agentic Visual Generation  \nHaozhe Wang 1 Weijia Feng3 Jinpeng Yu3 Che Liu4 Ping Nie2 Fangzhen Lin 1 Jiaming Liu3 ,B Ruihua Huang3 Jimmy Lin2 Wenhu Chen2 Cong Wei2 ,B  \n1Hong Kong University of Science and Technology 2University of Waterloo  \n3 Qwen Applications 4Imperial College London  \n􀂌 Project Page: [https://haozheh3.github](https://haozheh3.github) .io/SearchGen  \nAbstract  \nVisual generators excel at rendering, but they confidently fabricate what they do not know. User requests are unbounded, evolving, and deeply long-tailed: new characters, trending entities, post-cutoff events, etc. This world-knowledge bottleneck is structural: generators are trained on fixed corpora, but the visual world is open-ended. We construct SEARCHGEN-20K and SEARCHGEN-BENCH, 20,839 prompts spanning twelve failure categories and twenty-two domains, paired with a pre-executed multimodal SEARCHGEN-CORPUS-1M to facilitate offline, reproducible research. On SEARCHGEN-BENCH, frontier open generators score only 21–28 out of 100, a 40-point collapse invisible to existing benchmarks. The natural remedy to this knowledge bottleneck is to employ search tools, enabling agentic visual generation. But we found that naive search fails: it retrieves indiscriminately, injecting noise into prompts the generator already handles. Wetrace the root cause to a generator-specific, evolving knowledge boundary—the divide between what a generator can internalize through training and what must remain in external context—and show that this boundary, though hard to specify in advance, is discoverable through a teach-then-search co-training framework. Even a minimal recipe of this co-training produces monotonic improvement, laying the foundation for recursive self-improvement of visual generation that meets worldknowledge grounded requests. We release code, model and the full dataset, co-training corpus, and search corpus as a replayable harness for tool-augmented, world-knowledge-grounded visual generation.  \nFigure 1: Representative search-augmented generations from SEARCHGEN-20K, spanning all twelve failure categories identified from 20,840 production prompts. SEARCHGEN-20K captures the productionscale diverse user requests that demand the unbounded, evolving, and deeply long-tailed world knowledge.  \nB: Corresponding Authors.  \nPreprint.  \nFigure 2: Two paradigms for visual generation. (Left) Prompt rewriting relies on an LLM to expand the user query into a longer textual prompt, which is then passed directly to the generator. (Right) In contrast, our approach equips an agent with a search tool to retrieve relevant knowledge and visual references from a web corpus; the agent then organizes the retrieved evidence into multimodal context and provides it to the generator.  \n1 Introduction  \nAsk a frontier image generator for the mascot of the 2025 Osaka Expo: you get a polished, confident fabrication. Ask for a historically accurate Spartan phalanx: you get anachronistic armor rendered in exquisite detail. Modern generators produce complex scenes with precise lighting and coherent structure, saturating on standard benchmarks [29, 6, 14], yet they still fail on a substantial class of real-world user requests (Figure 1) . These failures reflect a world-knowledge bottleneck rather than a visual-synthesis bottleneck. Visual generators are trained on fixed corpora with inherent knowledge cutoffs, whereas user requests are unbounded, evolving, and deeply long-tailed: new characters, regional cultural symbols, niche typography, historical artifacts, and recent events. The world knowledge bottleneck motivates external knowledge access as a natural complement to visual generation, analogous to illustrators consulting references before depicting unfamiliar concepts. To systematically study this problem, we analyze over 20,000 user prompts from production-level AIGC pla","cbCaicq4NYF0nEpn","https://ap.wps.com/l/cbCaicq4NYF0nEpn","pdf",8199809,4,1,32,"English","en",105,"# Abstract\n# Introduction\n## Problem: world-knowledge bottleneck\n## Dataset and benchmarks: SEARCHGEN-20K and SEARCHGEN-BENCH\n## Agentic search and why naive search fails\n## Knowledge boundary and teach-then-search co-training","[{\"question\":\"What core limitation causes modern visual generators to fail on real-world user requests?\",\"answer\":\"They suffer a structural world-knowledge bottleneck: training relies on fixed corpora with knowledge cutoffs, while user requests are open-ended and deeply long-tailed.\"},{\"question\":\"What are SEARCHGEN-20K and SEARCHGEN-BENCH used for?\",\"answer\":\"They provide world-knowledge-grounded prompts spanning multiple domains and failure modes, enabling automated assessment and revealing evaluation gaps in existing benchmarks.\"},{\"question\":\"Why does naive agentic search degrade visual generation quality?\",\"answer\":\"Blindly triggering search injects noise into prompts that the generator can already handle, so the approach fails to respect the generator-specific evolving knowledge boundary.\"}]",1784191615,81,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"search-beyond-what-can-be-taught-evolving-the-knowledge-boundary-in-agentic-visual-generation","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/search-beyond-what-can-be-taught-evolving-the-knowledge-boundary-in-agentic-visual-generation/83946/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-26","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What core limitation causes modern visual generators to fail on real-world user requests?","Question",{"text":75,"@type":76},"They suffer a structural world-knowledge bottleneck: training relies on fixed corpora with knowledge cutoffs, while user requests are open-ended and deeply long-tailed.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What are SEARCHGEN-20K and SEARCHGEN-BENCH used for?",{"text":80,"@type":76},"They provide world-knowledge-grounded prompts spanning multiple domains and failure modes, enabling automated assessment and revealing evaluation gaps in existing benchmarks.",{"name":82,"@type":73,"acceptedAnswer":83},"Why does naive agentic search degrade visual generation quality?",{"text":84,"@type":76},"Blindly triggering search injects noise into prompts that the generator can already handle, so the approach fails to respect the generator-specific evolving knowledge boundary.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]