[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86570-en":3,"doc-seo-86570-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86570,1099514068365,"Aurelia","https://ap-avatar.wpscdn.com/avatar/10000253d8d9f28188e?_k=1776742907772140068",8,"Research & Report","GEIS A Generation–Evaluation–Improvement Loop of Agent Skills for Long-Form Article Generation","Long-form article generation is challenging for large language models due to long context, long instructions, and long outputs. Existing multi-agent pipelines often entangle capabilities in prompts and fixed procedures, limiting inspection, reuse, and iterative improvement. GEIS (Generation–Evaluation–Improvement loop of agent Skills) uses named declarative skills for Wikipedia-style long-form writing, evidence collection, diagram rendering, PDF-aware pairwise evaluation, and rule-level improvement. Experiments on 20 featured-article topics show 8.0-point gains over a default writer and superiority to STORM in structural and content quality.","arXiv :2607 . 1 1503v 1 [ cs .CL] 13 Jul 2026  \nGEIS: A Generation–Evaluation–Improvement Loop of Agent Skills for Long-Form Article  \nGeneration  \nJiale Zhang 1 , Juntao Hu2 , and Zhijian Ou 1 ,2⋆  \n1 Speech Processing and Machine Intelligence (SPMI) Lab, Tsinghua University, China  \n2 TasiTech Co. , Ltd. , China  \nAbstract. Long-form article generation remains difficult for large language models because it combines long context, long instructions, and long outputs. Existing multi-agent pipelines such as STORM improve information coverage by simulating role-specialized agents, but their capabilities are often entangled in prompts and fixed procedures, making them hard to inspect, reuse, or iteratively improve. This paper presents GEIS (Generation–Evaluation–Improvement loop of agent Skills), a loop of named and declarative skills for Wikipedia-style long-form article generation. Implemented and evaluated in Tasi Harness, GEIS composes skills for article writing, browser-based evidence and image collection, diagram rendering, PDF-aware pairwise evaluation, and rule-level skill improvement. Its core writing skill follows Request, Plan, Draft, Audit, Refine, and Deliver; the pairwise evaluation skill produces structured quality reports; and the improvement skill maps recurrent findings into permanent patches to the writing skill in our 20-topic experiment. We evaluate GEIS on 20 Wikipedia Featured Article topics. Under the same generation backend, GEIS improves over the Tasi Harness default writer by 8.0 points on a 100-point PDF quality rubric and outperforms STORM on the two comparable writing dimensions, structural quality and content quality. In the 20-topic improvement experiment, the patched writingskill raises the average score from 82.90 to 86.95, with 17 out of 20 topics improved and the gain mainly coming from content quality. These results show that long-form generation can be reframed from a fixed workflow into an inspectable, modular, and evaluation-guided improvement loop.  \nKeywords: Large language models · Long-form generation · agent skills  \n· Evaluation-guided improvement · LLM-as-a-judge  \n1 Introduction  \nLarge language models (LLMs) have become strong generators of short answers, summaries, code snippets, and conversational responses [1] . However, many real writing tasks are not short-form generation problems. Wikipedia-style article  \n⋆ Corresponding author ([ozj@tsinghua.edu.cn](ozj@tsinghua.edu.cn)). This work is supported by the National Science and Technology Major Project (2023ZD0121401) .  \n2 J. Zhang et al.  \nwriting, technical report preparation, proposal drafting, and analytical document creation require a system to collect information, organize it into a coherent structure, maintain factual consistency across sections, and produce a polished deliverable. We refer to such tasks as L3 long-form generation tasks because they jointly involve long context, long instructions, and long outputs, a combination emphasized in recent long-output and long-context generation studies [8,9] .  \nThe L3 setting exposes several weaknesses of single-pass LLM generation. First, long contexts can cause information loss and attention imbalance, includingthe well-known “lost in the middle” phenomenon [2] . Second, long articles require several objectives to be optimized simultaneously: coverage, structure, coherence, factual reliability, style, and delivery readiness. A single autoregressive pass has no explicit mid-course quality gate. Third, long outputs make hallucination, repetition, and unsupported factual claims more likely, especially when generation must integrate many retrieved facts [3, 17 , 18] . When a paragraph-level error appears early, it can be amplified by later sections or become hard to detect after the document has been assembled.  \nA natural solution is to decompose writing into multiple stages or multiple agents. STORM [4], for example, simulates Wikipedia writers and experts to explore a topic fr","cbCaiaUv4jpGVuqY","https://ap.wps.com/l/cbCaiaUv4jpGVuqY","pdf",441899,3,1,15,"English","en",105,"# Introduction\n## Challenges in L3 long-form generation\n## Decomposing writing into stages or agents\n## Agent-skill paradigm and controllability\n## GEIS organization: generation, evaluation, improvement","[{\"question\":\"What problem does GEIS target in long-form article generation?\",\"answer\":\"GEIS targets difficulties in long-form generation where models must handle long context, long instructions, and long outputs while maintaining coverage, structure, coherence, factual reliability, and deliverable quality.\"},{\"question\":\"How does GEIS differ from earlier multi-agent pipelines like STORM?\",\"answer\":\"GEIS emphasizes explicit, modular agent skills where each skill focuses on a specific capability (writing, retrieval, diagram generation, evaluation) and evaluation feedback is converted into reusable, improved writing rules.\"},{\"question\":\"What evidence and evaluation methods does GEIS use?\",\"answer\":\"GEIS delegates evidence and image collection to browser-based skills, generates diagrams via a dedicated diagram skill, and performs symmetric PDF-aware pairwise evaluation using a pdf-comparison-evaluation skill.\"}]",1784212698,38,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"geis-a-generationevaluationimprovement-loop-of-agent-skills-for-long-form-article-generation","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/geis-a-generationevaluationimprovement-loop-of-agent-skills-for-long-form-article-generation/86570/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-25","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does GEIS target in long-form article generation?","Question",{"text":75,"@type":76},"GEIS targets difficulties in long-form generation where models must handle long context, long instructions, and long outputs while maintaining coverage, structure, coherence, factual reliability, and deliverable quality.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does GEIS differ from earlier multi-agent pipelines like STORM?",{"text":80,"@type":76},"GEIS emphasizes explicit, modular agent skills where each skill focuses on a specific capability (writing, retrieval, diagram generation, evaluation) and evaluation feedback is converted into reusable, improved writing rules.",{"name":82,"@type":73,"acceptedAnswer":83},"What evidence and evaluation methods does GEIS use?",{"text":84,"@type":76},"GEIS delegates evidence and image collection to browser-based skills, generates diagrams via a dedicated diagram skill, and performs symmetric PDF-aware pairwise evaluation using a pdf-comparison-evaluation skill.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]