[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84849-en":3,"doc-seo-84849-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},84849,8796095461564,"Liam","https://ap-avatar.wpscdn.com/davatar_155a257f0dc6eb9ab79c44ca47cae57d",8,"Research & Report","FirstResearch: Auditable Question Formation for LLM Scientific Discovery Agents","LLM systems increasingly support scientific discovery through ideation, literature synthesis, experiment planning, and report generation, yet the first proposed research question can remain hard to audit. FirstResearch introduces a first-principles framework for LLM agents, centered on a structured Research Question Certificate that records primitives, assumptions, a mechanism model, tensions, a falsifiable hypothesis, a minimal decisive test, and a failure update rule. On a ten-topic benchmark, it outperforms prompt-level baselines under a DeepSeek-blind-judge protocol, with preserved rankings under Gemini rescoring and strong ablation evidence for the certificate-centric design.","arXiv :2607 .05682v 1 [ cs .AI] 6 Jul 2026  \nFirstResearch: Auditable Question Formation for LLM Scientific Discovery Agents  \nYufeng Wang [louiswang524@gmail. com](louiswang524@gmail. com)  \nJuly 8, 2026  \nAbstract  \nLLM systems for scientific discovery increasingly assist with ideation, literature synthesis, experiment planning, and report generation, but the first research question they propose can remain difficult to audit: it may sound plausible without exposing the mechanism, falsifier, or assumption that a scientist should inspect. We introduce FirstResearch, a first-principles research-question formation framework for scientific LLM agents whose core artifact is a structured Research Question Certificate. The certificate records primitive definitions, assumptions, a mechanism model, a tension or contradiction, a falsifiable hypothesis, a minimal decisive test, and a failure update rule, making the proposed question inspectable before downstream execution. On ten LLM-agent research topics, FirstResearch outperforms controlled prompt-level baselines inspired by AI co-scientist, Agent Laboratory, and AI Scientist-v2 under a primary DeepSeek-blind-judge protocol. AGemini-2.5-Flash independent-judge rescore of the same 40 baseline packages preserves the system-level ranking, with FirstResearch scoring 4.86/5 versus 4.38/5 for the strongest baseline and Pearson agreement of 0.865 on average score. A one-repeat ablation checkpoint further suggests that the certificate-centered core is the strongest component: certificate-only scoring reaches 4.90/5 under DeepSeek and 4.88/5 under Gemini, while removing certificates drops below 1/5 under both judges. These results are preliminary and use LLM judges rather than human domain experts, but they support a narrow scientific-discovery claim: explicit derivation constraints are a promising mechanism for making LLM-generated scientific questions more auditable. Code, prompts, saved outputs, and reproduction scripts are available at [https://github.com/louiswang524/FirstResearch](https://github.com/louiswang524/FirstResearch).  \n1 Introduction  \nLanguage models are becoming scientific collaborators: they brainstorm hypotheses, synthesize literature, plan experiments, analyze data, and draft reports. Recent autonomous research systems show how far this agenda can go. The AI Scientist generates ideas, executes experiments, writes manuscripts, and uses an automated reviewer for evaluation (Lu et al. , 2024); Agent Laboratory structures research assistance into literature review, experimentation, and report writing (Schmidgall et al. , 2025); AI co-scientist uses generation, debate, ranking, and evolution for hypothesis generation (Gottweis et al. , 2025); and AI Scientist-v2 extends autonomous discovery with agentic tree search and workshop-level manuscript generation (Yamada et al. , 2025) . These systems show impressive breadth, but breadth does not guarantee that the first research question is mechanistic, falsifiable, or traceable enough for scientific review.  \nThe central challenge in LLM-assisted scientific ideation is not only finding a topic that sounds plausible, but forming a question that exposesa mechanism. A weak research question can be implemented and written up, yet still fail scientifically because it tests an underspecified gap, a vague improvement claim, or a metric without a clear falsifying observation. This problem matters for human-AI collaboration in science: a scientist needs to know which assumptions the agent made, what observation would reject the proposed hypothesis, and how a negative result should update the research direction. Without an explicit derivation record, it is difficult to tell whether an agent has discovered a real tension or merely imitated the surface form of a research proposal.  \nWe propose FirstResearch, a research-question generation framework that treats derivation as a first-class artifact. Given a topic, FirstResearch first defines prim","cbCain7FL0YEogV9","https://ap.wps.com/l/cbCain7FL0YEogV9","pdf",226819,4,1,16,"English","en",105,"# Introduction\n## Background: LLMs as scientific collaborators\n## Problem: mechanism, falsification, and traceability gaps\n## FirstResearch framework and certificate artifact\n## Evaluation claim and contributions","[{\"question\":\"What problem does FirstResearch address in LLM-assisted scientific discovery?\",\"answer\":\"FirstResearch targets research questions that sound plausible but lack exposed mechanisms, falsifiers, or explicit assumptions, making scientific audit and failure-driven updating difficult.\"},{\"question\":\"What is the Research Question Certificate, and what does it record?\",\"answer\":\"The certificate is a structured artifact that records primitive definitions, assumptions, a mechanism model, tensions or contradictions, a falsifiable hypothesis, a minimal decisive test, expected observations, and a failure update rule.\"},{\"question\":\"How is FirstResearch evaluated against baseline methods?\",\"answer\":\"It is tested on a ten-topic benchmark using LLM-judge protocols, where it outperforms controlled prompt-level baselines under a primary DeepSeek-blind-judge and keeps the system-level ranking when the saved packages are rescored by Gemini-2.5-Flash.\"}]",1784198792,40,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"firstresearch-auditable-question-formation-for-llm-scientific-discovery-agents","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/firstresearch-auditable-question-formation-for-llm-scientific-discovery-agents/84849/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-23","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does FirstResearch address in LLM-assisted scientific discovery?","Question",{"text":75,"@type":76},"FirstResearch targets research questions that sound plausible but lack exposed mechanisms, falsifiers, or explicit assumptions, making scientific audit and failure-driven updating difficult.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What is the Research Question Certificate, and what does it record?",{"text":80,"@type":76},"The certificate is a structured artifact that records primitive definitions, assumptions, a mechanism model, tensions or contradictions, a falsifiable hypothesis, a minimal decisive test, expected observations, and a failure update rule.",{"name":82,"@type":73,"acceptedAnswer":83},"How is FirstResearch evaluated against baseline methods?",{"text":84,"@type":76},"It is tested on a ten-topic benchmark using LLM-judge protocols, where it outperforms controlled prompt-level baselines under a primary DeepSeek-blind-judge and keeps the system-level ranking when the saved packages are rescored by Gemini-2.5-Flash.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,119,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":29,"slug":118},7,"Healthcare","healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]