[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83464-en":3,"doc-seo-83464-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83464,1099513958762,"Logic","https://ap-avatar.wpscdn.com/avatar/1000023916a998db790?x-image-process=image/resize,m_fixed,w_180,h_180&k=1784791008015729253",8,"Research & Report","Testing Frontier Large Language Models’ Physics Literacy in Parallel Physical Worlds","Current large-language-model (LLM) physics benchmarks often reward answer accuracy, but they do not separate genuine reasoning from recall of familiar problem patterns, leaving uncertainty about where a model’s reasoning fails. A four-stage diagnostic is introduced to assess induction, formulation, prediction, and review inside unfamiliar physics frameworks. The approach uses locked pre-registrations, fresh stage sessions, dual-LLM judging, and a human-audit pathway, applied to three parallel physics worlds.","arXiv :2607 .00276v1 [ cs .LG] 30 Jun 2026  \nTesting Frontier Large Language Models’ Physics Literacy in Parallel Physical Worlds  \nDong Zhang [dongzhanghz@gmail. com](dongzhanghz@gmail. com)  \nAbstract  \nCurrent large-language-model (LLM) physics benchmarks are usually scored by answer accuracy, which cannot distinguish genuine reasoning from recall of familiar problem patterns and reveals little about where a model’s reasoning breaks down. We introduce an auditable four-stage diagnostic that evaluates whether an LLM can reason inside an unfamiliar physics framework through induction, formulation, prediction, and review. The diagnostic combines locked pre-registrations, fresh sessions between stages, dual-LLM judging, and a human-audit pathway, and we apply it to three parallel physics worlds: a single-equation counterfactual world (F = mv), a historical framework (Aristotelian mechanics), and a four-domain counterfactual world (Decay World) . Across Claude Opus 4.7, GPT-5.5, and Gemini 3.1 Pro, the three worlds yield composite PASS rates are 6/15, 6/15, and 0/15 respectively (content ∧ structural for F = mv and Aristotelian, content axis only for Decay World where the structural axis is out of scope) . The most pointed empirical pattern is a qualitative-versus-quantitative asymmetry: in Decay World, models almost never predict the wrong direction of change, but frequently compute the wrong ratio by slipping back to standard-physics relations. The protocol also surfaces two methodology findings: LLMjudge reliability does not transfer across frameworks, and Stage 4 self-review is weak in every framework, with the model’s own review wrongly reporting no earlier error in at least two-thirds of the trials that actually contained one. We release the full prompts, responses, verdicts, and audit records.  \n1 Introduction  \nWhether large language models (LLMs) can reason is one of the central debates in AI research. In the literature, reasoning is usually defined through tasks that cannot be answered in a single step, such as multistep arithmetic, geometric proofs, and long-horizon planning, where the model must produce intermediate steps that can be checked independently. Wei et al. (2022b) introduced Chain-of-Thought (CoT) prompting in 2022: supplying a few “question, intermediate steps, answer” examples in the prompt raised PaLM- 540B’s accuracy on GSM8K (Cobbe et al., 2021) from 17.9% to 56.9% . Interestingly, Kojima et al. later found that examples were not even necessary: the single phrase “Let’s think step by step” raised zero-shot InstructGPT-175B from 10.4% to 40.7%(Kojima et al., 2022) . Together with Brown et al. (2020)’s GPT-3 scaling observation and Wei et al. (2022a)’s emergence hypothesis, these results supported the optimistic view that scale plus prompting can produce emergent reasoning. On the other hand, this optimistic view has been contested by Schaeffer et al., who argued that much of the observed emergence is an artifact of the evaluation metric (Schaeffer et al., 2023) . Mirzadeh et al.’s GSM-Symbolic showed that frontier models are highly sensitive to symbolic perturbations of the problem text, with accuracy in the GSM-NoOp variant dropping by more than 60 percentage points (Mirzadeh et al., 2024) . This suggests that surface symbol matching contributes far more than true multi-step derivation. Huang and Chang’s survey places this debate on a spectrum, from “CoT is real reasoning” to “CoT is template retrieval from training data” (Huang & Chang, 2023) . Whether LLMs perform genuine reasoning thus remains an open question.  \nA more ambitious question goes beyond reasoning: can AI do scientific research? We sort existing “AI for science” work into three levels of autonomy, all of which remain anchored to a framework supplied in advance. Low autonomy: AI acts as an execution tool within an established theory or experimental pipeline, carrying  \nout optimization, automation, parameter search, or data analysis. Bo","cbCaiqjzuvq0Lrn8","https://ap.wps.com/l/cbCaiqjzuvq0Lrn8","pdf",905328,2,1,37,"English","en",105,"# Introduction\n## Reasoning in LLMs and evaluation debates\n## Autonomy levels in AI-for-science\n## Framework constraints across autonomy levels","[{\"question\":\"Why are standard LLM physics benchmarks insufficient for evaluating reasoning?\",\"answer\":\"They typically score answer accuracy and cannot distinguish genuine reasoning from recall of familiar solution patterns, which obscures where reasoning breaks down.\"},{\"question\":\"How does the proposed diagnostic evaluate physics literacy across unfamiliar frameworks?\",\"answer\":\"It uses an auditable four-stage process—induction, formulation, prediction, and review—combining locked pre-registrations, fresh sessions between stages, dual-LLM judging, and a human-audit pathway.\"},{\"question\":\"What key empirical pattern emerges in the Decay World results?\",\"answer\":\"Models almost never predict the wrong direction of change, but they frequently compute the wrong ratio by reverting to standard-physics relations.\"}]",1784188147,93,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"testing-frontier-large-language-models-physics-literacy-in-parallel-physical-worlds","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/testing-frontier-large-language-models-physics-literacy-in-parallel-physical-worlds/83464/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-25","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why are standard LLM physics benchmarks insufficient for evaluating reasoning?","Question",{"text":75,"@type":76},"They typically score answer accuracy and cannot distinguish genuine reasoning from recall of familiar solution patterns, which obscures where reasoning breaks down.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the proposed diagnostic evaluate physics literacy across unfamiliar frameworks?",{"text":80,"@type":76},"It uses an auditable four-stage process—induction, formulation, prediction, and review—combining locked pre-registrations, fresh sessions between stages, dual-LLM judging, and a human-audit pathway.",{"name":82,"@type":73,"acceptedAnswer":83},"What key empirical pattern emerges in the Decay World results?",{"text":84,"@type":76},"Models almost never predict the wrong direction of change, but they frequently compute the wrong ratio by reverting to standard-physics relations.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]