[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-160516-en":3,"doc-seo-160516-105":31,"detail-sidebar-cat-0-en-105":93},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":28,"seo_description":14,"update_tm":29,"read_time":30},160516,1099523882367,"Jordan Avery","https://ap-avatar.wpscdn.com/davatar_9964176cb1d06d4a9deccf72a44ae3dc",8,"Research & Report","Evaluating the Logical Reasoning Ability of ChatGPT and GPT-4","Harnessing logical reasoning ability is a comprehensive natural language understanding endeavor. With the release of Generative Pretrained Transformer 4 (GPT-4), highlighted as “advanced” at reasoning tasks, the work evaluates GPT-4 and ChatGPT across multiple logical reasoning benchmarks including LogiQA, ReClor, and AR-LSAT. It tests multi-choice reading comprehension and natural language inference, and builds an out-of-distribution dataset to assess robustness and performance drop-offs.","Evaluating the Logical Reasoning Ability of ChatGPT and GPT-4  \nHanmeng Liu  \nWestlake University [liuhanmeng@westlake.edu.cn](liuhanmeng@westlake.edu.cn)  \nRuoxi Ning  \nZhejiang University [ruoxining@zju.edu.cn](ruoxining@zju.edu.cn)  \nZhiyang Teng  \nNanyang Technological University [zhiyang.teng@ntu.edu.sg](zhiyang.teng@ntu.edu.sg)  \nJian Liu  \nFudan University [jianliu17@fudan.edu.cn](jianliu17@fudan.edu.cn)  \nQiji Zhou and Yue Zhang 􀀃  \nWestlake University zhouqiji, [zhangyue@westlake.edu.cn](zhangyue@westlake.edu.cn)  \narXiv :2304 .03439v3 [ cs .CL] 5 May 2023  \nAbstract  \nHarnessing logical reasoning ability is a comprehensive natural language understanding endeavor. With the release of Generative Pretrained Transformer 4 (GPT-4), highlighted as \"advanced\" at reasoning tasks, we are eager to learn the GPT-4 performance on various logical reasoning tasks. This report analyses multiple logical reasoning datasets, with popular benchmarks like LogiQA and ReClor, and newly-released datasets like AR-LSAT. We test the multi-choice reading comprehension and natural language inference tasks with benchmarks requiring logical reasoning. We further construct a logical reasoning out-ofdistribution dataset to investigate the robustness of ChatGPT and GPT-4 . We also make a performance comparison between ChatGPT and GPT-4 . Experiment results show that ChatGPT performs signiﬁcantly better than the RoBERTa ﬁne-tuning method on most logical reasoning benchmarks. With early access to the GPT-4 API we are able to conduct intense experiments on the GPT-4 model. The results show GPT-4 yields even higher performance on most logical reasoning datasets. Among benchmarks, ChatGPT and GPT-4 do relatively well on well-known datasets like LogiQA and ReClor. However, the performance drops signiﬁcantly when handling newly released and out-of-distribution datasets. Logical reasoning remains challenging for ChatGPT and GPT-4, especially on outof-distribution and natural language inference datasets. We release the prompt-style logical reasoning datasets as a benchmark suite and name it LogiEval.  \n1 Introduction  \nLogical reasoning is essential to human intelligence, and incorporating logical reasoning abilities into natural language understanding (NLU) systems has been an active research interest from  \nY􀀃 ue Zhang is the corresponding author  \nthe beginning of artiﬁcial intelligence (Cresswell, 1973) (Kowalski, 1979) (Iwa´nska, 1993) . Researchers have been exploring various approaches to achieve this goal, including rule-based methods, symbolic systems (MacCartney and Manning, 2007a), ﬁne-tuning large language models (Wanget al., 2018), and combining both neural and symbolic approaches (Li and Srikumar, 2019) .  \nIn the traditional logical and semantic approach, computational linguists developed symbolic systems utilizing First-Order-Logic (FOL) or Natural Logic (MacCartney and Manning, 2007a) to tackle fundamental inference tasks. Rule-based models struggle to unravel problems like the RTE challenge (Dagan et al., 2005) with hand-crafted rules and theorem provers. Formal logic reasoning adopted by early researchers came up with symbolic systems and hand-crafted rules, where knowledge was represented explicitly using formal logic or other symbolic representations. With rules, the systems can process deduction operations. However, these approaches face challenges in handling ambiguity and scalability. They are brittle when dealing with real-world natural language data.  \nThe era of neural network models sees the rise of large-scale NLI datasets as popular benchmarks. For example, the SNLI (Bowman et al., 2015) and the Multi-genre NLI (MNLI) (Williamset al., 2018a) datasets are created through crowdsourcing, featuring an immense data size and broad coverage. They catalyze the development of models with better representation abilities and become the go-to benchmark for natural language understanding research. The giant leap in model performance comes wi","cbCaiorvQ2qiW638","https://ap.wps.com/l/cbCaiorvQ2qiW638","pdf",1222216,6,1,11,"English","en",105,"# Introduction\n## Background: symbolic and neural approaches\n## Benchmarks and datasets for logical reasoning\n## Evaluation of ChatGPT and GPT-4","[{\"question\":\"What tasks and datasets are used to evaluate ChatGPT and GPT-4?\",\"answer\":\"The evaluation covers multi-choice reading comprehension and natural language inference using benchmarks such as LogiQA and ReClor, plus the newly released AR-LSAT dataset.\"},{\"question\":\"Why is an out-of-distribution logical reasoning dataset included?\",\"answer\":\"An out-of-distribution dataset is constructed to investigate robustness and reveal how performance changes beyond known benchmark distributions.\"},{\"question\":\"How do ChatGPT and GPT-4 perform on well-known versus newly released datasets?\",\"answer\":\"They do relatively well on well-known datasets like LogiQA and ReClor, but performance drops significantly on newly released and out-of-distribution datasets.\"}]","Evaluating the Logical Reasoning Ability of ChatGPT and GPT-4 | PDF",1788064722,28,{"code":4,"msg":32,"data":33},"ok",{"site_id":25,"language":24,"slug":34,"title":13,"keywords":35,"description":14,"schema_data":36,"social_meta":88,"head_meta":90,"extra_data":92,"updated_unix":29},"evaluating-the-logical-reasoning-ability-of-chatgpt-and-gpt-4","",{"@graph":37,"@context":87},[38,55,70],{"@type":39,"itemListElement":40},"BreadcrumbList",[41,45,49,52],{"item":42,"name":43,"@type":44,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":46,"name":47,"@type":44,"position":48},"https://docshare.wps.com/document/","Document",2,{"item":50,"name":12,"@type":44,"position":51},"https://docshare.wps.com/document/research-report/",3,{"item":53,"name":13,"@type":44,"position":54},"https://docshare.wps.com/document/evaluating-the-logical-reasoning-ability-of-chatgpt-and-gpt-4/160516/",4,{"url":53,"name":13,"@type":56,"author":57,"headline":13,"publisher":59,"fileFormat":62,"inLanguage":24,"description":14,"dateModified":63,"datePublished":64,"encodingFormat":62,"isAccessibleForFree":65,"interactionStatistic":66},"DigitalDocument",{"name":9,"@type":58},"Person",{"url":42,"name":60,"@type":61},"DocShare","Organization","application/pdf","2026-09-05","2026-08-30",true,{"@type":67,"interactionType":68,"userInteractionCount":20},"InteractionCounter",{"@type":69},"ViewAction",{"@type":71,"mainEntity":72},"FAQPage",[73,79,83],{"name":74,"@type":75,"acceptedAnswer":76},"What tasks and datasets are used to evaluate ChatGPT and GPT-4?","Question",{"text":77,"@type":78},"The evaluation covers multi-choice reading comprehension and natural language inference using benchmarks such as LogiQA and ReClor, plus the newly released AR-LSAT dataset.","Answer",{"name":80,"@type":75,"acceptedAnswer":81},"Why is an out-of-distribution logical reasoning dataset included?",{"text":82,"@type":78},"An out-of-distribution dataset is constructed to investigate robustness and reveal how performance changes beyond known benchmark distributions.",{"name":84,"@type":75,"acceptedAnswer":85},"How do ChatGPT and GPT-4 perform on well-known versus newly released datasets?",{"text":86,"@type":78},"They do relatively well on well-known datasets like LogiQA and ReClor, but performance drops significantly on newly released and out-of-distribution datasets.","https://schema.org",{"og:url":53,"og:type":89,"og:title":13,"og:site_name":60,"og:description":14},"article",{"robots":91,"canonical":53},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":94},[95,99,103,107,112,116,121,124,129,132,136],{"id":21,"doc_module":4,"doc_module_name":47,"category_name":96,"show_sort_weight":97,"slug":98},"Story & Novel",90,"story-novel",{"id":48,"doc_module":4,"doc_module_name":47,"category_name":100,"show_sort_weight":101,"slug":102},"Literature",80,"literature",{"id":54,"doc_module":4,"doc_module_name":47,"category_name":104,"show_sort_weight":105,"slug":106},"Exam",70,"exam",{"id":108,"doc_module":4,"doc_module_name":47,"category_name":109,"show_sort_weight":110,"slug":111},5,"Comic",60,"comic",{"id":20,"doc_module":4,"doc_module_name":47,"category_name":113,"show_sort_weight":114,"slug":115},"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":47,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":47,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":125,"doc_module":4,"doc_module_name":47,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":47,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":133,"doc_module":4,"doc_module_name":47,"category_name":134,"show_sort_weight":133,"slug":135},10,"Lifestyle","lifestyle",{"id":137,"doc_module":4,"doc_module_name":47,"category_name":138,"show_sort_weight":108,"slug":139},19,"General","general"]