[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-151698-en":3,"doc-seo-151698-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},151698,4398048950312,"Violet","https://ap-avatar.wpscdn.com/avatar/400002538284de19e3c?_k=1778320343897328908",8,"Research & Report","SPHINX - A Synthetic Environment for Visual Perception and Reasoning","SPHINX is a synthetic environment targeting core cognitive primitives in visual perception and reasoning. It procedurally generates puzzles from motifs, tiles, charts, icons, and geometric primitives, with verifiable ground-truth solutions for precise evaluation and scalable dataset creation. The benchmark spans 25 task types covering symmetry detection, geometric transformations, spatial reasoning, chart interpretation, and sequence prediction. Results on recent LVLMs show limited accuracy versus human performance, while RLVR substantially improves accuracy and transfers gains to external visual reasoning benchmarks.","SPHINX: A Synthetic Environment for Visual Perception and Reasoning  \nMd Tanvirul Alam Rochester Institute of Technology Rochester, NY, USA  \n[ma8235@rit.edu](ma8235@rit.edu)  \nSaksham Aggarwal Rochester Institute of Technology Rochester, NY, USA  \n[sxavse@rit.edu](sxavse@rit.edu)  \narXiv :2511 .208 14v2 [ cs .CV] 5 Apr 2026  \nJustin Yang Chae University of Washington Seattle, WA, USA  \n[jchae3@uw.edu](jchae3@uw.edu)  \nAbstract  \nWe present SPHINX, a synthetic environment for visual perception and reasoning that targets core cognitive primitives. SPHINX procedurally generates puzzles using motifs, tiles, charts, icons, and geometric primitives, each paired with verifiable ground-truth solutions, enabling both precise evaluation and large-scale dataset construction. The benchmark covers 25 task types spanning symmetry detection, geometric transformations, spatial reasoning, chart interpretation, and sequence prediction. Evaluating recent large vision–language models (LVLMs) shows that even state-of-the-art GPT-5 attains only 51. 1% accuracy, well below human performance. Finally, we demonstrate that reinforcement learning with verifiable rewards (RLVR) substantially improves model accuracy on these tasks and yields gains on external visual reasoning benchmarks, highlighting its promise for advancing multimodal reasoning. Project page, code, and dataset available at [https:](https:)//[maveryn.github.io/sphinx/](maveryn.github.io/sphinx/).  \n1. Introduction  \nLarge language models (LLMs) have recently demonstrated striking advances in reasoning, achieving gold medal level performance at the International Mathematical Olympiad [8] and strong results across mathematics, logical reasoning, and coding [14, 20, 23, 60, 65] . Because reasoning is a core component of human intelligence, it has become a central benchmark for progress toward Artificial General Intelligence (AGI) [19] . Techniques such as Chain-of-Thought prompting [56], testtime compute scaling [23], and post-training strategies such as rule-based reinforcement learning in DeepSeek-  \nNidhi Rastogi Rochester Institute of Technology  \nRochester, NY, USA  \n[nxrvse@rit.edu](nxrvse@rit.edu)  \n Human (75 . 8%)  GPT-5 Mini (47 . 1%)  \n GPT-5 (51 . 1%)  Qwen3-VL-235B (41 . 3%)  \nGeometric Reasoning  \nFigure 1 . Radar plot shows accuracies (%) achieved by LVLMs and by humans on the broad categories of SPHINX.  \nR1 have further improved model performance, helping mitigate reward hacking [20] and allowing more robust generalization across domains [2, 21, 62] .  \nIn contrast to the rapid progress of LLMs, large visionlanguage models (LVLMs) remain far less capable of visual reasoning [11, 33, 45, 70] . Unlike text-based systems that can leverage structured prompts and posttraining strategies, LVLMs must jointly parse visual inputs and integrate them with language, a substantially more complex challenge [6, 18, 20, 54, 62] . Current models often fail to construct coherent reasoning chainsand stumble on tasks trivial to humans [66] . Although reinforcement learning has been applied to strengthen LVLMs [28, 40], progress is constrained by benchmarks that emphasize perception over reasoning, such as referring to expression comprehension or math-with-diagram  \ndatasets, where models frequently reduce visual inputs to text and rely on language reasoning [63, 72] .  \nMore recently, several works have begun to investigate abstract visual reasoning (AVR) in LVLMs [6, 12, 24, 26, 33, 63], yet these efforts still fall short of systematically evaluating core perceptual primitives such as symmetry detection, mental rotation, and structured pattern matching. Cognitive science has long established that these abilities underpin fluid intelligence and matrix reasoning [7, 16, 41, 47], implying that practical machinelearning evaluation must directly target such primitives through controlled tasks that disentangle perception from abstraction. To address this gap, we introduce SPHINX, a synthetic envir","cbCaidvRBuzd23Js","https://ap.wps.com/l/cbCaidvRBuzd23Js","pdf",8505221,1,44,"English","en",105,"# Introduction\n## SPHINX Design\n### Design Principles","[{\"question\":\"What is SPHINX and what problem does it target?\",\"answer\":\"SPHINX is a synthetic environment for visual perception and reasoning, designed to evaluate core cognitive primitives such as symmetry detection and spatial transformations.\"},{\"question\":\"How does SPHINX generate tasks and supervision?\",\"answer\":\"It procedurally generates puzzles using motifs, tiles, charts, icons, and geometric primitives, and pairs each instance with deterministic verifiable ground-truth solutions.\"},{\"question\":\"How do large vision-language models perform on SPHINX and what improves results?\",\"answer\":\"Evaluations show state-of-the-art LVLMs achieve far below human accuracy; reinforcement learning with verifiable rewards (RLVR) substantially improves accuracy and yields gains on external visual reasoning benchmarks.\"}]","SPHINX - A Synthetic Environment for Visual Perception and Reasoning | PDF",1787847545,111,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"sphinx-a-synthetic-environment-for-visual-perception-and-reasoning","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/sphinx-a-synthetic-environment-for-visual-perception-and-reasoning/151698/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-09-04","2026-08-27",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What is SPHINX and what problem does it target?","Question",{"text":76,"@type":77},"SPHINX is a synthetic environment for visual perception and reasoning, designed to evaluate core cognitive primitives such as symmetry detection and spatial transformations.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How does SPHINX generate tasks and supervision?",{"text":81,"@type":77},"It procedurally generates puzzles using motifs, tiles, charts, icons, and geometric primitives, and pairs each instance with deterministic verifiable ground-truth solutions.",{"name":83,"@type":74,"acceptedAnswer":84},"How do large vision-language models perform on SPHINX and what improves results?",{"text":85,"@type":77},"Evaluations show state-of-the-art LVLMs achieve far below human accuracy; reinforcement learning with verifiable rewards (RLVR) substantially improves accuracy and yields gains on external visual reasoning benchmarks.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,121,124,129,132,136],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":46,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":46,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":46,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":133,"doc_module":4,"doc_module_name":46,"category_name":134,"show_sort_weight":133,"slug":135},10,"Lifestyle","lifestyle",{"id":137,"doc_module":4,"doc_module_name":46,"category_name":138,"show_sort_weight":107,"slug":139},19,"General","general"]