[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82082-en":3,"doc-seo-82082-105":29,"detail-sidebar-cat-0-en-105":83},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},82082,1374391975076,"Riley","https://ap-avatar.wpscdn.com/avatar/14000253ca4ec9f6853?x-image-process=image/resize,m_fixed,w_180,h_180&k=1783305029341752051",8,"Research & Report","Ortho2CAD 3D CAD Generation from Orthographic Drawings Using Vision Language Models","Ortho2CAD addresses the gap between 2D engineering orthographic drawings and editable 3D CAD by generating CadQuery code from orthographic projection images. The work explains how raster drawings block automated extraction, motivating vision language models conditioned on standardized multi-view inputs. It introduces scalable data synthesis using pythonOCC to produce large open datasets with dashed hidden lines and bounding-box dimensions. Two training regimes are investigated: supervised fine-tuning with code supervision on DeepCAD and reinforcement learning with geometric rewards on Fusion 360 reconstruction. Experiments show improved 3D reconstruction accuracy and code validity across settings.","ORTHO2CAD: 3D CAD GENERATION FROM ORTHOGRAPHIC DRAWINGS USING  \nVISION LANGUAGE MODELS  \nAditya Joglekar Amit Regmi Kenji Shimada Levent Burak Kara∗  \nDepartment of Mechanical Engineering  \nCarnegie Mellon University  \nPittsburgh, PA, 15213, USA  \n∗ Address all correspondences [to lkara@andrew.cmu.edu](to lkara@andrew.cmu.edu)  \ndownstream engineering workflows such as manufacturing planning, simulation and cost estimation. Yet, much of the design intent in practice is communicated through 2D engineering technical drawings, specifically multi-view orthographic that provide a standardized way to encode topology and dimensions. While these drawings may be authored digitally (often in vector form), the format most frequently exchanged in real manufacturing workflows is often an image-based (raster) drawing due to ease of sharing, quality assurance and intellectual-property protection because image drawings are not editable [1] . This prevalence of raster drawings creates a major impediment to 3D reconstruction automation because unlike vector formats that provide direct scripted access to geometric and semantic entities, raster drawings typically require human inspection to extract the information. Together, these realities create a persistent disconnect between how designs are specified and communicated (often as rasterized 2D drawings) and how they must be reconstructed as editable 3D CAD by designers and engineers for subsequent applications.  \nRecent progress in large language models (LLMs) and vision language models (VLMs) has enabled direct generation of CAD programs from text descriptions, perspective images and point clouds exemplified by works such as [2],[3] and [4] . They produce CadQuery 1 code and evaluate the resulting solids using geometric similarity metrics such as intersection-over-union (IoU) . While these approaches demonstrate the CadQuery codegenerating abilities of LLMs and VLMs, none of them tackle the task of converting 2D engineering drawings to CadQuery code. 2D engineering drawings with orthographic projections of a part are the standard for part design communication and a prime input modality designers deal with for 3D model creation. Hence, we focus on orthographic projection images as input to a VLM for  \n1CadQuery is a Python based parametric CAD scripting library  \ngenerating CadQuery code and thus creating editable 3D CAD models.  \nData synthesis. Due to the unavailability of large-scale engineering drawing data, we create a scalable data generation python code using the open source pythonOCC library to synthesize orthographic drawings with three views and critical bounding box dimension annotations from existing CAD repositories. We generate these drawings for the DeepCAD [5] and Fusion 360 Reconstruction [6] datasets and publish the data generation pipeline for further use by the community. Zhang et al. [7] release a 2D orthographic drawings dataset constructed from the Fusion 360 database and using FreeCAD [8], but it only consists of simple 2981 samples and we could not find their exact generation pipeline on any public platform. Zhang et al. [9] use the ABC dataset [10] and FreeCAD [8] for creating close to 70k 2D orthographic drawings but the drawings do not have key dimensions and do not have dashed hidden lines, thus not following the standard conventions. We present one of the largest open-source resources (more than 150k samples) of orthographic drawings, and importantly publish an automatic generation python code for orthographic drawings with dashed hidden lines and bounding box dimensions derived from public CAD datasets, enabling systematic study of drawing-to-editable-CAD generation under controlled, reproducible conditions.  \nWith this data foundation, two complementary learning regimes are investigated: supervised finetuning and reinforcement learning, reflecting the practical reality that code supervision is uneven across datasets.  \nSupervised finetuning on DeepCAD dataset. On ","cbCaie73jCvN4zZk","https://ap.wps.com/l/cbCaie73jCvN4zZk","pdf",4016628,1,14,"English","en",105,"# Motivation and Problem Setting\n# Data Synthesis for Orthographic Drawings\n# Learning Regimes\n## Supervised Fine-Tuning on DeepCAD\n## Reinforcement Learning on Fusion 360\n# Contributions and Results","[{\"question\":\"What training strategies does Ortho2CAD use for different datasets?\",\"answer\":\"It applies supervised fine-tuning on DeepCAD where ground-truth CadQuery programs exist, and uses reinforcement learning on Fusion 360 when aligned ground-truth CadQuery codes are not directly available. RL optimizes using geometric similarity rewards computed from executing generated code.\"}]",1784178124,35,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":78,"head_meta":80,"extra_data":82,"updated_unix":27},"ortho2cad-3d-cad-generation-from-orthographic-drawings-using-vision-language-models","",{"@graph":35,"@context":77},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/ortho2cad-3d-cad-generation-from-orthographic-drawings-using-vision-language-models/82082/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71],{"name":72,"@type":73,"acceptedAnswer":74},"What training strategies does Ortho2CAD use for different datasets?","Question",{"text":75,"@type":76},"It applies supervised fine-tuning on DeepCAD where ground-truth CadQuery programs exist, and uses reinforcement learning on Fusion 360 when aligned ground-truth CadQuery codes are not directly available. RL optimizes using geometric similarity rewards computed from executing generated code.","Answer","https://schema.org",{"og:url":51,"og:type":79,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":81,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":84},[85,89,93,97,102,107,112,115,120,123,127],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":86,"show_sort_weight":87,"slug":88},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":90,"show_sort_weight":91,"slug":92},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Exam",70,"exam",{"id":98,"doc_module":4,"doc_module_name":45,"category_name":99,"show_sort_weight":100,"slug":101},5,"Comic",60,"comic",{"id":103,"doc_module":4,"doc_module_name":45,"category_name":104,"show_sort_weight":105,"slug":106},6,"Technology",50,"technology",{"id":108,"doc_module":4,"doc_module_name":45,"category_name":109,"show_sort_weight":110,"slug":111},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":113,"slug":114},30,"research-report",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},9,"Religion & Spirituality",20,"religion-spirituality",{"id":118,"doc_module":4,"doc_module_name":45,"category_name":121,"show_sort_weight":118,"slug":122},"World Cup","world-cup",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":124,"slug":126},10,"Lifestyle","lifestyle",{"id":128,"doc_module":4,"doc_module_name":45,"category_name":129,"show_sort_weight":98,"slug":130},19,"General","general"]