[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83703-en":3,"doc-seo-83703-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83703,4398048949847,"Eliana","https://ap-avatar.wpscdn.com/avatar/400002536579ef2da7f?_k=1778318612642679267",8,"Research & Report","OpenGlass Sensing-Computing Split Architecture for Local MLLM-Driven Real-Time Visual Assistance","OpenGlass is an open-source, privacy-oriented local-first system for low-latency multimodal visual assistance, designed primarily for blind and low-vision users. It addresses cloud delays and privacy risks by splitting sensing and computing: an ESP32-based glasses unit captures first-person visual context, while a nearby consumer device runs local MLLM inference and produces spoken feedback over local wireless. Evaluations measure response quality, query-to-audio latency, safety-aware abstention, and auditable logs. Under real ESP32 Wi‑Fi capture, OpenGlass achieves 993 ms median user-to-audio latency with resized payloads and 1625 ms with raw 1280×720 payloads, with 97.5% and 93.3% of trials completing under 2 s.","OpenGlass: A Sensing-Computing Split Architecture for Local MLLM-Driven Real-Time Visual Assistance  \nMengzhang Li1 Yuan Yao1,2  \n1 Shanghai Qizhi Institute, 2 College of AI, Tsinghua University  \nCorrespondence: [yaoyuanthu@gmail.com](yaoyuanthu@gmail.com)  \n[https://github.com/OpenSQZ/OpenGlass](https://github.com/OpenSQZ/OpenGlass)  \narXiv :2607 .032 13v 1 [ cs .CV] 3 Jul 2026  \nAbstract  \nWe present OpenGlass, an open-source, privacy-oriented, local-first system for lowlatency multimodal visual assistance, with a primary focus on blind and low-vision users. Cloud MLLM assistants offer strong visual understanding, but often require uploading firstperson visual data and can suffer multi-second network delays; wearable glasses are ideal for sensing, but cannot host large models under tight compute and power budgets. OpenGlass addresses this gap with a sensing-computing split: an ESP32-based glasses-side unit captures visual context, while a nearby consumergrade device performs local MLLM inference and local speech output, reducing cloud reliance and keeping raw egocentric visual data on user-controlled devices by default. We evaluate response quality, query-ready-to-audio latency, safety-aware abstention, and auditable logs. Under real ESP32 Wi-Fi capture, OpenGlass reaches 993 ms median user-to-audio latency with resized payloads and 1625 ms with raw 1280×720 payloads; 97.5% and 93.3% of trials fall below 2 s, respectively. OpenGlass is a user-initiated visual-assistance reference platform for obstacle/hazard awareness, sign/object queries, and image-quality selfchecking, rather than a certified navigation aid. We release source code, hardware instructions, prompts, evaluation data, and logs.  \n1 Introduction  \nMultimodal foundation models have demonstrated strong visual understanding, yet deploying such capability in assistive settings still faces a fundamental tension between capability and interaction quality. Assistive applications for blind and lowvision (BLV) users exemplify this gap: a wearable glasses system must capture the scene, interpret the user request, and deliver spoken guidance quickly enough for the user to act; otherwise, responses become stale in dynamic tasks such as item recog-  \nnition and text reading (Bigham et al., 2010) . Today, many widely used visual assistance workflows and recent GPT class assistants are deployed as cloud services (Huang et al., 2025), which can provide strong model capability but often incur multisecond delays and network jitter under real wireless conditions. Cloud-only deployment also raises privacy concerns for egocentric assistive cameras, since first-person images may contain bystanders, private spaces, screens, documents, and other sensitive visual context. At the same time, wearable devices are ideal for continuous sensing, yet they typically have limited compute and power budgetsand therefore cannot host large multimodal models.  \nFrom a deployment perspective, existing compute platforms can be grouped into several categories. Cloud clusters offer the highest compute capacity but depend on network availability and introduce variable transfer delays. Local desktop workstations can run large models with stable throughput, but they are not naturally portable for everyday assistive use. Portable consumer devices such as laptops, tablets, and phones provide a practical middle ground for running models locally near the user. Wearable devices such as watches and glasses are effective for sensing, but are constrained in memory, compute, and battery. In addition, many users understandably prefer not to continuously upload first-person camera images or spoken interaction content to external servers, since these streams may include private spaces, bystanders, screens, documents, and other sensitive context. A local-first design therefore serves not only as a latency optimization, but also as a privacy-oriented deployment choice: raw visual inputs and generated spoken feedback can","cbCaioCPxQRohJnM","https://ap.wps.com/l/cbCaioCPxQRohJnM","pdf",2264515,5,1,11,"English","en",105,"# Introduction\n## Capability–Interaction Tension in Visual Assistance\n## Deployment Trade-offs: Cloud, Desktop, and Wearables\n## Local-First Privacy and Latency Rationale\n## OpenGlass System Overview\n## System Evaluation and Interaction Patterns","[{\"question\":\"What problem does OpenGlass target for visual assistance users?\",\"answer\":\"OpenGlass targets the gap between strong multimodal visual understanding and the need for fast, interactive spoken guidance in real time, especially for blind and low-vision users. It also addresses privacy risks from uploading first-person visual streams to cloud services.\"},{\"question\":\"How does OpenGlass reduce both latency and privacy exposure?\",\"answer\":\"It uses a sensing-computing split architecture: ESP32 glasses-side sensing captures visual context, while a nearby consumer device performs local MLLM inference and local speech output over Wi‑Fi or a hotspot. This keeps raw egocentric visual data on user-controlled devices by default and avoids cloud multi-second delays.\"},{\"question\":\"What latency performance does OpenGlass report under real ESP32 Wi‑Fi capture?\",\"answer\":\"OpenGlass reports a 993 ms median user-to-audio latency with resized payloads and 1625 ms with raw 1280×720 payloads. It also reports 97.5% and 93.3% of trials finishing below 2 seconds for the respective payload settings.\"}]",1784189839,28,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"openglass-sensing-computing-split-architecture-for-local-mllm-driven-real-time-visual-assistance","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/openglass-sensing-computing-split-architecture-for-local-mllm-driven-real-time-visual-assistance/83703/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-27","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What problem does OpenGlass target for visual assistance users?","Question",{"text":76,"@type":77},"OpenGlass targets the gap between strong multimodal visual understanding and the need for fast, interactive spoken guidance in real time, especially for blind and low-vision users. It also addresses privacy risks from uploading first-person visual streams to cloud services.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How does OpenGlass reduce both latency and privacy exposure?",{"text":81,"@type":77},"It uses a sensing-computing split architecture: ESP32 glasses-side sensing captures visual context, while a nearby consumer device performs local MLLM inference and local speech output over Wi‑Fi or a hotspot. This keeps raw egocentric visual data on user-controlled devices by default and avoids cloud multi-second delays.",{"name":83,"@type":74,"acceptedAnswer":84},"What latency performance does OpenGlass report under real ESP32 Wi‑Fi capture?",{"text":85,"@type":77},"OpenGlass reports a 993 ms median user-to-audio latency with resized payloads and 1625 ms with raw 1280×720 payloads. It also reports 97.5% and 93.3% of trials finishing below 2 seconds for the respective payload settings.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":20,"slug":138},19,"General","general"]