[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-188349-en":3,"doc-seo-188349-105":30,"detail-sidebar-cat-1-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":11,"category_id":12,"category_name":13,"doc_title":14,"doc_description":15,"doc_content":16,"file_id":17,"file_url":18,"file_type":19,"file_size":20,"view_count":11,"is_deleted":4,"is_public":11,"is_downloadable":11,"audit_status":11,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":15,"update_tm":28,"read_time":29},188349,5909890329169,"McQueen","https://ap-avatar.wpscdn.com/davatar_9964176cb1d06d4a9deccf72a44ae3dc",1,158,"General","MAI-Thinking-1","This document details the architecture and performance of large language models, specifically focusing on the MAI-Base-1 model. It presents a comparative analysis of various model configurations, including different numbers of layers, hidden units, and feed-forward network (FFN) dimensions. A key aspect highlighted is the use of Sparse Mixture-of-Experts (MoE) layers, with specific configurations like (8/512) experts, and their impact on model size and computational efficiency. The document also includes performance metrics across different evaluation categories such as General, STEM, Math, Code, and Multilingual tasks, comparing models with MoE layers in every layer versus those with shared experts. Detailed tables showcase the EGFLOPs and EGTime for various model sizes, illustrating trade-offs between computational cost and performance. Furthermore, the document outlines the diverse data sources used for training and evaluation, ranging from internal code projects and STEM problem sets to public community discussions and commissioned vendor data, emphasizing the multi-domain expertise of these models. The architectural diagram illustrates a typical transformer block with normalization, attention mechanisms (local/global), and dense FFNs, alongside the Sparse MoE component, suggesting a hybrid approach to model design.","| Model | Active | Total | Layers | Hidden | FFN | Down\u003Cbr>Proj | Expert\u003Cbr>FFN | Top-k/Experts | KV/Q |\n| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |\n| L12 | 365M | 3.9B | 12 | 1024 | 2048 | 512 | 1536 | 8/512 | 8/16 |\n| L18 | 760M | 13B | 18 | 1536 | 3072 | 768 | 2304 | 8/512 | 8/16 |\n| L24 | 1.5B | 30B | 24 | 2048 | 4096 | 1024 | 3072 | 8/512 | 8/24 |\n| L30 | 2.6B | 58B | 30 | 2560 | 5120 | 1280 | 3840 | 8/512 | 8/32 |\n| L36 | 4.0B | 100B | 36 | 3072 | 6144 | 1536 | 4608 | 8/512 | 8/32 |\n| L42 | 6.1B | 159B | 42 | 3584 | 7168 | 1792 | 5376 | 8/512 | 8/40 |\n| L66 | 21.7B | 615B | 66 | 5632 | 11264 | 2816 | 8448 | 8/512 | 8/64 |\n| L78 | 35.6B | 1015B | 78 | 6656 | 13312 | 3328 | 9984 | 8/512 | 8/80 |\n| MAI-Base-1 | 34.7B | 962B | 78 | 6656 | 13312 | 3072 | 10240 | 8/512 | 8/80 |\n\n\n| Eval Category | MoE every layer (8/384) |  | MoE every layer (7+1 shared/384) |  |\n| --- | --- | --- | --- | --- |\n|  | EGFLOPs | EGTime | EGFLOPs | EGTime |\n| General | 0.88 | 0.69 | 0.99 | 0.78 |\n| STEM | 0.93 | 0.73 | 1.04 | 0.82 |\n| Math | 0.94 | 0.74 | 1.06 | 0.84 |\n| Code | 0.94 | 0.74 | 1.03 | 0.82 |\n| Multilingual | 0.94 | 0.75 | 1.05 | 0.82 |\n| Weighted Average | 0.94 | 0.73 | 1.03 | 0.82 |\n\n| Category | Evaluation Data | Source |\n| --- | --- | --- |\n| Coding | Microsoft code and pull requests Human-AI coding sessions | Internal private projects |\n|  |  | Internal testing of previous models |\n| STEM | Worked solutions to graduate-level STEM problems | Commissioned from vendor |\n| Math | Worked solutions to advanced math problems | Commissioned from vendor |\n| General knowledge | Online Community Discussions Human-AI interactions\u003Cbr>Pyramidal trivia\u003Cbr>Hard trivia questions | Public web forums |\n|  |  | Internal testing of previous models |\n|  |  | Publicly available databases deduplicated against training data |\n|  |  | Commissioned from vendor |\n| Multilingual | Multilingual Human-AI interactions | Internal testing of previous models |","cbCaielGJV6IJttK","https://ap.wps.com/l/cbCaielGJV6IJttK","pdf",3317663,109,"English","en",105,"# Model Architectures\n## Performance Evaluation\n## Data Sources","[{\"question\":\"What is the MAI-Base-1 model?\",\"answer\":\"The MAI-Base-1 model is a large language model with 78 layers, 35.6 billion active parameters, and a total of 962 billion parameters. It features 6656 hidden units and a 13312 FFN dimension.\"},{\"question\":\"How does the Sparse Mixture-of-Experts (MoE) affect model performance?\",\"answer\":\"The document suggests that MoE layers, such as the (8/512) configuration (8 active experts out of 512 total), are incorporated into model architectures to potentially enhance efficiency and performance across various tasks, as indicated by the comparative evaluation tables.\"},{\"question\":\"What types of data are used for evaluating these large language models?\",\"answer\":\"The models are evaluated using diverse datasets covering Coding (Microsoft code and pull requests), STEM (graduate-level STEM problems), Math (advanced math problems), General knowledge (online discussions, trivia), and Multilingual tasks (human-AI interactions).\"}]","MAI-Thinking-1 | PDF",1788388848,38,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":14,"keywords":34,"description":15,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"mai-thinking-1","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":11},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/template/","Template",2,{"item":49,"name":13,"@type":43,"position":50},"https://docshare.wps.com/template/general/",3,{"item":52,"name":14,"@type":43,"position":53},"https://docshare.wps.com/template/mai-thinking-1/188349/",4,{"url":52,"name":14,"@type":55,"author":56,"headline":14,"publisher":58,"fileFormat":61,"inLanguage":23,"description":15,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-09-06","2026-09-02",true,{"@type":66,"interactionType":67,"userInteractionCount":11},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What is the MAI-Base-1 model?","Question",{"text":76,"@type":77},"The MAI-Base-1 model is a large language model with 78 layers, 35.6 billion active parameters, and a total of 962 billion parameters. It features 6656 hidden units and a 13312 FFN dimension.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How does the Sparse Mixture-of-Experts (MoE) affect model performance?",{"text":81,"@type":77},"The document suggests that MoE layers, such as the (8/512) configuration (8 active experts out of 512 total), are incorporated into model architectures to potentially enhance efficiency and performance across various tasks, as indicated by the comparative evaluation tables.",{"name":83,"@type":74,"acceptedAnswer":84},"What types of data are used for evaluating these large language models?",{"text":85,"@type":77},"The models are evaluated using diverse datasets covering Coding (Microsoft code and pull requests), STEM (graduate-level STEM problems), Math (advanced math problems), General knowledge (online discussions, trivia), and Multilingual tasks (human-AI interactions).","https://schema.org",{"og:url":52,"og:type":88,"og:title":14,"og:site_name":59,"og:description":15},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":93},[94,99,104,109,114,119,124,129,134],{"id":95,"doc_module":11,"doc_module_name":46,"category_name":96,"show_sort_weight":97,"slug":98},11,"Presentations",90,"presentations",{"id":100,"doc_module":11,"doc_module_name":46,"category_name":101,"show_sort_weight":102,"slug":103},12,"Resumes",80,"resumes",{"id":105,"doc_module":11,"doc_module_name":46,"category_name":106,"show_sort_weight":107,"slug":108},14,"Invoices",70,"invoices",{"id":110,"doc_module":11,"doc_module_name":46,"category_name":111,"show_sort_weight":112,"slug":113},15,"Posters",60,"posters",{"id":115,"doc_module":11,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},16,"Social Media",50,"social-media",{"id":120,"doc_module":11,"doc_module_name":46,"category_name":121,"show_sort_weight":122,"slug":123},17,"Forms",40,"forms",{"id":125,"doc_module":11,"doc_module_name":46,"category_name":126,"show_sort_weight":127,"slug":128},18,"Letters",30,"letters",{"id":130,"doc_module":11,"doc_module_name":46,"category_name":131,"show_sort_weight":132,"slug":133},21,"Paper Templates",5,"papers-templates",{"id":12,"doc_module":11,"doc_module_name":46,"category_name":13,"show_sort_weight":4,"slug":135},"general-158"]