[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-seo-190904-105":3,"detail-sidebar-cat-1-en-105":81,"doc-detail-190904-en":127},{"code":4,"msg":5,"data":6},0,"ok",{"site_id":7,"language":8,"slug":9,"title":10,"keywords":11,"description":12,"schema_data":13,"social_meta":74,"head_meta":76,"extra_data":78,"updated_unix":80},105,"en","gemini-model-report-performance-overview","Gemini model report - performance overview","","Gemini model report compares Gemini Ultra, Gemini Pro, and Gemini Nano against a range of established benchmarks across reasoning, multimodal understanding, mathematics, code generation, and reading comprehension. The content details model characteristics, such as TPU scalability for Ultra, latency-optimized cost performance for Pro, and on-device deployment with two Nano variants and 4-bit quantization. Benchmark tables cover MMLU, GSM8K, MATH, BIG-Bench-Hard, HumanEval, Natural2Code, DROP, HellaSwag, and WMT23.",{"@graph":14,"@context":73},[15,34,56],{"@type":16,"itemListElement":17},"BreadcrumbList",[18,23,27,31],{"item":19,"name":20,"@type":21,"position":22},"https://docshare.wps.com","Home","ListItem",1,{"item":24,"name":25,"@type":21,"position":26},"https://docshare.wps.com/template/","Template",2,{"item":28,"name":29,"@type":21,"position":30},"https://docshare.wps.com/template/presentations/","Presentations",3,{"item":32,"name":10,"@type":21,"position":33},"https://docshare.wps.com/template/gemini-model-report-performance-overview/190904/",4,{"url":32,"name":10,"@type":35,"image":36,"author":41,"headline":10,"publisher":44,"fileFormat":47,"inLanguage":8,"description":12,"dateModified":48,"datePublished":49,"encodingFormat":47,"isAccessibleForFree":50,"interactionStatistic":51},"DigitalDocument",{"url":37,"@type":38,"width":39,"height":40},"https://docshare.wps.com/thumbnails/gemini-model-report-performance-overview/190904.png","ImageObject",442,249,{"name":42,"@type":43},"Asher","Person",{"url":19,"name":45,"@type":46},"DocShare","Organization","application/pdf","2026-10-04","2026-09-03",true,{"@type":52,"interactionType":53,"userInteractionCount":55},"InteractionCounter",{"@type":54},"ViewAction",8,{"@type":57,"mainEntity":58},"FAQPage",[59,65,69],{"name":60,"@type":61,"acceptedAnswer":62},"What are the key differences between Gemini Ultra, Gemini Pro, and Gemini Nano?","Question",{"text":63,"@type":64},"Gemini Ultra targets state-of-the-art performance and is efficiently serveable at scale on TPU accelerators. Gemini Pro is optimized for cost and latency with strong reasoning and broad multimodal capabilities. Gemini Nano is designed to run on-device, with Nano-1 and Nano-2 variants using distillation from larger Gemini models and 4-bit quantization for deployment.","Answer",{"name":66,"@type":61,"acceptedAnswer":67},"Which benchmarks are included in the document for evaluating model performance?",{"text":68,"@type":64},"The document includes results for MMLU, GSM8K, MATH, BIG-Bench-Hard, HumanEval, Natural2Code, DROP, HellaSwag, and WMT23, along with additional tables for normalized accuracy and translation scoring across resource levels.",{"name":70,"@type":61,"acceptedAnswer":71},"How does the Nano deployment approach work according to the report?",{"text":72,"@type":64},"The report states that two Nano versions are trained (Nano-1 with 1.8B parameters and Nano-2 with 3.25B parameters) for different memory device constraints. It also notes that Nano is trained via distilling from larger Gemini models and is 4-bit quantized for deployment.","https://schema.org",{"og:url":32,"og:type":75,"og:title":10,"og:site_name":45,"og:description":12},"article",{"robots":77,"canonical":32},"index,follow",{"doc_id":79,"site_id":7},190904,1788405486,{"code":4,"msg":82,"data":83},"success",[84,88,93,98,103,108,113,118,123],{"id":85,"doc_module":22,"doc_module_name":25,"category_name":29,"show_sort_weight":86,"slug":87},11,90,"presentations",{"id":89,"doc_module":22,"doc_module_name":25,"category_name":90,"show_sort_weight":91,"slug":92},12,"Resumes",80,"resumes",{"id":94,"doc_module":22,"doc_module_name":25,"category_name":95,"show_sort_weight":96,"slug":97},14,"Invoices",70,"invoices",{"id":99,"doc_module":22,"doc_module_name":25,"category_name":100,"show_sort_weight":101,"slug":102},15,"Posters",60,"posters",{"id":104,"doc_module":22,"doc_module_name":25,"category_name":105,"show_sort_weight":106,"slug":107},16,"Social Media",50,"social-media",{"id":109,"doc_module":22,"doc_module_name":25,"category_name":110,"show_sort_weight":111,"slug":112},17,"Forms",40,"forms",{"id":114,"doc_module":22,"doc_module_name":25,"category_name":115,"show_sort_weight":116,"slug":117},18,"Letters",30,"letters",{"id":119,"doc_module":22,"doc_module_name":25,"category_name":120,"show_sort_weight":121,"slug":122},21,"Paper Templates",5,"papers-templates",{"id":124,"doc_module":22,"doc_module_name":25,"category_name":125,"show_sort_weight":4,"slug":126},158,"General","general-158",{"code":4,"msg":82,"data":128},{"doc_id":79,"user_id":129,"nickname":42,"user_avatar":130,"doc_module":22,"category_id":85,"category_name":29,"doc_title":10,"doc_description":12,"doc_content":131,"file_id":132,"file_url":133,"file_type":134,"file_size":135,"view_count":55,"is_deleted":4,"is_public":22,"is_downloadable":22,"audit_status":22,"page_count":86,"language":136,"language_code":8,"site_id":7,"html_lang":8,"table_of_contents":137,"faqs":138,"seo_title":139,"seo_description":12,"update_tm":80,"read_time":140},687197207639,"https://ap-avatar.wpscdn.com/davatar_a8503ba1806abce46bf441b54a3ca4cd","| Model size | Model description |\n| --- | --- |\n| Ultra | Our most capable model that delivers state-of-the-art performance across a wide range of highly complex tasks, including reasoning and multimodal tasks. It is efficiently serveable at scale on TPU accelerators due to the Gemini architecture. |\n| Pro | A performance-optimized model in terms of cost as well as latency that delivers significant performance across a wide range of tasks. This model exhibits strong reasoning performance and broad multimodal capabilities. |\n| Nano | Our most efficient model, designed to run on-device. We trained two versions of Nano, with 1.8B (Nano-1) and 3.25B (Nano-2) parameters, targeting low and high memory devices respectively. It is trained by distilling from larger Gemini models. It is 4-bit quantized for deployment and provides best-in-class performance. |\n\n|  | Gemini Ultra | Gemini Pro | GPT-4 | GPT-3.5 | PaLM 2-L | Claude 2 | Inflection-2 | Grok 1 | LLAMA-2 |\n| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |\n| MMLU\u003Cbr>Multiple-choice questions in 57 subjects (professional & academic)\u003Cbr>(Hendrycks et al., 2021a) | 90.04%\u003Cbr>CoT@32∗\u003Cbr>83.7%\u003Cbr>5-shot | 79.13%\u003Cbr>CoT@8∗\u003Cbr>71.8%\u003Cbr>5-shot | 87.29%\u003Cbr>CoT@32 (via API ∗∗ )\u003Cbr>86.4%\u003Cbr>5-shot (reported) | 70%\u003Cbr>5-shot | 78.4%\u003Cbr>5-shot | 78.5%\u003Cbr>5-shot CoT | 79.6%\u003Cbr>5-shot | 73.0%\u003Cbr>5-shot | 68.0%∗∗∗ |\n| GSM8K\u003Cbr>Grade-school math\u003Cbr>(Cobbe et al., 2021) | 94.4%\u003Cbr>Maj1@32 | 86.5%\u003Cbr>Maj1@32 | 92.0%\u003Cbr>SFT &\u003Cbr>5-shot CoT | 57.1%\u003Cbr>5-shot | 80.0%\u003Cbr>5-shot | 88.0%\u003Cbr>0-shot | 81.4%\u003Cbr>8-shot | 62.9%\u003Cbr>8-shot | 56.8%\u003Cbr>5-shot |\n| MATH\u003Cbr>Math problems across\u003Cbr>5 difficulty levels &\u003Cbr>7 subdisciplines (Hendrycks et al., 2021b) | 53.2%\u003Cbr>4-shot | 32.6%\u003Cbr>4-shot | 52.9%\u003Cbr>4-shot\u003Cbr>(via API ∗∗ )\u003Cbr>50.3%\u003Cbr>(Zheng et al., 2023) | 34.1%\u003Cbr>4-shot (via API ∗∗ ) | 34.4%\u003Cbr>4-shot | — | 34.8% | 23.9%\u003Cbr>4-shot | 13.5%\u003Cbr>4-shot |\n| BIG-Bench-Hard\u003Cbr>Subset of hard BIG-bench tasks written as CoT problems\u003Cbr>(Srivastava et al., 2022) | 83.6%\u003Cbr>3-shot | 75.0%\u003Cbr>3-shot | 83.1%\u003Cbr>3-shot (via API ∗∗ ) | 66.6%\u003Cbr>3-shot (via API ∗∗ ) | 77.7%\u003Cbr>3-shot | — | — | — | 51.2%\u003Cbr>3-shot |\n| HumanEval\u003Cbr>Python coding tasks\u003Cbr>(Chen et al., 2021) | 74.4%\u003Cbr>0-shot (PT∗∗∗∗ ) | 67.7%\u003Cbr>0-shot (PT∗∗∗∗ ) | 67.0%\u003Cbr>0-shot (reported) | 48.1%\u003Cbr>0-shot | — | 70.0%\u003Cbr>0-shot | 44.5%\u003Cbr>0-shot | 63.2%\u003Cbr>0-shot | 29.9%\u003Cbr>0-shot |\n| Natural2Code\u003Cbr>Python code generation.(New held-out set with no leakage on web) | 74.9%\u003Cbr>0-shot | 69.6%\u003Cbr>0-shot | 73.9%\u003Cbr>0-shot (via API ∗∗ ) | 62.3%\u003Cbr>0-shot (via API ∗∗ ) | — | — | — | — | — |\n| DROP\u003Cbr>Reading comprehension & arithmetic.\u003Cbr>(metric: F1-score)\u003Cbr>(Dua et al., 2019) | 82.4\u003Cbr>Variable shots | 74.1 Variable shots | 80.9 3-shot (reported) | 64.1\u003Cbr>3-shot | 82.0 Variable shots | — | — | — | — |\n| HellaSwag\u003Cbr>(validation set) Common-sense multiple choice questions (Zellers et al., 2019) | 87.8%\u003Cbr>10-shot | 84.7%\u003Cbr>10-shot | 95.3%\u003Cbr>10-shot (reported) | 85.5%\u003Cbr>10-shot | 86.8%\u003Cbr>10-shot | — | 89.0%\u003Cbr>10-shot | — | 80.0%∗∗∗ |\n| WMT23\u003Cbr>Machine translation (metric: BLEURT)\u003Cbr>(Tom et al., 2023) | 74.4\u003Cbr>1-shot (PT∗∗∗∗ ) | 71.7\u003Cbr>1-shot | 73.8 1-shot (via API ∗∗ ) | — | 72.7\u003Cbr>1-shot | — | — | — | — |\n\n\n| | |\n| --- | --- |\n\n|  | Gemini Nano 1\u003Cbr>accuracy normalized by Pro |  | Gemini Nano 2\u003Cbr>accuracy normalized by Pro |  |\n| --- | --- | --- | --- | --- |\n| BoolQ | 71.6 | 0.81 | 79.3 | 0.90 |\n| TydiQA (GoldP) | 68.9 | 0.85 | 74.2 | 0.91 |\n| NaturalQuestions (Retrieved) | 38.6 | 0.69 | 46.5 | 0.83 |\n| NaturalQuestions (Closed-book) | 18.8 | 0.43 | 24.8 | 0.56 |\n| BIG-Bench-Hard (3-shot) | 34.8 | 0.47 | 42.4 | 0.58 |\n| MBPP | 20.0 | 0.33 | 27.2 | 0.45 |\n| MATH (4-shot) | 13.5 | 0.41 | 22.8 | 0.70 |\n| MMLU (5-shot) | 45.9 | 0.64 | 55.8 | 0.78 |\n\n\n| WMT 23\u003Cbr>(Avg BLEURT) | Gemini Ultra | Gemini Pro | Gemini Nano 2 | Gemini Nano 1 | GPT-4 | PaLM 2-L |\n| --- | --- | --- | --- | --- | --- | --- |\n| High Resource | 74.2 | 71.7 | 67.7 ","cbCairFCshaazyIf","https://ap.wps.com/l/cbCairFCshaazyIf","pdf",27110234,"English","# Gemini model tiers\n## Ultra, Pro, Nano descriptions\n## Benchmark results across tasks","[{\"question\":\"What are the key differences between Gemini Ultra, Gemini Pro, and Gemini Nano?\",\"answer\":\"Gemini Ultra targets state-of-the-art performance and is efficiently serveable at scale on TPU accelerators. Gemini Pro is optimized for cost and latency with strong reasoning and broad multimodal capabilities. Gemini Nano is designed to run on-device, with Nano-1 and Nano-2 variants using distillation from larger Gemini models and 4-bit quantization for deployment.\"},{\"question\":\"Which benchmarks are included in the document for evaluating model performance?\",\"answer\":\"The document includes results for MMLU, GSM8K, MATH, BIG-Bench-Hard, HumanEval, Natural2Code, DROP, HellaSwag, and WMT23, along with additional tables for normalized accuracy and translation scoring across resource levels.\"},{\"question\":\"How does the Nano deployment approach work according to the report?\",\"answer\":\"The report states that two Nano versions are trained (Nano-1 with 1.8B parameters and Nano-2 with 3.25B parameters) for different memory device constraints. It also notes that Nano is trained via distilling from larger Gemini models and is 4-bit quantized for deployment.\"}]","Gemini model report - performance overview | PDF",31]