https://deepmind.google/models/gemini/pro/ Build with our next generation AI systems Explore models chevron_right Gemini Our most intelligent AI models [nav__models__] 2.5 Pro [nav__models__] 2.5 Flash [nav__models__] 2.0 Flash-Lite Learn more Gemma Lightweight, state-of-the-art open models [nav__models__] Gemma 3 [nav__models__] Gemma 3n [nav__models__] ShieldGemma 2 Learn more Generative models Image, music and video generation models [nav__models__] Imagen [nav__models__] Lyria [nav__models__] Veo Experiments AI prototypes and experiments [nav__models__] Project Astra [nav__models__] Project Mariner [nav__models__] Gemini Diffusion Our latest AI breakthroughs and updates from the lab Explore research chevron_right Projects Explore some of the biggest AI innovations Learn more Publications Read a selection of our recent papers Learn more News Discover the latest updates from our lab Learn more Unlocking a new era of discovery with AI Explore science chevron_right AI for biology [nav__science_] AlphaFold [nav__science_] AlphaMissense [nav__science_] AlphaProteo AI for climate and sustainability [nav__science_] WeatherNext AI for mathematics and computer science [nav__science_] AlphaEvolve [nav__science_] AlphaProof [nav__science_] AlphaGeometry AI for physics and chemistry [nav__science_] GNoME [nav__science_] Fusion [nav__science_] AlphaQubit AI transparency [nav__science_] SynthID Our mission is to build AI responsibly to benefit humanity About Google DeepMind chevron_right News Discover our latest AI breakthroughs, projects, and updates Learn more Careers We're looking for people who want to make a real, positive impact on the world Learn more Milestones For over 20 years, Google has worked to make AI helpful for everyone Learn more Education We work to make AI more accessible to the next generation Learn more Responsibility Ensuring AI safety through proactive security, even against evolving threats Learn more The Podcast Uncover the extraordinary ways AI is transforming our world Learn more Models Research Science About Build with our next generation AI systems Explore models chevron_right Gemini Our most intelligent AI models [nav__models__] 2.5 Pro [nav__models__] 2.5 Flash [nav__models__] 2.0 Flash-Lite Learn more Gemma Lightweight, state-of-the-art open models [nav__models__] Gemma 3 [nav__models__] Gemma 3n [nav__models__] ShieldGemma 2 Learn more Generative models Image, music and video generation models [nav__models__] Imagen [nav__models__] Lyria [nav__models__] Veo Experiments AI prototypes and experiments [nav__models__] Project Astra [nav__models__] Project Mariner [nav__models__] Gemini Diffusion Our latest AI breakthroughs and updates from the lab Explore research chevron_right Projects Explore some of the biggest AI innovations Learn more Publications Read a selection of our recent papers Learn more News Discover the latest updates from our lab Learn more Unlocking a new era of discovery with AI Explore science chevron_right AI for biology [nav__science_] AlphaFold [nav__science_] AlphaMissense [nav__science_] AlphaProteo AI for climate and sustainability [nav__science_] WeatherNext AI for mathematics and computer science [nav__science_] AlphaEvolve [nav__science_] AlphaProof [nav__science_] AlphaGeometry AI for physics and chemistry [nav__science_] GNoME [nav__science_] Fusion [nav__science_] AlphaQubit AI transparency [nav__science_] SynthID Our mission is to build AI responsibly to benefit humanity About Google DeepMind chevron_right News Discover our latest AI breakthroughs, projects, and updates Learn more Careers We're looking for people who want to make a real, positive impact on the world Learn more Milestones For over 20 years, Google has worked to make AI helpful for everyone Learn more Education We work to make AI more accessible to the next generation Learn more Responsibility Ensuring AI safety through proactive security, even against evolving threats Learn more The Podcast Uncover the extraordinary ways AI is transforming our world Learn more Models Build with our next generation AI systems Explore models chevron_right Gemini Our most intelligent AI models [nav__models__] 2.5 Pro [nav__models__] 2.5 Flash [nav__models__] 2.0 Flash-Lite Learn more Gemma Lightweight, state-of-the-art open models [nav__models__] Gemma 3 [nav__models__] Gemma 3n [nav__models__] ShieldGemma 2 Learn more Generative models Image, music and video generation models [nav__models__] Imagen [nav__models__] Lyria [nav__models__] Veo Experiments AI prototypes and experiments [nav__models__] Project Astra [nav__models__] Project Mariner [nav__models__] Gemini Diffusion Research Our latest AI breakthroughs and updates from the lab Explore research chevron_right Projects Explore some of the biggest AI innovations Learn more Publications Read a selection of our recent papers Learn more News Discover the latest updates from our lab Learn more Science Unlocking a new era of discovery with AI Explore science chevron_right AI for biology [nav__science_] AlphaFold [nav__science_] AlphaMissense [nav__science_] AlphaProteo AI for climate and sustainability [nav__science_] WeatherNext AI for mathematics and computer science [nav__science_] AlphaEvolve [nav__science_] AlphaProof [nav__science_] AlphaGeometry AI for physics and chemistry [nav__science_] GNoME [nav__science_] Fusion [nav__science_] AlphaQubit AI transparency [nav__science_] SynthID About Our mission is to build AI responsibly to benefit humanity About Google DeepMind chevron_right News Discover our latest AI breakthroughs, projects, and updates Learn more Careers We're looking for people who want to make a real, positive impact on the world Learn more Milestones For over 20 years, Google has worked to make AI helpful for everyone Learn more Education We work to make AI more accessible to the next generation Learn more Responsibility Ensuring AI safety through proactive security, even against evolving threats Learn more The Podcast Uncover the extraordinary ways AI is transforming our world Learn more Try Google AI Studio # Try Gemini # Google DeepMind Google AI Learn about all of our AI Google DeepMind Explore the frontier of AI Google Labs Try our AI experiments Google Research Explore our research Gemini app Chat with Gemini Google AI Studio Build with our next-gen AI models Models Research Science About # Try Google AI Studio # Try Gemini Gemini Our most intelligent AI models Chat with Gemini Try in Google AI Studio [aKGkv9hI3gEBsmPkLux9v9qMfZd] Models Gemini 2.5: Our most intelligent models are getting even better [h9YL1McSoifYFJy1L7hYhdkOcNY] Models Gemini 2.5 Pro Preview: even better coding performance [SSbxSkG2rVlx0EBDPkef4DlizlS] Models Build rich, interactive web apps with an updated Gemini 2.5 Pro [SpF9hIf_C3KSo-xElVbNF768qUJtb4TNoXTczjyxdH9BSUNnNaedlC7QYq6d9C8YGVnEfSjDdvC3hR4p81UijHFgLoqWqqRPQ70lpnJB50OAAot0iw] Gemini 2.5 models are capable of reasoning through their thoughts before responding, resulting in enhanced performance and improved accuracy. * What's new * Models * Hands-on * Performance * Safety * Build What's new Access the latest preview of Gemini 2.5 Pro We're introducing an upgraded preview of Gemini 2.5 Pro, our most intelligent model yet. Try in Google AI Studio [zLWgkcGrRi_ETSVXGfmgssJ1V8hAo45YsNNeLVoJ3X88h3L3pz9Gm20wjm7RDrdtKoCo5v5thnPU_rVVbz8nsPqth_ZSE3Qpvq4wGQlCNv1qMIsRzQ] Deep Think We're making Gemini 2.5 Pro even better by introducing an enhanced reasoning mode called Deep Think. Learn more Pause video Play video Native audio Converse in more expressive ways with native audio outputs that capture the subtle nuances of how we speak. Learn more Unmute video Mute video Pause video Play video An even better 2.5 Flash Improved across key benchmarks for reasoning, multimodality, code and long context while getting even more efficient. Try in Google AI Studio Pause video Play video Model family Gemini 2.5 builds on the best of Gemini -- with native multimodality and a long context window. * [lLQG1sGhmjTxlX0uvPdyB_fzSJrhcd4uCLcJzAMZww4NtQsTshFhdTKDqv] Preview 2.5 Pro Best for coding and complex prompts Learn more * [UfKWV8Zdl5SMqIrVS4PKHKaW1VbVBUWfLNmQalnL-0JopkepklxiAuGzWn] Preview 2.5 Flash Best for fast performance on complex tasks Learn more * [xzKpRoyzgbXdm11EtNY2GL0WyjLdktmLwt5VublbcRkuG8VkklgtBrVibc] General availability 2.0 Flash-Lite Best for cost-efficient performance Learn more Hands-on with 2.5 Pro See how Gemini 2.5 Pro uses its reasoning capabilities to create interactive simulations and do advanced coding. [SKA8kprwmNx0jLuWpDSW92dKgMFMKvhnLoQFDBwyy2bV9sF66oz63wxUkhC5u8afr5Ct4J] [hqdefault] Watch Make an interactive animation See how Gemini 2.5 Pro uses its reasoning capabilities to create an interactive animation of "cosmic fish" with a simple prompt. [Uey7LzpJXJAi6ypT9AmMC9S_yK-jL7PvoZi8z8d49UkPUF-bccFFQHC6-jKW5Opse10HPg] [hqdefault] Watch Create your own dinosaur game Watch Gemini 2.5 Pro create an endless runner game, using executable code from a single line prompt. [dBysJaamRHrlXjFJxdD00z15fO7VBbIknHoar-BXyex8pEWdHWmD101kAx1CNs-dkrdIUr] [hqdefault] Watch Code a fractal visualization See how Gemini 2.5 Pro creates a simulation of intricate fractal patterns to explore a Mandelbrot set. [IRPqYnoHkdkEAqqz--k9rmoeu7cOsWB4sUo3-Cjn3dGac3V9S4S1mhLNief4rUmZe119y9] [hqdefault] Watch Plot interactive economic data Watch Gemini 2.5 Pro use its reasoning capabilities to create an interactive bubble chart to visualize economic and health indicators over time. [DOLiWOw9PzulyTnrd7E8xQ0ecWLwJU3YbaqZvdT-SZrOG2Kfio_McNij1I3bxqnnWxmFgT] [hqdefault] Watch Animate complex behavior See how Gemini 2.5 Pro creates an interactive Javascript animation of colorful boids inside a spinning hexagon. [zGWm9SxkW69eqZ1ZArzuWgFabqDjB_gzwPR5FCRLIW4kOycXldqr8ux1JrTkc5U1eXJFGy] [hqdefault] Watch Code particle simulations Watch Gemini 2.5 Pro use its reasoning capabilities to create an interactive simulation of a reflection nebula. Performance Gemini 2.5 is state-of-the-art across a wide range of benchmarks. [KCVQJ3Q9X2phZ22j9HDrWJLJzy8cBqGvq-dZ21Gkk9OwrH7a0AfVFAWWd4m2w9Zj7w4rHPCM6k6ug5OWfYTvHGjZ3dh18Wh0llfjNOjc6KxCIyH4VJs] Benchmarks Gemini 2.5 Pro demonstrates significantly improved performance across a wide range of benchmarks. Gemini 2.5 OpenAI Claude Grok 3 Benchmark Pro Preview OpenAI o4-mini Opus 4 Beta DeepSeek 06-05 o3 High High 32k Extended R1 05-28 Thinking thinking thinking $/1M tokens $1.25 Input price (no $2.50 > $10.00 $1.10 $15.00 $3.00 $0.55 caching) 200k tokens $10.00 Output price $/1M tokens $15.00 > $40.00 $4.40 $75.00 $15.00 $2.19 200k tokens Reasoning & knowledge Humanity's 21.6% 20.3% 14.3% 10.7% -- 14.0%* Last Exam (no tools) Science GPQA single 86.4% 83.3% 81.4% 79.6% 80.2% 81.0% diamond attempt multiple -- -- -- 83.3% 84.6% -- attempts Mathematics single 88.0% 88.9% 92.7% 75.5% 77.3% 87.5% AIME 2025 attempt multiple -- -- -- 90.0% 93.3% -- attempts Code generation LiveCodeBench single 69.0% 72.0% 75.8% 51.1% -- 70.5% (UI: 1/1/ attempt 2025-5/1/ 2025) Code editing 82.2% 79.6% 72.0% 72.0% 53.3% Aider diff-fenced diff diff diff diff 71.6% Polyglot Agentic coding single 59.6% -- -- 72.5% -- -- SWE-bench attempt Verified multiple 67.2% 69.1% 68.1% 79.4% -- 57.6% attempts Factuality 54.0% 48.6% 19.3% -- 43.6% 27.8% SimpleQA Factuality FACTS 87.8% 69.6% 62.1% 77.7% 74.8% -- grounding Visual single no MM reasoning attempt 82.0% 82.9% 81.6% 76.5% 76.0% support MMMU multiple -- -- -- -- 78.0% no MM attempts support Image understanding 67.2% -- -- -- -- no MM Vibe-Eval support (Reka) Video no MM understanding 83.6% -- -- -- -- support VideoMMMU Long context 128k MRCR v2 (average) 58.0% 57.1% 36.3% -- 34.0% -- (8-needle) 1M 16.4% no no no no no (pointwise) support support support support support Multilingual performance 89.2% -- -- -- -- -- Global MMLU (Lite) Methodology Gemini results: All Gemini scores are pass @1."Single attempt" settings allow no majority voting or parallel test-time compute; "multiple attempts" settings allow test-time selection of the candidate answer. They are all run with the AI Studio API for the model-id gemini-2.5-pro-preview-06-05 with default sampling settings. To reduce variance, we average over multiple trials for smaller benchmarks. Aider Polyglot score is the pass rate average of 3 trials. Vibe-Eval results are reported using Gemini as a judge. Non-Gemini results: All the results for non-Gemini models are sourced from providers' self reported numbers unless mentioned otherwise below. All SWE-bench Verified numbers follow official provider reports, using different scaffoldings and infrastructure. Google's scaffolding for "multiple attempts" for SWE-Bench includes drawing multiple trajectories and re-scoring them using model's own judgement. Thinking vs not-thinking: For Claude 4 results are reported for the reasoning model where available (HLE, LCB, Aider). For Grok-3 all results come with extended reasoning except for SimpleQA (based on xAI reports) and Aider. For OpenAI models high level of reasoning is shown where results are available (except for GPQA, AIME 2025, SWE-Bench, FACTS, MMMU). Single attempt vs multiple attempts: When two numbers are reported for the same eval higher number uses majority voting with n=64 for Grok models and internal scoring with parallel test time compute for Anthropic models. Result sources: Where provider numbers are not available we report numbers from leaderboards reporting results on these benchmarks: Humanity's Last Exam results are sourced from https://agi.safe.ai/ and https://scale.com/leaderboard/humanitys_last_exam, AIME 2025 numbers are sourced from https://matharena.ai/. LiveCodeBench results are from https://livecodebench.github.io/leaderboard.html (1/1/2025 - 5/1/2025 in the UI), Aider Polyglot numbers come from https:// aider.chat/docs/leaderboards/. FACTS come from https://www.kaggle.com /benchmarks/google/facts-grounding. For MRCR v2 which is not publicly available yet we include 128k results as a cumulative score to ensure they can be comparable with other models and a pointwise value for 1M context window to show the capability of the model at full length. The methodology has changed in this table vs previously published results for MRCR v2 as we have decided to focus on a harder, 8-needle version of the benchmark going forward. API costs are sourced from providers' website and are current as of June 5th. * indicates evaluated on text problems only (without images) Building responsibly in the agentic era As we develop these new technologies, we recognize the responsibility it entails, and aim to prioritize safety and security in all our efforts. Learn more [en8USUshUuTe_aFMln_4oBIGpAeDseY0aIDY90NQ40ihXZASHQZuUCE6sPsOyDorfXYgjGMggSKMh49zjXQu7VlNYhUP84_97fLUjZaxlLowh0Pn] For developers Gemini's advanced thinking, native multimodality and massive context window empowers developers to build next-generation experiences. Start building [md_zZiipNmxFRr6-amI3KcaGrhtA0rWSuHvCU7EsyYkvu3uFJJ_iW1-nbipeYkA25mRVeGGhihhnIFKa_JpStPJwUUsPP4GYnunDKathmvYGIHqU-A] Developer ecosystem Build with cutting-edge generative AI models and tools to make AI helpful for everyone. [gemini_overview_google-ai-studio__icon] Google AI Studio Build with the latest models from Google DeepMind [gemini_overview_gemini-api__icon] Gemini API Easily integrate Google's most capable AI model to your apps Accessing our latest AI models We want developers to gain access to our models as quickly as possible. We're making these available through Google AI Studio. Sign in to Google AI Studio Get the latest updates Sign up for news on the latest innovations from Google DeepMind. Email address [ ] Please enter a valid email (e.g., "name@example.com") I accept Google's Terms and Conditions and acknowledge that my information will be used in accordance with Google's Privacy Policy. Sign up Follow us footer__x footer__instagram footer__youtube footer__linkedin footer__github Build AI responsibly to benefit humanity Models Build with our next generation AI systems # Gemini # Gemma # Veo # Imagen # Lyria Science Unlocking a new era of discovery with AI # AlphaFold # SynthID # WeatherNext Learn more About News Careers Research Responsibility & Safety Sign up for updates on our latest innovations I accept Google's Terms and Conditions and acknowledge that my information will be used in accordance with Google's Privacy Policy. [ ] chevron_right Please enter a valid email (e.g., "name@example.com") About Google Google products Privacy Terms Manage cookies Gemini Pro Gemini Pro Preview Gemini 2.5 Pro Best for coding and complex prompts Try in Google AI Studio [HTs_nyUe37i_bc2SU2dz5EH283AdGn5JR_YMMk7BfE_LZT_omu7OT5IQG6f9r3iqcRef4layxxRhXLuGZ-Cibaz8Fuy4U-mQEoZdhQgYA--qQuKiR24] Gemini 2.5 Pro is our most advanced model yet, excelling at coding and complex prompts. Pro performance * [FuUQuVu6jpSS6k5vE3lQy1Z2U_vr3x1ikx0_WgfWMe_bS9byPbqs4Y1afo] Enhanced reasoning State-of-the-art in key math and science benchmarks. * [1LFgm7CAbuqmBskCmTVuo7zBHpa9UwbCJAZN4lWs2WfFGMwID9QmqGyo_O] Advanced coding Easily generate code for web development tasks. * [PVyBgeY20xCZfbtqOE2lDP7YBB80jnXmgu9AYMM6SB_dse9q63Bdebwprr] Natively multimodal Understands input across text, audio, images and video. * [VUNQxDBiDi-DqY2DpsZ4evg1UXpfoEJ1ySyvlUcIyPbPlFS0tMVf8rIrjT] Long context Explore vast datasets with a 1-million token context window. --------------------------------------------------------------------- Deep Think We're making Gemini 2.5 Pro even better by introducing an enhanced reasoning mode called Deep Think. [hqdefault] Watch It uses our latest cutting edge research in reasoning - including parallel thinking techniques - resulting in incredible performance. [aPsyn7kd4tkP7hyXEgC30K4CnKFmK5VkQOC4MCTN4CWmuqfaiuRraKts] [8soduZM_U7zg28P8Wf4h8R1oiTkwYDxww_SQ5HzqAZHsiMeCcPYSwUz8] [RGbTJVWEpwh59ufm_zkklMGJ6kAonYI9oY-1Xpt3-2I6b7M1T3dqB2wH] Methodology All Gemini results come from our runs. USAMO 2025: https:// matharena.ai. LiveCodeBench V6: * o3 High: Internal runs since numbers are not available in official leaderboard, o4-mini High: https://livecodebench.github.io/leaderboard.html (2/1/2025-5/1/2025). MMMU: Self reported by OpenAI --------------------------------------------------------------------- Preview Native audio Converse in more expressive ways with native audio outputs that capture the subtle nuances of how we speak. Seamlessly switch between 24 languages, all with the same voice. Try in Google AI Studio Unmute video Mute video Pause video Play video --------------------------------------------------------------------- Vibe-coding nature with 2.5 Pro Images transformed into code-based representations of its natural behavior. [hqdefault] Watch Hands-on with 2.5 Pro See how Gemini 2.5 Pro uses its reasoning capabilities to create interactive simulations and do advanced coding. Make an interactive animation See how Gemini 2.5 Pro uses its reasoning capabilities to create an interactive animation of "cosmic fish" with a simple prompt. [SKA8kprwmNx0jLuWpDSW92dKgMFMKvhnLoQFDBwyy2bV9sF66oz63wxUkhC5u8afr5Ct4JgnbXj8vlvHsl_6hh] Watch Create your own dinosaur game Watch Gemini 2.5 Pro create an endless runner game, using executable code from a single line prompt. [Uey7LzpJXJAi6ypT9AmMC9S_yK-jL7PvoZi8z8d49UkPUF-bccFFQHC6-jKW5Opse10HPgyefPMAW_zB1xSvvY] Watch Code a fractal visualization See how Gemini 2.5 Pro creates a simulation of intricate fractal patterns to explore a Mandelbrot set. [dBysJaamRHrlXjFJxdD00z15fO7VBbIknHoar-BXyex8pEWdHWmD101kAx1CNs-dkrdIUrmOnjxDlu8bDmzpKE] Watch Plot interactive economic data Watch Gemini 2.5 Pro use its reasoning capabilities to create an interactive bubble chart to visualize economic and health indicators over time. [IRPqYnoHkdkEAqqz--k9rmoeu7cOsWB4sUo3-Cjn3dGac3V9S4S1mhLNief4rUmZe119y9eIL2qiiYVltQ0jPQ] Watch Animate complex behavior See how Gemini 2.5 Pro creates an interactive Javascript animation of colorful boids inside a spinning hexagon. [DOLiWOw9PzulyTnrd7E8xQ0ecWLwJU3YbaqZvdT-SZrOG2Kfio_McNij1I3bxqnnWxmFgTINK6odEu_tfjo8pi] Watch Code particle simulations Watch Gemini 2.5 Pro use its reasoning capabilities to create an interactive simulation of a reflection nebula. [zGWm9SxkW69eqZ1ZArzuWgFabqDjB_gzwPR5FCRLIW4kOycXldqr8ux1JrTkc5U1eXJFGyH8A4QSlJiENT02XN] Watch Benchmarks Gemini 2.5 Pro leads common benchmarks by meaningful margins. Gemini 2.5 OpenAI Claude Grok 3 Benchmark Pro Preview OpenAI o4-mini Opus 4 Beta DeepSeek 06-05 o3 High High 32k Extended R1 05-28 Thinking thinking thinking $/1M tokens $1.25 Input price (no $2.50 > $10.00 $1.10 $15.00 $3.00 $0.55 caching) 200k tokens $10.00 Output price $/1M tokens $15.00 > $40.00 $4.40 $75.00 $15.00 $2.19 200k tokens Reasoning & knowledge Humanity's 21.6% 20.3% 14.3% 10.7% -- 14.0%* Last Exam (no tools) Science GPQA single 86.4% 83.3% 81.4% 79.6% 80.2% 81.0% diamond attempt multiple -- -- -- 83.3% 84.6% -- attempts Mathematics single 88.0% 88.9% 92.7% 75.5% 77.3% 87.5% AIME 2025 attempt multiple -- -- -- 90.0% 93.3% -- attempts Code generation LiveCodeBench single 69.0% 72.0% 75.8% 51.1% -- 70.5% (UI: 1/1/ attempt 2025-5/1/ 2025) Code editing 82.2% 79.6% 72.0% 72.0% 53.3% Aider diff-fenced diff diff diff diff 71.6% Polyglot Agentic coding single 59.6% -- -- 72.5% -- -- SWE-bench attempt Verified multiple 67.2% 69.1% 68.1% 79.4% -- 57.6% attempts Factuality 54.0% 48.6% 19.3% -- 43.6% 27.8% SimpleQA Factuality FACTS 87.8% 69.6% 62.1% 77.7% 74.8% -- grounding Visual single no MM reasoning attempt 82.0% 82.9% 81.6% 76.5% 76.0% support MMMU multiple -- -- -- -- 78.0% no MM attempts support Image understanding 67.2% -- -- -- -- no MM Vibe-Eval support (Reka) Video no MM understanding 83.6% -- -- -- -- support VideoMMMU Long context 128k MRCR v2 (average) 58.0% 57.1% 36.3% -- 34.0% -- (8-needle) 1M 16.4% no no no no no (pointwise) support support support support support Multilingual performance 89.2% -- -- -- -- -- Global MMLU (Lite) Methodology Gemini results: All Gemini scores are pass @1."Single attempt" settings allow no majority voting or parallel test-time compute; "multiple attempts" settings allow test-time selection of the candidate answer. They are all run with the AI Studio API for the model-id gemini-2.5-pro-preview-06-05 with default sampling settings. To reduce variance, we average over multiple trials for smaller benchmarks. Aider Polyglot score is the pass rate average of 3 trials. Vibe-Eval results are reported using Gemini as a judge. Non-Gemini results: All the results for non-Gemini models are sourced from providers' self reported numbers unless mentioned otherwise below. All SWE-bench Verified numbers follow official provider reports, using different scaffoldings and infrastructure. Google's scaffolding for "multiple attempts" for SWE-Bench includes drawing multiple trajectories and re-scoring them using model's own judgement. Thinking vs not-thinking: For Claude 4 results are reported for the reasoning model where available (HLE, LCB, Aider). For Grok-3 all results come with extended reasoning except for SimpleQA (based on xAI reports) and Aider. For OpenAI models high level of reasoning is shown where results are available (except for GPQA, AIME 2025, SWE-Bench, FACTS, MMMU). Single attempt vs multiple attempts: When two numbers are reported for the same eval higher number uses majority voting with n=64 for Grok models and internal scoring with parallel test time compute for Anthropic models. Result sources: Where provider numbers are not available we report numbers from leaderboards reporting results on these benchmarks: Humanity's Last Exam results are sourced from https://agi.safe.ai/ and https://scale.com/leaderboard/humanitys_last_exam, AIME 2025 numbers are sourced from https://matharena.ai/. LiveCodeBench results are from https://livecodebench.github.io/leaderboard.html (1/1/2025 - 5/1/2025 in the UI), Aider Polyglot numbers come from https:// aider.chat/docs/leaderboards/. FACTS come from https://www.kaggle.com /benchmarks/google/facts-grounding. For MRCR v2 which is not publicly available yet we include 128k results as a cumulative score to ensure they can be comparable with other models and a pointwise value for 1M context window to show the capability of the model at full length. The methodology has changed in this table vs previously published results for MRCR v2 as we have decided to focus on a harder, 8-needle version of the benchmark going forward. API costs are sourced from providers' website and are current as of June 5th. * indicates evaluated on text problems only (without images) Model information 2.5 Pro Model Card Model deployment status Preview Supported data types for input Text, Image, Video, Audio Supported data types for output Text Supported # tokens for input 1M Supported # tokens for output 64k Knowledge cutoff January 2025 Function calling Tool use Structured output Search as a tool Code execution Reasoning Best for Coding Complex prompts Google AI Studio Availability Gemini API Gemini App