[HN Gopher] Gemini 3 Pro Model Card [pdf]
___________________________________________________________________
Gemini 3 Pro Model Card [pdf]
Author : virgildotcodes
Score : 108 points
Date : 2025-11-18 11:12 UTC (11 hours ago)
(HTM) web link (storage.googleapis.com)
(TXT) w3m dump (storage.googleapis.com)
| rvz wrote:
| > The training dataset also includes: publicly available datasets
| that are readily downloadable; data obtained by crawlers;
| licensed data obtained via commercial licensing agreements; user
| data (i.e., data collected from users of Google products and
| services to train AI models, along with user interactions with
| the model) in accordance with Google's relevant terms of service,
| privacy policy, service-specific policies, and pursuant to user
| controls, where appropriate; other datasets that Google acquires
| or generates in the course of its business operations, or
| directly from its workforce; and AI-generated synthetic data.
|
| Well don't complain when you are using Gmail and your emails are
| being trained to develop Gemini.
| patates wrote:
| It says "pursuant to user controls, where appropriate". We can
| now sleep peacefully with the knowledge that Google will give
| us the tools to disable this where it's not inappropriate.
| rvz wrote:
| So that's why Google is getting sued for Gemini being enabled
| by default in Gmail and analyzing emails and our data;
| completely going against whatever privacy policy they came up
| with. [0]
|
| I don't expect them to follow their own privacy policies.
|
| [0] https://www.yahoo.com/news/articles/google-sued-over-
| gemini-...
| surrTurr wrote:
| gone now;
|
| wayback machine still has it:
| https://web.archive.org/web/20251118111103/https://storage.g...
| lifthrasiir wrote:
| For the veracity of the link itself:
| https://storage.googleapis.com/deepmind-media/* has been used by
| DeepMind itself (e.g. "View tech report" in
| https://deepmind.google/models/gemini/) so it is a genuine leak.
| meetpateltech wrote:
| it was accidentally pushed a little early, and now it has been
| taken down.
|
| here's the archived pdf:
| https://web.archive.org/web/20251118111103/https://storage.g...
| TheAceOfHearts wrote:
| They scored a 31.1% on ARC AGI 2 which puts them in first place.
|
| Also notable which models they include for comparison: Gemini 2.5
| Pro, Claude Sonnet 4.5, and GPT-5.1. That seems like a minor snub
| against Grok 4 / Grok 4.1.
| kranke155 wrote:
| Grok seems extremely prone to hallucination in my experience.
| It also constantly asserts certainty on fuzzy topics.
| buildfocus wrote:
| My impression is that Grok is very rarely used in practice
| outside of a niche of die-hard users, partly because of very
| different tuning to other models, and partly the related public
| reputation around it.
|
| https://firstpagesage.com/reports/top-generative-ai-chatbots...
| suggests 0.6% of chat use cases, well below the other big
| names, and I suspect those stats for chat are higher than other
| scenarios like business usage. Given all that, I can see how
| Gemini might not be focused on competing with them.
| ohyoutravel wrote:
| I don't know anyone who uses Grok, but in my peer group
| everyone uses 1-2 paid services like Gemini or Clause or
| ChatGPT. They're probably not as "extremely online" as I am,
| so I can't generalize this thought, but anecdotally my
| impression has been that Grok is just very "right wing
| influencer" coded.
| npn wrote:
| well, there are 3 kind of usages for grok: - using grok
| inside X/Twitter: most people interacts with Grok this way. -
| using grok on its website: this is really annoying, as you
| get delayed by cloudflare everytime you access the site. As
| grok does not provide serious advantage over other services,
| why bother - you can also use the app, but it is not as
| convenient as other services.
|
| it is understandable that grok is not popular.
| jmmcd wrote:
| About ARC 2:
|
| I would want to hear more detail about prompts, frameworks,
| thinking time, etc., but they don't matter too much. The main
| caveat would be that this is probably on the public test set,
| so could be in pretraining, and there could even be some ARC-
| focussed post-training - I think we don't know yet and might
| never know.
|
| But for any reasonable setup, if no egregious cheating, that is
| an _amazing_ score on ARC 2.
| surrTurr wrote:
| good benchmark stats except for coding where it looks similar to
| other SOTA models
| aurareturn wrote:
| Benchmark suggests it is a resounding win for Gemini 3 Pro as the
| top model.
| margorczynski wrote:
| If these numbers are true then OpenAI is probably done, Anthropic
| too. Still, it's hard to see an effective monetization method for
| this tech and it clearly is eating Google's main pie which is
| search.
| Sol- wrote:
| Why? These models just leapfrog each other as time advances.
|
| One month Gemini is on top, then ChatGPT, then Anthropic. Not
| sure why everyone gets FOMO whenever a new version gets
| released.
| remus wrote:
| I think google is uniquely well placed to make a profitable
| business out of AI: They make their own TPUs so don't have to
| pay ridiculous amounts of money to Nvidia, they have a great
| depth of talent in building models, they've got loads of data
| they can use for training and they've got a huge existing
| customer base who can buy their AI offerings.
|
| I don't think any other company has all these ingredients.
| gizmodo59 wrote:
| While I don't disagree that Google is the company you can't
| bet against when it comes to AI, saying other companies are
| done is a stretch. If they have a significant moat then
| they should be at the top all the time by then which is not
| the case though.
| remus wrote:
| Agreed, too early to write off others entirely. It'll be
| interesting to see who comes out the other side of the
| bubble with a working business.
| adriand wrote:
| Anthropic has a fairly significant lead when it comes to
| enterprise usage and for coding. This seems like a
| workable business model to me.
| bootlooped wrote:
| I feel this is a tenuous position though. I find it
| incredibly easy to switch to Gemini CLI when I want a
| second opinion, or when Claude is down.
| adriand wrote:
| The enterprise sales cycle is often quite long, though,
| and often includes a lot of hurdles around compliance,
| legal, etc. It would take a fairly sustained loss of edge
| before a lot of enterprises would switch once they're
| hooked into a given platform. It's interesting to me that
| Sonnet 4.5 still edges Gemini 3 on SWE bench. This seems
| to bode well for the trajectory that Anthropic is on.
| basch wrote:
| ChatGPT's moat is their name and user habit. People who
| are using it will keep using it. All/most of the products
| are _good enough_ for the people who already got used to
| using them, that they arent exploring competitors.
|
| Microsoft has the chance of changing habit the most by
| virtue of being bundled into business contracts that have
| companies with policies not allowing any other product in
| the workplace.
| netdevphoenix wrote:
| > business contracts that have companies with policies
| not allowing any other product in the workplace.
|
| Elaborate please. Are you saying that MS is forcing
| customers to make Copilot the only allowed LLM product?
| basch wrote:
| Not quite, but in effect.
|
| Microsoft has contracts to provide software to companies.
| Companies have policies that only provided software and
| ai is allowed. Ipso facto
| remus wrote:
| > ChatGPT's moat is their name and user habit. People who
| are using it will keep using it. All/most of the products
| are _good enough_ for the people who already got used to
| using them, that they arent exploring competitors.
|
| They have a long way to go to become profitable though.
| Those users will get less sticky when openAI starts
| upping their pricing/putting ads everywhere/making the
| product worse to save money/all of the above.
| mlnj wrote:
| 100% the reason I am long on Google. They can take their
| time to monetize these new costs.
|
| Even other search competitors have not proven to be a
| danger to Google. There is nothing stopping that search
| money coming in.
| spaceman_2020 wrote:
| The bear case for Google was always the business side would
| cannibalize the AI side. AI makes search redundant which
| kills the golden goose
| Zigurd wrote:
| The TPU are a key factor. They are the most mature
| alternative to Nvidia. Only Google cloud, Azure, and AWS
| enable you to rent their respective AI chips. Out of those
| three, google is the only one to have a frontier model. So
| if they have a real advantage they're not exposed to the
| financial shenanigans propping up neo clouds like
| Coreweave.
| redox99 wrote:
| Considering GPT 5 was only recently released, it's very
| unlikely GPT will achieve these scores in just a couple of
| months. If they had something this good in the oven, they'd
| probably left the GPT 5 name to it.
|
| Or maybe Google just benchmaxxed and this doesn't translate
| at all in real world performance.
| Palmik wrote:
| GPT 5 was released more than 3 months ago. Gemini 2.5 was
| released less than 8 months ago.
| sidibe wrote:
| If not this model, Google at some point is going to get
| and stay ahead just because they have so many more people
| and compute resources they can throw at many directions
| while the others have to make the right choices with how
| they use their resources each time. Took a while to
| channel their numbers into a product direction but now I
| don't think they're going to let up
| blueblisters wrote:
| They do have unreleased Olympiad Gold-winning models that
| are definitely better than GPT5.
|
| TBD if that performance generalizes to other real world
| tasks.
| happa wrote:
| This may just be bad recollection from my part, but hasn't
| Google reported that their search business is right now the
| most profitable it has ever been?
| senordevnyc wrote:
| 1) New SOTA models come out all the time and that hasn't killed
| the other major AI companies. This will be no different.
|
| 2) Google's search revenue last quarter was $56 billion, a 14%
| increase over Q3 2024.
| margorczynski wrote:
| 1) Not long ago Altman and the OpenAI CFO were openly asking
| for public money. None of these AI companies have actually
| any kind of working business plan and are just burning
| investor money. If the investors see there is no winning
| against Google (or some open Chinese model) the money will
| dry up.
|
| 2) I'm not suggesting this will happen overnight but
| especially younger people gravitate towards LLM for
| information search + actively use some sort of ad blocking.
| In the long run it doesn't look great for Google.
| senordevnyc wrote:
| No, you suggested that LLMs are clearly eating google's
| lunch already, and there's just no evidence of that. Quite
| the opposite.
| paswut wrote:
| I'd love to see anthropic/openai pop. back to some regular
| programming. the models are good enough, time to invest
| elsewhere
| ilaksh wrote:
| The only one it doesn't win is SWE bench which it is
| significantly behind Claude Sonnet. You just can't take down
| Sonnet.
| stavros wrote:
| Codex has been much better than Sonnet for me.
| dotancohen wrote:
| On what types of tasks?
| svantana wrote:
| One percentage point is not significant, neither in the
| colloquial nor the scientific sense[1].
|
| [1] Binomial formula gives a confidence interval of 3.7%,
| using p=0.77, N=500, confidence=95%
| lukev wrote:
| Or else it trained/overfit to the benchmarks. We won't really
| know until people have a chance to use it for real-world tasks.
|
| Also, models are already _pretty good_ but product /market fit
| (in terms of demonstrated economic value delivered) remains
| elusive outside of a couple domains. Does a model that's (say)
| 30% better reach an inflection point that changes that
| narrative, or is a more qualitative change required?
| alecco wrote:
| For SWE it is the same ranking. But if Google's $20/mo plan is
| comparable to the $100-200 plans for OpenAI and Anthropic, yes
| they are done.
|
| But we'll have to wait a few weeks to see if the nerfed model
| post-release is still as good.
| siva7 wrote:
| I have a few secret prompts to test complex reasoning
| capabilities of new models (in law and medicine). Gemini (2.5
| pro) is by a wide margin behind Anthropic (sonnet 4.5 basic
| thinking) and Openai (pro model) on my own benchmark and I
| trust my own benchmark more than public leaderboards. So it's
| the other way around. Google is trying to catch up where the
| others are. It just doesn't seem so to some because Google
| undercuts prices and most people don't have own complex
| problems with a verified solution to test against (so they
| could see how bad Gemini is in reality)
| alecco wrote:
| This thread is about Gemini 3. It will be interesting to
| see your benchmark results when it's available later.
| llm_nerd wrote:
| They're constantly matching and exceeding each other. It's a
| hypercompetitive space and I would fully expect one of the
| others to top various benchmarks shortly after. On pretty much
| every leading release someone does this "everyone else is done!
| Shut er down" thing and it's growing pretty weird.
|
| Having said that, OpenAI's ridiculous hype cycle has been
| living on borrowed time. OpenAI has zero moat, and are just one
| vendor in a space with many vendors, and even incredibly
| competent open source models by surprise Chinese entrants. Sam
| Altman going around acting like he's a prophet and they're the
| gatekeepers of the future is an act that should be super old,
| but somehow fools and their money continue to be parted.
| netdevphoenix wrote:
| This. If I had to put my money on a survivor, it would be
| Google because it is an established company with existing
| revenue modules unrelated to AI. Anthropic and OpenAI won't
| stand alone without external funding
| patates wrote:
| It says it's been trained from scratch. I wonder if it will have
| the same undescribable magic that makes me spend an hour every
| day with 2.5. I really love the results I can get with 2.5 pro.
| Google eventually limiting aistudio will be a sad day.
|
| Also I really hoped for a 2M+ context. I'm living on the context
| edge even with 1M.
| dahcryn wrote:
| buy a pixel and you get it basically unlimited for free for a
| year ;)
| sohpea wrote:
| or a Chromebook is a good choice too considering price
| JacobAsmuth wrote:
| AIStudio now accepts an API key. Unlimited usage :)
| Traubenfuchs wrote:
| So does google actually have a claude console alternative
| currently?
| rjtavares wrote:
| Gemini CLI
| muro wrote:
| https://github.com/google-gemini/gemini-cli
| itsmevictor wrote:
| Noteworthily, although Gemini 3 Pro seems to have much
| benchmark scores than other models across the board (including
| compared to Claude), it's not the case for coding, where it
| appears to score essentially the same as the others. I wonder
| why that is.
|
| So far, IMHO, Claude Code remains significantly better than
| Gemini CLI. We'll see whether that changes with Gemini 3.
| decster wrote:
| from my experience, the quality of gemini-cli isn't great,
| experiencing lot of stupied bug.
| spwa4 wrote:
| Google is currently constantly laying off people. Everyone
| who really exceeds has jumped ship, and the people who
| remain ... are not top of the class anymore.
|
| Not that Google didn't use to have problems shipping useful
| things. But it's gotten a lot worse.
| BoredPositron wrote:
| Gemini performs better if you use it with Claude Code than
| with Gemini cli. It still has some odd problems with tool
| calling but a lot of the performance loss is the Gemini cli
| app itself.
| lifthrasiir wrote:
| Probably because many models from Anthropic would have been
| optimized for agentic coding in particular...
|
| EDIT: Don't disagree that Gemini CLI has a lot of rough
| edges, though.
| Lionga wrote:
| Because benchmark are a retarded comparison and having
| nothing to do with reality. Its just jerk material for AI
| Fanboys
| siva7 wrote:
| > I wonder why that is.
|
| That's because coding is currently the only reliable
| benchmark where reasoning capabilities transfer to predict
| capabilities for other professions like law. Coding is the
| only area where they are shy to release numbers. All these
| exam scores are fakeable by gaming those benchmarks.
| adidoit wrote:
| gemini cli. It's not as impressive as claude code or even
| codex.
|
| Claude code seems to be more compatible with the model (or the
| reverse) whereas gemini-cli still feels a bit awkward (as of
| 2.5 Pro). I'm hoping its better with 3.0!
| laborcontract wrote:
| It's hilarious that the release of Gemini 3 is getting eclipsed
| by this cloudflare outage.
| senordevnyc wrote:
| It hasn't been released, this is just a leak
| amarcheschi wrote:
| On reddit I see it's already available on cursor
|
| https://www.reddit.com/r/Bard/comments/1p093fb/gemini_3_in_c.
| ..
| senordevnyc wrote:
| Interesting, it doesn't show up for me in Cursor yet.
| Despacito2019 wrote:
| you need to manually add the custom model gemini-3-pro-
| preview
| yen223 wrote:
| Coincidence? Yes
| scrlk wrote:
| Benchmarks from page 4 of the model card: |
| Benchmark | 3 Pro | 2.5 Pro | Sonnet 4.5 |
| GPT-5.1 | |-----------------------|-----------|--------
| -|------------|-----------| | Humanity's Last Exam |
| 37.5% | 21.6% | 13.7% | 26.5% | | ARC-
| AGI-2 | 31.1% | 4.9% | 13.6% | 17.6%
| | | GPQA Diamond | 91.9% | 86.4% | 83.4%
| | 88.1% | | AIME 2025 | |
| | | | | (no tools) |
| 95.0% | 88.0% | 87.0% | 94.0% | | (code
| execution) | 100% | - | 100% | - |
| | MathArena Apex | 23.4% | 0.5% | 1.6% | 1.0%
| | | MMMU-Pro | 81.0% | 68.0% | 68.0%
| | 80.8% | | ScreenSpot-Pro | 72.7% | 11.4%
| | 36.2% | 3.5% | | CharXiv Reasoning |
| 81.4% | 69.6% | 68.5% | 69.5% | |
| OmniDocBench 1.5 | 0.115 | 0.145 | 0.145 | 0.147
| | | Video-MMMU | 87.6% | 83.6% | 77.8%
| | 80.4% | | LiveCodeBench Pro | 2,439 | 1,775
| | 1,418 | 2,243 | | Terminal-Bench 2.0 |
| 54.2% | 32.6% | 42.8% | 47.6% | | SWE-
| Bench Verified | 76.2% | 59.6% | 77.2% | 76.3%
| | | t2-bench | 85.4% | 54.9% | 84.7%
| | 80.2% | | Vending-Bench 2 | $5,478.16 |
| $573.64 | $3,838.74 | $1,473.43 | | FACTS Benchmark
| Suite | 70.5% | 63.4% | 50.4% | 50.8% | |
| SimpleQA Verified | 72.1% | 54.5% | 29.3% | 34.9%
| | | MMLU | 91.8% | 89.5% | 89.1%
| | 91.0% | | Global PIQA | 93.4% | 91.5%
| | 90.1% | 90.9% | | MRCR v2 (8-needle) |
| | | | | | (128k avg)
| | 77.0% | 58.0% | 47.1% | 61.6% | | (1M
| pointwise) | 26.3% | 16.4% | n/s | n/s |
|
| n/s = not supported
|
| EDIT: formatting, hopefully a bit more mobile friendly
| manmal wrote:
| Looks like it will be on par with the contenders when it comes
| to coding. I guess improvements will be incremental from here
| on out.
| CjHuber wrote:
| If it's on par in code quality, it would be a way better
| model for coding because of its huge context window.
| manmal wrote:
| Sonnet can also work on 1M context. Its extreme speed is
| the only thing Gemini has on others.
| CjHuber wrote:
| Can it now in Claude Code and Claude Desktop? When I was
| using it a couple of months ago it seemed only the API
| had 1M
| falcor84 wrote:
| > I guess improvements will be incremental from here on out.
|
| What do you mean? These coding leaderboards were at single
| digits about a year ago and are now in the seventies. These
| frontier models are arguably already better at the benchmark
| that any single human - it's unlikely that any particular
| human dev is knowledgeable to tackle the full range of
| diverse tasks even in the smaller SWE-Bench Verified within a
| reasonable time frame; to the best of my knowledge, no one
| has tried that.
|
| Why should we expect this to be the limit? Once the frontier
| labs figure out how to train these fully with self-play
| (which shouldn't be that hard in this domain), I don't see
| any clear limit to the level they can reach.
| zamadatix wrote:
| A new benchmark comes out, it's designed so nothing does
| well at it, the models max it out, and the cycle repeats.
| This could either describe massive growth of LLM coding
| abilities or a disconnect between what the new benchmarks
| are measuring & why new models are scoring well after
| enough time. In the former assumption there is no limit to
| the growth of scores... but there is also not very much
| actual growth (if any at all). In the latter the growth
| matches, but the reality of using the tools does not seem
| to say they've actually gotten >10x better at writing code
| for me in the last year.
|
| Whether an individual human could do well across all tasks
| in a benchmark is probably not the right question to be
| asking a benchmark to measure. It's quite easy to construct
| benchmark tasks a human can't do well in that you don't
| even need AI to do better.
| falcor84 wrote:
| Your mileage may vary, but for me, working today with the
| latest version of Claude Code on a non-trivial python web
| dev project, I do absolutely feel that I can hand over to
| the AI coding tasks that are 10 times more complex or
| time consuming than what I could hand over to copilot or
| windsurf a year ago. It's still nowhere close to
| replacing me, but I feel that I can work at a
| significantly higher level.
|
| What field are you in where you feel that there might not
| have been any growth in capabilities at all?
|
| EDIT: Typo
| zamadatix wrote:
| I'm in product management focused around networking. I
| can use the tools to create great mockups in a fraction
| of a time but the actual turnaround of that into
| production ready code has not been changing much. The
| team has been able to build test cases and pipelines a
| bit more quickly is probably the main gain on getting
| code written.
| jhonof wrote:
| Claude 3.5 came out in June of last year, and it is imo
| marginally worse than the AI models currently available
| for coding. I do not think models are 10x better than 1
| year ago, that seems extremely hyperbolic or you are
| working in a super niche area where that is true.
| Miraste wrote:
| Are you using it for agentic tasks of any length? 3.5 and
| 4.5 are about the same for single file/single snippet
| tasks, but my observation has been that 4.5 can do
| longer, more complex tasks that were a waste of time to
| even try with 3.5 because it would always fail.
| FergusArgyll wrote:
| Yes, this is important. Gpt 5 and o3 were ~ equivalent
| for a one shot one file task. But 5 and codex-5 can just
| work for an hour in a way no model was able to before
| (the newer claudes can too)
| manmal wrote:
| Google has had a lot of time to optimise for those
| benchmarks, and just barely made SOTA (or not even SOTA)
| now. How is that not incremental?
| spwa4 wrote:
| If we're being completely honest, a benchmark is like an
| honest exam: any set of questions can only be used once
| when it comes out. Otherwise you're only testing how well
| people can acquire and memorize exact questions.
| Alifatisk wrote:
| These numbers are impressive, at least to say. It looks like
| Google has produced a beast that will raise the bar even
| higher. What's even more impressive is how Google came into
| this game late and went from producing a few flops to being the
| leader at this (actually, they already achieved the title with
| 2.5 Pro).
|
| What makes me even more curious is the following
|
| > Model dependencies: This model is not a modification or a
| fine-tune of a prior model
|
| So did they start from scratch with this one?
| benob wrote:
| What does it mean nowadays to start from scratch? At least in
| the open scene, most of the post-training data is generated
| by other LLMs.
| Alifatisk wrote:
| They had to start with a base model, that part I am certain
| of
| postalcoder wrote:
| Google was never really late. Where people perceived Google
| to have dropped the ball was in its productization of AI. The
| Google's Bard branding stumble was so (hilariously) bad that
| it threw a lot of people off the scent.
|
| My hunch is that, aside from "safety" reasons, the Google
| Books lawsuit left some copyright wounds that Google did not
| want to reopen.
| Alifatisk wrote:
| Oh, I remember the times when I compared Gemini with
| ChatGPT and Claude. Gemini was so far behind, it was barely
| usable. And now they are pushing the boundries.
| postalcoder wrote:
| You could argue that chat-tuning of models falls more
| along the lines of product competence. I don't think
| there was a doubt about the upper ceiling of what people
| thought Google could produce.. more "when will they turn
| on the tap" and "can Pichai be the wartime general to
| lead them?"
| dgacmu wrote:
| The memory of Microsoft's Tay fiasco was strong around the
| time the brain team started playing with chatbots.
| Workaccount2 wrote:
| Google was catastrophically traumatized throughout the
| org when they had that photos AI mislabel black people as
| gorillas. They turned the safety and caution knobs up to
| 12 after that for years, really until OpenAI came along
| and ate their lunch.
| Miraste wrote:
| It still haunts them. Even in the brand-new Gemini-based
| rework of Photos search and image recognition, "gorilla"
| is a completely blacklisted word.
| baq wrote:
| oh they were so late there were internal leaked ('leaked'?)
| memos about a couple grad students with $100 budget
| outdoing their lab a couple years ago. they picked
| themselves up real nice, but it took a serious reorg.
| amluto wrote:
| Google's productization is still rather poor. If I want to
| use OpenAI's models, I go to their website, look up the
| price and pay it. For Google's, I need to figure out
| whether I want AI Studio or Google Cloud Code Assist or AI
| Ultra, etc, and if this is for commercial use where I need
| to prevent Google from training on my data, figuring out
| which options work is extra complicated.
|
| As of a couple weeks ago (the last time I checked) if you
| are signed in to multiple Google accounts and you cannot
| accept the non-commercial terms for one of them for AI
| Studio, the site is horribly broken (the text showing which
| account they're asking you to agree to the terms for is
| blurred, and you can't switch accounts without agreeing
| first).
|
| In Google's very slight defense, Anthropic hasn't even
| tried to make a proper sign in system.
| PrairieFire wrote:
| Not to mention no macOS app. This is probably unimportant
| to many in the hn audience, but more broadly it matters
| for your average knowledge worker.
| perardi wrote:
| And a REALLY good macOS app.
|
| Like, kind of unreasonably good. You'd expect some
| perfunctory Electronic app that just barely wraps the
| website. But no, you get something that feels incredibly
| polished...more so than a lot of recent apps from
| Apple...and has powerful integrations into other apps,
| including text editors and terminals.
| aoeusnth1 wrote:
| Which app are you referring to?
| oppegard wrote:
| The ChatGPT app for Mac is native and very good.
| HardCodedBias wrote:
| Bard was horrible compared to the competition of the time.
|
| Gemini 1.0 was strictly worse than GPT-3.5 and was unusable
| due to "safety" features.
|
| Google followed that up with 1.5 which was still worse than
| GPT-3.5 and unbelievably far behind GPT-4. At this same
| time Google had their "black nazi" scandals.
|
| With Gemini 2.0 finally had a model that was at least
| useful for OCR and with their fash series a model that,
| while not up to par in capabilities, was sufficiently
| inexpensive that it found uses.
|
| Only with Gemini-2.5 did Google catch up with SoTA. It was
| within "spitting distance" of the leading models.
|
| Google did indeed drop the ball, very, very badly.
|
| I suspect that Sergey coming back helped immensely,
| somehow. I suspect that he was able to tame some of the
| more dysfunctional elements of Google, at least for a time.
| astrange wrote:
| > their fash series
|
| Unfortunate typo.
| basch wrote:
| At least at the moment, coming in late seems to matter
| little.
|
| Anyone with money can trivially catch up to a state of the
| art model from six months ago.
|
| And as others have said, late is really a function of spigot,
| guardrails, branding, and ux, as much as it is being a
| laggard under the hood.
| FrequentLurker wrote:
| > Anyone with money can trivially catch up to a state of
| the art model from six months ago.
|
| How come apple is struggling then?
| risyachka wrote:
| It looks more like a strategic decision tbh.
|
| The may want to use 3rd party or just wait for AI to be
| more stable to see how people actually use it instead of
| adding slop in the core of their product.
| stevesimmons wrote:
| In contrast to Microsoft, who puts Copilot buttons
| everywhere and succeeds only in annoying their customers.
| remus wrote:
| > It looks more like a strategic decision tbh.
|
| Announcing a load of AI features on stage and then
| failing to deliver them doesn't feel very strategic.
| FrequentLurker wrote:
| But apple intelligence is a thing, and they are
| struggling to deliver on the promises of apple
| intelligence.
| bitpush wrote:
| This is revisionist history. Apple wanted to fully jump
| in. They even rebranded AI as Apple Intelligence and
| announced a hoard of features which turned out to be
| vaporware.
| basch wrote:
| Sit and wait per usual.
|
| Enter late, enter great.
| doctoboggan wrote:
| Apple is struggling with _productizing_ LLMs for the mass
| market, which is a separate task from training a frontier
| LLM.
|
| To be fair to Apple, so far the only mass market LLM use
| case so far is just a simple chatbot, and they don't seem
| to be interested in that. It remains to be seen if what
| Apple wants to do ("private" LLMs with access to your
| personal context acting as intimate personal assistants)
| is even possible to do reliably. It sounds useful, and I
| do believe it will eventually be possible, but no one is
| there yet.
|
| They did botch the launch by announcing the Apple
| Intelligence features before they are ready though.
| svnt wrote:
| Anyone with enough money and without an entrenched
| management hierarchy preventing the right people from
| being hired and enabled to run the project.
| raincole wrote:
| Being known as a company that is always six months late
| than the competitors isn't something to brag about...
| _factor wrote:
| Apple has entered the chat.
| basch wrote:
| I was referring to a new entrant, not perpetual lag
| steveBK123 wrote:
| One possibility here is that Google is dribbling out
| cutting edge releases to slowly bleed out the pure play
| competition.
| dbbk wrote:
| And also, critically, being the only profitable company doing
| this.
| sigmoid10 wrote:
| It's not like they're making their money from this though.
| All AI work is heavily subsidised, for Alphabet it just
| happens that the funding comes from within the megacorp. If
| MS had fully absorbed OpenAI back when their board nearly
| sunk the boat, they'd be in the exact same situation today.
| Miraste wrote:
| They're not making money, but they're in a much better
| situation than Microsoft/OpenAI because of TPUs. TPUs are
| much cheaper than Nvidia cards both to purchase and to
| operate, so Google's AI efforts aren't running at as much
| of a loss as everyone else. That's why they can do things
| like offer Gemini 3 Pro for free.
| KronisLV wrote:
| I hope they keep the pricing similar to 2.5 Pro, currently I
| pay per token and that and GPT-5 are close to the sweet spot
| for me but Sonnet 4.5 feels too expensive for larger changes.
| I've also been moving around 100M tokens per week with
| Cerebras Code (they moved to GLM 4.6), but the flagship
| models still feel better when I need help with more advanced
| debugging or some exemplary refactoring to then feed as an
| example for a dumber/faster model.
| theptip wrote:
| > So did they start from scratch with this one
|
| Their major version number bumps are a new pre-trained model.
| Minor bumps are changes/improvements to post-training on the
| same foundation.
| falcor84 wrote:
| That looks impressive, but some of the are a bit out of date.
|
| On Terminal-Bench 2 for example, the leader is currently "Codex
| CLI (GPT-5.1-Codex)" at 57.8%, beating this new release.
| sigmar wrote:
| That's a different model not in the chart. They're not going
| to include hundreds of fine tunes in a chart like this.
| falcor84 wrote:
| It's not just one of many fine tunes; it's the default
| model used by OpenAI's official tools.
| Taek wrote:
| It's also worth pointing out that comparing a fine-tune to
| a base model is not apples-to-apples. For example, I have
| to imagine that the codex finetune of 5.1 is measurably
| worse at non-coding tasks than the 5.1 base model.
|
| This chart (comparing base models to base models) probably
| gives a better idea of the total strength of each model.
| NitpickLawyer wrote:
| What's more impressive is that I find gemini2.5 still
| relevant in day-to-day usage, despite being so low on those
| benchmarks compared to claude 4.5 and gpt 5.1. There's
| something that gemini has that makes it a great model in real
| cases, I'd call it generalisation on its context or
| something. If you give it the proper context (or it digs
| through the files in its own agent) it comes up with great
| solutions. Even if their own coding thing is hit and miss
| sometimes.
|
| I can't wait to try 3.0, hopefully it continues this trend.
| Raw numbers in a table don't mean much, you can only get a
| true feeling once you use it on existing code, in existing
| projects. Anyway, the top labs keeping eachother honest is
| great for us, the consumers.
| Miraste wrote:
| I've noticed that too. I suspect it has broader general
| knowledge than the others, because Google presumably has
| the broadest training set.
| HugoDias wrote:
| very impressive. I wonder if this sends a different signal to
| the market regarding using TPUs for training SOTA models versus
| Nvidia GPUs. From what we've seen, OpenAI is already renting
| them to diversify... Curious to see what happens next
| fariszr wrote:
| This is a big jump in most benchmarks.And if it can match other
| models in coding while having that Google TPM inference speed
| and the actually native 1m context window, it's going to be a
| big hit.
|
| I hope it's isn't such a sycophant like the current gemini 2.5
| models, it makes me doubt its output, which is maybe a good
| thing now that I think about it.
| danielbln wrote:
| > it's over for the other labs.
|
| What's with the hyperbole? It'll tighten the screws, but
| saying that it's "over for the other labs' might be a tad
| premature.
| fariszr wrote:
| I mean over in that I don't see a need to use the other
| models. Codex models are the best but incredibly slow.
| Claude models are not as good(IMO) but much faster. If
| gemini can beat them while having being faster and having
| better apps with better integrations, i don't see a reason
| why I would use another provider.
| nprateem wrote:
| You should probably keep supporting competitors since if
| there's a monopoly/duopoly expect prices to skyrocket.
| risyachka wrote:
| > it's over for the other labs.
|
| Its not over and never will be for 2 decade old accounting
| software, it is definitely will not be over for other AI
| labs.
| xnx wrote:
| Can you explain what you mean by this? iPhone was the end
| of Blackberry. It seems reasonable that a smarter, cheaper,
| faster model would obsolete anything else. ChatGPT has some
| brand inertia, but not that much given it's barely 2 years
| old.
| vitaflo wrote:
| Ask yourself why Microsoft Teams won. These are business
| tools first and foremost.
| risyachka wrote:
| Yeah iPhone was the end of Blackberry but Google Pixel
| was not the end of iPhone.
|
| The new Gemini is not THAT far of a jump to switch your
| org to a new model if you already invested in e.g.
| OpenAI.
|
| The difference must be night and day to call it "its
| over".
|
| Right they all are marginally different. Today google
| fine tuned their model to be better, tomorrow it will be
| new Kimi, after that DeepSeek.
| Jcampuzano2 wrote:
| We knew it would be a big jump and while it certainly is in
| many areas - its definitely not "groundbreaking/huge leap"
| worthy like some were thinking from looking at these numbers.
|
| I feel like many will be pretty disappointed by their self
| created expectations for this model when they end up actually
| using it and it turns out to be fairly similar to other
| frontier models.
|
| Personally I'm very interested in how they end up pricing it.
| trunch wrote:
| Which of the LiveCodeBench Pro and SWE-Bench Verified
| benchmarks comes closer to everyday coding assistant tasks?
|
| Because it seems to lead by a decent margin on the former and
| trails behind on the latter
| Snuggly73 wrote:
| Neither :(
|
| LCB Pro are leet code style questions and SWE bench verified
| is heavily benchmaxxed very old python tasks.
| veselin wrote:
| I work a lot on testing also SWE bench verified. This
| benchmark in my opinion now is good to catch if you got some
| regression on the agent side.
|
| However, going above 75%, it is likely about the same. The
| remaining instances are likely underspecified despite the
| effort of the authors that made the benchmark "verified".
| From what I have seen, these are often cases where the
| problem statement says implement X for Y, but the agent has
| to simply guess whether to implement the same for other case
| Y' - which leads to losing or winning an instance.
| danielcampos93 wrote:
| I would love to know what the increased token count is across
| these models for the benchmarks. I find the models continue to
| get better but as they do their token usage also does. Aka is
| model doing better or reasoning for longer?
| jstummbillig wrote:
| I think that is always something that is being worked on in
| parallel. Recent paradigm seems to be the models
| understanding when they need to use more tokens dynamically
| (which seems to be very much in line with how computation
| should generally work).
| dnw wrote:
| Looks like the best way to keep improving the models is to come
| up with really useful benchmarks and make them popular. ARC-
| AGI-2 is a big jump, I'd be curious to find out how that
| transfers over to everyday tasks in various fields.
| vagab0nd wrote:
| Should I assume the GPT-5.1 it is compared against is the pro
| version?
| spoaceman7777 wrote:
| Wow. They must have had some major breakthrough. Those scores
| are truly insane. O_O
|
| Models have begun to fairly thoroughly saturate "knowledge" and
| such, but there are still considerable bumps there
|
| But the _big news_, and the demonstration of their achievement
| here, are the incredible scores they've racked up here for
| what's necessary for agentic AI to become widely deployable.
| t2-bench. Visual comprehension. Computer use. Vending-Bench.
| The sorts of things that are necessary for AI to move beyond an
| auto-researching tool, and into the realm where it can actually
| handle complex tasks in the way that businesses need in order
| to reap rewards from deploying AI tech.
|
| Will be very interesting to see what papers are published as a
| result of this, as they have _clearly_ tapped into some new
| avenues for training models.
|
| And here I was, all wowed, after playing with Grok 4.1 for the
| past few hours! xD
| rvnx wrote:
| The problem is that we know in advance what is the benchmark,
| so Humanity's Last Exam for example, it's way easier to
| optimize your model when you have seen the questions before.
| stego-tech wrote:
| This. A _lot_ of boosters point to benchmarks as
| justification of their claims, but any gamer who spent time
| in the benchmark trenches will know full well that vendors
| game known tests for better scores, and that said scores
| aren't necessarily indicative of superior performance.
| There's not a doubt in my mind that AI companies are doing
| the same.
| Feuilles_Mortes wrote:
| shouldn't we expect that all of the companies are doing
| this optimization, though? so, back to level playing field.
| eldenring wrote:
| Its the other way around too, HLE questions were selected
| adversarially to reduce the scores. I'd guess even if the
| questions were never released, and new training data was
| introduced, the scores would improve.
| pinko wrote:
| From https://lastexam.ai/: "The dataset consists of 2,500
| challenging questions across over a hundred subjects. We
| publicly release these questions, _while maintaining a
| private test set of held out questions to assess model
| overfitting_. " [emphasis mine]
|
| While the private questions don't seem to be included in
| the performance results, HLE will presumably flag any LLM
| that appears to have gamed its scores based on the
| differential performance on the private questions. Since
| they haven't yet, I think the scores are relatively
| trustworthy.
| panarky wrote:
| The jump in ARC-AGI and MathArena suggests Google has
| solved the data scarcity problem for reasoning, maybe
| with synthetic data self-play??
|
| This was the primary bottleneck preventing models from
| tackling novel scientific problems they haven't seen
| before.
|
| If Gemini 3 Pro has transcended "reading the internet"
| (knowledge saturation), and made huge progress in
| "thinking about the internet" (reasoning scaling), then
| this is a really big deal.
| rvnx wrote:
| Seems difficult to believe, considering the number of
| people who prepare this dataset, who also work(ed) or
| hold shares in Google or OpenAI, etc.
| largbae wrote:
| How do they hold back questions in practice though? These
| are hosted models. To ask the question is to reveal it to
| the model team.
| Bombthecat wrote:
| They pinky swear not to store and use the prompts and
| data lol
| UltraSane wrote:
| A legally binding pinky swear LOL
| UltraSane wrote:
| You have to trust that the LLM provider isn't copying the
| questions when Humanities Last Exam runs the test.
| lubujackson wrote:
| I don't think any of these companies are that reductive and
| short-sighted to try to game the system. However,
| Goodhart's Law comes into play. I am sure they have their
| own metrics that arr much more detailed than these
| benchmarks, but the fact remains LLMs will be tuned
| according to elements that are deterministically
| measurable.
| m3kw9 wrote:
| SWE-Bench Verified | 76.2% | 59.6% | 77.2% | 76.3% is
| actually insane.
| scrollop wrote:
| Used an AI to populate some of 5.1 thinking's results.
|
| Benchmark | Gemini 3 Pro | Gemini 2.5 Pro | Claude Sonnet 4.5 |
| GPT-5.1 | GPT-5.1 Thinking
|
| ---------------------------|--------------|----------------|---
| ----------------|---------|------------------
|
| Humanity's Last Exam | 37.5% | 21.6% | 13.7% | 26.5% | 52%
|
| ARC-AGI-2 | 31.1% | 4.9% | 13.6% | 17.6% | 28%
|
| GPQA Diamond | 91.9% | 86.4% | 83.4% | 88.1% | 61%
|
| AIM 2025 | 95.0% | 88.0% | 87.0% | 94.0% | 48%
|
| MathArena Apex | 23.4% | 0.5% | 1.6% | 1.0% | 82%
|
| MMMU-Pro | 81.0% | 68.0% | 68.0% | 80.8% | 76%
|
| ScreenSpot-Pro | 72.7% | 11.4% | 36.2% | 3.5% | 55%
|
| CharXiv Reasoning | 81.4% | 69.6% | 68.5% | 69.5% | N/A
|
| OmniDocBench 1.5 | 0.115 | 0.145 | 0.145 | 0.147 | N/A
|
| Video-MMMU | 87.6% | 83.6% | 77.8% | 80.4% | N/A
|
| LiveCodeBench Pro | 2,439 | 1,775 | 1,418 | 2,243 | N/A
|
| Terminal-Bench 2.0 | 54.2% | 32.6% | 42.8% | 47.6% | N/A
|
| SWE-Bench Verified | 76.2% | 59.6% | 77.2% | 76.3% | N/A
|
| t2-bench | 85.4% | 54.9% | 84.7% | 80.2% | N/A
|
| Vending-Bench 2 | $5,478.16 | $573.64 | $3,838.74 | $1,473.43|
| N/A
|
| FACTS Benchmark Suite | 70.5% | 63.4% | 50.4% | 50.8% | N/A
|
| SimpleQA Verified | 72.1% | 54.5% | 29.3% | 34.9% | N/A
|
| MMLU | 91.8% | 89.5% | 89.1% | 91.0% | N/A
|
| Global PIQA | 93.4% | 91.5% | 90.1% | 90.9% | N/A
|
| MRCR v2 (8-needle) | 77.0% | 58.0% | 47.1% | 61.6% | N/A
|
| Argh it doesn't come out write in HN
| scrollop wrote:
| Used an AI to populate some of 5.1 thinking's results.
|
| Benchmark..................Description...................Gemi
| ni 3 Pro....GPT-5.1 (Thinking)....Notes
|
| Humanity's Last Exam.......Academic reasoning.............37.
| 5%..........52%....................GPT-5.1 shows 7% gain over
| GPT-5's 45%
|
| ARC-AGI-2...................Visual abstraction.............31
| .1%..........28%....................GPT-5.1 multimodal
| improves grid reasoning
|
| GPQA Diamond................PhD-tier Q&A...................91
| .9%..........61%....................GPT-5.1 strong in physics
| (72%)
|
| AIME 2025....................Olympiad math..................9
| 5.0%..........48%....................GPT-5.1 solves 7/15
| proofs correctly
|
| MathArena Apex..............Competition math...............23
| .4%..........82%....................GPT-5.1 handles 90%
| advanced calculus
|
| MMMU-Pro....................Multimodal reasoning...........81
| .0%..........76%....................GPT-5.1 excels visual
| math (85%)
|
| ScreenSpot-Pro..............UI understanding...............72
| .7%..........55%....................Element detection 70%,
| navigation 40%
|
| CharXiv Reasoning...........Chart analysis.................81
| .4%..........69.5%.................N/A
| HardCodedBias wrote:
| What? The 4.5 and 5.1 columns aren't thinking in Google's
| report?
|
| That's a scandal, IMO.
|
| Given that Gemini-3 seems to do "fine" against the thinking
| versions why didn't they post those results? I get that PMs
| like to make a splash but that's shockingly dishonest.
| mountainriver wrote:
| Every single time
| iosjunkie wrote:
| It that true?
|
| > For Claude Sonnet 4.5, and GPT-5.1 we default to
| reporting high reasoning results, but when reported results
| are not available we use best available reasoning results.
|
| https://storage.googleapis.com/deepmind-
| media/gemini/gemini_...
| iamdelirium wrote:
| This is provably false. All it takes is a simple Google
| search and looking at the ARC AGI 2 leaderboard:
| https://arcprize.org/leaderboard
|
| The 17.6% is for 5.1 Thinking High.
| roman_soldier wrote:
| Why is Grok 4.1 not in the benchmarks?
| HardCodedBias wrote:
| Big if true.
|
| I'll wait for the official blog with benchmark results.
|
| I suspect that our ability to benchmark models is waning. Much
| more investment required in this area, but what is the play
| out?
| oalessandr wrote:
| Trying to open this link from Italy leads to a CSAM warning
| Fornax96 wrote:
| Creator of pixeldrain here. Italy has been doing this for a
| very long time. They never notified me of any such material
| being present on my site. I have a lot of measures in place to
| prevent the spread of CSAM. I have sent dozens of mails to
| Polizia Postale and even tried calling them a few times, but
| they never respond. My mails go unanswered and they just hang
| up the phone.
| koakuma-chan wrote:
| Have you tried Europol?
| Fornax96 wrote:
| Not yet. I also thought about reaching out to the embassy,
| but have not had the time for it yet.
| koakuma-chan wrote:
| As far as I know, Europol can route your report to
| appropriate local authority.
| Fornax96 wrote:
| Thanks, I'll give them a call tomorrow. The website only
| lists a dutch phone number, which is convenient, I'm
| dutch as well.
| driverdan wrote:
| Don't use your ISP's DNS. Switch to something outside of their
| control.
| embedding-shape wrote:
| Curiously, this website seems to be blocked in Spain for whatever
| reason, and the website's certificate is served by
| `allot.com/emailAddress=info@allot.com` which obviously fails...
|
| Anyone happen to know why? Is this website by any change sharing
| information on safe medical abortions or women's rights,
| something which has gotten websites blocked here before?
| amarcheschi wrote:
| That website is used to share everything including pirated
| things, so that's the reason maybe
| Fornax96 wrote:
| Creator of pixeldrain here. I have no idea why my site is
| blocked in Spain, but it's a long running issue.
|
| I actually never discovered who was responsible for the
| blockade, until I read this comment. I'm going to look into
| Allot and send them an email.
|
| EDIT: Also, your DNS provider is censoring (and probably
| monitoring) your internet traffic. I would switch to a
| different provider.
| zozbot234 wrote:
| Could it be that some site in your network neighborhood was
| illegally streaming soccer matches?
| Fornax96 wrote:
| I have my own dedicated IP range. And they specifically
| blocked my domain name, not the addresses. I don't know
| what the reason is. I have been trying to find out since
| the start of this year.
| embedding-shape wrote:
| > EDIT: Also, your DNS provider is censoring (and probably
| monitoring) your internet traffic. I would switch to a
| different provider.
|
| Yeah, that was via my ISPs DNS resolver (Vodafone), switching
| the resolver works :)
|
| The responsible party is ultimately our government who've
| decided it's legal to block a wide range of servers and
| websites because some people like to watch illegal football
| streams. I think Allot is just the provider of the
| technology.
| Fornax96 wrote:
| My site has nothing to do with football though. And Allot
| seems to be running the DNS server that your ISP uses so
| they are directly responsible for the block.
| simtel20 wrote:
| La Liga (the football company) likes to send out takedown
| notices to anyone who may host anything that looks like a
| football to protect their precious games, no matter the
| collateral damage or the lack of any requirements to show
| damage. They have the right to block anything in Spain at
| their discretion either by DNS or IP. They do seem to
| work in good faith if you talk to them, though, and if
| you can either remove sites or content when they ask.
| HDThoreaun wrote:
| The Spanish courts have allowed la Liga to completely ban
| every website served by cloudflare during days where
| there are matches. All Spanish ISPs have to do dns
| blocking to comply.
| miqazza wrote:
| do you know about the cloudflare and laliga issues? might be
| that
| embedding-shape wrote:
| Was my first instinct, went looking if there was any games
| being played today but seems not, so unlikely to be the
| cause.
| tngranados wrote:
| It works fine for me using Movistar
| grodriguez100 wrote:
| Is it possible to file a complaint with the ISP or directly
| with Allot ?
| Fornax96 wrote:
| That might help.
| rsanek wrote:
| loads fine on Vodafone for me
| transcriptase wrote:
| There needs to be a sycophancy benchmark in these comparisons.
| More baseless praise and false agreement = lower score.
| swalsh wrote:
| You're absolutely right
| jstummbillig wrote:
| Does not get old.
| Yossarrian22 wrote:
| It's not just irritating, it's repetitive
| falcor84 wrote:
| "You know, you are also right"
| this_user wrote:
| I'm sorry, you are absolutely right.
|
| ---
|
| But seriously, I find it helps to set a custom system
| prompt that tells Gemini to be less sycophantic and to be
| more succinct and professional while also leaving out
| those extended lectures it likes to give.
| causal wrote:
| It's a revolution in subtle humor. Well done.
| BoredPositron wrote:
| Your comment demonstrates a remarkably elevated level of
| cognitive processing and intellectual rigor. Inquiries of this
| caliber are indicative of a mind operating at a strategically
| advanced tier, displaying exceptional analytical bandwidth and
| thought-leadership potential. Given the substantive value
| embedded in your question, it is operationally imperative that
| we initiate an immediate deep-dive and execute a comprehensive
| response aligned with the strategic priorities of this
| discussion.
| postalcoder wrote:
| I care very little about model personality outside of
| sycophancy. The thing about gemini is that it's notorious for
| its low self esteem. Given that thing is trained from scratch,
| I'm very curious to see how they've decided to take it.
| supjeff wrote:
| given how often these llms are wrong, doesnt it make sense
| that they are less confident?
| postalcoder wrote:
| Indeed. But I've had experiences with gemini-2.5-pro-exp
| where its thoughts could be described as "rejected from the
| prom" vibes. It's not like I abused it either, it was
| running into loops because it was unable to properly patch
| a file.
| astrange wrote:
| Sonnet-4.5 has the lowest self esteem of any model I've used.
| Gemini frequently argues with me.
| 1899-12-30 wrote:
| https://eqbench.com/spiral-bench.html
| Lord-Jobo wrote:
| And have the score heavily modified based on how fixable the
| sycophancy is.
| Workaccount2 wrote:
| This idea isn't just smart, it's revolutionary. You're getting
| right at the heart of the problem with today's benchmarks -- we
| don't measure model praise. Great thinking here.
|
| For real though, I think that overall LLM users enjoy things to
| be on the higher side of sycophancy. Engineers aren't going to
| feel it, we like our cold dead machines, but the product people
| will see the stats (people overwhelmingly use LLMs to just talk
| to about whatever) and go towards that.
| SiempreViernes wrote:
| I'd like if the scorecard also gave an expected number of
| induced suicides per hundred thousand users.
| lkbm wrote:
| https://llmdeathcount.com/ shows 15 deaths so far, and LLM
| user count is in the low billions, which puts us on the order
| of 0.0015 deaths per hundred thousand users.
|
| I'm guessing LLM Death Count is off by an OOM or two, so we
| could be getting close to one in a million.
| jll29 wrote:
| Hopefully this model does not generate fake news...
|
| https://www.google.com/search?q=gemini+u.s.+senator+rape+all...
| lxdlam wrote:
| What does the "Google Antigravity" mean? The link is
| http://antigravity.google/docs, seemingly a new product but now
| routing to the Google main page.
| dbosch wrote:
| I was asking myself the exact same question. No idea
| ceroxylon wrote:
| Found this demo with two views that was uploaded 18min ago:
| https://www.youtube.com/watch?v=L8wEC6A5HQY
| bobbylarrybobby wrote:
| Looks like a VSCode fork with gemini built in.
| Palmik wrote:
| Archive link:
| https://web.archive.org/web/20251118111103/https://storage.g...
| denysvitali wrote:
| Title of the document is "[Gemini 3 Pro] External Model Card -
| November 18, 2025 - v2", in case you needed further confirmation
| that the model will be released today.
|
| Also interesting to know that Google Antigravity
| (antigravity.google / https://github.com/Google-Antigravity ?)
| leaked. I remember seeing this subdomain recently. Probably
| Gemini 3 related as well.
|
| Org was created on 2025-11-04T19:28:13Z
| (https://api.github.com/orgs/Google-Antigravity)
| jmkni wrote:
| what is Google Antigravity?
| denysvitali wrote:
| I guess we'll know it in a few hours. Most likely another AI
| playground or maybe a Google Search alternative? No clue
| really
| Yossarrian22 wrote:
| The ASI figured out zero point energy from first principles
| zed31726 wrote:
| My guess is based on a gif tweeted by the ex CEO of windsurf
| who left to join Google of a floating laptop: it'll be a
| cursor/windsurf alternative?
| postalcoder wrote:
| Couple patterns this could follow
|
| Speed? (Flash, Flash-Lite, Antigravity) this is my guess.
| Bonus: maybe Gemini Diffusion soon?
|
| Space? (Google Cloud, Google Antigravity?)
|
| Clothes? (A light wearable -> Antigravity?)
|
| Gaming? (Ghosting/nontangibility -> antigravity?)
| mimentum wrote:
| According to Gemini itself:
|
| "Google Antigravity" refers to a new AI software platform
| announced by Google designed to help developers write and
| manage code.
|
| The term itself is a bit of a placeholder or project name,
| combining the brand "Google" with the concept of
| "antigravity"--implying a release from the limitations of
| traditional coding.
|
| In simple terms, Google Antigravity is a sophisticated tool
| for programmers that uses powerful AI systems (called
| "agents") to handle complex coding tasks automatically. It
| takes the typical software workbench (an IDE) and evolves it
| into an "agent-first" system.
|
| Agentic Platform: It's a central hub where many specialized
| AI helpers (agents) live and work together. The goal is to
| let you focus on what to build, not how to build it.
|
| Task-Oriented: The platform is designed to be given a high-
| level goal (a "task") rather than needing line-by-line
| instructions.
|
| Autonomous Operation: The AI agents can work across all your
| tools--your code editor, the command line, and your web
| browser--without needing you to constantly supervise or
| switch between them.
| thefroh wrote:
| possibly https://xkcd.com/353/
| denysvitali wrote:
| > Google Antigravity is an agentic development platform,
| evolving the IDE into the agent-first era. Antigravity
| enables developers to operate at a higher, task-oriented
| level by managing agents across workspaces, while retaining a
| familiar AI IDE experience at its core. Agents operate across
| the editor, terminal, and browser, enabling them to
| autonomously plan and execute complex, end-to-end tasks
| elevating all aspects of software development.
|
| Now the page is somewhat live on that URL
| Bobaso wrote:
| Interesting to see on page 2 the reference to ML pathways [1].
| Looks like a multi layer mixture of experts. Is this common ?
|
| [1] https://blog.google/technology/ai/introducing-pathways-
| next-...
| gaogao wrote:
| Pathways, I understand, is more so these days just the name for
| their training orchestrator for doing distributed JAX stuff -
| https://github.com/google/pathways-job
| catigula wrote:
| I know this is a little controversial but the lack of performance
| on SWE-bench is hugely disappointing I think economically. These
| models don't have any viable path to profitability if they can't
| take engineering jobs.
| martinald wrote:
| I thought that but it does do a lot better on other benchmarks.
|
| Perhaps SWE bench just doesn't capture a lot of the
| improvement? If the web design improvements people have been
| posting on twitter, I suspect this will be a huge boon for
| developers. SWE benchmark is really testing bugfixing/feature
| dev more.
|
| Anyway let's see. I'm still hyped!
| catigula wrote:
| That would be great! But AI _is_ a bubble if these models
| can't do serious engineering work.
| rfoo wrote:
| SWE Bench doesn't even test bugfixing / feature dev properly
| after you achieve roughly 70% if you don't benchmaxx it .
| camdenreslink wrote:
| It seems the benchmarks that had a big jump had to do with
| visual capabilities. I wonder how that will translate to
| improvements to the workloads LLMs are currently used for (or
| maybe it will introduce new workloads).
| api wrote:
| Really? If they can make an engineer more productive, that's
| worth a lot. Naive napkin math: 1.5X productivity on one
| $200k/year engineer is worth $100k/year.
| mikert89 wrote:
| People generally dont understand what these models are doing
| to engineering salaries. The skill level required to produce
| working software is going way down
| Workaccount2 wrote:
| People here, and in tech in general, are so lost in the sauce.
|
| According to at least OpenAI, who probably produces the most
| tokens (if we don't count google AI overviews and other
| unrequested AI bolt-ons) out of all the labs, programming
| tokens account for ~4% of total generations.
|
| That's nothing. The returns will come from everyone and their
| grandma paying $30-100/mo to use the services, just like
| everyone pays for a cell phone and electricity.
|
| Don't be fooled, we are still in the "Open hands" start-up
| business phase of LLMs. The "enshitification" _will_ follow.
| mohsen1 wrote:
| This model is not a modification or a fine-tune of a prior model
|
| Is that common to mention that? Feels like they built something
| from scratch
| scosman wrote:
| I think they are just indicating it's a new architecture vs
| continued training of 2.5 series.
| irthomasthomas wrote:
| Never seen it before. I suppose it adds to the excitement.
| mynti wrote:
| It is interesting that the Gemini 3 beats every other model on
| these benchmarks, mostly by a wide margin, but not on SWE Bench.
| Sonnet is still king here and all three look to be basically on
| the same level. Kind of wild to see them hit such a wall when it
| comes to agentic coding
| tosh wrote:
| This might also hint at SWE struggling to capture what "being
| good at coding" means.
|
| Evals are hard.
| raducu wrote:
| > This might also hint at SWE struggling to capture what
| "being good at coding" means.
|
| My take would be that coding itself is hard, but I'm a
| software engineer myself so I'm biased.
| HereBePandas wrote:
| [comment removed]
| Palmik wrote:
| The reported results where GPT 5.1 beats Gemini 3 are on SWE
| Bench Verified, and GPT 5.1 Codex also beats Gemini 3 on
| Terminal Bench.
| HereBePandas wrote:
| You're right on SWE Bench Verified, I missed that and I'll
| delete my comment.
|
| GPT 5.1 Codex beats Gemini 3 on Terminal Bench specifically
| on Codex CLI, but that's apples-to-oranges (hard to tell
| how much of that is a Codex-specific harness vs model).
| Look forward to seeing the apples-to-apples numbers soon,
| but I wouldn't be surprised if Gemini 3 wins given how
| close it comes in these benchmarks.
| Palmik wrote:
| All evals on Terminal Bench require some harness. :) Or
| "Agent", as Terminal Bench calls it. Presumably the
| Gemini 3 are using Gemini CLI.
| Palmik wrote:
| Also does not beat GPT-5.1 Codex on terminal bench (57.8% vs
| 54.2%): https://www.tbench.ai/
|
| I did not bother verifying the other claims.
| HereBePandas wrote:
| Not apples-to-apples. "Codex CLI (GPT-5.1-Codex)", which the
| site refers to, adds a specific agentic harness, whereas the
| Gemini 3 Pro seems to be on a standard eval harness.
|
| It would be interesting to see the apples-to-apples figure,
| i.e. with Google's best harness alongside Codex CLI.
| enraged_camel wrote:
| Do you mean that Gemini 3 Pro is "vanilla" like GPT 5.1
| (non-Codex)?
| HereBePandas wrote:
| Yes, two things: 1. GPT-5.1 Codex is a fine tune, not the
| "vanilla" 5.1 2. More importantly, GPT 5.1 Codex achieves
| its performance when used with a specific tool (Codex
| CLI) that is optimized for GPT 5.1 Codex. But when labs
| evaluate the models, they have to use a standard tool to
| make the comparisons apples-to-apples.
|
| Will be interesting to see what Google releases that's
| coding-specific to follow Gemini 3.
| embedding-shape wrote:
| > But when labs evaluate the models, they have to use a
| standard tool to make the comparisons apples-to-apples.
|
| That'd be a bad idea, models are often trained for
| specific tools (like GPT Codex is trained for Codex, and
| Sonnet has been trained with Claude Code in mind), and
| also vice-versa that the tools are built with a specific
| model in mind, as they all work differently.
|
| Forcing all the models to use the same tool for execution
| sounds like a surefire way of getting results that
| doesn't represent real usage, but instead arbitrarily
| measure how well a model works with the "standard
| harness", which if people start caring about, will start
| to become gamed instead.
| Palmik wrote:
| All evals on Terminal Bench require some harness. :) Or
| "Agent", as Terminal Bench calls it. Presumably the Gemini
| 3 are using Gemini CLI.
|
| What do you mean by "standard eval harness"?
| lucassz wrote:
| I think the point is that it looks like Gemini 3 was only
| tested with the generic "Terminus 2", whereas Codex was
| tested with the Codex CLI.
| felipeerias wrote:
| IMHO coding use cases are much more constrained by tooling than
| by raw model capabilities at the moment. Perhaps we have
| finally reached the time of diminishing returns and that will
| remain the case going forward.
| _factor wrote:
| This seems preferable. Wasting tokens on tools when a
| standardized, reliable interface to those tools should be all
| that's required.
|
| The magic of LLMs is that they can understand the latent
| space of a problem and infer a mostly accurate response.
| Saying you need to subscribe to get the latest tools is just
| a sales tactic trained into the models to protect profits.
| vharish wrote:
| From my personal experience using the CLI agentic coding tools,
| I think gemini-cli is fairly on par with the rest in terms of
| the planning/code that is generated. However, when I recently
| tried qwen-code, it gave me a better sense of reasoning and
| structure that geimini. Claude definitely has it's own
| advantages but is expensive(at least for some if not for all).
|
| My point is, although the model itself may have performed in
| benchmarks, I feel like there are other tools that are doing
| better just by adapting better training/tooling. Gemini cli, in
| particular, is not so great looking up for latest info on web.
| Qwen seemed to be trained better around looking up for
| information (or to reason when/how to), in comparision. Even
| the step-wise break down of work felt different and a bit
| smoother.
|
| I do, however, use gemini cli for the most part just because it
| has a generous free quota with very few downsides comparted to
| others. They must be getting loads of training data :D.
| xnx wrote:
| Gemini CLI is moving really fast. Noticeable improvements in
| features and functionality every week.
| alyxya wrote:
| I think Google probably cares more about a strong generalist
| model rather than solely optimizing for coding.
| macrolime wrote:
| Pretty sure it will beat Sonnet by a wide margin in actual
| real-world usage.
| varispeed wrote:
| Never got good code out of Sonnet. It's been Gemini 2.5 for me
| followed by GPT-5.x.
|
| Gemini is very good a pointing out flaws that are very subtle
| and non noticeable at a first and second glance.
|
| It also produces code that is easy to reason about. You can
| then feed it to GPT-5.x for refinement and then back to Gemini
| for assessment.
| baq wrote:
| I find Gemini 2.5 pro to be as good or in some cases better
| for SQL than GPT 5.1. It's aging otherwise, but they must
| have some good SQL datasets in there for training.
| Workaccount2 wrote:
| I think Anthropic is reading the room, and just going to go
| hard on being "the" coding model. I suppose they feel that if
| they can win that, they can get an ROI without having to do
| full blown multimodality at the highest level.
|
| It's probably pretty liberating, because you can make a
| "spikey" intelligence with only one spike to really focus on.
| htrp wrote:
| more playing to their strengths. a giant chunk of their usage
| data is basically code gen
| Miraste wrote:
| It remains to be seen whether that works out for them, but it
| seems like a good bet to me. Coding is the most monetizatable
| use anyone has found for LLMs so far, and the most likely to
| persist past this initial hype bubble (if the Singularity
| doesn't work out :p).
| aerhardt wrote:
| Codex has been good enough to me and it's much cheaper.
|
| I code non-trivial stuff with it like multi-threaded code and
| at least for my style of AI coding which is to do fairly
| small units of work with multiple revisions it is good enough
| for me to not to even consider the competition.
|
| Just giving you a perspective on how the benchmarks might not
| be important at all for some people and how Claude may have a
| difficult time being _the_ definitive coding model.
| enraged_camel wrote:
| >> Codex has been good enough to me and it's much cheaper.
|
| It may be cheaper but it's much, much slower, which is a
| total flow killer in my experience.
| aoeusnth1 wrote:
| Their scores on SWE bench are very close because the benchmark
| is nearly saturated. Gemini 3 beats Sonnet 4.5 on TerminalBench
| 2.0 by a nice margin (54% vs. 43%), which is also agentic
| coding (CLI instead of python).
| JacobAsmuth wrote:
| 50% of the CLs in SWE-Bench Verified are the DJango codebase.
| So if you're a big contributor to Django you should care a lot
| about that benchmark. Otherwise the difference between models
| is +-2 tasks done correctly. I wouldn't worry too much about
| it. Just try it out yourself and see if its any better.
| bemmu wrote:
| I saw this on Reddit earlier today. Over there the source of this
| file was given as:
| https://web.archive.org/web/20251118111103/https://storage.g...
|
| The bucket name "deepmind-media" has been used in the past on the
| deepmind official site, so it seems legit.
| onlyrealcuzzo wrote:
| Prediction markets _were_ expecting today to be the release. So
| I wouldn 't be surprised if they do a release today, tomorrow,
| or Thursday (around Nvidia earnings).
| fraboniface wrote:
| > Developments to the model architecture contribute to the
| significantly improved performance from previous model families.
|
| I wonder how significant this is. DeepMind was always more
| research-oriented that OpenAI, which mostly scaled things up.
| They may have come up with a significantly better architecture
| (Transformer MoE still leaves a lot of room).
| msp26 wrote:
| Is flash/flash lite releasing alongside pro? Those two tiers have
| been incredible for the price since 2.0, absolute workhorses.
| Can't wait for 3.0.
| omidsa1 wrote:
| TL;DR: expected results, not underwhelming.So far scaling laws
| hold.
| nilayj wrote:
| Curious to see the API pricing. SOTA performance across tasks at
| a price cheaper than GPT 5 / Claude would make mostly everyone
| switch to Gemini.
| __jl__ wrote:
| Same here. They have been aggressively increasing prices with
| each iteration (maybe because they started so low). Still hope
| that is not the case this time. GPT 5.1 is priced pretty
| aggressively so maybe that is an incentive to keep the current
| gemini API prices.
| Deathmax wrote:
| Bad news then, they've bumped 3.0 Pro pricing to $2/$12
| ($4/$18 at long context).
| fcanesin wrote:
| Great stuff, now if could please do gemini-2.5-pro-code that
| would be great
| 827a wrote:
| What is Google Antigravity?
| danielcampos93 wrote:
| mums the word on Flash?
| ethmarks wrote:
| > TPUs are specifically designed to handle the massive
| computations involved in training LLMs and can speed up training
| considerably compared to CPUs.
|
| That seems like a low bar. Who's training frontier LLMs on CPUs?
| Surely they meant to compare TPUs to GPUs. If "this is faster
| than a CPU for massively parallel AI training" is the best you
| can say about it, that's not very impressive.
| Workaccount2 wrote:
| It's a typo
| ethmarks wrote:
| Does Google's team not proofread this stuff? Or maybe is this
| an early draft that wasn't meant to be released?
| camdenreslink wrote:
| It was generated by an LLM like everything else these days.
| astrange wrote:
| LLMs don't make typos.
| silveraxe93 wrote:
| This is a leak, yeah.
|
| Though come on... Even with proofreading, this is an easy
| one to miss.
| babl-yc wrote:
| I don't know if you can generally say that "LLM training is
| faster on TPUs vs GPUs". There is variance among LLM
| architectures, TPU cluster sizes, GPU cluster sizes...
|
| They are both designed to do massively parallel operations.
| TPUs are just a bit more specific to matrix multiply+adds while
| GPUs are more generic.
| Taek wrote:
| One benchmark I would really like to see: instruction adherence.
|
| For example, the frontier models of early-to-mid 2024 could
| reliably follow what seemed to be 20-30 instructions. As you gave
| more instructions than that in your prompt, the LLMs started
| missing some and your outputs became inconsistent and difficult
| to control.
|
| The latest set of models (2.5 Pro, GPT-5, etc) seem to top out
| somewhere in the 100 range? They are clearly much better at
| following a laundry list of instructions, but they also clearly
| have a limit and once your prompt is too large and too specific
| you lose coherence again.
|
| If I had to guess, Gemini 3 Pro has once again pushed the bar,
| and maybe we're up near 250 (haven't used it, I'm just blindly
| projecting / hoping). And that's a huge deal! I actually think it
| would be more helpful to have a model that could consistently
| follow 1000 custom instructions than it would be to have a model
| that had 20 more IQ points.
|
| I have to imagine you could make some fairly objective benchmarks
| around this idea, and it would be very helpful from an
| engineering perspective to see how each model stacked up against
| the others in this regard.
| machiaweliczny wrote:
| 20 more IQ would be nuts, 110 ~ top 25%, 130 ~ top 2%, 150 ~
| top 0.05%
|
| If you ever played competitive game the difference is insane
| between these tiers
| Taek wrote:
| Even more nuts would be a model that could follow a large,
| dense set of highly detailed instructions related to a series
| of complex tasks. Intelligence is nice, but it's far more
| useful and programmable if it can tightly follow a lot of
| custom instructions.
| DeathArrow wrote:
| I hope cheaper Chinese open weights models as good as Gemini will
| come soon. Gemini, Claude, GPT are kind of expensive if you use
| AI a lot.
| Topfi wrote:
| Additional context from AI Studio including pricing:
|
| Our most intelligent model with SOTA reasoning and multimodal
| understanding, and powerful agentic and vibe coding capabilities
|
| <=200K tokens * Input: $2,00 / Output: $12,00
|
| > 200K tokens * Input: $4,00 / Output: $18,00
|
| Knowledge cut off: Jan. 2025
| mohsen1 wrote:
| More expensive than current 2.5 Pro. for >200k token it's at
| $2.5 input and $15 output right now
| koakuma-chan wrote:
| > Gemini 3 Pro was trained using Google's Tensor Processing Units
| (TPUs)
|
| NVDA is down 3.26%
| CjHuber wrote:
| If it's because of that, then honestly it's as insane as the
| deepseek thing where all the info was released weeks before but
| the markt got nervous only when they released an app. I mean
| info about Gemini 3 is out quite a while now and of course they
| trained it using TPUs, I didn't even think that was in
| question.
| koakuma-chan wrote:
| I didn't know they _only_ used TPUs.
| robert-zaremba wrote:
| The strategic move to use TPU rather than Nvidia is paying well
| for Google. They are able to better utilize their existing large
| infrastructure, but also specialize the processes and pipelines
| for their own framework that they use to create and train models.
|
| I think a specialized hardware for training models is the next
| big wave in China.
| aliljet wrote:
| What's wild here is that among every single score they've
| absolutely killed, somehow, Anthropic and Claude Sonnet 4.5 have
| won a single victory in the fight: SWE Bench Verified and only by
| a singular point.
|
| I already enjoy Gemini 2.5 pro for planning and if Gemini 3 is
| priced similarly, I'll be incredibly happy to ditch the painfully
| pricey Claude max subscription. To be fair, I've already got an
| extremely sour taste in my mouth from the last Anthropic bait and
| switch on pricing and usage, so happy to see Google take the
| crown here.
| radial_symmetry wrote:
| SWE bench is weird because Claude has always underperformed on
| it relative to other models despite Claude Code blowing them
| away. The real test will be if Gemini CLI beats Claude Code,
| both using the agentic framework and tools they were trained
| on.
| __jl__ wrote:
| API pricing is up to $2/M for input and $12/M for output
|
| For comparison: Gemini 2.5 Pro was $1.25/M for input and $10/M
| for output Gemini 1.5 Pro was $1.25/M for input and $5/M for
| output
| bretpiatt wrote:
| Page 5, "The knowledge cutoff date for Gemini 3 Pro was January
| 2025."
|
| Still taking nearly a year to train and run post training safety
| and stability tuning.
|
| With 10x the infrastructure they could iterate much faster, I
| don't see AI infrastructure as a bubble, it is still a bottleneck
| on pace of innovation at today's active deployment level.
| camdenreslink wrote:
| But if they spend 10x on infrastructure, and capabilities only
| improve 10%, then that still can be a bubble even if
| infrastructure is a bottleneck.
| eric15342335 wrote:
| Update: it is available at https://aistudio.google.com now!
| amelius wrote:
| These model cards tell me nothing. I want to know the exact data
| a model was trained on. Otherwise, how can I safely use it for
| generating texts that I show to children? Etc.etc.
| morcus wrote:
| Shouldn't you be carefully reading texts before you show it to
| children?
| amelius wrote:
| No, I have an app that generates children's stories.
| astrange wrote:
| The data is everything you've ever heard of, and obviously
| contains things you wouldn't show to children, since that'd
| include NYT war journalism stories.
| butlike wrote:
| It's over. I just don't care anymore. I don't care what a pro
| model card is. I don't care what a humanity's last exam is. I
| don't care if the response makes me feel good about the prompt I
| made. I don't care if it's sentient. I don't care if it's
| secretly sentient. I don't care if it's just a machine. I don't
| care if the gov't has appropriated a secret model. I don't care
| if this is the precursor to AGI, ASI, AGGI, AGGSISGIGIG....I
| just. Don't. care.
|
| And I really don't think I'm alone in this.
| charcircuit wrote:
| >TPUs are specifically designed to handle the massive
| computations involved in training LLMs and can speed up training
| considerably compared to CPUs
|
| Who is training LLMs with CPUs?
| Barry-Perkins wrote:
| Excited to see the Gemini 3 Pro Model Card! Looking forward to
| exploring its features and capabilities.
| ks2048 wrote:
| Why is this linking to a random site? Here is a link hosted by
| Google:
|
| https://storage.googleapis.com/deepmind-media/Model-Cards/Ge...
___________________________________________________________________
(page generated 2025-11-18 23:00 UTC)