[HN Gopher] Google calls Gemma 3 the most powerful AI model you ...
___________________________________________________________________
Google calls Gemma 3 the most powerful AI model you can run on one
GPU
Author : gmays
Score : 80 points
Date : 2025-03-20 18:32 UTC (4 hours ago)
(HTM) web link (www.theverge.com)
(TXT) w3m dump (www.theverge.com)
| CamperBob2 wrote:
| Technically, the 1.58-bit Unsloth quant of DeepSeek R1 runs on a
| single GPU+128GB of system RAM. It performs amazingly well, but
| you'd better not be in a hurry.
| bestouff wrote:
| You probably meant 128GB
| CamperBob2 wrote:
| Edited. TBF, I _did_ say it was slow...
| refulgentis wrote:
| Have you tried the 0.000000001 bit quant? IIUC it's a single
| param so youd get much faster speeds.
| jfim wrote:
| What's the recommended way to run LLMs these days?
|
| Ollama seems to work with DeepSeek R1 with enough memory using
| an older CPU but it's around 1 token/second on my desktop.
| ForTheKidz wrote:
| ...until this coming tuesday? ...let's talk value.
|
| EDIT: I do feel like a fool, thank you.
| refulgentis wrote:
| Good news: it's free. Infinite value.
| rafaelmn wrote:
| Plenty of free things with no value
| 2OEH8eoCRo0 wrote:
| There's nothin more expensive than free
| impure wrote:
| It's a 27B model, I highly doubt that.
| xnx wrote:
| Which part do you doubt? That it is the most powerful? That it
| runs on a single GPU?
| Alifatisk wrote:
| That it is the most powerful
| 7734128 wrote:
| The claim is not that it's the most powerful, it is that
| it's the most powerful model that fit on one card.
| warkdarrior wrote:
| What's a better model to run on one GPU?
| int_19h wrote:
| QwQ-32
| snovv_crash wrote:
| If I ask qwq32 anything that is even slightly complicated
| it will ramble until it exceeds the context window, then
| forget my question. Q4k which is all that fits (with
| context) on a 3090.
|
| Gemma3 27B gives me a rapid 1shot response, and actually
| works really well for the type of rubber duck brainstorming
| partner I often need.
| simonw wrote:
| What's a better model that can run on a single GPU?
| idonotknowwhy wrote:
| It depends what you're trying to do.
|
| Coding - Mistral-Small-2503 or Qwen2.5-32b-Coder Reasoning -
| QwQ-32b Writing - Gemma-3-27b is good at this. etc
| simonw wrote:
| Right, this thread is about which models are better than
| Gemma-3-27B.
|
| I'm a fan of Mistral Small 3 personally but I've not spent
| enough time with it, Gemma and the new Mistral Small 3.1 to
| have an opinion of which of those is the "best" model.
|
| The best indicator of model quality I can find right now is
| still https://lmarena.ai/?leaderboard=
|
| Gemma 3 27B holds an impressive 8th place right now, second
| highest non-proprietary model after DeekSeek R1 (at 6th).
|
| QwQ-32B is 12th. Weirdly I couldn't find either of the
| Mistral Small 3 models on there.
| RandyRanderson wrote:
| So says Gemma 3.
| ChrisArchitect wrote:
| Google post from last week:
| https://blog.google/technology/developers/gemma-3/
| cwoolfe wrote:
| Apparently it can also pray. Seriously, I asked it for biblical
| advice about a tough situation today and it said it was praying
| for me. XD
| rurp wrote:
| This reminds me a recent chat I had with Claude, trying to
| identify what looked like an unusual fossil. The responses
| included things along the lines of "What a neat find!" or
| "That's fascinating! I'd love it if you could share more
| details". The statements were normal nice things to hear from a
| friend, but I found them pretty off-putting coming from a
| computer that of course couldn't care less and isn't a living
| thing I have a relationship with.
|
| This sort of thing worries me quite a bit. The modern internet
| has already sparked an awful lot of pseudo or para-social
| relationships through social media, OnlyFans and the like, with
| serious mental health and social cohesion costs. I think we're
| heading into a world where a lot of the remaining normal
| healthy social behavior gets subsumed by LLMs pretending to be
| your friend or romantic interest.
| thijson wrote:
| https://www.youtube.com/watch?v=LghsLs3DYUs
|
| Snip Snip was pleasant to talk to until the very end.
| hengheng wrote:
| Trying to find the silver lining in this makes me God's
| advocate I guess?
|
| I was able to reflect a lot on my upbringing by reading
| reddit threads. Advice columns, relationships, parenting
| advice, just dealing with people. It was great to finally
| have a normalized, standardized world view to bounce my own
| concepts off. It was like an advice column in an old
| magazine, but infinitely big. In my early 20s I must have
| spent entire days on there.
|
| I guess LLMs are the modern, ultra personalized version of
| that. Internet average, westernized culture, infinite supply,
| instantly. Just add water and enjoy a normal view of the
| world, no matter your surroundings or how you grew up. This
| is going to help out so many kids.
|
| And they're not evil yet. Host your own LLMs before you tell
| them your secrets, people.
| lamuswawir wrote:
| New term, "God's advocate"!
| nemothekid wrote:
| > _I think we 're heading into a world where a lot of the
| remaining normal healthy social behavior gets subsumed by
| LLMs pretending to be your friend or romantic interest._
|
| This is already happening.
|
| https://x.com/zymillyyy/status/1902181493553733941
| FridgeSeal wrote:
| Wow, Twitch parasocial relationships have _nothing_ on
| this.
| cube00 wrote:
| After giving me continuous wrong answers ChatGPT decided it
| would try allow me to indulge it in a "learning opportunity"
| instead.
|
| _I completely understand your frustration, and I genuinely
| appreciate you pushing me to be more accurate. You clearly
| know your way around Rust, and I should have been more
| precise from the start._
|
| _If you've already figured out the best approach, I'd love
| to hear it! Otherwise, I'm happy to keep digging until we
| find the exact method that works for your case. Either way, I
| appreciate the learning opportunity._
| k8sToGo wrote:
| I asked Claude for advice and it said something about it being
| heartbreaking.
| harvey9 wrote:
| Douglas Adams' Electric Monk finally available to the public!
| mdp2021 wrote:
| Which reveals strong reasons to suspect strong "parroting"
| qualities. "Parroting" qualities that should have been fought
| in implementation since day 0.
| Obscurity4340 wrote:
| Its fun to pick it apart sometimes and get it to correct
| itself but you often would never know unless you had direct
| or deep cut knowledge to interrogate it from
| SubiculumCode wrote:
| I'm not sure how much people parrot on the daily.
| KoolKat23 wrote:
| And henceforth, I'll equate that statement as a pithy one GPU
| level effort from whoever offers it.
|
| "Make sure you write atleast a Jetson Orin Nano Super level
| message in the condolence card"
| odysseus wrote:
| Does it run on the severed floor?
| butterlettuce wrote:
| I love how this show (2022) is _just_ heavily emanating into
| pop culture.
|
| Stoked for the season 2 finale today. It'll be like watching
| the Super Bowl.
| siva7 wrote:
| what show?
| alas44 wrote:
| Severance
| jprd wrote:
| Severance
| fortyseven wrote:
| Severance
| vsgherzi wrote:
| actually the finale is on friday not today!
| zeroq wrote:
| I call it the biggest bs since I had my supper.
| grej wrote:
| It lasted until Mistral released 3.1 Small a week later. Such is
| the pace of AI...
| nirav72 wrote:
| yeah, I can't even keep up these days. So now mainly focus on
| what I can run locally via Ollama.
| simonw wrote:
| Gemma 3 is on Ollama now https://ollama.com/library/gemma3
| but surprisingly they don't have Mistral 3.1 yet.
|
| I've managed to run Mistral 3.1 on my laptop using MLX, notes
| here https://simonwillison.net/2025/Mar/17/mistral-small-31/
| pram wrote:
| Gemma 3 is a lot better at writing for sure, compared to 2, but
| the big improvement is I can actually use a 32k+ context window
| and not have it start flipping out with random garbage.
| williamDafoe wrote:
| Does anyone use GoogleAI? For an AI Company with an AI Ceo using
| AI language translation, I think their actual GPT products are
| all terrible and have a terrible rep. And who wants their private
| conversation shipped back to google for spying?
| KoolKat23 wrote:
| Gemini 2.0 Flash et al are excellent. Think you need to try
| them again.
| deepsquirrelnet wrote:
| I tried a lot of models on openrouter recently, and I have to
| say that I found Gemini 2.0 flash to be surprisingly useful.
|
| I'd never used one of Google's proprietary models before
| that, but it really hits a sweet spot in the quality vs
| latency space right now.
| LeoPanthera wrote:
| I use Gemini (Advanced) over ChatGPT. Google's privacy issues
| are concerning, but no moreso than giving my conversations to
| OpenAI.
|
| In my experience, Gemini simply works better.
| timmg wrote:
| I'm wondering how small of a model can be "generally intelligent"
| (as in LLM intelligent, not AGI). Like there must be a size too
| small to hold "all the information" in.
|
| And I also wonder at what point we'll see specialized small
| models. Like if I want help coding, it's probably ok if the model
| doesn't know who directed "Jaws". I suspect that is the future:
| many small, specialized models.
|
| But maybe training compute will just get to the point where we
| can run a full-featured model on our desktop (or phone)?
| idonotknowwhy wrote:
| > Like there must be a size too small to hold "all the
| information" in.
|
| We're already there. If you running a Mistral-Large-2411 and
| Mistral-Small-2409 locally, you'll find the larger model is
| able to recall more specific details about works of fiction.
| And Deepseek-R1 is aware of a lot more.
|
| Then you ask one of the Qwen2.5 coding models, and they won't
| even be aware of it, because they're:
|
| > small, specialized models.
|
| > But maybe training compute will just get to the point where
| we can run a full-featured model on our desktop (or phone)?
|
| Training time compute won't allow the model to do anything out
| of distribution. You can test this yourself if you run one of
| the "R1 Distill" models. Eg. If you run the Qwen R1 distill and
| ask it about niche fiction, no matter how long you let it
| <think> for, it can't tell you something the original Qwen
| didn't know.
| LeoPanthera wrote:
| Maybe Llama 3.3 70B doesn't count as running on "one GPU", but it
| certainly runs just fine on one Mac, and in my tests it's far
| better at holding onto concepts over a longer conversation than
| Gemma 3 is, which starts getting confused after about 4000
| tokens.
___________________________________________________________________
(page generated 2025-03-20 23:00 UTC)