[HN Gopher] Google calls Gemma 3 the most powerful AI model you ...
       ___________________________________________________________________
        
       Google calls Gemma 3 the most powerful AI model you can run on one
       GPU
        
       Author : gmays
       Score  : 80 points
       Date   : 2025-03-20 18:32 UTC (4 hours ago)
        
 (HTM) web link (www.theverge.com)
 (TXT) w3m dump (www.theverge.com)
        
       | CamperBob2 wrote:
       | Technically, the 1.58-bit Unsloth quant of DeepSeek R1 runs on a
       | single GPU+128GB of system RAM. It performs amazingly well, but
       | you'd better not be in a hurry.
        
         | bestouff wrote:
         | You probably meant 128GB
        
           | CamperBob2 wrote:
           | Edited. TBF, I _did_ say it was slow...
        
         | refulgentis wrote:
         | Have you tried the 0.000000001 bit quant? IIUC it's a single
         | param so youd get much faster speeds.
        
         | jfim wrote:
         | What's the recommended way to run LLMs these days?
         | 
         | Ollama seems to work with DeepSeek R1 with enough memory using
         | an older CPU but it's around 1 token/second on my desktop.
        
       | ForTheKidz wrote:
       | ...until this coming tuesday? ...let's talk value.
       | 
       | EDIT: I do feel like a fool, thank you.
        
         | refulgentis wrote:
         | Good news: it's free. Infinite value.
        
           | rafaelmn wrote:
           | Plenty of free things with no value
        
             | 2OEH8eoCRo0 wrote:
             | There's nothin more expensive than free
        
       | impure wrote:
       | It's a 27B model, I highly doubt that.
        
         | xnx wrote:
         | Which part do you doubt? That it is the most powerful? That it
         | runs on a single GPU?
        
           | Alifatisk wrote:
           | That it is the most powerful
        
             | 7734128 wrote:
             | The claim is not that it's the most powerful, it is that
             | it's the most powerful model that fit on one card.
        
         | warkdarrior wrote:
         | What's a better model to run on one GPU?
        
           | int_19h wrote:
           | QwQ-32
        
             | snovv_crash wrote:
             | If I ask qwq32 anything that is even slightly complicated
             | it will ramble until it exceeds the context window, then
             | forget my question. Q4k which is all that fits (with
             | context) on a 3090.
             | 
             | Gemma3 27B gives me a rapid 1shot response, and actually
             | works really well for the type of rubber duck brainstorming
             | partner I often need.
        
         | simonw wrote:
         | What's a better model that can run on a single GPU?
        
           | idonotknowwhy wrote:
           | It depends what you're trying to do.
           | 
           | Coding - Mistral-Small-2503 or Qwen2.5-32b-Coder Reasoning -
           | QwQ-32b Writing - Gemma-3-27b is good at this. etc
        
             | simonw wrote:
             | Right, this thread is about which models are better than
             | Gemma-3-27B.
             | 
             | I'm a fan of Mistral Small 3 personally but I've not spent
             | enough time with it, Gemma and the new Mistral Small 3.1 to
             | have an opinion of which of those is the "best" model.
             | 
             | The best indicator of model quality I can find right now is
             | still https://lmarena.ai/?leaderboard=
             | 
             | Gemma 3 27B holds an impressive 8th place right now, second
             | highest non-proprietary model after DeekSeek R1 (at 6th).
             | 
             | QwQ-32B is 12th. Weirdly I couldn't find either of the
             | Mistral Small 3 models on there.
        
       | RandyRanderson wrote:
       | So says Gemma 3.
        
       | ChrisArchitect wrote:
       | Google post from last week:
       | https://blog.google/technology/developers/gemma-3/
        
       | cwoolfe wrote:
       | Apparently it can also pray. Seriously, I asked it for biblical
       | advice about a tough situation today and it said it was praying
       | for me. XD
        
         | rurp wrote:
         | This reminds me a recent chat I had with Claude, trying to
         | identify what looked like an unusual fossil. The responses
         | included things along the lines of "What a neat find!" or
         | "That's fascinating! I'd love it if you could share more
         | details". The statements were normal nice things to hear from a
         | friend, but I found them pretty off-putting coming from a
         | computer that of course couldn't care less and isn't a living
         | thing I have a relationship with.
         | 
         | This sort of thing worries me quite a bit. The modern internet
         | has already sparked an awful lot of pseudo or para-social
         | relationships through social media, OnlyFans and the like, with
         | serious mental health and social cohesion costs. I think we're
         | heading into a world where a lot of the remaining normal
         | healthy social behavior gets subsumed by LLMs pretending to be
         | your friend or romantic interest.
        
           | thijson wrote:
           | https://www.youtube.com/watch?v=LghsLs3DYUs
           | 
           | Snip Snip was pleasant to talk to until the very end.
        
           | hengheng wrote:
           | Trying to find the silver lining in this makes me God's
           | advocate I guess?
           | 
           | I was able to reflect a lot on my upbringing by reading
           | reddit threads. Advice columns, relationships, parenting
           | advice, just dealing with people. It was great to finally
           | have a normalized, standardized world view to bounce my own
           | concepts off. It was like an advice column in an old
           | magazine, but infinitely big. In my early 20s I must have
           | spent entire days on there.
           | 
           | I guess LLMs are the modern, ultra personalized version of
           | that. Internet average, westernized culture, infinite supply,
           | instantly. Just add water and enjoy a normal view of the
           | world, no matter your surroundings or how you grew up. This
           | is going to help out so many kids.
           | 
           | And they're not evil yet. Host your own LLMs before you tell
           | them your secrets, people.
        
             | lamuswawir wrote:
             | New term, "God's advocate"!
        
           | nemothekid wrote:
           | > _I think we 're heading into a world where a lot of the
           | remaining normal healthy social behavior gets subsumed by
           | LLMs pretending to be your friend or romantic interest._
           | 
           | This is already happening.
           | 
           | https://x.com/zymillyyy/status/1902181493553733941
        
             | FridgeSeal wrote:
             | Wow, Twitch parasocial relationships have _nothing_ on
             | this.
        
           | cube00 wrote:
           | After giving me continuous wrong answers ChatGPT decided it
           | would try allow me to indulge it in a "learning opportunity"
           | instead.
           | 
           |  _I completely understand your frustration, and I genuinely
           | appreciate you pushing me to be more accurate. You clearly
           | know your way around Rust, and I should have been more
           | precise from the start._
           | 
           |  _If you've already figured out the best approach, I'd love
           | to hear it! Otherwise, I'm happy to keep digging until we
           | find the exact method that works for your case. Either way, I
           | appreciate the learning opportunity._
        
         | k8sToGo wrote:
         | I asked Claude for advice and it said something about it being
         | heartbreaking.
        
         | harvey9 wrote:
         | Douglas Adams' Electric Monk finally available to the public!
        
         | mdp2021 wrote:
         | Which reveals strong reasons to suspect strong "parroting"
         | qualities. "Parroting" qualities that should have been fought
         | in implementation since day 0.
        
           | Obscurity4340 wrote:
           | Its fun to pick it apart sometimes and get it to correct
           | itself but you often would never know unless you had direct
           | or deep cut knowledge to interrogate it from
        
           | SubiculumCode wrote:
           | I'm not sure how much people parrot on the daily.
        
         | KoolKat23 wrote:
         | And henceforth, I'll equate that statement as a pithy one GPU
         | level effort from whoever offers it.
         | 
         | "Make sure you write atleast a Jetson Orin Nano Super level
         | message in the condolence card"
        
       | odysseus wrote:
       | Does it run on the severed floor?
        
         | butterlettuce wrote:
         | I love how this show (2022) is _just_ heavily emanating into
         | pop culture.
         | 
         | Stoked for the season 2 finale today. It'll be like watching
         | the Super Bowl.
        
           | siva7 wrote:
           | what show?
        
             | alas44 wrote:
             | Severance
        
             | jprd wrote:
             | Severance
        
             | fortyseven wrote:
             | Severance
        
           | vsgherzi wrote:
           | actually the finale is on friday not today!
        
       | zeroq wrote:
       | I call it the biggest bs since I had my supper.
        
       | grej wrote:
       | It lasted until Mistral released 3.1 Small a week later. Such is
       | the pace of AI...
        
         | nirav72 wrote:
         | yeah, I can't even keep up these days. So now mainly focus on
         | what I can run locally via Ollama.
        
           | simonw wrote:
           | Gemma 3 is on Ollama now https://ollama.com/library/gemma3
           | but surprisingly they don't have Mistral 3.1 yet.
           | 
           | I've managed to run Mistral 3.1 on my laptop using MLX, notes
           | here https://simonwillison.net/2025/Mar/17/mistral-small-31/
        
       | pram wrote:
       | Gemma 3 is a lot better at writing for sure, compared to 2, but
       | the big improvement is I can actually use a 32k+ context window
       | and not have it start flipping out with random garbage.
        
       | williamDafoe wrote:
       | Does anyone use GoogleAI? For an AI Company with an AI Ceo using
       | AI language translation, I think their actual GPT products are
       | all terrible and have a terrible rep. And who wants their private
       | conversation shipped back to google for spying?
        
         | KoolKat23 wrote:
         | Gemini 2.0 Flash et al are excellent. Think you need to try
         | them again.
        
           | deepsquirrelnet wrote:
           | I tried a lot of models on openrouter recently, and I have to
           | say that I found Gemini 2.0 flash to be surprisingly useful.
           | 
           | I'd never used one of Google's proprietary models before
           | that, but it really hits a sweet spot in the quality vs
           | latency space right now.
        
         | LeoPanthera wrote:
         | I use Gemini (Advanced) over ChatGPT. Google's privacy issues
         | are concerning, but no moreso than giving my conversations to
         | OpenAI.
         | 
         | In my experience, Gemini simply works better.
        
       | timmg wrote:
       | I'm wondering how small of a model can be "generally intelligent"
       | (as in LLM intelligent, not AGI). Like there must be a size too
       | small to hold "all the information" in.
       | 
       | And I also wonder at what point we'll see specialized small
       | models. Like if I want help coding, it's probably ok if the model
       | doesn't know who directed "Jaws". I suspect that is the future:
       | many small, specialized models.
       | 
       | But maybe training compute will just get to the point where we
       | can run a full-featured model on our desktop (or phone)?
        
         | idonotknowwhy wrote:
         | > Like there must be a size too small to hold "all the
         | information" in.
         | 
         | We're already there. If you running a Mistral-Large-2411 and
         | Mistral-Small-2409 locally, you'll find the larger model is
         | able to recall more specific details about works of fiction.
         | And Deepseek-R1 is aware of a lot more.
         | 
         | Then you ask one of the Qwen2.5 coding models, and they won't
         | even be aware of it, because they're:
         | 
         | > small, specialized models.
         | 
         | > But maybe training compute will just get to the point where
         | we can run a full-featured model on our desktop (or phone)?
         | 
         | Training time compute won't allow the model to do anything out
         | of distribution. You can test this yourself if you run one of
         | the "R1 Distill" models. Eg. If you run the Qwen R1 distill and
         | ask it about niche fiction, no matter how long you let it
         | <think> for, it can't tell you something the original Qwen
         | didn't know.
        
       | LeoPanthera wrote:
       | Maybe Llama 3.3 70B doesn't count as running on "one GPU", but it
       | certainly runs just fine on one Mac, and in my tests it's far
       | better at holding onto concepts over a longer conversation than
       | Gemma 3 is, which starts getting confused after about 4000
       | tokens.
        
       ___________________________________________________________________
       (page generated 2025-03-20 23:00 UTC)