[HN Gopher] Can I run AI locally?
       ___________________________________________________________________
        
       Can I run AI locally?
        
       Author : ricardbejarano
       Score  : 764 points
       Date   : 2026-03-13 12:46 UTC (10 hours ago)
        
 (HTM) web link (www.canirun.ai)
 (TXT) w3m dump (www.canirun.ai)
        
       | John23832 wrote:
       | RTX Pro 6000 is a glaring omission.
        
         | schaefer wrote:
         | No Nvidia Spark workstation is another omission.
        
         | embedding-shape wrote:
         | Yeah, that's weird, seems it has later models, and earlier, but
         | specifically not Pro 6000? Also, based on my experience, the
         | given numbers seems to be at least one magnitude off, which
         | seems like a lot, when I use the approx values for a Pro 6000
         | (96GB VRAM + 1792 GB/s)
        
       | sxates wrote:
       | Cool thing!
       | 
       | A couple suggestions:
       | 
       | 1. I have an M3 Ultra with 256GB of memory, but the options list
       | only goes up to 192GB. The M3 Ultra supports up to 512GB. 2. It'd
       | be great if I could flip this around and choose a model, and then
       | see the performance for all the different processors. Would help
       | making buying decisions!
        
         | utopcell wrote:
         | Unfortunately, Apple retired the 512GiB models.
        
           | ProllyInfamous wrote:
           | Sure, but _those already sold still exist_.
        
       | GrayShade wrote:
       | This feels a bit pessimistic. Qwen 3.5 35B-A3B runs at 38 t/s tg
       | with llama.cpp (mmap enabled) on my Radeon 6800 XT.
        
         | Aurornis wrote:
         | At what quantization and with what size context window?
        
           | GrayShade wrote:
           | Looks like it's a bit slower today. Running llama.cpp b8192
           | Vulkan.
           | 
           | $ ./llama-cli
           | unsloth_Qwen3.5-35B-A3B-GGUF_Qwen3.5-35B-A3B-UD-Q4_K_XL.gguf
           | -c 65536 -p "Hello"
           | 
           | [snip 73 lines]
           | 
           | [ Prompt: 86,6 t/s | Generation: 34,8 t/s ]
           | 
           | $ ./llama-cli
           | unsloth_Qwen3.5-35B-A3B-GGUF_Qwen3.5-35B-A3B-UD-Q4_K_XL.gguf
           | -c 262144 -p "Hello"
           | 
           | [snip 128 lines]
           | 
           | [ Prompt: 78,3 t/s | Generation: 30,9 t/s ]
           | 
           | I suspect the ROCm build will be faster, but it doesn't work
           | out of the box for me.
        
       | phelm wrote:
       | This is awesome, it would be great to cross reference some
       | intelligence benchmarks so that I can understand the trade off
       | between RAM consumption, token rate and how good the model is
        
       | S4phyre wrote:
       | Oh how cool. Always wanted to have a tool like this.
        
       | adithyassekhar wrote:
       | This just reminded me of this
       | https://www.systemrequirementslab.com/cyri.
       | 
       | Not sure if it still works.
        
       | twampss wrote:
       | Is this just llmfit but a web version of it?
       | 
       | https://github.com/AlexsJones/llmfit
        
         | deanc wrote:
         | Yes. But llmfit is far more useful as it detects your system
         | resources.
        
           | dgrin91 wrote:
           | Honestly I was surprised about this. It accurately got my GPU
           | and specs without asking for any permissions. I didnt realize
           | I was exposing this info.
        
             | dekhn wrote:
             | How could it not? That information is always available to
             | userspace.
        
               | bityard wrote:
               | "Available to userspace" is a much different thing than
               | "available to every website that wants it, even in
               | private mode".
               | 
               | I too was a little surprised by this. My browser
               | (Vivladi) makes a big deal about how privacy-conscious
               | they are, but apparently browser fingerprinting is not on
               | their radar.
        
               | swiftcoder wrote:
               | It's pretty hard to avoid GPU fingerprinting if you have
               | webgl/webgpu enabled
        
               | dekhn wrote:
               | We switched to talking about llmfit in this subthread, it
               | runs as native code.
        
             | rithdmc wrote:
             | Do you mean the OPs website? Mine's way off.
             | 
             | > Estimates based on browser APIs. Actual specs may vary
        
             | spudlyo wrote:
             | I run LibreWolf, which is configured to ask me before a
             | site can use WebGL, which is commonly used for
             | fingerprinting. I got the popup on this site, so I assume
             | that's how they're doing it.
        
             | johnisgood wrote:
             | Why were you surprised?
             | 
             | You can check out here how it does that:
             | https://github.com/AlexsJones/llmfit/blob/main/llmfit-
             | core/s...
             | 
             | To detect NVIDIA GPUs, for example:
             | https://github.com/AlexsJones/llmfit/blob/main/llmfit-
             | core/s...
             | 
             | In this case it just runs the command "nvidia-smi".
             | 
             | Note: llmfit is not web-based.
        
           | Someone1234 wrote:
           | I feel like they both solve different issues well:
           | 
           | - If you already HAVE a computer and are looking for models:
           | LLMFit
           | 
           | - If you are looking to BUY a computer/hardware, and want to
           | compare/contrast for local LLM usage: This
           | 
           | You cannot exactly run LLMFit on hardware you don't have.
        
         | rootusrootus wrote:
         | That's super handy, thanks for sharing the link. Way more
         | useful than the web site this post is about, to be honest.
         | 
         | It looks like I can run more local LLMs than I thought, I'll
         | have to give some of those a try. I have decent memory (96GB)
         | but my M2 Max MBP is a few years old now and I figured it would
         | be getting inadequate for the latest models. But llmfit thinks
         | it's a really good fit for the vast majority of them.
         | Interesting!
        
           | hrmtst93837 wrote:
           | Your hardware can run a good range of local models, but keep
           | an eye on quantization since 4-bit models trade off some
           | accuracy, especially with longer context or tougher tasks.
           | Thermal throttling is also an issue, since even Apple silicon
           | can slow down when all cores are pushed for a while, so
           | sustained performance might not match benchmark numbers.
        
       | mrdependable wrote:
       | This is great, I've been trying to figure this stuff out
       | recently.
       | 
       | One thing I do wonder is what sort of solutions there are for
       | running your own model, but using it from a different machine. I
       | don't necessarily want to run the model on the machine I'm also
       | working from.
        
         | cortesoft wrote:
         | Ollama runs a web server that you use to interact with the
         | models: https://docs.ollama.com/quickstart
         | 
         | You can also use the kubernetes operator to run them on a
         | cluster: https://ollama-operator.ayaka.io/pages/en/
        
         | rebolek wrote:
         | ssh?
        
       | g_br_l wrote:
       | could you add raspi to the list to see which ridiculously small
       | models it can run?
        
       | vova_hn2 wrote:
       | It says "RAM - unknown", but doesn't give me an option to specify
       | how much RAM I have. Why?
        
       | charcircuit wrote:
       | On mobile it does not show the name of the model in favor of the
       | other stats.
        
       | debatem1 wrote:
       | For me the "can run" filter says "S/A/B" but lists S, A, B, and C
       | and the "tight fit" filter says "C/D" but lists F.
       | 
       | Just FYI.
        
       | metalliqaz wrote:
       | Hugging Face can already do this for you (with much more up-to-
       | date list of available models). Also LM Studio. However they
       | don't attempt to estimate tok/sec, so that's a cool feature.
       | However I don't really trust those numbers that much because it
       | is not incorporating information about the CPU, etc. True GPU
       | offload isn't often possible on consumer PC hardware. Also there
       | are different quants available that make a big difference.
        
       | havaloc wrote:
       | Missing the A18 Neo! :)
        
       | arjie wrote:
       | Cool website. The one that I'd really like to see there is the
       | RTX 6000 Pro Blackwell 96 GB, though.
        
       | ge96 wrote:
       | Raspberry pi? Say 4B with 4GB of ram.
       | 
       | I also want to run vision like Yocto and basic LLM with TTS/STT
        
         | boutell wrote:
         | I've been trying to get speech to text to work with a
         | reasonable vocabulary on pis for a while. It's tough. All the
         | modern models just need more GPU than is available
        
           | ge96 wrote:
           | Whispr?
           | 
           | For wakewords I have used pico rhino voice
           | 
           | I want to use these I2S breakout mics
        
           | meatmanek wrote:
           | For ASR/STT on a budget, you want
           | https://huggingface.co/nvidia/parakeet-tdt-0.6b-v3 - it works
           | great on CPU.
           | 
           | I haven't tried on a raspberry pi, but on Intel it uses a
           | little less than 1s of CPU time per second of audio. Using
           | https://github.com/NVIDIA-
           | NeMo/NeMo/blob/main/examples/asr/a... for chunked streaming
           | inference, it takes 6 cores to process audio ~5x faster than
           | realtime. I expect with all cores on a Pi 4 or 5, you'd
           | probably be able to at least keep up with realtime.
           | 
           | (Batch inference, where you give it the whole audio file up
           | front, is slightly more efficient, since chunked streaming
           | inference is basically running batch inference on overlapping
           | windows of audio.)
           | 
           | EDIT: there are also the multitalker-parakeet-
           | streaming-0.6b-v1 and nemotron-speech-streaming-en-0.6b
           | models, which have similar resource requirements but are
           | built for true streaming inference instead of chunked
           | inference. In my tests, these are slightly less accurate. In
           | particular, they seem to completely omit any sentence at the
           | beginning or end of a stream that was partially cut off.
        
       | LeifCarrotson wrote:
       | This lacks a whole lot of mobile GPUs. It also does not
       | understand that you can share CPU memory with the GPU, or perform
       | various KV cache offloading strategies to work around memory
       | limits.
       | 
       | It says I have an Arc 750 with 2 GB of shared RAM, because that's
       | the GPU that renders my browser...but I actually have an RTX1000
       | Ada with 6 GB of GDDR6. It's kind of like an RTX 4050 (not listed
       | in the dropdowns) with lower thermal limits. I also have 64 GB of
       | LPDDR5 main memory.
       | 
       | It works - Qwen3 Coder Next, Devstral Small, Qwen3.5 4B, and
       | others can run locally on my laptop in near real-time. They're
       | not quite as good as the latest models, and I've tried some
       | bigger ones (up to 24GB, it produces tokens about half as fast as
       | I can type...which is disappointingly slow) that are slower but
       | smarter.
       | 
       | But I don't run out of tokens.
        
       | sshagent wrote:
       | I don't see my beloved 5060ti. looks great though
        
       | carra wrote:
       | Having the rating of how well the model will run for you is cool.
       | I miss to also have some rating of the model capabilities (even
       | if this is tricky). There are way too many to choose. And just
       | looking at the parameter number or the used memory is not always
       | a good indication of actual performance.
        
       | jrmg wrote:
       | Is there a reliable guide somewhere to setting up local AI for
       | coding (please don't say 'just Google it' - that just results in
       | a morass of AI slop/SEO pages with out of date, non-self-
       | consistent, incorrect or impossible instructions).
       | 
       | I'd like to be able to use a local model (which one?) to power
       | Copilot in vscode, and run coding agent(s) (not general purpose
       | OpenClaw-like agents) on my M2 MacBook. I know it'll be slow.
       | 
       | I suspect this is actually fairly easy to set up - if you know
       | how.
        
         | AstroBen wrote:
         | Ollama or LM Studio are very simple to setup.
         | 
         | You're probably not going to get anything working well as an
         | agent on an M2 MacBook, but smaller models do surprisingly well
         | for focused autocomplete. Maybe the Qwen3.5 9B model would run
         | decently on your system?
        
           | jrmg wrote:
           | Right - setting up LM studio is not hard. But how do I
           | connect LM Studio to Copilot, or set up an agent?
        
             | brcmthrowaway wrote:
             | Basically LM Studio has a server that serves models over
             | HTTP (localhost). Configure/enable the server and connect
             | OpenCode to it.
             | 
             | Try this article https://advanced-stack.com/fields-
             | notes/qwen35-opencode-lm-s...
             | 
             | I'm looking for an alternative to OpenCode though, I can
             | barely see the UI.
        
               | AstroBen wrote:
               | Codex also supports configuring an alternative API for
               | the model, you could try that:
               | https://unsloth.ai/docs/basics/codex#openai-codex-cli-
               | tutori...
        
             | AstroBen wrote:
             | It looks like Copilot has direct support for Ollama if
             | you're willing to set that up:
             | https://docs.ollama.com/integrations/vscode
             | 
             | For LM Studio under server settings you can start a local
             | server that has an OpenAI-compatible API. You'd need to
             | point Copilot to that. I don't use Copilot so not sure of
             | the exact steps there
        
             | NortySpock wrote:
             | I tried the Zed editor and it picked up Ollama with almost
             | no fiddling, so that has allowed me to run Qwen3.5:9B just
             | by tweaking the ollama settings (which had a few dumb
             | defaults, I thought, like assuming I wanted to run 3 LLMs
             | in parallel, initially disabling Flash Attention, and
             | having a very short context window...).
             | 
             | Having a second pair of "eyes" to read a log error and dig
             | into relevant code is super handy for getting ideas
             | flowing.
        
         | chatmasta wrote:
         | Any time I google something on this topic, the results are
         | useful but also out of date, because this space is moving so
         | absurdly fast.
        
         | randusername wrote:
         | Personally I'd start with llamafile [0] then move to compiling
         | your own llama.cpp.
         | 
         | It's not as bad as you might think to compile llama.cpp for
         | your target architecture and spin up an OpenAI compatible API
         | endpoint. It even downloads the models for you.
         | 
         | [0]: https://github.com/mozilla-ai/llamafile
        
       | AstroBen wrote:
       | This doesn't look accurate to me. I have an RX9070 and I've been
       | messing around with Qwen 3.5 35B-A3B. According to this site I
       | can't even run it, yet I'm getting 32tok/s ^.-
        
         | misnome wrote:
         | It seems to be missing a whole load of the quantized Qwen
         | models, Qwen3.5:122b works fine in the 96GB GH200 (a machine
         | that is also missing here....)
        
         | mongrelion wrote:
         | Which quantization are you running and what context size?
         | 32tok/s for that model on that card sounds pretty good to me!
        
       | unfirehose wrote:
       | if you do, would you still want to collect data in a single pane
       | of glass? see my open source repo for aggregating harness data
       | from multiple machine learning model harnesses & models into a
       | single place to discover what you are working on & spending time
       | & money. there is plans for a scrobble feature like last.fm but
       | for agent research & code development & execution.
       | 
       | https://github.com/russellballestrini/unfirehose-nextjs-logg...
       | 
       | thanks, I'll check for comments, feel free to fork but if you
       | want to contribute you'll have to find me off of github, I
       | develop privately on my own self hosted gitlab server. good luck
       | & God bless.
        
       | varispeed wrote:
       | Does it make any sense? I tried few models at 128GB and it's all
       | pretty much rubbish. Yes they do give coherent answers, sometimes
       | they are even correct, but most of the time it is just plain
       | wrong. I find it massive waste of time.
        
         | boutell wrote:
         | I'm not sure how long ago you tried it, but look at Qwen 3.5
         | 32b on a fast machine. Usually best to shut off thinking if
         | you're not doing tool use.
        
         | mongrelion wrote:
         | Apparently there is a whole science behind running models. I
         | have seen the instructions that unsloth publishes for their
         | quants and depending on the model they'll tweak things like the
         | temperature, top k, etc.
         | 
         | The size of the quantization you chose also makes a difference.
         | 
         | The GPU driver also plays an important role.
         | 
         | What was your approach? What software did you use to run the
         | models?
        
       | orthoxerox wrote:
       | For some reason it doesn't react to changing the RAM amount in
       | the combo box at the top. If I open this on my Ryzen AI Max 395+
       | with 32 GB of unified memory, it thinks nothing will fit because
       | I've set it up to reserve 512MB of RAM for the GPU.
        
         | bityard wrote:
         | Yeah, this site is iffy at best. I didn't even see Strix Halo
         | on the list, but I selected 128GB and bumped up the memory
         | bandwidth. It says gpt-oss-120b "barely runs" at ~2 t/s.
         | 
         | In reality, gpt-oss-120b fits great on the machine with plenty
         | of room to spare and easily runs inference north of 50 t/s
         | depending on context.
        
       | kylehotchkiss wrote:
       | My Mac mini rocks qwen2.5 14b at a lightning fast 11/tokens a
       | second. Which is actually good enough for the long term data
       | processing I make it spend all day doing. It doesn't lock up the
       | machine or prevent its primary purpose as webserver from being
       | fulfilled.
        
       | freediddy wrote:
       | i think the perplexity is more important than tokens per second.
       | tokens per second is relatively useless in my opinion. there is
       | nothing worse than getting bad results returned to you very
       | quickly and confidently.
       | 
       | ive been working with quite a few open weight models for the last
       | year and especially for things like images, models from 6 months
       | would return garbage data quickly, but these days qwen 3.5 is
       | incredible, even the 9b model.
        
         | sroussey wrote:
         | No, getting bad results slowly is much worse. Bad results
         | quickly and you can make adjustments.
         | 
         | But yes, if there is a choice I want quality over speed. At
         | same quality, I definitely want speed.
        
       | meatmanek wrote:
       | This seems to be estimating based on memory bandwidth / size of
       | model, which is a really good estimate for dense models, but MoE
       | models like GPT-OSS-20b don't involve the entire model for every
       | token, so they can produce more tokens/second on the same
       | hardware. GPT-OSS-20B has 3.6B active parameters, so it should
       | perform similarly to a 3-4B dense model, while requiring enough
       | VRAM to fit the whole 20B model.
       | 
       | (In terms of intelligence, they tend to score similarly to a
       | dense model that's as big as the geometric mean of the full model
       | size and the active parameters, i.e. for GPT-OSS-20B, it's
       | roughly as smart as a sqrt(20b*3.6b) [?] 8.5b dense model, but
       | produces tokens 2x faster.)
        
         | lambda wrote:
         | Yeah, I looked up some models I have actually run locally on my
         | Strix Halo laptop, and its saying I should have much lower
         | performance than I actually have on models I've tested.
         | 
         | For MoE models, it should be using the active parameters in
         | memory bandwidth computation, not the total parameters.
        
         | littlestymaar wrote:
         | While your remark is valid, there's two small inaccuracies
         | here:
         | 
         | > GPT-OSS-20B has 3.6B active parameters, so it should perform
         | similarly to a 3-4B dense model, while requiring enough VRAM to
         | fit the whole 20B model.
         | 
         | First, the token generation speed is going to be comparable,
         | but not the prefil speed (context processing is going to be
         | much slower on a big MoE than on a small dense model).
         | 
         | Second, without speculative decoding, it is correct to say that
         | a small dense model and a bigger MoE with the same amount of
         | active parameters are going to be roughly as fast. But if you
         | use a small dense model you will see token generation
         | performance improvements with speculative decoding (up to x3
         | the speed), whereas you probably won't gain much from
         | speculative decoding on a MoE model (because two consecutive
         | tokens won't trigger the same "experts", so you'd need to load
         | more weight to the compute units, using more bandwidth).
        
           | lambda wrote:
           | So, this is all true, but this calculation isn't that
           | nuanced. It's trying to get you into a ballpark range, and
           | based on my usage on my real hardware (if I put in my specs,
           | since it's not in their hardware list), the results are
           | fairly close to my real experience if I compensate for the
           | issue where it's calculating based on total params instead of
           | active.
           | 
           | So by doing so, this calculator is telling you that you
           | should be running entirely dense models, and sparse MoE
           | models that maybe both faster and perform better are not
           | recommended.
        
             | littlestymaar wrote:
             | I agree, and I even started my response expressing my
             | agreement with the whole point.
             | 
             | But since this is a tech forum, I assumed some people would
             | be interested by the correction on the details that were
             | wrong.
        
         | pbronez wrote:
         | The docs page addresses this:
         | 
         | > A Mixture of Experts model splits its parameters into groups
         | called "experts." On each token, only a few experts are active
         | -- for example, Mixtral 8x7B has 46.7B total parameters but
         | only activates ~12.9B per token. This means you get the quality
         | of a larger model with the speed of a smaller one. The
         | tradeoff: the full model still needs to fit in memory, even
         | though only part of it runs at inference time.
         | 
         | > A dense model activates all its parameters for every token --
         | what you see is what you get. A MoE model has more total
         | parameters but only uses a subset per token. Dense models are
         | simpler and more predictable in terms of memory/speed. MoE
         | models can punch above their weight in quality but need more
         | VRAM than their active parameter count suggests.
         | 
         | https://www.canirun.ai/docs
        
           | lambda wrote:
           | It discusses it, and they have data showing that they know
           | the number of active parameters on an MoE model, but they
           | don't seem to use that in their calculation. It gives me
           | answers far lower than my real-world usage on my setup; its
           | calculation lines up fairly well for if I were trying to run
           | a dense model of that size. Or, if I increase my memory
           | bandwidth in the calculator by a factor of 10 or so which is
           | the ratio between active and total parameters in the model, I
           | get results that are much closer to real world usage.
        
         | tommy_axle wrote:
         | I'm guessing this is also calculating based on the full context
         | size that the model supports but depending on your use case it
         | will be misleading. Even on a small consumer card with Qwen 3
         | 30B-A3B you probably don't need 128K context depending on what
         | you're doing so a smaller context and some tensor overrides
         | will help. llama.cpp's llama-fit-params is helpful in those
         | cases.
        
       | nilslindemann wrote:
       | 1. More title attributes please ("S 16 A 7 B 7 C 0 D 4 F 34",
       | huh?)
       | 
       | 2. Add a 150% size bonus to your site.
       | 
       | Otherwise, cool site, bookmarked.
        
       | amelius wrote:
       | Why isn't there some kind of benchmark score in the list?
        
       | amelius wrote:
       | What is this S/A/B/C/etc. ranking? Is anyone else using it?
        
         | vikramkr wrote:
         | Just a tier list I think
        
         | relaxing wrote:
         | Apparently S being a level above A comes from Japanese grading.
         | I've been confused by that, too.
        
           | swiftcoder wrote:
           | It's very common in Japanese-developed video games as well
        
       | tcbrah wrote:
       | tbh i stopped caring about "can i run X locally" a while ago. for
       | anything where quality matters (scripting, code, complex
       | reasoning) the local models are just not there yet compared to
       | API. where local shines is specific narrow tasks - TTS,
       | embeddings, whisper for STT, stuff like that. trying to run a 70b
       | model at 3 tok/s on your gaming GPU when you could just hit an
       | API for like $0.002/req feels like a weird flex IMO
        
         | hatthew wrote:
         | For me and probably many other people, local has nothing to do
         | with cost and everything to do with privacy
        
           | tcbrah wrote:
           | genuine question - what are you working on that needs that
           | level of privacy? outside of NSFW stuff most API providers
           | arent doing anything with your prompts
        
             | hatthew wrote:
             | I would answer that, but it's private :)
             | 
             | I can think of several reasons: corporate policy, personal
             | principles, NSFW stuff, illegal stuff
        
       | sdingi wrote:
       | When running models on my phone - either through the web browser
       | or via an app - is there any chance it uses the phone's NPU, or
       | will these be GPU only?
       | 
       | I don't really understand how the interface to the NPU chip looks
       | from the perspective of a non-system caller, if it exists at all.
       | This is a Samsung device but I am wondering about the general
       | principle.
        
       | amelius wrote:
       | It would be great if something like this was built into ollama,
       | so you could easily list available models based on your current
       | hardware setup, from the CLI.
        
         | rootusrootus wrote:
         | Someone linked to llmfit. That would be a great tool to
         | integrate with ollama. Just highlight the one you want and tell
         | it to install.
         | 
         | Quick, someone go vibe code that.
        
           | dugidugout wrote:
           | The latest level of abstraction! You just release your ideas
           | half baked in some internet connected box and wake up with
           | products! Yahoo! Onwards into the Gestell!
        
       | am17an wrote:
       | You can still run larger MoE models using expert weight off-
       | loading to the CPU for token generation. They are by and large
       | useable, I get ~50 toks/second on a kimi linear 48B (3B active)
       | model on a potato PC + a 3090
        
       | brcmthrowaway wrote:
       | If anyone hasn't tried Qwen3.5 on Apple Silicon, I highly suggest
       | you to! Claude level performance on local hardware. If the Qwen
       | team didn't get fired, I would be bullish on Local LLM.
        
       | golem14 wrote:
       | Has anyone actually built anything with this tool?
       | 
       | The website says that code export is not working yet.
       | 
       | That's a very strange way to advertise yourself.
        
       | cafed00d wrote:
       | Open with multiple browsers (safari vs chrome) to get more
       | "accurate + glanceable" rankings.
       | 
       | Its using WebGPU as a proxy to estimate system resource. Chrome
       | tends to leverage as much resources (Compute + Memory) as the OS
       | makes available. Safari tends to be more efficient.
       | 
       | Maybe this was obvious to everyone else. But its worth re-
       | iterating for those of us skimmers of HN :)
        
       | ryandrake wrote:
       | Missing RTX A4000 20GB from the GPU list.
        
       | mark_l_watson wrote:
       | I have spent a HUGE amount of time the last two years
       | experimenting with local models.
       | 
       | A few lessons learned:
       | 
       | 1. small models like the new qwen3.5:9b can be fantastic for
       | local tool use, information extraction, and many other embedded
       | applications.
       | 
       | 2. For coding tools, just use Google Antigravity and gemini-cli,
       | or, Anthropic Claude, or...
       | 
       | Now to be clear, I have spent perhaps 100 hours in the last year
       | configuring local models for coding using Emacs, Claude Code
       | (configured for local), etc. However, I am retired and this time
       | was a lot of fun for me: lot's of efforts trying to maximize
       | local only results. I don't recommend it for others.
       | 
       | I do recommend getting very good at using embedded local models
       | in small practical applications. Sweet spot.
        
         | nine_k wrote:
         | What kind of hardware did you use? I suppose that a 8GB gaming
         | GPU and a Mac Pro with 512 GB unified RAM give quite different
         | results, both formally being local.
        
           | fzzzy wrote:
           | A Mac Pro with 512 gb unified ram does not exist.
        
             | nine_k wrote:
             | Mac Studio Ultra, my bad. The 512 GB option existed up
             | until March 2026:
             | https://macdailynews.com/2026/03/06/apple-
             | drops-512gb-m3-ult...
        
         | manmal wrote:
         | What about running e.g. Qwen3.5 128B on a rented RTX Pro 6000?
        
           | girvo wrote:
           | IMO you're better off using qwen3.5-plus through the model
           | studio coding plan, but ymmv
        
         | kylehotchkiss wrote:
         | I've been really interested in the difference between 3.5 9b
         | and 14b for information extraction. Is there a discernible
         | difference in quality of capability?
        
         | johnmaguire wrote:
         | I'd love to know how you fit smaller models into your workflow.
         | I have an M4 Macbook Pro w/ 128GB RAM and while I have toyed
         | with some models via ollama, I haven't really found a nice
         | workflow for them yet.
        
           | philipkglass wrote:
           | It really depends on the tasks you have to perform. I am
           | using specialized OCR models running locally to extract page
           | layout information and text from scanned legal documents. The
           | quality isn't perfect, but it is _really_ good compared to
           | desktop /server OCR software that I formerly used that cost
           | hundreds or thousands of dollars for a license. If you have
           | similar needs and the time to try just one model, start with
           | GLM-OCR.
           | 
           | If you want a general knowledge model for answering questions
           | or a coding agent, nothing you can run on your MacBook will
           | come close to the frontier models. It's going to be
           | frustrating if you try to use local models that way. But
           | there are a lot of useful applications for local-sized models
           | when it comes to interpreting and transforming unstructured
           | data.
        
             | mandeepj wrote:
             | > I formerly used that cost hundreds or thousands of
             | dollars for a license
             | 
             | Azure Doc Intelligence charges $1.50 for 1000 pages. Was
             | that an annual/recurring license?
             | 
             | Would you mind sharing your OCR model? I'm using Azure for
             | now, as I want to focus on building the functionality
             | first, but would later opt for a local model.
        
               | philipkglass wrote:
               | I took a long break from document processing after
               | working on it heavily 20 years ago. The tools I used
               | before were ABBYY FineReader and PrimeOCR. I haven't
               | tried any of the commercial cloud based solutions. I'm
               | currently using GLM-OCR, Chandra OCR, and Apple's
               | LiveText in conjunction with each other (plus custom code
               | for glue functionality and downstream processing).
               | 
               | Try just GLM-OCR if you want to get started quickly. It
               | has good layout recognition quality, good text
               | recognition quality, and they actually tested it on Apple
               | Silicon laptops. It works easily out-of-the-box without
               | the yak shaving I encountered with some other models.
               | Chandra is even more accurate on text but its layout
               | bounding boxes are worse and it runs very slowly unless
               | you can set up batched inference with vLLM on CUDA. (I
               | tried to get batching to run with vllm-mlx so it could
               | work entirely on macOS, but a day spent shaving the yak
               | with Claude Opus's help went nowhere.)
               | 
               | If you just want to transcribe documents, you can also
               | try end-to-end models like olmOCR 2. I need pipeline
               | models that expose inner details of document layout
               | because I need to segment and restructure page contents
               | for further processing. The end-to-end models just
               | "magically" turn page scans into complete Markdown or
               | HTML documents, which is more convenient for some uses
               | but not mine.
        
               | D-Machine wrote:
               | These are some really great explicit examples and links,
               | much appreciated.
        
           | saltwounds wrote:
           | I use Raycast and connect it to LM Studio to run text clean
           | up and summaries often. The models are small enough I keep
           | them in memory more often than not
        
           | Bluecobra wrote:
           | I didn't realize that you can get 128GB of memory in a
           | notebook, that is impressive!
        
             | AzN1337c0d3r wrote:
             | Most workstation class laptops (i.e. Lenovo P-series, Dell
             | Precision) have 4 DIMM slots and you can get them with 256
             | GB (at least, before the current RAM shortages).
             | 
             | There's also the Ryzen AI Max+ 395 that has 128GB unified
             | in laptop form factor.
             | 
             | Only Apple has the unique dynamic allocation though.
        
               | the_pwner224 wrote:
               | Yep, I have a 13" gaming tablet with the 128 GB AMD Strix
               | Halo chip (Ryzen AI Max+ 395, what a name). Asus ROG Flow
               | Z13. It's a beast; the performance is totally
               | disproportionate to its size & form factor.
               | 
               | I'm not sure what exactly you're referring to with "Only
               | Apple has the unique dynamic allocation though." On Strix
               | Halo you set the fixed VRAM size to 512 MB in the BIOS,
               | and you set a few Linux kernel params that enable dynamic
               | allocation to whatever limit you want (I'm using 110 GB
               | max at the moment). LLMs can use up to that much when
               | loaded, but it's shared fully dynamically with regular
               | RAM and is instantly available for regular system use
               | when you unload the LLM.
        
               | wilkystyle wrote:
               | What operating system are you using? I was looking at
               | this exact machine as a potential next upgrade.
        
               | the_pwner224 wrote:
               | Arch with KDE, it works perfectly out of the box.
               | 
               | I configured/disabled RGB lighting in Windows before
               | wiping and the settings carried over to Linux. On Arch,
               | install & enable power-profiles-daemon and you can switch
               | between quiet/balanced/performance fan & TDP profiles. It
               | uses the same profiles & fan curves as the options in
               | Asus's Windows software. KDE has native integration for
               | this in the GUI in the battery menu. You don't need to
               | install asus-linux or rog-control-center.
               | 
               | For local AI: set VRAM size to 512 MB in the BIOS, add
               | these kernel params:
               | 
               | ttm.pages_limit=31457280 ttm.page_pool_size=31457280
               | amd_iommu=off
               | 
               | Pages are 4 KiB each, so 120 GiB = 120 x 1024^3 / 4096 =
               | 31457280
               | 
               | To check that it worked: sudo dmesg | grep
               | "amdgpu.*memory" will report two values. VRAM is what's
               | set in BIOS (minimum static allocation). GTT is the
               | maximum dynamic quota. The default is 48 GB of GTT. So if
               | you're running small models you actually don't even need
               | to do anything, it'll just work out of the box.
               | 
               | LM Studio worked out of the box with no setup, just
               | download the appimage and run it. For Ollama you just
               | `pacman -S ollama-rocm` and `systemctl enable --now
               | ollama`, then it works. I recently got ComfyUI set up to
               | run image gen & 3d gen models and that was also very
               | easy, took <10 minutes.
               | 
               | I can't believe this machine is still going for $2,800
               | with 128 GB. It's an incredible value.
        
               | xnzakg wrote:
               | You may wanna see if openrgb isn't able to configure the
               | RGB. Could even do some fun stuff like changing the color
               | once done with a training run or something
        
               | lambda wrote:
               | > Only Apple has the unique dynamic allocation though.
               | 
               | What do you mean? On Linux I can dynamically allocate
               | memory between CPU and GPU. Just have to set a few kernel
               | parameters to set the max allowable allocation to the
               | GPU, and set the BIOS to the minimum amount of dedicated
               | graphics memory.
        
               | AzN1337c0d3r wrote:
               | Maybe things have changed but the last time I looked at
               | this, it was only max 96GB to the GPU. And it isn't
               | dynamic in the sense you still have to tweak the kernel
               | parameters, which require a reboot.
               | 
               | Apple has none of this.
        
               | the_pwner224 wrote:
               | Strix Halo you can get _at least_ 120 GB to the GPU (out
               | of 128 GB total), I 'm using this configuration.
               | 
               | Setting the kernel params is a one-time initial setup
               | thing. You have 128 GB of RAM, set it to 120 or whatever
               | as the max VRAM. The LLM will use as much as it needs and
               | the rest of the system will use as much it needs. Fully
               | dynamic with real-time allocation of resources. Honestly
               | I literally haven't even thought of it after setting
               | those kernel args a while ago.
               | 
               | So: "options ttm.pages_limit=31457280
               | ttm.page_pool_size=31457280", reboot, and that's
               | literally all you have to do.
               | 
               | Oh and even that is only needed because the AMD driver
               | defaults it to something like 35-48 GB max VRAM
               | allocation. It is fully dynamic out of the box, you're
               | only configuring the max VRAM quota with those params.
               | I'm not sure why they choice that number for the default.
        
               | lambda wrote:
               | You do have to set the kernel parameters once to set the
               | max GPU allocation, I have it set to 110 GiB, and you
               | have to set a BIOS setting to set the minimum GPU
               | allocation, I have it set to 512 MiB. Once you've set
               | those up, it's dynamic within those constraints, with no
               | more reboots required.
               | 
               | On Windows, I think you're right, it's max 96 GiB to the
               | GPU and it requires a reboot to change it.
        
             | lambda wrote:
             | I've got a 128 GiB unified memory Ryzen Ai Max+ 395 (aka
             | Strix Halo) laptop.
             | 
             | Trying to run LLM models somehow makes 128 GiB of memory
             | feel incredibly tight. I'm frequently getting OOMs when I'm
             | running models that are pushing the limits of what this can
             | fit, I need to leave more memory free for system memory
             | than I was expecting. I was expecting to be able to run
             | models of up to ~100 GiB quantized, leaving 28 GiB for
             | system memory, but it turns out I need to leave more room
             | for context and overhead. ~80 GiB quantized seems like a
             | better max limit when trying not running on a headless
             | system so I'm running a desktop environment, browser, IDE,
             | compilers, etc in addition to the model.
             | 
             | And memory bandwidth limitations for running the models is
             | real! 10B active parameters at 4-6 bit quants feels usable
             | but slow, much more than that and it really starts to feel
             | sluggish.
             | 
             | So this can fit models like Qwen3.5-122B-A10B but it's not
             | the speediest and I had to use a smaller quant than
             | expected. Qwen3-Coder-Next (80B/3B active) feels quite on
             | speed, though not quite as smart. Still trying out models,
             | Nemotron-3-Super-120B-A12B just came out, but looks like
             | it'll be a bit slower than Qwen3.5 while not offering up
             | any more performance, though I do really like that they
             | have been transparent in releasing most of its training
             | data.
        
               | zozbot234 wrote:
               | There's been some very recent ongoing work in some local
               | AI frameworks on enabling mmap by default, which can
               | potentially obviate some RAM-driven limitations
               | especially for sparse MoE models. Running with mmap and
               | too little RAM will then still come with severe slowdowns
               | since read-only model parameters will have to be shuttled
               | in from storage as they're needed, but for hardware with
               | fast enough storage and especially for models that
               | "almost" fit in the RAM filesystem cache, this can be a
               | huge unblock at negligible cost. Especially if it
               | potentially enables further unblocks via adding extra
               | swap for K-V cache and long context.
        
           | echelon wrote:
           | Shouldn't we prioritize large scale open weights and open
           | source cloud infra?
           | 
           | An OpenRunPod with decent usage might encourage more non-
           | leading labs to dump foundation models into the commons. We
           | just need infra to run it. Distilling them down to desktop is
           | a fool's errand. They're meant to run on DC compute.
           | 
           | I'm fine with running everything in the cloud as long as we
           | own the software infra and the weights.
           | 
           | This is conceivably the only way we could catch up to Claude
           | Code is to have the Chinese start releasing their best coding
           | models and for them to get significant traction with
           | companies calling out to hosted versions. Otherwise, we're
           | going to be stuck in a take off scenario with no bridge.
        
             | girvo wrote:
             | I run Qwen3.5-plus through Alibaba's coding plan (Model
             | Studio): incredibly cheap, pretty fast, and decent. I can't
             | compare it to the highest released weight one though.
        
               | singpolyma3 wrote:
               | Is that https://www.alibabacloud.com/help/en/model-
               | studio/coding-pla... ? I was a bit confused that it seems
               | to be sized in requests not tokens
        
           | tempaccount5050 wrote:
           | Not OP but I had an XML file with inconsistent formatting for
           | album releases. I wanted to extract YouTube links from it,
           | but the formatting was different from album to album. Nothing
           | you could regex or filter manually. I shoved it all into a
           | DB, looked up the album, then gave the xml to a local LLM and
           | said "give me the song/YouTube pairs from this DB entry".
           | Worked like a charm.
        
         | sdrinf wrote:
         | Just want to echo the recommendation for qwen3.5:9b. This is a
         | smol, thinking, agentic tool-using, text-image multimodal
         | creature, with very good internal chains of thought. CoT can be
         | sometimes excessive, but it leads to very stable decision-
         | making process, even across very large contexts -something we
         | haven't seen models of this size before.
         | 
         | What's also new here, is VRAM-context size trade-off: for 25%
         | of it's attention network, they use the regular KV cache for
         | global coherency, but for 75% they use a new KV cache with
         | linear(!!!!) memory-token-context size expansion! which means,
         | eg ~100K token -> 1.5gb VRAM use -meaning for the first time
         | you can do extremely long conversations / document processing
         | with eg a 3060.
         | 
         | Strong, strong recommend.
        
           | steve_adams_86 wrote:
           | I've been building a harness for qwen3.5:9b lately (to better
           | understand how to create agentic tools/have fun) and I'm not
           | going to use it instead of Opus 4.6 for my day job but it's
           | remarkably useful for small tasks. And more than snappy
           | enough on my equipment. It's a fun model to experiment with.
           | I was previously using an old model from Meta and the
           | contrast in capability is pretty crazy.
           | 
           | I like the idea of finding practical uses for it, but so far
           | haven't managed to be creative enough. I'm so accustomed to
           | using these things for programming.
        
           | kingo55 wrote:
           | How's it compare in quality with larger models in the same
           | series? E.g 122b?
        
           | ggsp wrote:
           | How much difference are you seeing between standard and Q4
           | versions in terms of degradation, and is it constant across
           | tasks or more noticeable in some vs others?
        
             | rnewme wrote:
             | Less than expected, search for unsloths recent benchmark
        
           | dsr_ wrote:
           | Correction: not thinking, not a creature.
           | 
           | If it was a creature I would feel some sorrow when I killed
           | it.
           | 
           | If you are feeling sorrow when you reboot a machine running
           | an LLM, get to a psychiatrist ASAP.
        
           | threecheese wrote:
           | You can really see the limitations of qwen3.5:9b in reasoning
           | traces- it's fascinating. When a question "goes bad",
           | sometimes the thinking tokens are WILD - it's like watching
           | the Poirot after a head injury.
           | 
           | Example: "what is the air speed velocity of a swallow?" -
           | qwen knew it was a Monty Python gag, but couldnt and didnt
           | figure out which one.
        
         | cyanydeez wrote:
         | Cline (https://marketplace.visualstudio.com/items?itemName=saou
         | driz...) in vscode, inside a code-server run within docker
         | (https://docs.linuxserver.io/images/docker-code-server/) using
         | lmstudio (https://lmstudio.ai/) to access unsloth models
         | (https://unsloth.ai/docs/get-started/unsloth-model-catalog)
         | speficially (https://unsloth.ai/docs/models/qwen3-coder-next)
         | appears to be right at the edge of productivity, as long as you
         | realize what complexity means when issuing tasks.
        
         | dataflow wrote:
         | Thanks for sharing this, it's super helpful. I have a question
         | if you don't mind: I want a model that I can feed, say, my
         | entire email mailbox to, so that I can ask it questions later.
         | (Just the text content, which I can clean and preprocess
         | offline for its use.) Have any offline models you've dealt with
         | seemed suitable for that sort of use case, with that volume of
         | content?
        
           | perbu wrote:
           | Prompt injection is a problem if your agent has access to
           | anything.
           | 
           | The local models are quite weak here.
        
             | dataflow wrote:
             | Security is not a concern for the purpose of my question
             | here, please ignore that for now. I'm just looking for text
             | summary and search functionality here, not looking to give
             | it full system access and let it loose on my computer or
             | network. I can easily set up VM/sandboxing/airgapping/etc.
             | as needed.
             | 
             | My question is really just about what can handle that
             | volume of data (ideally, with the quoted
             | sections/duplications/etc. that come with email chains) and
             | still produce useful (textual) output.
        
         | adamkittelson wrote:
         | Anecdotal but for some reason I had a pretty bad time with
         | qwen3.5 locally for tool usage. I've been using GPT-OSS-120B
         | successfully and switched to qwen so that I could process
         | images as well (I'm using this for a discord chat bot).
         | 
         | Everything worked fine on GPT but Qwen as often as not
         | preferred to pretend to call a tool and not actually call it.
         | After much aggravation I wound up just setting my bot / llama
         | swap to use gpt for chat and only load up qwen when someone
         | posts an image and just process / respond to the image with
         | qwen and pop back over to gpt when the next chat comes in.
        
           | GorbachevyChase wrote:
           | You are responsible for the dead internet theory.
        
         | dhblumenfeld1 wrote:
         | Have you found that using a frontier model for planning and
         | small local model for writing code to be a solid workflow? Been
         | wanting to experiment with relying less on Claude Code/Codex
         | and more on local models.
        
         | eek2121 wrote:
         | Qwen is actually really good at code as well. I used
         | qwen3-coder-next a while back and it was every bit as good as
         | claude code in the use cases I tested it in. Both made the same
         | amount of mistakes, and both did a good job of the rest.
        
         | sakesun wrote:
         | Becoming a retired builder is the ultimate bliss.
        
         | chrisweekly wrote:
         | Thanks for this, Mark. And for your website and books and
         | generosity of spirit. Signal in the noise. Have an awesome
         | weekend!
        
       | andy_ppp wrote:
       | Is it correct that there's zero improvement in performance
       | between M4 (+Pro/Max) and M5 (+Pro/Max) the data looks identical.
       | Also the memory does not seem to improve performance on larger
       | models when I thought it would have?
       | 
       | Love the idea though!
       | 
       | EDIT: Okay the whole thing is nonsense and just some rough
       | guesswork or asking an LLM to estimate the values. You should
       | have real data (I'm sure people here can help) and put ESTIMATE
       | next to any of the combinations you are guessing.
        
         | GeekyBear wrote:
         | > Is it correct that there's zero improvement in performance
         | between M4 (+Pro/Max) and M5 (+Pro/Max)
         | 
         | Preliminary testing did not come to that conclusion.
         | 
         | > Apple's New M5 Max Changes the Local AI Story
         | 
         | https://www.youtube.com/watch?v=XGe7ldwFLSE
        
           | lostmsu wrote:
           | From the video: 4.4k is "almost" 4x times 1.8k because 4.4k
           | has "number 4" in the beginning, and the other one - number
           | 1.
           | 
           | For the lazy: that's less then 3x: 1.8 * 3 = 5.4
        
             | andy_ppp wrote:
             | It's not even the largest part, just prefill so I think
             | maybe M5 Max is 30% faster overall. Still pretty good I
             | think but the 4x nonsense is just marketing!
        
       | mkagenius wrote:
       | Literally made the same app, 2 weeks back -
       | https://news.ycombinator.com/item?id=47171499
        
         | mongrelion wrote:
         | What front-end framework did you use? I find the UI so visually
         | appealing
        
           | mkagenius wrote:
           | Thanks. I actually used Google AI Studio for this. Prompted
           | with my color choices and let it do the rest, turned out
           | pretty good.
        
       | zitterbewegung wrote:
       | The M4 Ultra doesn't exist and there is more credible rumors for
       | an M5 Ultra. I wouldn't put a projection like that without
       | highlighting that this processor doesn't exist yet.
        
       | rcarmo wrote:
       | This is kind of bogus since some of the S and A tier models are
       | pretty useless for reasoning or tool calls and can't run with any
       | sizable system prompt... it seems to be solely based on tokens
       | per second?
        
       | polyterative wrote:
       | awesome, needed this
        
       | tristor wrote:
       | This does not seem accurate based on my recently received M5 Max
       | 128GB MBP. I think there's some estimates/guesswork involved, and
       | it's also discounting that you can move the memory divider on
       | Unified Memory devices like Apple Silicon and AMD AI Max 395+.
        
       | bheadmaster wrote:
       | Missing 5060 Ti 16GB
        
       | tencentshill wrote:
       | Missing laptop versions of all these chips.
        
       | mopierotti wrote:
       | This (+ llmfit) are great attempts, but I've been generally
       | frustrated by how it feels so hard to find any sort of guidance
       | about what I would expect to be the most straightforward/common
       | question:
       | 
       | "What is the highest-quality model that I can run on my hardware,
       | with tok/s greater than <x>, and context limit greater than <y>"
       | 
       | (My personal approach has just devolved into guess-and-check,
       | which is time consuming.) When using TFA/llmfit, I am immediately
       | skeptical because I already know that Qwen 3.5 27B Q6 @ 100k
       | context works great on my machine, but it's buried behind
       | relatively obsolete suggestions like the Qwen 2.5 series.
       | 
       | I'm assuming this is because the tok/s is much higher, but I
       | don't really get much marginal utility out of tok/s speeds beyond
       | ~50 t/s, and there's no way to sort results by quality.
        
         | J_Shelby_J wrote:
         | It's a hard problem. I've been working on it for the better
         | part of a year.
         | 
         | Well, granted my project is trying to do this in a way that
         | works across multiple devices and supports multiple models to
         | find the best "quality" and the best allocation. And this puts
         | an exponential over the project.
         | 
         | But "quality" is the hard part. In this case I'm just choosing
         | the largest quants.
        
           | mopierotti wrote:
           | Supporting all the various devices does sound quite
           | challenging.
           | 
           | I wouldn't expect a perfect single measurement of "quality"
           | to exist, but it seems like it could be approximated enough
           | to at least be directionally useful. (e.g. comparing
           | subsequent releases of the same model family)
        
         | downrightmike wrote:
         | LLMs are just special purpose calculators, as opposed to normal
         | calculators which just do numbers and MUST be accurate. There
         | aren't very good ways of knowing what you want because the
         | people making the models can't read your mind and have
         | different goals
        
         | comboy wrote:
         | What is the $/Mtok that would make you choose your time vs
         | savings of running stuff locally?
         | 
         | Just to be clear, it may sound like a snarky comment but I'm
         | really curious from you or others how do you see it. I mean
         | there are some batches long running tasks where ignoring
         | electricity it's kind of free but usually local generation is
         | slower (and worse quality) and we all kind of want some stuff
         | to get done.
         | 
         | Or is it not about the cost at all, just about not pushing your
         | data into the clouds.
        
           | wilkystyle wrote:
           | For me it's a combination of privacy and wanting to be able
           | to experiment as much as I want without limits. I'd happily
           | take something that is 80% as good as SOTA but I can run it
           | locally 24/7. I don't think there's anything out there yet
           | that would _100%_ obviate my desire to at least
           | _occasionally_ fall back to e.g. Claude, but I think most of
           | it could be done locally if I had infinite tokens to throw at
           | it.
        
           | mopierotti wrote:
           | Good question. I agree with what I think you're implying,
           | which is that local generation is not the right choice if you
           | want to maximize results per time/$ spent. In my experience,
           | hosted models like Claude Opus 4.6 are just so effective that
           | it's hard to justify using much else.
           | 
           | Nevertheless, I spend a lot of time with local models because
           | of:
           | 
           | 1. Pure engineering/academic curiosity. It's a blast to
           | experiment with low-level settings/finetunes/lora's/etc. (I
           | have a Cog Sci/ML/software eng background.)
           | 
           | 2. I prefer not to share my data with 3rd party services, and
           | it's also nice to not have to worry too much about
           | accidentally pasting sensitive data into prompts (like
           | personal health notes), or if I'm wasting $ with silly
           | experiments, or if I'm accidentally poisoning some stateful
           | cross-session 'memories' linked to an account.
           | 
           | 3. It's nice to be able solve simple tasks without having to
           | reason about any external 'side-effects' outside my machine.
        
           | phillmv wrote:
           | i can think of some tasks (classification, structured info
           | extraction) that i _imagine_ even small meh models could do
           | quite well at
           | 
           | on data i would never ever want to upload to any vendor if i
           | can avoid it
        
         | 0xbadcafebee wrote:
         | Too generic question. Gotta be more specific:
         | "what is the best open weight model for high-quality coding
         | that fits in 8GB VRAM and 32GB system RAM with t/s >= 30 and
         | context >= 32768" -> Qwen2.5-Coder-7B-Instruct
         | "what is the best open weight model for research w/web search
         | that fits in 24GB VRAM and 32GB system RAM with t/s >= 60 and
         | context >= 400k" -> Qwen3-30B-A3B-Instruct-2507
         | "what is the best open weight embedding model for RAG on a
         | collection of 100,000 documents that fits in 40GB VRAM and
         | 128GB system RAM with t/s >= 50 and context >= 200k" ->
         | Qwen3-Embedding-8B
         | 
         | Specific models & sizes for specific use cases on specific
         | hardware at specific speeds.
        
       | SXX wrote:
       | Sorry if already been answered, but will there be a metric for
       | latency aka time to first token?
       | 
       | Since I considered buying M3 Ultra and feel like it the most
       | often discussed regarding using Apple hardware for runninh local
       | LLMs. Where speed might be okay, but prompt processing can take
       | ages.
        
         | teaearlgraycold wrote:
         | Wait for the M5 Ultra. It will get the 4x prompt processing
         | speeds from the rest of the M5 product line. I hear rumors it
         | will be released this year.
        
       | tkfoss wrote:
       | Nice UI, but crap data, probably llm generated.
        
       | anigbrowl wrote:
       | Useful tool, although some of the dark grey text is dark that I
       | had to squint to make it out against the background.
        
       | lagrange77 wrote:
       | Finally! I've been waiting for something like this.
        
       | mmaunder wrote:
       | OP can you please make it not as dark and slightly larger. Super
       | useful otherwise. Qwen 3.5 9B is going to get a lot of love out
       | of this.
        
         | ProllyInfamous wrote:
         | I'm not usually one to whine, but agreed; additionally, add
         | contrast to the modifiers (e.g. processor select). First thing
         | I did when I visited was scale the website to 150%
         | 
         | Super impressive comparisons, and correlates with my perception
         | having three seperate generations of GPU (from your list
         | pulldown). Thanks for including the "old AMD" Polaris chipsets,
         | which are actually _still much faster_ than lower-spec Apple
         | silicon. I have Ollama3.1 on a VEGA64 and it really is _twice
         | as fast as an M2Pro_...
         | 
         | ----
         | 
         | For anybody that thinks installing a local LLM is complicated:
         | it's not (so long as you have more than one computer, don't
         | tinker on your primary workhorse). I am a blue collar
         | electrician (admittedly: geeky); no more difficult than
         | installing linux. I used an online LLM to help me install both
         | =D
        
         | ricardbejarano wrote:
         | OP here, it's not mine though!
        
         | aanet wrote:
         | +1
         | 
         | The website is super useful. That theme though... low-contrast
         | text on too-dark theme is, uh, barely readable for me.
        
       | reactordev wrote:
       | This shows no models work with my hardware but that's furthest
       | from the truth as I'm running Qwen3.5...
       | 
       | This isn't nearly complete.
        
         | kennywinker wrote:
         | Well... don't keep us guessing -what hardware? And which size
         | qwen3.5?
        
       | azmenak wrote:
       | From my personal testing, running various agentic tasks with a
       | bunch of tool calls on an M4 Max 128GB, I've found that running
       | quantized versions of larger models to produce the best results
       | which this site completely ignores.
       | 
       | Currently, Nemotron 3 Super using Unsloth's UD Q4_K_XL quant is
       | running nearly everything I do locally (replacing Qwen3.5 122b)
        
       | bearjaws wrote:
       | So many people have vibe coded these websites, they are posted to
       | Reddit near daily.
        
       | kuon wrote:
       | I have amd 9700 and it is not listed while it is great llm
       | hardware because it has 32Gb for a reasonable price. I tried
       | doing "custom" but it didn't seem to work.
       | 
       | The tool is very nice though.
        
       | ipunchghosts wrote:
       | What is S? Also, NVIDIA RTX 4500 Ada is missing.
        
       | fraywing wrote:
       | This is amazing. Still waiting for the "Medusa" class AMD chips
       | to build my own AI machine.
        
       | kpw94 wrote:
       | People complaining about how hard to get simple answer is don't
       | appreciate the complexity in figuring out optimal models...
       | 
       | There's so many knobs to tweak, it's a non trivial problem
       | 
       | - Average/median length of your Prompts
       | 
       | - prompt eval speed (tok/s)
       | 
       | - token generation speed (tok/s)
       | 
       | - Image/media encoding speed for vision tasks
       | 
       | - Total amount of RAM
       | 
       | - Max bandwidth of ram (ddr4, ddr5, etc.?)
       | 
       | - Total amount of VRAM
       | 
       | - "-ngl" (amount of layers offloaded to GPU)
       | 
       | - Context size needed (you may need sub 16k for OCR tasks for
       | instance)
       | 
       | - Size of billion parameters
       | 
       | - Size of active billion parameters for MoE
       | 
       | - Acceptable level of Perplexity for your use case(s)
       | 
       | - How aggressive Quantization you're willing to accept (to
       | maintain low enough perplexity)
       | 
       | - even finer grain knobs: temperature, penalties etc.
       | 
       | Also, Tok/s as a metric isn't enough then because there's:
       | 
       | - thinking vs non-thinking: which mode do you need?
       | 
       | - models that are much more "chatty" than others in the same area
       | (i remember testing few models that max out my modest desktop
       | specs, qwen 2.5 non-thinking was so much faster than equivalent
       | ministral non-thinking even though they had equivalent tok/s...
       | Qwen would respond to the point quickly)
       | 
       | At the end, final questions are: are you satisfied with how long
       | getting an answer took? and was the answer good enough?
       | 
       | The same exercise with paid APIs exists too, obviously less knobs
       | but depending on your use case, there's still differences between
       | providers and models. You can abstract away a lot of the knobs ,
       | just add "are you satisfied with how much it cost" on top of the
       | other 2 questions
        
       | paxys wrote:
       | I wish creators of local model inference tools (LM Studio, Ollama
       | etc.) would release these numbers publicly, because you can be
       | sure they are sitting on a large dataset of real-world
       | performance.
        
       | gopalv wrote:
       | Chrome runs Gemini Nano if you flip a few feature flags on [1].
       | 
       | The model is not great, but it was the "least amount of setup"
       | LLM I could run on someone else's machine.
       | 
       | Including structured output, but has a tiny context window I
       | could use.
       | 
       | [1] - https://notmysock.org/code/voice-gemini-prompt.html
        
       | vednig wrote:
       | Our work at DoShare is a lot of this stuff we've been on it for 2
       | years
        
       | sidchilling wrote:
       | I have been trying to run Qwen Coder models (8B at 4bit) on my M3
       | Pro 18GB behind Ollama and connecting codex CLI to it. The tool
       | usage seems practically zero, like it returns the tool call in
       | text JSON and codex CLI doesn't run the tool (just displays the
       | tool call in text). Has anyone succeeded in doing something like
       | this? What am I missing?
        
         | MikeNotThePope wrote:
         | I have the same hardware. Been curious about trying it with
         | Opencode.
        
         | mongrelion wrote:
         | It might be that the system prompt sent by codex is not optimal
         | for that model. Try with open code and see if your results
         | improve
        
       | starkeeper wrote:
       | This is awesome!!!
       | 
       | Could you please add title="explanation" over each selected item
       | at the top. For example, when I choose my video card the ram
       | changes... I'm not sure if the RAM selection is GPU RAM? The GRAM
       | was already listed with the graphics card. SO I choose 96GB which
       | is my Main memory? And the GB/s I am assuming it's GPU -> CPU
       | speed?
        
       | pants2 wrote:
       | This really highlights the impracticality of local models:
       | 
       | My $3k Macbook can run `GPT-OSS 20B` at ~16 tok/s according to
       | this guide.
       | 
       | Or I can run `GPT-OSS 120B` (a 6X larger model) at 360 tok/s (30X
       | faster) on Groq at $0.60/Mtok output tokens.
       | 
       | To generate $3k worth of output tokens on my local Mac at that
       | pricing it would have to run 10 years continuously without
       | stopping.
       | 
       | There's virtually no economic break-even to running local models,
       | and no advantage in intelligence or speed. The only thing you
       | really get is privacy and offline access.
        
         | danny_codes wrote:
         | A million tokens is like 5 minutes of inference for heavy
         | coding use.
        
           | girvo wrote:
           | At work I regularly hit my 7.5mil tokens per hour limit one
           | of our tools has, and have to switch model of tool, and I'm
           | not even really a remotely heavy user. I think people don't
           | realise how many tokens get burned with CoT and tool calls
           | these days
           | 
           | At 7.5mil per hour hard limit, 84 days to hit the
           | grandparents $3k
           | 
           | That said local models really are slow still, or fast enough
           | and not that great
        
         | xandrius wrote:
         | You're saying it as if privacy was worthless? Also not many
         | people would consider the price of buying a macbook and put it
         | strictly towards running a local model.
         | 
         | Instead if you wanted to get a macbook anyway, you get to run
         | local models for free on top. Very different story.
        
           | pants2 wrote:
           | The privacy angle is not that interesting to me.
           | 
           | - You can find inference providers with whatever privacy
           | terms you're looking for
           | 
           | - If you're using LLMs with real data (let's say handling
           | GMail) then Google has your data anyway so might as well use
           | Gemini API
           | 
           | - Even if you're a hardcore roll-your-own-mail-server type,
           | you probably still use a hosted search engine and have gotten
           | comfortable with their privacy terms
           | 
           | Also on cost the point is you can use an API that's many
           | times smarter and faster for a rounding error in cost
           | compared to your Mac. So why bother with local except for the
           | cool factor?
        
       | amdivia wrote:
       | I found this to be inaccurate, I can run OSS GPT 120B (4 bit
       | quant) on my 5090 and 64 ram system with around 40 t/s. Yet here
       | the site claims it won't work
        
       | Readerium wrote:
       | Qwen 3.5 4B is the goat then
        
       | ThrowawayTestr wrote:
       | For image generation or even video generation, local models are
       | totally feasible. I can generate a 5 second clip with wan 2.2 in
       | about 30 minutes on my 3060 12G. Plus, I have full control on the
       | loras used.
        
       | dzink wrote:
       | This would be wonderful if it is accurate - instead of
       | guesstimating, let people report their actual findings. I can
       | confirm GLM 4.7 is possible on M1 Max and it can do nice
       | comprehensive answers (albeit at 12 min an answer) locally. You
       | can also easily do Mistral7B and OSS 20B and others. Structure it
       | as a way to report accruals, similarly to Levels.xyz for
       | salaries, instead of guestimating.
        
       | nicklo wrote:
       | the animation of the model name text when opening the detail view
       | is so smooth and delightful
        
       | torginus wrote:
       | Huh, I never knew my browser just volunteers my exact hardware
       | specs to any website without so much as even notifying me about
       | it.
        
         | Jaxan wrote:
         | It doesn't really. The website thinks I'm on a iPhone 19 pro,
         | although I'm actually on a iPhone SE 1st gen. So it's off by
         | roughly a decade.
        
           | torginus wrote:
           | Maybe that's one of Safari's numerous 'quirks' our frontend
           | devs keep bitching about.
           | 
           | Which in this case Im thankful that Apple isn't too keen on
           | following standards like these.
        
           | weikju wrote:
           | > on a iPhone 19 pro
           | 
           | I wish the website could tell us how life is like in 2027!
        
         | ebbi wrote:
         | I thought that's how airlines do the whole trickery around
         | having different pricing if you access the site from Windows or
         | Mac...
        
         | DanielHB wrote:
         | This stuff is used a lot in browser fingerprinting for tracking
         | purposes. More privacy-focused browsers usually feed randomized
         | info.
        
       | comrade1234 wrote:
       | I can't tell at a glance what this page is showing, but I am
       | curious about the licenses on the various models that let me run
       | it locally and make money off it. Awhile ago only deepseek let
       | you do that - not sure now.
        
         | mind_heist wrote:
         | nice, this is an interesting idea. Can you elaborate on the
         | licensing issue ? how do you get blocked for using the models
         | commercially ?
        
           | comrade1234 wrote:
           | Just read the license agreement. Last time I looked into this
           | the only model I could run locally and do what I want was
           | deepseek. I think it was the MIT license. The others had
           | various restrictions that just didn't make it worth it.
           | 
           | I stopped researching this because buying the hardware to run
           | deepseek full model just isn't practical right now. Our
           | customers will have to be happy with us sending data to
           | OpenAI/deepseek/etc if they want to use those features.
        
             | singpolyma3 wrote:
             | qwen3.5 is just apache
        
       | urba_ wrote:
       | Man, I wonder when there will be AI server farms made from iCloud
       | locked jailbroken iPhone 16s with backported MacOS
        
       | Akuehne wrote:
       | Can we get some of the ancient Nvidia Teslas, like the p40 added?
        
       | dirk94018 wrote:
       | We wrote the linuxtoaster inference engine, toasted, and are
       | getting 400 prefill, 100 gen on a M4 Max w 128GB RAM on
       | Qwen3-next-coder 6bit, 8bit runs too. KV caching means it feels
       | snappy in chat mode. Local can work. For pro work, programming,
       | I'd still prefer SOTA models, or GLM 4.7 via Cerebras.
        
       | 0xbadcafebee wrote:
       | Couple thoughts:
       | 
       | - The t/s estimation per machine is off. Some of these models run
       | generation at twice the speed listed (I just checked on a couple
       | macs & an AMD laptop). I guess there's no way around that, but
       | some sort of sliding scale might be better.
       | 
       | - Ollama vs Llama.cpp vs others produce different results. I can
       | run gpt-oss 20b with Ollama on a 16GB Mac, but it fails with "out
       | of memory" with the latest llama.cpp (regardless of param tuning,
       | using their mxfp4). Otoh, when llama.cpp does work, you can
       | usually tweak it to be faster, if you learn the secret arts (like
       | offloading only specific MoE tensors). So the t/s rating is even
       | more subjective than just the hardware.
       | 
       | - It's great that they list speed and size per-quant, but that
       | needs to be a filter for the main list. It might be "16 t/s" at
       | Q4, but if it's a small model you need higher quant (Q5/6/8) to
       | not lose quality, so the advertised t/s should be one of those
       | 
       | - Why is there an initial section which is all "performs poorly",
       | and then "all models" below it shows a ton of models that perform
       | well?
        
       | TheCapn wrote:
       | @OP are you the creator? Could you add my GPU to the list?
       | 
       | Radeon VII
       | 
       | https://www.amd.com/en/support/downloads/drivers.html/graphi...
        
       | sand500 wrote:
       | How does it have details for M4 ultra?
        
       | ementally wrote:
       | In mobile section it is missing Tensor chips (used by Google
       | Pixel devices).
        
       | johneth wrote:
       | Re: the design of the site. Please use higher contrast colours,
       | especially the barely visible grey text on black background. It's
       | annoying to try to read.
        
       | nazbasho wrote:
       | its perfect
        
       | adamhsn wrote:
       | Cool project!!
       | 
       | It would be useful to filter which model to use based on the
       | objective or usage (i.e., for data extraction vs. coding).
       | 
       | Also, just looking at VRAM kind of misses that a lot of CPU
       | memory can be shared with the GPU via layer offloading. I think
       | there is ultimately a need for a native client, like a CPU/GPU
       | benchmark, to figure out how the model will actually perform more
       | precisely.
        
       ___________________________________________________________________
       (page generated 2026-03-13 23:00 UTC)