[HN Gopher] Running Qwen3 on your macbook, using MLX, to vibe co...
       ___________________________________________________________________
        
       Running Qwen3 on your macbook, using MLX, to vibe code for free
        
       Author : avetiszakharyan
       Score  : 243 points
       Date   : 2025-05-01 11:54 UTC (11 hours ago)
        
 (HTM) web link (localforge.dev)
 (TXT) w3m dump (localforge.dev)
        
       | avetiszakharyan wrote:
       | I'd thought to share this quick tutorial to get an actual
       | autonomous agent running on your local and doing some simple
       | tasks. Still in progress trying to figure ou right MLX settings
       | or proper model version to do it, but the framework around this
       | approach is solid, so i hought i'd share!
        
         | nottorp wrote:
         | Now how do you feed it an existing codebase as part of your
         | prompt? Does it even support that (prompt size etc).
        
           | avetiszakharyan wrote:
           | Ye you can just run it in a folder, and ask it to look
           | around, it can execute bash commands, do anything that Claude
           | Code can do. it will read all the codebase if it has to
        
             | pylotlight wrote:
             | Typically I'd use tools for that as context will be finite,
             | but I hear it does a decent job at tool calling too so
             | should see solid perf there.
        
               | avetiszakharyan wrote:
               | 8b was very bad at tool calling. 30b does an "ok" job but
               | requires the wrapper to do a lot of job, like for example
               | the model will be very random, either have multiple
               | tool_call ags, or have multiple tool calls in one tag
               | inconsistant. but "handle-able"
        
       | walthamstow wrote:
       | Looks good. I've been looking for a local-first AI-assisted IDE
       | to work with Google's Gemma 3 27B
       | 
       | I do think you should disclose that Localforge is your own
       | project though.
        
         | danw1979 wrote:
         | Personally, I assumed that a blog post on the domain
         | localforge.dev was written by the developers of localforge, but
         | I might be wrong.
        
           | SquareWheel wrote:
           | They likely mean that the submitter, avetiszakharyan, should
           | disclose their relationship to Localforge.
        
             | zarathustreal wrote:
             | Fascinating.. I wonder how much of the economy runs on
             | social proof
        
               | SquareWheel wrote:
               | It's not uncommon on HN! We frequently have people
               | chiming in as CEOs, insiders, and experts in various
               | fields without much proof. Generally, it hasn't been a
               | problem. Or at least I've not seen any examples of having
               | wool pulled over our eyes in this fashion.
        
           | walthamstow wrote:
           | Sure, if you already know what Localforge is before clicking.
        
             | avetiszakharyan wrote:
             | Where do i put that, in the blogpost or?
        
               | walthamstow wrote:
               | If you can still edit it, adding it to your first comment
               | is fine I would say. "Disclosure: I am the author of
               | Localforge" or similar.
        
               | avetiszakharyan wrote:
               | No thats the only thing I can't edit tbh :(
        
             | tasuki wrote:
             | I didn't know, and still assumed the blog post on
             | localforge.dev was written by the localforge.dev people.
             | Who else?
        
       | endlessvoid94 wrote:
       | I've found the local models useful for non-coding tasks, however
       | the 8B parameter models so far have proven lacking _enough_ for
       | coding tasks that I 'm waiting another few months for whatever
       | the Moore's law equivalent of LLM power is to catch up. Until
       | then, I'm sticking with Sonnet 3.7.
        
         | walthamstow wrote:
         | If you have a 32GB Mac then you should be able to run up to 27B
         | params, I have done so with Google's `gemma3:27b-it-qat`
        
           | endlessvoid94 wrote:
           | Hm, I've got an M2 air w/ 24GB. Running the 27B model was
           | _crawling_. Maybe I had something misconfigured.
        
             | 100721 wrote:
             | No, that sounds right. 24GB isn't enough to feasibly run
             | 27B parameters. The rule of thumb is approximately 1GB of
             | ram per billion parameters.
             | 
             | Someone in another comment on this post mentioned using one
             | of the micro models (Qwen 0.6B I think?) and having decent
             | results. Maybe you can try that and then progressively move
             | upwards?
             | 
             | EDIT: "Queen" -> "Qwen"
        
               | simonw wrote:
               | You also need to leave space for other apps. If you run a
               | 27B model on a 32GB machine you may find that you can't
               | productively run other apps.
               | 
               | I have 64GB and I can only just fit a bunch of Firefox
               | and VS Code windows at the same time as running a 27B
               | model.
        
               | brandall10 wrote:
               | That rule of thumb is only related to 8 bit quants at low
               | context. The default for ollama is 4 bit, which puts it
               | roughly about 14GB.
               | 
               | The vast majority of people run between 4-6 bit depending
               | on system capability. The extra accuracy above 6 tends to
               | not be worth it relative to the performance hit.
        
             | redman25 wrote:
             | I think only 2/3 of ram is allocated to be available to the
             | gpu, so like 14gb which is probably not enough to run even
             | Q4 quant.
        
               | tstrimple wrote:
               | This is configurable by the way.
               | 
               | sudo sysctl iogpu.wired_limit_mb=12345
        
           | alkh wrote:
           | How much RAM was it taking during inference?
        
             | walthamstow wrote:
             | 15.4GB during inference according to Activity Monitor
        
               | alkh wrote:
               | Oh, nice, that's actually not bad at all. Thanks, will
               | give it a try on my 36Gb Mac
        
       | ttoinou wrote:
       | Great thank you. Side topic : anyone knows a way to have a
       | centralized proxy to all LLMs services, online or local, that
       | lets our services connect to it and we manage access to LLMs only
       | once there ? And also records calls to LLM. Would make the whole
       | UX of switching LLMs weekly easier, we would only reconfigure the
       | proxy. I know only LiteLLM that can do that but its record of all
       | LLMs calls is a bit clunky to use properly
        
         | mnholt wrote:
         | I've been looking for this for my team but haven't found it.
         | Providers like OpenAI and Anthropic offer admin token to manage
         | team accounts and you look hook into Ollama or another self
         | managed service for local AI.
         | 
         | Seems like a great way to roll out AI to a medium sized team
         | where a very small team can coordinate access to the best
         | available tools so the entire team doesn't need to keep pace at
         | the current break-neck speed.
        
         | ramesh31 wrote:
         | https://openrouter.ai/
        
         | Havoc wrote:
         | Litellm is definitely your best bet. For recording - you can
         | probably vibe code a proxy in front of it that mitms it and
         | dumps the request into whatever format you need
        
           | rcarmo wrote:
           | Litellm can log stuff pretty well on its own.
        
         | tidbeck wrote:
         | Could you maybe make use of Simon Willsons [LLM
         | lib/app](https://github.com/simonw/llm)? It has great LLM
         | support (just pass in the model to use) and records everything
         | by default.
        
           | simonw wrote:
           | The one feature missing from LLM core for this right now is
           | serving models over an HTTP OpenAI-compatible local server.
           | There's a plugin you can try for that here though:
           | https://github.com/irthomasthomas/llm-model-gateway
        
         | calebkaiser wrote:
         | I'm a maintainer of Opik, an open source LLM eval/observability
         | framework. If you use something like LiteLLM or OpenRouter to
         | handle the proxying of requests, Opik basically provides an
         | out-of-the-box recording layer via its integrations with both:
         | 
         | https://github.com/comet-ml/opik
        
       | chuckadams wrote:
       | Anyone know of a setup, perhaps with MCP, where I can get my
       | local LLM to work in tandem on tasks, compress context, or
       | otherwise act in concert with the cloud agent I'm using with
       | Augment/Cursor/whatever? It seems silly that my shiny new M3 box
       | just renders the UI while the cloud LLM alone refactors my
       | codebase, I feel they could negotiate the tasks between
       | themselves somehow.
        
         | _joel wrote:
         | There's a few Ollama-MCP bridge servers already (from a quick
         | search, also interested myself):
         | 
         | ollama-mcp-bridge: A TypeScript implementation that "connects
         | local LLMs (via Ollama) to Model Context Protocol (MCP)
         | servers. This bridge allows open-source models to use the same
         | tools and capabilities as Claude, enabling powerful local AI
         | assistants"
         | 
         | simple-mcp-ollama-bridge: A more lightweight bridge connecting
         | "Model Context Protocol (MCP) servers to OpenAI-compatible LLMs
         | like Ollama"
         | 
         | rawveg/ollama-mcp: "An MCP server for Ollama that enables
         | seamless integration between Ollama's local LLM models and MCP-
         | compatible applications like Claude Desktop"
         | 
         | How you route would be an interesting challenge, presumably
         | could just tell it to use the mcp for certain tasks, thereby
         | offloading locally.
        
         | rcarmo wrote:
         | I've been toying with Visual Studio Code's MCP and agent
         | support and gotten it to offload things like reference searches
         | and targeted web crawling (look up module X on git repo Y via
         | this URL pattern that the MCP server goes, fetches and parses).
         | 
         | I started by giving it a reference Python MCP server and asking
         | it to modify the code to do that. Now I have 3-4 tools that
         | give me reproducible results.
        
       | crazymoka wrote:
       | Why do you need mlx? Like your blog post by you never explain why
       | things need to be used.
       | 
       | Why isn't using localforge enough as it ties into models?
        
         | turnsout wrote:
         | I believe mlx will allow you to run the models marginally
         | faster (per a recent blog post by @simonw)
        
           | simonw wrote:
           | Yeah, you don't necessarily _need_ it but it 's optimized for
           | Apple Silicon and in my experience feels like it gives
           | slightly better performance than GGUFs. I really need to
           | formally measure that so I'm not just running on vibes!
        
             | indigodaddy wrote:
             | I for one, am willing to just trust you bro ;)
        
               | turnsout wrote:
               | Yeah I'll go with Simon's vibes over most people's
               | measurements!
        
         | freeone3000 wrote:
         | mlx is an alternative model format to GGUF. It executes
         | natively on apple silicon using Apple's AI accelerator, rather
         | than through GGUF as a compute shader(!). It's faster and uses
         | fewer resources on Apple devices.
        
         | avetiszakharyan wrote:
         | I was just trying to make sure is maximally performant, and did
         | it with MLX because i am running on mac hardware and wanted to
         | be able to run 30b in reasonable time so it can actually
         | autonomously code something. Otherwise there are many ways to
         | do it!
        
           | p0w3n3d wrote:
           | It would be nice if you mentioned it's about apple silicon,
           | and not apple intel computers. They're still ubiquitous
           | nowadays
        
             | Tokumei-no-hito wrote:
             | we're on the 4th generation of silicon now
        
       | omneity wrote:
       | I'm using Qwen3-30B-A3B locally and it's very impressive. Feels
       | like the GPT-4 killer we were waiting for for two years. I'm
       | getting 70 tok/s on an M3 Max, which is pushing it into the "very
       | usable" quadrant.
       | 
       | What was even more impressive is the 0.6B model which made the
       | sub 1B actually useful for non-trivial tasks.
       | 
       | Overall very impressed. I am evaluating how it can integrate with
       | my current setup and will probably report somewhere about that.
        
         | mtw wrote:
         | how much RAM do you have? I want to compare with my local setup
         | (M4 Pro)
        
           | omneity wrote:
           | 128GB but it's not using much.
           | 
           | I'm running Q4 and it's taking 17.94 GB VRAM with 4k context
           | window, 20GB with 32k tokens.
        
             | A4ET8a8uTh0_v2 wrote:
             | I am not a mac person, but I am debating buying one for the
             | unified ram now that the prices seem to be inching down. Is
             | it painful to set up? The general responses I seem to get
             | range from "It is takes zero effort" to "It was a major
             | hassle to set everything up."
        
               | dghlsakjg wrote:
               | Read the article you are commenting on. It is a how to
               | that answers your exact question. It takes 4 commands in
               | the terminal.
        
               | bloqs wrote:
               | The article should answer your question. Or do you mean
               | setting up a Mac for use as a Linux or windows user
        
               | A4ET8a8uTh0_v2 wrote:
               | << Or do you mean setting up a Mac for use as a Linux or
               | windows user
               | 
               | This part, yes. I assume the setting a complete
               | environment is a little more involved than the 4 commands
               | sibling is also refers to.
        
               | tough wrote:
               | I mean that's more like old habits no? MacOS is pretty
               | easy to setup. if all your target apps are available on
               | it, you shouldn't have much of a problem.
               | 
               | As a Windows/MacOS/Linux dweller kinto is a godsend so I
               | can have macos keyboard (but you could have linux or
               | windows by default) on all OSes https://kinto.sh/
        
               | Shraal wrote:
               | It's relatively easy. macOS is easier to set up than
               | Linux, but it will always depend on your specific needs
               | and environment.
               | 
               | E.g.: I go a little bit overboard for the average macOS
               | user:
               | 
               | - custom system- and app-specific keyboard mappings
               | (ultra-modifier on caps-lock; custom tabbing-key-
               | modifier) via Karabiner Elements
               | 
               | - custom trackpad mappings via BetterTouchTool
               | 
               | - custom Time Machine schedule and backup logic; you can
               | vibe-code your install script once and re-use it in the
               | future; just make it idempotent
               | 
               | - custom quake-like Terminal via iTerm
               | 
               | - shell customizations
               | 
               | - custom Alfred workflows
               | 
               | - etc.
               | 
               | If all you need is just a sensible package manager and
               | the terminal to get started, just set up Time Machine
               | with default settings, Homebrew, your shell, and
               | optionally iTerm2, and you're good to go. Other
               | noteworthy power-user tools:
               | 
               | - Hammerspoon
               | 
               | - Syncthing / Resilio Sync
               | 
               | - Arq. Naturally, the usual backup tools also run on
               | macOS: Borg, Kopia, etc.
               | 
               | - Affinity suite for image processing
               | 
               | - Keyshape for animations on web, mobile, etc.
        
               | simonw wrote:
               | LM Studio and Ollama are both very low complexity ways to
               | get local LLMs running on a Mac.
               | 
               | As a Python person I've found uv + MLX to be pretty
               | painless on a Mac too.
        
               | MR4D wrote:
               | You can use the method in this tutorial or you can
               | download LM Studio and run it.
               | 
               | The latter is super easy. Just download the model (thru
               | the GUI) and go.
        
               | avetiszakharyan wrote:
               | Honestly, it is quite a hastle, took me 2 hours BUT. if
               | you just take the whole article text and paste that to
               | gemini-2.5-pro and give your circumstance, i think it
               | will give you specific steps for your case and it should
               | be trivial from that moment on
        
               | PhilippGille wrote:
               | > I am not a mac person, but I am debating buying one for
               | the unified ram
               | 
               | Soon some AMD Ryzen AI Max PCs will be available, with
               | unified memory as well. For example the Framework Desktop
               | with up to 128 GB, shared with the iGPU:
               | 
               | - Product: https://frame.work/us/en/desktop?tab=overview
               | 
               | - Video, discussing 70B LLMs at around 3m:50s :
               | https://youtu.be/zI6ZQls54Ms
        
               | A4ET8a8uTh0_v2 wrote:
               | oh boy... i was genuinely hoping someone less pricy would
               | enter this market
               | 
               | edit: ok.. i am excited.
        
           | dust42 wrote:
           | I have a MBP M1 Max 64GB and I get 40t/s with llama.cpp and
           | unsloth q4_k_m on the 30B A3B model. I always use /nothink
           | and Temperature=0.7, TopP=0.8, TopK=20, and MinP=0 - these
           | are the settings recommended for Qwen3 and they make a big
           | difference. With the default settings from llama-server it
           | will always run into an endless loop.
           | 
           | The quality of the output is decent, just keep in mind it is
           | only a 30B model. It also translates really well from french
           | to german and vice versa, much better than Google translate.
           | 
           | Edit: for comparision, Qwen2.5-coder 32B q4 is around
           | 12-14t/s on this M1 which is too slow for me. I usually used
           | the Qwen2.5-coder 17B at around 30t/s for simple tasks. Qwen3
           | 30B is imho better and faster.
           | 
           | [1] parameters for Qwen3:
           | https://huggingface.co/Qwen/Qwen3-30B-A3B
           | 
           | [2] unsloth quant:
           | https://huggingface.co/unsloth/Qwen3-30B-A3B-GGUF
           | 
           | [3] llama.cpp: https://github.com/ggml-org/llama.cpp
        
         | UK-Al05 wrote:
         | It's fits entirely in my 7900xtx memory. But tbh i've been
         | disappointed with programming ability so far.
         | 
         | It's using 20GB of memory according to ollama.
        
         | c0brac0bra wrote:
         | What tasks have you found the 0.6B model useful for? The
         | hallucination that's apparent during its thinking process put
         | up a big red flag for me.
         | 
         | Conversely, the 4B model actually seemed to work really well
         | and gave results comparable to Gemini 2.0 Flash (at least in my
         | simple tests).
        
           | SparkyMcUnicorn wrote:
           | You can use 0.6B for speculative decoding on the larger
           | models. It'll speed up 32B, but slows down 30B-A3B
           | dramatically.
        
           | omneity wrote:
           | It's okay for extracting simple things like addresses, or for
           | formatting text with some input data, like a more advanced
           | form of mail merge.
           | 
           | I haven't evaled these tasks so YMMV. I'm exploring other
           | possibilities as well. I suspect it might be decent at
           | autocomplete, and it's small enough one could consider
           | finetuning it on a codebase.
        
         | jasonjmcghee wrote:
         | Importantly they note that using a draft model screws it up and
         | this was my experience. I was initially impressed, then started
         | seeing problems, but after disabling my draft model it started
         | working much better. Very cool stuff- it's fast too as you
         | note.
         | 
         | The /think and /no_think commands are very convenient.
        
           | marcalc wrote:
           | What do you mean by draft model? And how would one disable
           | it? Cheers
        
             | _neil wrote:
             | A draft model is something that you would explicitly
             | enable. It uses a smaller model to speculatively generate
             | next tokens, in theory speeding up generation.
             | 
             | Here's the LM Studio docs on it:
             | https://lmstudio.ai/docs/app/advanced/speculative-decoding
        
           | woadwarrior01 wrote:
           | That should not be the case. Speculative decoding is trading
           | off compute for memory bandwidth. The model's output is
           | guaranteed to be the same, with or without it. Perhaps
           | there's a bug in the implementation that you're using.
        
         | avetiszakharyan wrote:
         | I went form bottom up, started with 4B, then 8B, then 30B, and
         | 30B was the only one that started to "use tools". Other models
         | were saying they will use but never did, or didnt notice all
         | the tools. I think anything above 30b would atually be able to
         | go full GPT on a task. 30b does it, but a bit.. meh
        
         | TuxSH wrote:
         | Personally, I'm getting 15 tok/s on both RTX 3060 and my
         | MacBook Air M4 (w/ 32GB, but 24 should suffice), with the
         | default config from LMStudio.
         | 
         | Which I find even more impressive, considering the 3060 is the
         | most used GPU (on Steam) and that M4 Air and future SoCs
         | are/will be commonplace too.
         | 
         | (Q4_K_M with filesize=18GB)
        
         | tomr75 wrote:
         | I'm getting 56 with mlx and lmstudio. How 76?
        
         | anon373839 wrote:
         | One of the most interesting things about that model is its
         | excellent score on the RAG confabulations (hallucination)
         | leaderboard. It's the 3rd best model overall, beating all
         | OpenAI models, for example. I wonder what Alibaba did to
         | achieve that.
         | 
         | https://github.com/lechmazur/confabulations
        
       | maille wrote:
       | I have a Windows PC with a GTX 5070 (12GB) any chance to run it?
        
         | simonw wrote:
         | I expect that will run Qwen 3 8B quite happily, and I've found
         | that to be a surprisingly capable model for its size.
        
         | UK-Al05 wrote:
         | The 30B one requires 20 GB of memory for me. But some of the
         | lower parameters one should be ok
        
           | avetiszakharyan wrote:
           | For me it was peaking at 35GB even when using
        
       | api wrote:
       | I'm really impressed and also very interested to see models I can
       | run on my MacBook Pro start to generate results close to large
       | hosted "frontier" models, and do so with what I _assume_ are far
       | fewer parameters.
       | 
       | I wonder how far this can go?
        
         | simonw wrote:
         | It's been a solid trend for the last two years: I've not
         | upgraded my laptop in the time and the quality of results I'm
         | getting from local models on that same machine has continued to
         | rise.
         | 
         | My hunch is that there's still some remaining optimization
         | fruit to be harvested but I expect we may be nearing a plateau.
         | I may have to upgrade from 64GB of RAM this year.
        
           | api wrote:
           | Seeing diffusion language models mature and get better will
           | be interesting. They can be much, much faster on less
           | hardware.
        
       | at0mic22 wrote:
       | Is there a way to achieve the same with ollama?
        
         | simonw wrote:
         | Yes, Ollama has Qwen 3 and it works great on a Mac. It may be
         | slightly slower than MLX since Ollama hasn't integrated that
         | (Apple Silicon optimized) library yet, but Ollama models still
         | use the Mac's GPU.
         | 
         | https://ollama.com/library/qwen3
        
           | avetiszakharyan wrote:
           | Yes, i did that but its not apple silicon optimized so it was
           | taking forever for 30b models. So its ok, but its not
           | fantastic
        
             | spmurrayzzz wrote:
             | You can just use llama.cpp instead (which is what ollama is
             | using under the hood via bindings). Just need to make sure
             | youre using commit `d3bd719` or newer. I normally use this
             | with nvidia/cuda, but tested on my mbp and havent had any
             | speed issues thus far.
             | 
             | Alternatively, LMStudio has MLX support you can use as
             | well.
        
       | freeone3000 wrote:
       | There needs to be more mention about the requirement of setting
       | the model-name correctly. For this tutorial to be executed top-
       | to-bottom, the model name must be "mlx-
       | community/Qwen3-30B-A3B-8bit". Other model names will result in a
       | 404 -- rightly so, as this is used to determine which model is
       | executed in mlx_lm.serve!
        
       | xnx wrote:
       | It's very cool that useful models can be run on single personal
       | computers at all. For coding, your time is very valuable, and I'd
       | never want to use anything less than the best. I'm happy to pay
       | pennies to use a frontier model with a huge context model and
       | great speed.
        
         | chipsrafferty wrote:
         | This is mostly for one of 4 reasons:
         | 
         | 1. Sovereignty over data, your outputs can't be stolen or
         | trained on
         | 
         | 2. Just for fun / learning / experiment on
         | 
         | 3. Avoid detection that you're using AI
         | 
         | 4. No Internet connection, in the woods at your cabin or
         | something
        
           | marcalc wrote:
           | This is my key points too. I love the power of having search
           | engine on my laptop.
        
           | biker142541 wrote:
           | Agreed. It's definitely been fun playing locally, learning,
           | fine tuning, etc, but these models just don't quite cut it
           | for serious development tasks (yet, and assuming none of the
           | above considerations apply). I haven't found better than
           | Gemini 2.5 for my work so far.
        
       | rickydroll wrote:
       | I understand why people use the Mac for their local LLM work. I
       | can't bring myself to spend any money on Apple products. I need
       | to find an alternative platform that runs under Linux, and
       | preferably, since I would run this remotely from my work laptop.
       | I would also want to find some way to modulate the power
       | consumption to turn it off automatically when I'm idle.
        
         | telotortium wrote:
         | Entirely due to the unified RAM between CPU and GPU in Apple
         | Silicon. Laptops otherwise almost never have a GPU with
         | sufficient RAM for LLMs.
        
           | rickydroll wrote:
           | Should have been clearer. I was thinking of a dedicated in-
           | house LLM server I could use from different laptops.
        
         | lreeves wrote:
         | The new AMD chips in the Framework laptops would be a good
         | candidate and I think you can get 96GB RAM in them. Also if the
         | LLM software is idle (like llama.cpp or ollama) there is
         | negligible extra power consumption.
        
           | organsnyder wrote:
           | I preordered a Framework Desktop with 128GB RAM for exactly
           | this reason. Apparently under Linux it's possible to assign
           | >100GB to the GPU.
        
         | badsectoracula wrote:
         | If you don't mind going through the eldritchian horror that is
         | building ROCm from source[0], Qwen_Qwen3-30B-A3B-Q6_K (6bit
         | quantization of the LLM mentioned in the article which in
         | practice shouldn't be much different) works decently fast on a
         | RX 7900 XTX using koboldcpp and llama.cpp. And by "decently
         | fast" i mean "it writes faster i can read".
         | 
         | If you're on Debian AFAIK AMD is paying someone to experience
         | the pain in your place, so that is an option if you're building
         | something from scratch, but my openSUSE Tumbleweed installation
         | predates the existence of llama.cpp by a few years and i'm not
         | subjecting myself to the horror that is Python projects
         | (mis)managed by AI developers[1] :-P.
         | 
         | EDIT: my mistake, ROCm isn't needed (or actually, supported) by
         | koboldcpp, it uses Vulkan. ROCm is available via a fork. Still,
         | with Vulkan it is fast too.
         | 
         | [0] ...and more than once as after some OS upgrade it might
         | break, like mine
         | 
         | [1] ok, i did it once, because recently i wanted to try out
         | some tool someone wrote that relied on some AI stuff and i was
         | too stubborn to give up - i had to install Python _from source_
         | on a Debian docker container because some dependency 2-3 layers
         | deep didn 't compile with a newer _minor version_ release of
         | Python. It convinced me to thank yet again to thank Georgi
         | Gerganov for making AI-related tooling that enables people to
         | stick with C++
        
           | rationably wrote:
           | If you are on Debian, ROCm is already packaged in Debian 13
           | (Trixie).
           | 
           | llama.cpp can be built using Debian-supplied libraries with
           | ROCm backend enabled.
        
             | badsectoracula wrote:
             | Yeah, as i wrote "if you're on Debian AFAIK AMD is paying
             | someone to experience the pain in your place" :-).
             | 
             | I used to use Debian at the past but when i was about to
             | install my current OS i already had the openSUSE Tumbleweed
             | installer in a USB so i went with that. Ultimately i just
             | needed "a Linux" and didn't care which. I do end up
             | building more stuff from source than when i used Debian but
             | TBH the only time that annoyed me was with ROCm because it
             | is broken into 2983847283 pieces, many of them have their
             | own flags for the same stuff, some claim they allow to
             | install them anywhere but in practice can only work via the
             | default in "/opt", and a bunch of them have their own
             | special snowflake build process (including one that
             | downloads some random stuff via a script through the build
             | process - IIRC a Gentoo packager made a bug report about it
             | to remove the need to download stuff, but i'm not sure if
             | it has been addressed or not).
             | 
             | If i was doing a fresh OS install i'd probably go with
             | Gentoo - it packages ROCm like Debian, but AFAICT (i
             | haven't tried it) it also provides some tools for you to
             | make bespoke patches to packages you install that survive
             | updates and i'd like to do some customizations on stuff i
             | install.
        
           | rickydroll wrote:
           | Yesterday I was successively using olama installed qwen3:32b
           | and drove it using Simon Willison's llm tool
           | (https://llm.datasette.io/en/stable/). Using CPU only, it ran
           | (if you can call moving at the speed of a walker running) and
           | sucked up almost all of my 32 GB ram.
           | 
           | My laptop has dual (and dueling) graphics chips, Intel and
           | Quadro K1200M with 4 GB of RAM. I will need to learn more
           | about LLM setup, so maybe I can torture myself getting the
           | Nvidia driver working on Linux and experiment with that.
        
       | joejoo wrote:
       | What's the difference between using MLX and MPS?
        
         | Tokumei-no-hito wrote:
         | i think MPS is the term for the APIs apple exposes to control
         | their GPUs and MLX is a machine learning framework optimized
         | for using MPS.
        
       | kamranjon wrote:
       | Just wanted to give a shout out to MLX and MLX-LM - I've been
       | using it to fine-tune Gemma 3 models locally and it's a
       | surprisingly well put together library and set of tools from the
       | Apple devs.
        
       | croemer wrote:
       | Site seems to have been struck with the HN hug of death
        
       | avetiszakharyan wrote:
       | I just wana say i got it to make a snake game! :D for free
        
       | nico wrote:
       | Very cool to see this and glad to discover localforge. Question
       | about localforge, can I combine two agents to do something like:
       | pass an image to a multimodal agent to provide html/css for it,
       | and another to code the rest?
       | 
       | In the post I saw there's gemma3 (multimodal) and qwen3 (not
       | multimodal). Could they be used as above?
       | 
       | How does localforge know when to route a prompt to which agent?
       | 
       | Thank you
        
         | avetiszakharyan wrote:
         | you can combine agents in 2 ways, you can constantly swap
         | agents during one conversation, or you can have the in separate
         | conversations and collaborate. I was even thinking 2 agents can
         | work on 2 separate git clones, and then do PR's to each other.
         | I also like using code to do the image parsing and css, adn
         | then gemini to do the coding. I tried using gemma and qwen for
         | "real stuff" but its more of a, simple stuff only, if i really
         | need output, id' rather spend money for now. Hopefully to
         | change soon. as for rotuing, localforge does NOT know. you
         | choose the agent, and it will loop inside that agent forever.
         | Like, the way it works is that unless agent decides to talk to
         | a user, it will forever be doing function calls and "talking to
         | functions", as one agent. The only routing happens this way.
         | there is main model and there is Expert model. main model knows
         | to ask expert model (see system promp), when its stuck. so for
         | any rouing to happen 1) system prompt needs to mention it 2) a
         | routing to another model should be a function call. that way
         | model knows how to ask another model for something
        
           | nico wrote:
           | Great insights, thank you for the extended and detailed
           | answer, I'll have to try it out
        
       | artdigital wrote:
       | Cool, but running qwen3 and doing a ls tool call is not "vibe
       | coding", this reads more like a lazy ad for localforge
       | 
       | I doubt it can perform well with actual autonomous tasks like
       | reading multiple files, navigating dirs and figuring out where to
       | make edits. That's at least what I would understand under "vibe
       | coding"
        
         | 85392_school wrote:
         | You should try it. It's trained for tool calling and thinks
         | before taking action.
        
         | avetiszakharyan wrote:
         | Definitely try it, it can navigate files search for stuff, run
         | bash commands, and while 30b is a bit cranky it gets the job
         | done (much worse then i would get when i plug in gpt-4.1, but
         | its still not bad, Kudos o qwen. As for localforge, it really
         | is a vibe coding tool, just like claude or codex, but with the
         | possibility to plug more than just one provider. What's wrong
         | with that?
        
           | tough wrote:
           | They're just pointing out how most -educational- content is
           | actually marketing in disguise, which is fine, but also fine
           | to acknowledge i guess, even if a bit snarkily
        
             | avetiszakharyan wrote:
             | Well its an oss project, free, I kind of didnt see it that
             | way i guess, that something thats given for free is bad-
             | tone to market in any possible way. Iguess from my
             | standpoint its more of a, I just want to show this thing to
             | people, because I am proud of it as a personal project, and
             | it brings me joy to just, put it out there And since if you
             | "just put it out there" it will sink to the bottom of the
             | HN pit, why not get a bit more creative.
        
         | datpuz wrote:
         | What *is* vibe coding? And how can I stop hearing about it for
         | the rest of my life?
        
           | krashidov wrote:
           | The real definition of vibe coding is coding with just an
           | LLM. Never looking at what code it outputs, never doing
           | manual edits. Just iterating with the LLM and always pressing
           | approve.
           | 
           | It is a viable way of making software. People have made
           | working software with it. It will likely only ever be more
           | prevalent but might be renamed to just plain old making apps.
        
           | baq wrote:
           | 'If it works it ships'
        
       | paul7986 wrote:
       | Forgive me I am just digesting the term "vibe coding," which
       | doesn't seem like coding at all? It's just typing into your AI's
       | text prompt and describing it to do xyz and then keep making
       | edits til the AI has a working prototype for what you seek. Is
       | that a correct assumption?
        
         | prophesi wrote:
         | Karpathy's original tweet defining "vibe coding":
         | 
         | https://x.com/karpathy/status/1886192184808149383
        
           | paul7986 wrote:
           | So it's not coding ... it's talking to a LLM via voice or
           | chat and have it code for you. Then ask it to change/edit
           | things and then review the code some or just run an error
           | check so the LLM fixes the error and your done.
           | 
           | And so people who are vibe coding are getting paid multiple
           | six figure salaries .... that's not sustainable anyone at any
           | age and in any country can vibe code.
           | 
           | Looks like we are embracing the demise of our skill-sets,
           | careers and livelihoods quickly!
        
             | colesantiago wrote:
             | Do not fear, there will be new jobs available from AI.
        
               | jimbokun wrote:
               | And AI will do those too.
        
               | abc_lisper wrote:
               | lol
        
               | paul7986 wrote:
               | Indeed it will DOGE all those jobs too
        
             | harvey9 wrote:
             | Are people who have no other programming skills really
             | landing well paid jobs with this? I would like to imagine
             | that as a step towards an Iain M Banks future but
             | realistically I'm more likely to see you all at the Skynet
             | work camp.
        
             | jimbokun wrote:
             | Well yes now you are up to speed on current developments.
        
               | paul7986 wrote:
               | lol thank you my LLM is now more advanced :)
        
       | desireco42 wrote:
       | You can just use Ollama and have a bunch of models, some are good
       | for planning, some are for executing tasks... this sounds more
       | complex then it should be or maybe I am lazy and want everything
       | neatly sorted.
       | 
       | I have models on external drive because Apple and through Ollama
       | server they interact really well with Cline or Roo code or even
       | Bolt, but I found Bolt really not working well.
        
         | desireco42 wrote:
         | To add, you can use so called, abliterated models that are
         | stripped of censorship for example. Much better experience
         | sometimes.
        
       | seanhunter wrote:
       | You can already do this, with qwen or (which I use) deepscaler
       | using aider and ollama. This is just an advert for localforge.
        
       | jononor wrote:
       | Running models locally is starting to get interesting now.
       | Especially the 30B-A3B version seems like a promising direction,
       | though it is still out of reach on 16 GB VRAM (quite accessible).
       | Hoping for new Nvidia RTX cards with 24/32 GB VRAM. Seems that we
       | might get to GPT4-ish levels within a few years? Which is useful
       | for a bunch of tasks.
        
         | avetiszakharyan wrote:
         | I think we are just tiny bit away of being able to really
         | "code" with ai, locally. Because even if it would be on
         | gemini2.5 level, since its free, you can make it self prompt a
         | bit more and eventually solve any problem. if i could ran 200b
         | or if 30b wouldve been as good - it wouldve been enough
        
       | rcarmo wrote:
       | Coincidentally, I just managed to get Qwen3 to go into a loop by
       | using a fairly simple prompt:
       | 
       | "create a python decorator that uses a trie to do mqtt topic
       | routing"
       | 
       | phi4-reasoning works, but I think the code is buggy
       | 
       | phi4-mini-reasoning freaks out
       | 
       | qwen3:30b starts looping and forgets about the decorator
       | 
       | mistral-small gets straight to the point and the code seems sane
       | 
       | https://mastodon.social/@rcarmo/114433075043021470
       | 
       | I regularly use Copilot models, and they can manage this without
       | too many issues (Claude 3.7 and Gemini output usable code with
       | tests), but local models seem to not have the ability to do it
       | quite yet.
        
         | avetiszakharyan wrote:
         | Is there an additional system prompt before that? Or i can
         | repro with just this?
        
           | rcarmo wrote:
           | Just that. I purposefully used exactly the same thing I did
           | with Claude and Gemini to see how the models dealt with
           | ambiguity.
        
         | datpuz wrote:
         | I think your prompt is bad. Still impressive that Claude 3.7
         | handled your bad prompt, but qwen3 had no problem with this
         | prompt:
         | 
         | Create a Python decorator that registers functions as handlers
         | for MQTT topic patterns (including + and # wildcards).
         | Internally, use a trie to store the topic patterns and match
         | incoming topic strings to the correct handlers. Provide an
         | example showing how to register multiple handlers and dispatch
         | a message to the correct one based on an incoming topic.
        
           | rcarmo wrote:
           | I purposefully used exactly the same thing I did with Claude
           | and Gemini to see how the models dealt with ambiguity. It
           | shouldn't have degraded the chain of thought to the point
           | where it starts looping.
        
         | GaggiX wrote:
         | You should probably try a different quantization, have you try
         | UD-Q4_K_XL?
        
         | datpuz wrote:
         | Here's qwen-30b-a3b's response to your prompt when I worded it
         | better:
         | 
         | The prompt was:
         | 
         | "Create a Python decorator that registers functions as handlers
         | for MQTT topic patterns (including + and # wildcards).
         | Internally, use a trie to store the topic patterns and match
         | incoming topic strings to the correct handlers. Provide an
         | example showing how to register multiple handlers and dispatch
         | a message to the correct one based on an incoming topic."
         | 
         | https://pastebin.com/wefw7X2h
        
           | rcarmo wrote:
           | I went back and used your prompt, and it is still looping:
           | 
           | https://pastebin.com/VfmhCTFm
        
       | jedisct1 wrote:
       | Qwen3 is great, but not for writing code. Even after the recent
       | fixes, and with the recommended parameters, it gets often trapped
       | in a loop.
       | 
       | Qwen2.5-32B, Cogito-32B and GLM-32B remain the best options for
       | local coding agents, even though the recently released MiMo model
       | is also quite good for its size.
        
       | bitbasher wrote:
       | I've been vibe coding for over a decade. All you need is a decent
       | pair of headphones and a pot of coffee. It's not free, but it's
       | pretty cheap.
        
         | datpuz wrote:
         | That's just called coding
        
           | redcobra762 wrote:
           | *with good vibes
        
         | madduci wrote:
         | And the "for free" in the title excludes the electricity costs
        
         | thih9 wrote:
         | In case someone didn't see that yet, vibe coding has a recent
         | and more specific meaning.
         | 
         | https://en.m.wikipedia.org/wiki/Vibe_coding
         | 
         | > a programming paradigm dependent on artificial intelligence
         | (AI), where a person describes a problem in a few sentences as
         | a prompt to a large language model (LLM) tuned for coding.
         | 
         | > A key part of the definition of vibe coding is that the user
         | accepts code without full understanding.
        
       | 999900000999 wrote:
       | Very impressive, it doesn't need to be as good as the pay for
       | token models. For example I've probably spent at least $300 last
       | month on vibe coding, a big part of this is I want to know what
       | tools I'm going to end up competing with, and another is I got a
       | working implementation of one of my side projects, and then I
       | decided I wanted it to be rewritten in another programming
       | language.
       | 
       | Even if I chill out a bit here, a refurbished Nvidia laptop would
       | pay for itself within a year. I am a bit disappointed Ollama
       | can't handle the full flow yet, IE it could be a single command.
       | 
       | ollama code qwen3
        
         | _bin_ wrote:
         | I just tried it. It got stuck looping on a `cargo check` call
         | and literally wouldn't do anything else. No additional context,
         | just repeatedly spitting out the same tool call.
         | 
         | The problem is the best models barely clear the bar for some
         | stuff in terms of coherence and reliability; anything else just
         | isn't particularly usable.
        
           | 999900000999 wrote:
           | This happens when I'm using Claude Code too. Even the best
           | models need humans to get unstuck.
           | 
           | Fron what I've seen most of them are good at writing new code
           | from scratch.
           | 
           | Refactoring is very difficult.
        
             | _bin_ wrote:
             | I tried it 3-4 times before giving up and it did this every
             | single time. I checked the tool call output and it was
             | running cargo check appropriately. I think maybe the
             | 30b-scale models just aren't sufficient for typical
             | development.
             | 
             | You're generally correct though, that from-scratch gets
             | better results. This is a huge constraint of them: I don't
             | want a model that will write something its way. I've
             | already gone through my design and settled on the
             | style/principles/libraries I did for a reason; the bot
             | working terribly with that _is_ a major flaw and I don 't
             | see saying "let the bot do things its preferred way" as a
             | good answer. Some systems, things like latency matters, and
             | the bot's way just isn't good enough.
             | 
             | The vast majority of man-hours are maintaining and
             | extending code, not green-fielding new stuff. Vendors
             | should be hyper-focused on this, on compliance with user
             | directions, not with building something that makes a react
             | todo-list app marginally faster or better than competitors.
        
               | 999900000999 wrote:
               | If anything, it's a good sign that these tools are no
               | where close to replacing us.
               | 
               | I was trying to get postgres working with a project the
               | other day, and Claude decided that it was going to just
               | replace it with SQL lite when it couldn't get the build
               | to work.
               | 
               | All I want is "I don't know how to do this." But now
               | these tools would rather just do it wrong.
               | 
               | They also have a very very strong tendency to try and
               | force unoptimized solutions. You'll have 3 classes that
               | do the exact same thing with only minor variable
               | differences. Something a human would do in one class.
               | 
               | For my latest project I'm strongly tempted to just suck
               | it up and code the whole thing by hand.
        
       | pritambarhate wrote:
       | Just tried Qwen 3 32B 4bit and 8bit GGUF via Ollama today on
       | RunPod A40 GPUs. For general knowledge questions the results were
       | very underwhelming. Here are answers to 3 questions I asked to
       | the 8 bit GGUF model (skipping the thinking tokens to keep it
       | short):
       | 
       | Q: who won the cricket worldcup 2023?
       | 
       | A: The *2023 ICC Cricket World Cup* was won by *Australia*. They
       | defeated *New Zealand* in the final on *November 19, 2023*, at
       | the *M. A. Chidambaram Stadium in Chennai, India*.
       | 
       | ### Key Details: - *Final Score*: Australia scored *334/5* in
       | their 50 overs, while New Zealand reached *199/8* in their chase.
       | - *Man of the Match*: Pat Cummins (Australia) with 5 wickets for
       | 34 runs. - *Tournament Host*: India. - *Notable Performance*:
       | Australia's aggressive batting and disciplined bowling sealed the
       | victory, marking their *sixth Cricket World Cup title* (tying
       | them with West Indies for the most titles).
       | 
       | This was the 13th edition of the ICC Cricket World Cup (50-over
       | format). The T20 World Cup is a separate tournament, with India
       | winning the 2024 edition. Let me know if you need further
       | details!
       | 
       | total duration: 34.911890735s load duration: 65.034584ms prompt
       | eval count: 20 token(s) prompt eval duration: 305.460537ms prompt
       | eval rate: 65.47 tokens/s eval count: 511 token(s) eval duration:
       | 34.539660654s eval rate: 14.79 tokens/s
       | 
       | Comment: Wrong - https://www.espncricinfo.com/series/icc-cricket-
       | world-cup-20...
       | 
       | Q: give last 10 movies to win the oscar best picture award
       | 
       | Here is a clear and accurate list of the *last 10 Best Picture
       | winners* at the Academy Awards, from *2014 to 2023*, including
       | the ceremony year, the winning film, and the film's release year:
       | 
       | ---
       | 
       | 1. *2023 (95th Oscars)* - *Winner*: _Everything Everywhere All at
       | Once_ - *Release Year*: 2022
       | 
       | 2. *2022 (94th Oscars)* - *Winner*: _The Power of the Dog_ -
       | *Release Year*: 2021
       | 
       | 3. *2021 (93rd Oscars)* - *Winner*: _Nomadland_ - *Release Year*:
       | 2020
       | 
       | 4. *2020 (92nd Oscars)* - *Winner*: _Parasite_ - *Release Year*:
       | 2019
       | 
       | 5. *2019 (91st Oscars)* - *Winner*: _Green Book_ - *Release
       | Year*: 2018
       | 
       | 6. *2018 (90th Oscars)* - *Winner*: _The Shape of Water_ -
       | *Release Year*: 2017
       | 
       | 7. *2017 (89th Oscars)* - *Winner*: _Moonlight_ - *Release Year*:
       | 2016
       | 
       | 8. *2016 (88th Oscars)* - *Winner*: _Spotlight_ - *Release Year*:
       | 2015
       | 
       | 9. *2015 (87th Oscars)* - *Winner*: _12 Years a Slave_ - *Release
       | Year*: 2013
       | 
       | 10. *2014 (86th Oscars)* - *Winner*: _Argo_ - *Release Year*:
       | 2012
       | 
       | Comment: Answer number 2 ans 9 are wrong.
       | (https://en.wikipedia.org/wiki/Academy_Award_for_Best_Picture)
       | 
       | I would have expected it to get things which are such big events
       | right at least.
        
       | Tacite wrote:
       | Trying on Macbook Pro (M4) with 24 GB, the whole system freeze
       | after the first question.
        
         | thih9 wrote:
         | I'm also team 24GB.
         | 
         | Is anyone using a similar setup with 24GB RAM? Which model
         | would you recommend?
         | 
         | I only saw a sub-1B model mentioned in other comments, but that
         | one seems too small.
        
           | TMWNN wrote:
           | `ollama run qwen3:14b`
        
       | avetiszakharyan wrote:
       | Quick Video i made on this topic:
       | https://www.youtube.com/watch?v=-h_IZhOdAeU
        
       ___________________________________________________________________
       (page generated 2025-05-01 23:02 UTC)