[HN Gopher] Workhorse LLMs: Why Open Source Models Dominate Clos...
       ___________________________________________________________________
        
       Workhorse LLMs: Why Open Source Models Dominate Closed Source for
       Batch Tasks
        
       Author : cmogni1
       Score  : 95 points
       Date   : 2025-06-06 18:38 UTC (1 days ago)
        
 (HTM) web link (sutro.sh)
 (TXT) w3m dump (sutro.sh)
        
       | ramesh31 wrote:
       | Flash is just so obscenely cheap at this point it's hard to
       | justify the headache of self hosting though. Really only applies
       | to sensitive data IMO.
        
         | behnamoh wrote:
         | You're getting downvoted but what you said is true. The cost of
         | self-hosting (and achieving +70 tok/sec consistently across the
         | entire context window) has never been low enough to justify
         | open source as a viable competitor to proprietary models of
         | OpenAI, Google, and Anthropic.
        
           | grepfru_it wrote:
           | I am curious the need for 70 t/sec?
        
             | Aeolun wrote:
             | Waiting minutes for your call to succeed is too
             | frustrating?
        
               | ekianjo wrote:
               | Depends entirely on the use case. Not every LLM workflow
               | is a chatbot
        
               | jbellis wrote:
               | no, but if you're not latency sensitive you should
               | probably be using DeepSeek v3 (cheaper than flash,
               | significantly smarter)
        
               | lostmsu wrote:
               | What makes you believe DeepSeek is smarter than Flash
               | 2.5? It is lower on all leaderboards.
        
               | jbellis wrote:
               | you're right, I should clarify that I'm talking about no
               | thinking mode, otherwise flash goes from "a bit more
               | expensive than dsv3" to "10x more expensive"
        
             | cootsnuck wrote:
             | High concurrency voice AI systems.
        
         | jacob019 wrote:
         | That's true for Flash 2.0 at $0.40/mtok output. GPT-4.1-nano is
         | the same price and also surprisingly capable. I can spend real
         | money with 2.5 flash, with those $3.50/mtok thinking tokens,
         | worth it though. OP is an inference provider, so there may be
         | some bias. Open source can't compete on context length either,
         | nothing touches 2.5 flash for the price with long context--I've
         | experimented with this a lot for my agentic pricing system.
         | Open source models are improving, but they aren't really any
         | cheaper right now, R1 for example does quite well performance
         | wise, but it uses a LOT of tokens to get there, further
         | limiting the shorter context window. There's still value in the
         | open source models, each model has unique strengths and they're
         | advancing quickly, but the frontier labs are moving fast too
         | and have very compelling "workhorse" offers.
        
         | mkl wrote:
         | With tools like Ollama, self-hosting is easier than hosted. No
         | sign-up, no API keys, no permission to spend money, no worries
         | about data security, just an easy install then import a Python
         | library. Qwen2.5-VL 7B is proving useful even on a work laptop
         | with insufficient VRAM - I just leave it running over a night
         | or weekend and it's saving me dozens of hours of work (that I
         | then get to spend on other higher-value work).
        
           | mgraczyk wrote:
           | It does not take dozens of hours to get an API key for gemini
        
             | mkl wrote:
             | I never claimed that it did. Gemini would probably save me
             | the same dozens of hours, but come with ongoing costs and
             | additional starting up hurdles (some near insurmountable in
             | my organisation, like data security for some of what I'm
             | doing).
        
               | shmoogy wrote:
               | Gemini flash or any free LLM on openrouter would be
               | orders of magnitude faster and effectively free. Unless
               | you are concerned about privacy of the conversation -
               | it's really purely being able to say you did it locally.
               | 
               | I definitely do appreciate and believe in the value of
               | open source / open weight LLMs - but inference is so
               | cheap right now for non frontier models.
        
             | cortesoft wrote:
             | They weren't saying getting the api key would take that
             | long, just getting permission from their company to let
             | them do it.
        
           | genewitch wrote:
           | I got the 70b qwen llama distill, I have 24GB of vram.
           | 
           | I opened aider and gave a small prompt, roughly:
           | Implement a JavaScript 2048 game that exists as flat file(s)
           | and does not require a server, just the game HTML, CSS, and
           | js. Make it compatible with firefox, at least.
           | 
           | That's it. Several hours later, it finished. The game ran. It
           | was worth it because this was in the winter and it heated my
           | house a bit, yay. I think the resulting 1-shot output is on
           | my github.
           | 
           | I know it was in the training set, etc, but I wanted to see
           | how big of a hassle it was, if it would 1-shot with such a
           | small prompt, how long it would take.
           | 
           | Makes me want to try deepseek 671B, but I don't have any
           | machines with >1TB of memory.
           | 
           | I do take donations of hardware.
        
             | mechagodzilla wrote:
             | Buy a used workstation with 512GB of DDR4 RAM. It will
             | probably cost like $1-1.5k, and be able to run a Q4 version
             | of the full deepseek 671B models. I have a similar setup
             | with dual-socket 18 core Xeons (and 768GB of RAM, so it
             | cost about $2k), and can get about 1.5 tokens/sec on those
             | models. Being able to see the full thinking trace on the R1
             | models is awesome compared to the OpenAI models.
        
           | 3036e4 wrote:
           | If/when Corporate Legal approves a tool like Ollama for use
           | on company computers, yes. Might not require purchasing
           | anything, but there can still be red tape.
        
           | yb6677 wrote:
           | What do you ask it to do overnight or weekend?
        
         | xfalcox wrote:
         | You'd be surprised how often people in enterprise can be left
         | waiting months to get an API key approved for an LLM provider.
        
           | diggan wrote:
           | Are you saying that it's faster for them to get the hardware
           | to run the weights themselves? Otherwise I'm not sure what
           | the relevancy is.
        
             | ChromaticPanic wrote:
             | Yes some have existing infra
        
               | diggan wrote:
               | I'm having a somewhat hard time believing a corporation
               | where getting a API key for a LLM service is very
               | difficult, somehow has the (GPU) infrastructure already
               | running for doing the same thing themselves, unless they
               | happen to be a ML corporation, but I don't think we're
               | talking about those in this context.
        
               | oooyay wrote:
               | Nah this is definitely a real scenario. Getting access to
               | public models requires a lot of security review, but
               | proving through Bedrock is much more simple. I may be
               | spoiled in having worked for companies that have ML
               | departments and developer XP departments though.
        
               | diggan wrote:
               | Not sure Bedrock counts as self-hosting though, isn't it
               | a managed service Amazon provides?
               | 
               | > I may be spoiled in having worked for companies that
               | have ML
               | 
               | Sounds likely, yeah, how many companies have ML
               | departments today? DS departments seem common, but ML i'm
               | not too sure about
        
               | achierius wrote:
               | No, this is very real. One reason why this can happen: a
               | company has elaborate processes for protecting their
               | internal data from leaking, but otherwise lets engineers
               | do what they want with resources allocated to them.
        
             | pegasus wrote:
             | Unless they are already in the possession of such hardware
             | (like an M3 mac, for example).
        
         | cortesoft wrote:
         | There is a wide range of opinions on what should be considered
         | sensitive data. Many people would classify a vast majority of
         | their data as sensitive.
        
       | delichon wrote:
       | Pass the choices through, please. It's so context dependent that
       | I want a <dumber> and a <smarter> button, with units of $/M
       | tokens. And another setting to send a particular prompt to "[x]
       | batch" and email me with the answer later. For most things I'll
       | start dumb and fast, but switch to smart and slow when the going
       | gets rough.
        
       | jbellis wrote:
       | This is a useful analysis, but only as a first cut and sometimes
       | not even that -- Grok 3 mini and DeepSeek V3 are by far the least
       | expensive coding models that are worth trying for scenarios where
       | you do and don't care about the vendor training on your requests,
       | respectively. One of those is "open source" (by which he seems to
       | mean "open weights") but far too large to run locally.
       | 
       | [I guess that must be a useful market niche though, apparently
       | this is by a company selling batch compute on exactly those small
       | open weights models.]
       | 
       | The problem is the author is evaluating by dividing the
       | Artificial Analysis score by a blended cost per token, but most
       | tasks have an intelligence "floor" below which it doesn't matter
       | how cheap something is, it will never succeed. And when you strip
       | out the very high results from super cheap 4B OSS models the rest
       | are significantly outclassed by Flash 2.0 (not on his chart but
       | still worth considering) and 2.5, not to mention other models
       | that might be better in domain specific tasks like grok-3 mini
       | for code.
       | 
       | (Nobody should be using Haiku in 2025. The OpenAI mini models are
       | not as bad as Haiku in p/p and maybe there is a use case for
       | prefering one over Flash but if so I don't know what it is.)
        
         | dinosaurdynasty wrote:
         | DeepSeek has a lot of competing providers that at least state
         | they don't train on API data, OpenRouter lists a bunch of them:
         | https://openrouter.ai/deepseek/deepseek-chat-v3-0324/provide...
         | 
         | (This is a big advantage of open weight models; even if they're
         | too big to host yourself, if it's worth anything there's a lot
         | of competition for inference)
        
           | jbellis wrote:
           | and all of them are so much more expensive than OG deepseek
           | as to completely remove themselves from consideration
           | 
           | you should probably use grok 3 mini if you want "cheapest
           | model that is reasonably good at code"
        
       ___________________________________________________________________
       (page generated 2025-06-07 23:02 UTC)