[HN Gopher] Workhorse LLMs: Why Open Source Models Dominate Clos...
___________________________________________________________________
Workhorse LLMs: Why Open Source Models Dominate Closed Source for
Batch Tasks
Author : cmogni1
Score : 95 points
Date : 2025-06-06 18:38 UTC (1 days ago)
(HTM) web link (sutro.sh)
(TXT) w3m dump (sutro.sh)
| ramesh31 wrote:
| Flash is just so obscenely cheap at this point it's hard to
| justify the headache of self hosting though. Really only applies
| to sensitive data IMO.
| behnamoh wrote:
| You're getting downvoted but what you said is true. The cost of
| self-hosting (and achieving +70 tok/sec consistently across the
| entire context window) has never been low enough to justify
| open source as a viable competitor to proprietary models of
| OpenAI, Google, and Anthropic.
| grepfru_it wrote:
| I am curious the need for 70 t/sec?
| Aeolun wrote:
| Waiting minutes for your call to succeed is too
| frustrating?
| ekianjo wrote:
| Depends entirely on the use case. Not every LLM workflow
| is a chatbot
| jbellis wrote:
| no, but if you're not latency sensitive you should
| probably be using DeepSeek v3 (cheaper than flash,
| significantly smarter)
| lostmsu wrote:
| What makes you believe DeepSeek is smarter than Flash
| 2.5? It is lower on all leaderboards.
| jbellis wrote:
| you're right, I should clarify that I'm talking about no
| thinking mode, otherwise flash goes from "a bit more
| expensive than dsv3" to "10x more expensive"
| cootsnuck wrote:
| High concurrency voice AI systems.
| jacob019 wrote:
| That's true for Flash 2.0 at $0.40/mtok output. GPT-4.1-nano is
| the same price and also surprisingly capable. I can spend real
| money with 2.5 flash, with those $3.50/mtok thinking tokens,
| worth it though. OP is an inference provider, so there may be
| some bias. Open source can't compete on context length either,
| nothing touches 2.5 flash for the price with long context--I've
| experimented with this a lot for my agentic pricing system.
| Open source models are improving, but they aren't really any
| cheaper right now, R1 for example does quite well performance
| wise, but it uses a LOT of tokens to get there, further
| limiting the shorter context window. There's still value in the
| open source models, each model has unique strengths and they're
| advancing quickly, but the frontier labs are moving fast too
| and have very compelling "workhorse" offers.
| mkl wrote:
| With tools like Ollama, self-hosting is easier than hosted. No
| sign-up, no API keys, no permission to spend money, no worries
| about data security, just an easy install then import a Python
| library. Qwen2.5-VL 7B is proving useful even on a work laptop
| with insufficient VRAM - I just leave it running over a night
| or weekend and it's saving me dozens of hours of work (that I
| then get to spend on other higher-value work).
| mgraczyk wrote:
| It does not take dozens of hours to get an API key for gemini
| mkl wrote:
| I never claimed that it did. Gemini would probably save me
| the same dozens of hours, but come with ongoing costs and
| additional starting up hurdles (some near insurmountable in
| my organisation, like data security for some of what I'm
| doing).
| shmoogy wrote:
| Gemini flash or any free LLM on openrouter would be
| orders of magnitude faster and effectively free. Unless
| you are concerned about privacy of the conversation -
| it's really purely being able to say you did it locally.
|
| I definitely do appreciate and believe in the value of
| open source / open weight LLMs - but inference is so
| cheap right now for non frontier models.
| cortesoft wrote:
| They weren't saying getting the api key would take that
| long, just getting permission from their company to let
| them do it.
| genewitch wrote:
| I got the 70b qwen llama distill, I have 24GB of vram.
|
| I opened aider and gave a small prompt, roughly:
| Implement a JavaScript 2048 game that exists as flat file(s)
| and does not require a server, just the game HTML, CSS, and
| js. Make it compatible with firefox, at least.
|
| That's it. Several hours later, it finished. The game ran. It
| was worth it because this was in the winter and it heated my
| house a bit, yay. I think the resulting 1-shot output is on
| my github.
|
| I know it was in the training set, etc, but I wanted to see
| how big of a hassle it was, if it would 1-shot with such a
| small prompt, how long it would take.
|
| Makes me want to try deepseek 671B, but I don't have any
| machines with >1TB of memory.
|
| I do take donations of hardware.
| mechagodzilla wrote:
| Buy a used workstation with 512GB of DDR4 RAM. It will
| probably cost like $1-1.5k, and be able to run a Q4 version
| of the full deepseek 671B models. I have a similar setup
| with dual-socket 18 core Xeons (and 768GB of RAM, so it
| cost about $2k), and can get about 1.5 tokens/sec on those
| models. Being able to see the full thinking trace on the R1
| models is awesome compared to the OpenAI models.
| 3036e4 wrote:
| If/when Corporate Legal approves a tool like Ollama for use
| on company computers, yes. Might not require purchasing
| anything, but there can still be red tape.
| yb6677 wrote:
| What do you ask it to do overnight or weekend?
| xfalcox wrote:
| You'd be surprised how often people in enterprise can be left
| waiting months to get an API key approved for an LLM provider.
| diggan wrote:
| Are you saying that it's faster for them to get the hardware
| to run the weights themselves? Otherwise I'm not sure what
| the relevancy is.
| ChromaticPanic wrote:
| Yes some have existing infra
| diggan wrote:
| I'm having a somewhat hard time believing a corporation
| where getting a API key for a LLM service is very
| difficult, somehow has the (GPU) infrastructure already
| running for doing the same thing themselves, unless they
| happen to be a ML corporation, but I don't think we're
| talking about those in this context.
| oooyay wrote:
| Nah this is definitely a real scenario. Getting access to
| public models requires a lot of security review, but
| proving through Bedrock is much more simple. I may be
| spoiled in having worked for companies that have ML
| departments and developer XP departments though.
| diggan wrote:
| Not sure Bedrock counts as self-hosting though, isn't it
| a managed service Amazon provides?
|
| > I may be spoiled in having worked for companies that
| have ML
|
| Sounds likely, yeah, how many companies have ML
| departments today? DS departments seem common, but ML i'm
| not too sure about
| achierius wrote:
| No, this is very real. One reason why this can happen: a
| company has elaborate processes for protecting their
| internal data from leaking, but otherwise lets engineers
| do what they want with resources allocated to them.
| pegasus wrote:
| Unless they are already in the possession of such hardware
| (like an M3 mac, for example).
| cortesoft wrote:
| There is a wide range of opinions on what should be considered
| sensitive data. Many people would classify a vast majority of
| their data as sensitive.
| delichon wrote:
| Pass the choices through, please. It's so context dependent that
| I want a <dumber> and a <smarter> button, with units of $/M
| tokens. And another setting to send a particular prompt to "[x]
| batch" and email me with the answer later. For most things I'll
| start dumb and fast, but switch to smart and slow when the going
| gets rough.
| jbellis wrote:
| This is a useful analysis, but only as a first cut and sometimes
| not even that -- Grok 3 mini and DeepSeek V3 are by far the least
| expensive coding models that are worth trying for scenarios where
| you do and don't care about the vendor training on your requests,
| respectively. One of those is "open source" (by which he seems to
| mean "open weights") but far too large to run locally.
|
| [I guess that must be a useful market niche though, apparently
| this is by a company selling batch compute on exactly those small
| open weights models.]
|
| The problem is the author is evaluating by dividing the
| Artificial Analysis score by a blended cost per token, but most
| tasks have an intelligence "floor" below which it doesn't matter
| how cheap something is, it will never succeed. And when you strip
| out the very high results from super cheap 4B OSS models the rest
| are significantly outclassed by Flash 2.0 (not on his chart but
| still worth considering) and 2.5, not to mention other models
| that might be better in domain specific tasks like grok-3 mini
| for code.
|
| (Nobody should be using Haiku in 2025. The OpenAI mini models are
| not as bad as Haiku in p/p and maybe there is a use case for
| prefering one over Flash but if so I don't know what it is.)
| dinosaurdynasty wrote:
| DeepSeek has a lot of competing providers that at least state
| they don't train on API data, OpenRouter lists a bunch of them:
| https://openrouter.ai/deepseek/deepseek-chat-v3-0324/provide...
|
| (This is a big advantage of open weight models; even if they're
| too big to host yourself, if it's worth anything there's a lot
| of competition for inference)
| jbellis wrote:
| and all of them are so much more expensive than OG deepseek
| as to completely remove themselves from consideration
|
| you should probably use grok 3 mini if you want "cheapest
| model that is reasonably good at code"
___________________________________________________________________
(page generated 2025-06-07 23:02 UTC)