[HN Gopher] Qwen3.8 Max now ranked as the best overall model by ...
___________________________________________________________________
Qwen3.8 Max now ranked as the best overall model by agentic index
Author : apitman
Score : 504 points
Date : 2026-08-06 18:44 UTC (16 hours ago)
(HTM) web link (artificialanalysis.ai)
(TXT) w3m dump (artificialanalysis.ai)
| embedding-shape wrote:
| Strange that the page
| https://artificialanalysis.ai/agents/coding-agents doesn't even
| mention "Qwen" once if it's now the "best" according to one of
| their one index?
| scrlk wrote:
| Different benchmarks:
|
| > Artificial Analysis Agentic Index: Represents the weighted
| average of agentic capabilities benchmarks in the Artificial
| Analysis Intelligence Index (GDPval-AA v2, Tau3-Banking)
|
| > Artificial Analysis Coding Agent Index v1.3 incorporates 3
| benchmarks: DeepSWE, Terminal-Bench v2, and SWE-Atlas-QnA
|
| Qwen3.8 Max is 55.4 on the Agentic Index but hasn't been tested
| for the Coding Agent Index.
| apitman wrote:
| Looks like coding agent is model+harness. There are far fewer
| models represented on that page. I believe "agentic index" is
| still the metric to look at for coding performance. I could
| be wrong about that though.
| Bootvis wrote:
| Indeed, and this Qwen 3.8 max specific page:
|
| https://artificialanalysis.ai/models/qwen3-8-max
|
| Doesn't have the claim either. Clickbait?
| petu wrote:
| This page has it, scroll to "Intelligence" header (not the
| highlights one, but second on the page / with black square)
| and click "Agentic Index"
| Bootvis wrote:
| So the original link should be: https://artificialanalysis.
| ai/models/qwen3-8-max?intelligenc...
|
| Even then, this seems a much more marginal win than the
| headline suggested to me.
| amelius wrote:
| According to those graphs, Grok 4.5 appears to be the most
| cost-effective model.
| user43928 wrote:
| $0.05 per task, Intelligence Index score 52 -> GPT 5.6 Luna
| max
|
| $0.36 per task, Intelligence Index score 56 -> Grok 4.5 high
|
| $1.13 per task, Intelligence Index score 58 -> Qwen 3.8 Max
|
| $0.81 per task, Intelligence Index score 59 -> GPT 5.6 Sol
| xhigh
|
| $1.80 per task, Intelligence Index score 63 -> Opus 5 xhigh
| artemisart wrote:
| They didn't run all benchmarks. It's the best in AA agentic
| index (GDPval-AA v2, 3-Banking) but not coding index (DeepSWE
| which is missing, Terminal-Bench v2.1 they have 81% vs 90% for
| Sol, SWE-Atlas-QnA missing).
| moritzwarhier wrote:
| Does "artificial analysis" mean what it says? Dubious.
|
| But: I've been very impressed by the larger Qwen Models, and a
| brief try of Kimi also impressed me.
|
| A lingering sense of quality degradation when going deep
| remains.
|
| But that's not an accusation: they seem to be hitting the
| compute/quality tradeoff extremely well.
|
| And on-prem capability is simply irreplaceable.
|
| Apart from all the innovations that were driven by the strive
| for this optimization: quantization, "distilling" (without
| obvious mad-cows-disease)... I think China was an invaluable
| player in this progress. Intuitively, I'd even go so far to
| speculate that LLaMa wouldn't exist without the competition.
| syntaxing wrote:
| I am so excited for Qwen 3.8 27B. It's a shame how slow prefill
| (~3-400) is on a strix halo but it's such a good model for
| agentic tasks.
| tarr11 wrote:
| What type of agentic tasks are you using it for (eg how
| complex)?
| syntaxing wrote:
| For personal stuff, I use it with AnythingLLM. It replaced
| any Google search for me. For coding, I run opencode though I
| have been debating switching to Pi. I would argue it's at
| Sonnet 3 level.
| CamperBob2 wrote:
| How are you running it on a Strix Halo? The weights aren't out
| yet, are they?
| 13rac1 wrote:
| I interpret @syntaxing as meaning they are looking forward to
| running Qwen3.8-27B, but are frustrated by prefill times with
| other models, such as Qwen3.6-27B.
| syntaxing wrote:
| I meant Qwen3.6. Unsloth supposedly has early preview of the
| model and the VRAM requirement is the same so most people
| expect similar model size and type.
| LoganDark wrote:
| I find that 35B-A3B is much easier to run on my M4 Max (both
| prefill and generation)
| markasoftware wrote:
| It's well known 35b is much faster (on any hardware) and
| quite a bit dumber
| dofm wrote:
| This really very much depends on how you are using it, I
| think. If you intend to leave it to solve long context
| problems and write whole prototypes, the 27B is going to be
| much better.
|
| But if you are sort of pair-programming with the model, the
| speed obviously matters and I think then the 35B is
| acceptably smart, and when it's wrong it'll be wrong much
| more quickly. It seems very good on SQL and PHP, and I
| assume on typical JS and Python.
|
| I would rather work that way, so I hope they do produce a
| small MoE model.
| colingauvin wrote:
| Prefill is survivable if you cache well. But what kills me is
| the context. Qwen 27 needs a ton of room for KV Cache. I guess
| not an issue on a 128 GB Halo or Spark, but if you are running
| of consumer/prosumer GPUs it's miserable to be compacting every
| 120k tokens.
| drnick1 wrote:
| Why does an open weights model cost nearly the same as GPT5.6?
| $1.14 vs $1.23 on the cost index. Since you can't presumably run
| this on your own hardware given the model size and hence gain
| other things like privacy, I don't see any reason to move away
| from GPT at this rate.
| eli wrote:
| It's not enough that it's better?
|
| Many providers will host it and will compete on price. It also
| can't easily be taken away because one company (or one
| government) decides they don't want it around any more. People
| can fine-tune it for particular workloads.
| drnick1 wrote:
| > It's not enough that it's better?
|
| It's barely better, and barely cheaper, not really enough to
| challenge the status quo IMO. Half the price for basically
| the same performance would be a much stronger value
| proposition.
| ux266478 wrote:
| What status quo? Just look at Openrouter's rankings:
| https://openrouter.ai/rankings
|
| Things change radically month to month. Nobody is remotely
| close to capturing the market or having any kind of
| stability over time. People move around quite a lot, often
| to sidegrade within a generation. Just playing fly on the
| wall with discourse would be enough to tell you all of
| this, even without the data to back it up.
| eli wrote:
| That's got a significant selection bias. Claude and
| ChatGPT and Gemini and other subs do not go through
| openrouter.
| ux266478 wrote:
| Not really, because that's not a unique aspect of any of
| those. It's true of all subscription services (that I'm
| aware of), as well as all of the free models. The
| selection bias primarily will be against models which be
| an outlier in the difference between openrouter users and
| total users, which is a much harder position to argue for
| any given company except for maybe Twitter.
|
| You can argue there's a selection bias that openrouter
| users are less likely to display model loyalty, but it
| would still be a visible confounding factor if it was a
| statistically significant behavior. And it's not. Nor is
| there a visibly meaningful indication that people don't
| sidegrade between models. With every single data set,
| you're going to see that. You're also going to see it
| reflected in discourse, as I mentioned. Fact of the
| matter is there isn't a status quo in AI any more than
| there's a status quo in cars.
| SwellJoe wrote:
| If anybody has, it's DeepSeek. But, with the promised
| price hikes, I'm sure that'll change. I'm guessing
| they're raising prices not because they're not making a
| profit at those prices, but because they're running into
| capacity problems and need to slow down until they've got
| more or risk providing poor service. For now DeepSeek
| Flash is the best deal going for API usage and its
| popularity makes sense.
|
| Also, OpenRouter misses most of the usage of the US
| models, as most people are getting those from the vendor
| directly via subscriptions.
| Art9681 wrote:
| They cherrypicked benchmarks. The ONE weighed benchmark where
| is beats Opus5 by 0.1 points is what was linked because
| that's how propaganda works. The Agentic Index that includes
| the full benchmark suite has it in 5th place.
|
| Might as well use gpt-sol.
| iAMkenough wrote:
| The whole industry cherry picks benchmarks.
|
| I stopped paying attention to self-published benchmarks
| when Apple started using those non-sensical performance
| graphs with "relative performance" as a vertical axis when
| announcing a new chip.
| apitman wrote:
| For one thing, providers of open models can't arbitrarily
| increase their prices without facing competition.
| frereubu wrote:
| But given the extremely low cost of switching, why wouldn't
| you use the cheaper one if they're comparable?
| apitman wrote:
| As low as it is, switching between providers on OpenRouter
| is still lower.
|
| That said, it's a fair point. For me, it boils down to
| things covered here: https://earendil.com/posts/session-
| portability/
|
| Things like obscured reasoning traces.
| copperx wrote:
| Speed and reliability.
| Alpha3031 wrote:
| You said it yourself, model size and hardware. Big models cost
| more (good optimisation reduces things slightly, but they still
| need the hardware).
| efficax wrote:
| it's a big honking trillion some parameters model. it's not
| cheap to run
| jjice wrote:
| Qwen Max is their large model - over a trillion params. Similar
| to Kimi K3 in size. Qwen 3.8 27B is going to be more accessible
| to your own hardware. I'd say that Qwen Max is not approachable
| for the majority of people and companies to self-host.
| criley2 wrote:
| GPT5.6Sol completes the suite in 70M tokens, while Qwen3.8Max
| needs like 145M tokens. So this is a case where models like
| Qwen 3.8 and Kimi K3 use a lot more output (reasoning) tokens,
| go a good bit slower, so they can ultimately achieve a better
| intelligence score than if they went more quickly.
|
| There are a couple of frontiers (ok bad word, maybe categories)
| in open weight models.
|
| These Qwen 3.8 and Kimi K3 style models aren't trying to win on
| price, they're trying to compete on intelligence and
| capability.
|
| Models like Deepseek V4 Flash (updated this week) are $0.03 a
| task, or 50X cheaper than Qwen3.8/Kimi K3, and 100X cheaper
| than Fable, while offering stunning intelligence. That's a
| different frontier for competition, and perhaps one more
| interesting for someone who wants to see them compete on cost.
| ecocentrik wrote:
| Why should open weights correlate with cost? Cost correlates
| with the expense of running the model more than it does to the
| expense of developing the model.
| jazzyjackson wrote:
| Running a large model on rented GPU is still meaningfully more
| private than handing your chat logs over to FAGA
| TheCycoONE wrote:
| The acronym is new to me: Facebook, Anthropic, Google,
| openAi?
| benjiro29 wrote:
| _Why does an open weights model cost nearly the same as GPT5.6?
| $1.14 vs $1.23 on the cost index._
|
| What cost the most in API. Input, Cached Input, or Output.
| There you have your answer.
|
| Unfortunately, we have moved so much of the actual intelligence
| of models towards reasoning, what results in some models
| getting good scores, but this is because they are dumping a
| insane amount of reasoning tokens at the problem.
|
| So a mid priced model, with heavy reasoning output, cost the
| same as a expensive model, with medium reasoning output.
|
| Before the GPT Luna price drop of 80%, you actually had the
| same price if you used Luna High and Sol Low. With the
| difference that Sol Low was insane fast, and often way better
| code.
|
| https://deepswe.datacurve.ai/
|
| Do not look at the top score but more what is on the horizontal
| axis as you go down. Sol Medium is frankly, was the best
| performance for dollar, until that Luna price drop. I will even
| argue that despite the higher price, Sol Medium is still way
| better despite Luna Max being cheaper. Or Opus Low, one of the
| better values also.
|
| What do you notice? Is that those models all have a high
| intelligence start point for their low setting. So that means
| they do not rely as much on output tokens aka thinking.
| eli wrote:
| I believe it. It's extremely good at troubleshooting. I gave Qwen
| and Kimi K3 the same annoying, complicated, intermittent bug to
| track down. Kimi did a bit better in understanding the existing
| code, but Qwen built some diagnostic tools and did an excellent
| statistical analysis on the log data. Qwen got way closer to the
| truth.
|
| I'm very much looking forward to their forthcoming smaller model
| Qwen 3.8 releases. A version that can easily run locally would be
| great.
| comboy wrote:
| How CLI are you guys using for qwen and kimi?
| eli wrote:
| I use https://pi.dev/ which works fine out of the box but is
| fairly minimal and intended to be customized. There are many
| extensions.
|
| OpenCode or oh-my-pi might make more sense if you just want a
| batteries-included agent. You can also make Claude Code work
| with other models without too much work, but I think that's
| asking for headaches.
| trey-jones wrote:
| I used claude with GLM and it's easy to set up, just hard
| to find the documentation. No headaches really, unless you
| want to use it against multiple different APIs.
| Gooblebrai wrote:
| Is there any subscription of any kind for Qwen? Or via
| Pi.dev needs to be used with API credits?
| iAMkenough wrote:
| My first web search turned up this as the top result
| https://www.alibabacloud.com/help/en/model-studio/coding-
| pla...
| adamgoodapp wrote:
| Currently only these models are available: qwen3.7-plus
| (vision), qwen3.6-plus (vision), kimi-k2.5 (vision),
| glm-5, and MiniMax-M2.5
| lkt wrote:
| Opencode Go has Qwen 3.8 Max at $10/month
| Gooblebrai wrote:
| Found the usage limits on OpenCode Go quite poor tbh
| g58892881 wrote:
| pi
| bitexploder wrote:
| OpenRouter and Pi here now.
| MrDrMcCoy wrote:
| Hermes. It can be set up to delegate to OpenCode for better
| multitasking and context usage as well.
| thefourthchime wrote:
| Did you also try Opus 5 and 5.6 Sol?
| ghosty141 wrote:
| 5.6 sol was very impressive for me. I had a weird behavior
| while using Qt and I gave it a screenshot and my expectation
| of what should happen and it read the Qt sourcode and showed
| me that my issue was a bug (including link to the ticket).
| sscaryterry wrote:
| Opus 5 is just terrible
| delduca wrote:
| Go China!
| SwellJoe wrote:
| I find that surprising.
|
| I've been trying it on several projects and have found it's
| pretty sloppy. It leaves stuff broken, doesn't reliably write
| tests to check its own work unless explicitly prompted,
| misunderstands the assignment, etc.
|
| It is smart and reasonably quick but not reliable.
| dyauspitr wrote:
| It's because they're doing some sort of combined score of
| intelligence, speed and cost. On pure intelligence it doesn't
| even show up in the top 10.
| superfrank wrote:
| I've come to the same conclusion over and over with all of the
| Chinese models that have been claimed to be catching up with
| OpenAI's and Anthropic's frontier models (Deepseek 4, GLM 5.2,
| Kimi K3).
|
| At their best, I think they're closing in on Opus and GPT, but
| they're incredibly inconsistent and the variance in output
| quality is much higher than the best from any of the Anthropic
| or OpenAI models from the last few generations. The only way I
| can describe it is that it feels like a lack of intuition with
| the models which means I find my self needing to write longer
| prompts or have more back and forth to get them to do what I
| want from them.
|
| To give an example, I have a saved prompt that I use as a
| sanity check on some data I'm storing. It reads about 50 rows
| from a DB and matches them to the UI and makes sure the data is
| displaying correctly. I've been using this with GPT 5.5 and now
| 5.6 for a few months and running it a few times a week with no
| issue. Sometimes I'll run it multiple times in a single chat if
| I notice bad data (run it, fix thing, run again, fix another
| thing).
|
| I recently tried to switch to using Deepseek v4 (first flash
| and then pro) and while both did the task just fine, both would
| do things like change the response format from one message to
| another in the same chat or randomly decide to omit things it
| didn't think were relevant. At one point I ran the prompt,
| fixed some bad data, and then said "Okay, I fixed row 7, run
| {prompt} again" and so it decided to leave row 7 out of the
| response. A few times the first message would contain a table
| and then the next run in the same chat would contain the data
| in a bulleted list.
|
| None of those are major issues and all could be solved with a
| bit more rigor in my prompting, but for me it makes them harder
| to work with. Those examples are a bit trivial, I think they're
| the easiest way for me to illustrate the gaps I see with them.
| aliljet wrote:
| Is there a path to distill this model to do very specific things?
| Like a RAG strategy for a small (or even large) corpus?
| Alpha3031 wrote:
| Depends on what you want to do. Some task specific models can
| be trained with a few ten or hundred thousand training examples
| so you can use a bigger model to produce synthetic training
| examples and then fine tune a smaller student model. I think
| that's the usual process. Whether you'd get acceptable
| performance this way depends, as mentioned, on what you're
| trying to do and what you'd consider acceptable.
| teravor wrote:
| once you are able to get the full probability distributions per
| token you can distill it on specific domains. distilling
| without that isn't generally a good idea unless you have
| invested millions in the requisite infrastructure.
| brcmthrowaway wrote:
| Could someone like Apple be playing the long game - Good
| Enough(tm) intelligence will eventually fit in our pocket and
| homes?
| LPisGood wrote:
| Almost surely. Apple is extremely well positioned to take
| advantage of this over the next decade.
| colingauvin wrote:
| DS4 Flash Q2/Q4 mixed quant fits on a DGX Spark (a $4000 device
| which is not particularly unheard of expense for Apple
| customers), and is indistinguishable for me from Opus for my
| personal daily use/assistant benchmarks[0].
|
| [0]https://humanparadox.org/local-vs-frontier-benchmarks-for-
| my... - note here I tested Q8 but have found no difference at
| lower quant.
| dofm wrote:
| Indeed. I like using Macs mostly, and the bargain M1 Max MBP
| I am using for local LLMs is a fabulous experimentation
| platform and does loads of other stuff well, so I am in no
| rush, but if I reached the point of buying dedicated hardware
| for an LLM, I'd be looking at the DGX Spark machines.
| kyxsc wrote:
| Apple is already doing this... they worked with Gemini to
| distill the model into a smaller one that fits on your phone.
| If you have iOS 27 Beta, you're already using this
| notatoad wrote:
| sort of. they have a local model, it does some things. they
| also have significant cloud infrastructure backing it, and
| most tasks are going to be sent off to the cloud for
| processing, not be handled by the on-device model. Siri is
| not on-device by any stretch of the imagination.
| sirbor wrote:
| Qwen is the way to go
| dyauspitr wrote:
| It doesn't even show up in the raw intelligence index, so how
| could it possibly be the best?
| quirino wrote:
| A couple days ago they had published an overall score of 53 for
| this model, but that was removed and today it returned with a
| score of 56.
|
| I wasn't able to find an explanation from them. Anyone knows what
| happened?
| Art9681 wrote:
| A wire transfer happened.
| ignoramous wrote:
| The kind of distillation guaranteed to work.
| steve-atx-7600 wrote:
| curious about methodology. ive seen them post results for
| claude/codex when they only ran over benchmarks 3 times per
| model...
| onomojo wrote:
| Any benchmark showing Opus 5 as the best just loses credibility
| for me. Anyone who's actually used Opus 5 daily knows what I'm
| talking about.
| copperx wrote:
| I'm dumbfounded to see Opus 5 making SO MANY mistakes in coding
| simple stuff. Most times, Fable 5 comes out to be cheaper
| because it nails so many things much quicker than Opus 5.
| garciasn wrote:
| I have Fable plan and Opus implement. I haven't had any major
| issues working this way; however, Opus does seem plain
| fucking stupid compared to what I experienced with Sonnet
| previously.
| aenis wrote:
| I do the same, and generally have good results, but it does
| stupid things with gusto.
|
| I'd open a blog with "weird things Opus did". Today it
| launched a swarm of cpu-hogging processes to test if the
| widget showing machine and I/O load is rendering nicely and
| correctly. The test went fine, but it was no longer able to
| kill those processes since they were really effectively
| hogging the CPU _in various ways_ - being diligent, some of
| them were hogging CPU, some were murdering the SSD, some
| were pounding on the network adapters. Took me 30 mins to
| recover the machine to a working state without killing the
| meaningful, messy, in-flight sessions i had going on on
| other projects.
| petesergeant wrote:
| > however, Opus does seem plain fucking stupid
|
| Infuriatingly so, in a way I don't remember Opus 4.8 being,
| but maybe I've just been ruined by Fable 5.
| hbn wrote:
| I bought my first LLM subscription with Claude right
| before they gave access to Fable 5.
|
| I got so used to it, when they finally pulled access for
| me and I had to go back to Opus I felt like I was working
| with my hands tied.
|
| I finally know what those women with AI boyfriends felt
| like when their app updated and it won't dirty talk with
| them anymore.
| moffkalast wrote:
| Fable has spoiled us all.
| sscaryterry wrote:
| Not so sure, I'm sure Opus 5 is just shit.
| moffkalast wrote:
| Eh it's better than 4.8 in terms of what it can get done
| on a good day, it's just far more taxing to get it there.
|
| Like the Fable ban stunt, I wouldn't put it pass
| Anthropic to kneecap Opus deliberately to drive more
| people to their more expensive option.
| usef- wrote:
| Weird how different people's experiences are. If it's making
| _simple_ mistakes something must be wrong in your setup
| /context I assume? It's been solid for me, beyond the usual
| LLMisms that all models have. But I keep context pretty
| minimal.
| cromka wrote:
| Statements like this typically come from working on the
| same setup and context using different models. I actually
| have that very experience now; I work on something
| security-adjacent so Fable often drops out, at which point
| Opus behaves like its lobotomized half-sibling. Pardon me
| the language, but I can't find a better example to be
| honest.
| nimonian wrote:
| Agreed. Opus 5 is doing just fine, slightly better than
| 4.8. It's personality is insufferable, but I find myself
| catching fewer problems at code review. It generally
| understands my conventions and isn't so eager to accrue
| tech debt.
| efficax wrote:
| Every model that comes out comes with a bunch of people
| saying "this one is actually dumb they were smart before"
| and I don't really get it. The models since Opus 4.5 have
| all been basically the same to me. Sometimes they do the
| wrong thing, so you have to steer and stop and correct
| them. Leaving them to operate on their own in no-human-in-
| the-loop harnesses often gets bad results. But if you
| single thread it, and keep your work targeted (you have to
| know what you want the thing to do!), clear your context,
| the models will do what you ask pretty reliably.
| PacificSpecific wrote:
| Glad to see this comment as this has generally been my
| experience as well. I'm really curious to see why it's so
| infuriating for others. My best guess is I'm using it
| more conservatively than most other users in this thread.
| yeeeloit wrote:
| At this stage in the game almost none of the comments or
| articles on HN can be trusted, if you know what I mean...
| TacticalCoder wrote:
| > I'm dumbfounded to see Opus 5 making SO MANY mistakes in
| coding simple stuff.
|
| To me it's not so much the dumb mistakes (although there are
| some of those) but the ultra-verbose, mega-inefficient
| "solutions" to some problems / prompts.
|
| Stuff that "works" if you're the kind of person that
| considers slamming a semi-trailer at 200 mph into a door did,
| technically, result in the door being somehow "open".
|
| As it's supposed to be one of the most advanced model, I
| can't help but wonder if the solutions are that
| bad/verbose/inefficient because we're already in a loop of
| models being trained on sloppy-pasta from previous models.
| visarga wrote:
| Sent to solve one task, came back with half of it solved and 2
| more problems.
| capnjazz wrote:
| "One thing worth your attention", "Two things worth knowing",
| "One thing to eyeball"
| greenchair wrote:
| This is driving me crazy. opus 4.8 did not do this to me
| not (at least during pre-5.0 timeframe). Feels like the new
| cycle is one step forward, two steps back.
| FridgeSeal wrote:
| And one of them is always something just completely out of
| scope and the other is something obvious it missed.
|
| "One thing worth your attention, if you were to detonate a
| pipe bomb in your house, it would have a negative effect on
| your living room".
| ethin wrote:
| I don't use Claude code, just Claude web, and I get this
| all the time. Or (since I have it push me to actually
| think) it will ask me some question in our back-and-forth,
| and then right after it'll provide the answer. As a "hint".
| Like come on
| vunderba wrote:
| What really enrages me is the amount of effort it puts into
| justifying weaseling out of work. (THAT'S MY JOB!)
|
| It will do everything it can to defer or push it off, to the
| point where I've had to add multiple imperative directives to
| the AGENTS file telling it, in no uncertain terms, not to
| defer tasks under any circumstances.
| cyanydeez wrote:
| sounds like someone needs a local llm.
| vunderba wrote:
| Oh I do. Headless 128GB RAM machine serving llama.cpp
| with a number of local models that I use on a daily
| basis.
|
| * Qwen3-VL picks up new images in a NAS, auto captions
| and adds the text descriptions as a hidden EXIF layer
| into the image, which is used for fast search and
| organization in conjunction with a Qdrant vector
| database.
|
| * Gemma3:27b is used for personal translation work
| (mostly English and Chinese).
|
| * Some small 8b models (like llama3.1) for sentiment
| analysis on text.
|
| But haven't really tried using local LLMs in conjunction
| with agentic harnesses yet.
| cyanydeez wrote:
| recommend opencode w/qwen 35B or 27B with MTP.
|
| My secret sauce is to use LLAMAcpp's reasoning-budget and
| reasoning-message that trigger cut off to overthinking
| with a message that says to either us subagents or
| compress the context. opencode's dynamic context pruning
| plugin can get you pretty far into the stratosphere.
| vunderba wrote:
| > My secret sauce is to use LLAMAcpp's reasoning-budget
| and reasoning-message that trigger cut off to
| overthinking with a message that says to either us
| subagents or compress the context
|
| Thanks for the tip - I like this a lot. I remember having
| to do a lot of tweaking to curtail Qwen QwQ-32b when it
| would go down an endless psychotic recursive reasoning
| loops as part of its _" chain of reasoning."_
| smartbit wrote:
| Any experience with Sleev as replacement of DCP?
|
| See https://news.ycombinator.com/item?id=48883538 25 days
| ago
|
| > _The Sleev (the project has been renamed to make a
| startup) creator was shilling their project in the
| OpenCode Discord. That person is very convinced they have
| something that no one has ever built before. They focused
| on token reduction without any real evals for capability
| impacts.
|
| I'm generally against this context pruning without
| prompting or details. Sleev is very opaque about how it
| works and definitely will bust your cache._
| logicchains wrote:
| "As you requested, I've finished task X. Honestly, task X
| turned out to require task Y, which I haven't actually done.
| Task Y is the next step if you'd like to continue along this
| route."
| pornel wrote:
| This is the hard-won load-bearing quote.
| dr_dshiv wrote:
| Belt and braces all the way down
| paradox460 wrote:
| The shape of this problem is very heavy
| vardalab wrote:
| Yeah, I told it to save in its memory that I don't want to
| have any more word salad!
| cromka wrote:
| Agreed, it's extremely frustrating. It's the only model that
| actually makes me curse when talking to it, even knowing how
| counterproductive it is.
| moffkalast wrote:
| I'd certainly rank it at the very top of the want to kill
| yourself when using it benchmark. It outperforms everything
| else on that leaderboard.
|
| With weaker models you can sort of understand, they're trying
| their best and failing, but this thing just channels its
| immense inteligence into being as annoying as possible
| instead. I know it can do what I'm asking it to do, but it
| just finds a way to weasel out of it, or maybe just thinks
| for 10 minutes instead, then fixes one thing and breaks four
| additional ones.
| msp26 wrote:
| yep matches my experience completely
|
| But even fable has the annoying tendency to invent new
| jargon and produce an incomprehensible soup of text.
| cromka wrote:
| Is there any model that knows how to smooth an overly
| literary text over? I find Opus and Fable constantly
| decorate the documentation they write like a damn 19/20th
| century writer. We're working with IT stuff yet it writes
| like it's going to win some Pulitzer prize. It's that one
| thing I don't get why they can't train them to do
| properly: I have not encountered a model yet that sticks
| to the _current_ language of the domain it 's tasked
| with.
| msp26 wrote:
| Not sure how to fully fix this but I remember a session
| last week where I got so fed up mid way though reading a
| response that I used the following:
|
| "give me this again without jargon invented this session
| at high density
|
| and with a couple (maybe more or less) simple useful
| ascii diagrams underneath each design"
|
| The context is that I was discussing an experimental new
| idea for my video game review analysis product.
|
| Designs 1,2, and 3 were horrible: the model even
| suggested a rejection after the word soup so it would
| have been pointless to waste my fleeting time on Earth
| reading it.
|
| Otherwise, I generally really enjoyed using fable for
| bouncing ideas. It was an absolute joy to have this thing
| provide useful criticism, analyse sample data, and create
| prototypes so that I could elevate my understanding of
| the problem without stepping down from a pure
| intuition/design headspace.
|
| But I don't consider the purely model written code usable
| for a feature this important. I'll probably scrap it
| entirely and start from scratch with newfound
| understanding.
| hungryhobbit wrote:
| The cursing thing blows my mind. "User is upset? Let's make
| decisions even faster (ie. more wrong) because clearly that's
| what they want!"
|
| It's a simple switch to make: cursing = try harder instead of
| cursing = stop trying. Is it really impossible to train
| Claude that way?
| kloop wrote:
| It's training data might have a ton of examples of people
| hurrying and screwing more after being yelled at
| cromka wrote:
| Replying to myself, because I just bumped into these: https:/
| /www.reddit.com/r/claude/comments/1vfvdgz/anthropic_l... http
| s://www.reddit.com/r/ClaudeAI/comments/1vgpyni/my_opus_5...
|
| Especially the second one seems exactly like my experience.
| sscaryterry wrote:
| I've ditched Anthropic completely because of it. It makes me
| furious.
| bontaq wrote:
| It's an infuriating model
| nomel wrote:
| What's the clear best, that you see?
| petesergeant wrote:
| Fable 5
| drschwabe wrote:
| GPT 5.6 Sol
| kachnuv_ocasek wrote:
| GLM 5.2
| dgellow wrote:
| Hilarious to see only different responses
| ofjcihen wrote:
| Should have been "clear best and what do you do"
| enraged_camel wrote:
| It's my daily driver. I like it and find it noticeably better
| than Opus 4.8.
|
| After I started reading complaints about Opus 5, I gave Fable
| the task of evaluating a bunch of code Opus 4.8 had written and
| compare it to Opus 5's code. Fable ran a dynamic workflow and
| the scores came back 15-20% higher for Opus 5's code in terms
| of quality, correctness and readability/conciseness. I did not
| tell Fable which Opus wrote which code, and I turned off memory
| as well to ensure there was no pollution from that angle.
|
| My only complaint is that Opus 5's prose is annoying as hell. I
| wrote a custom skill for it for concise debriefs and it has
| been working pretty well for me.
| agopaul wrote:
| > My only complaint is that Opus 5's prose is annoying as
| hell. I wrote a custom skill for it for concise debriefs and
| it has been working pretty well for me.
|
| I'm doing the same right now, and I've found that asking for
| "simple English" works most of the times, although not
| always.
|
| Did you find better wording that works consistently?
| fellowniusmonk wrote:
| I have some internal tests I use for areas where one particular
| solution/paradigm is dominant but worse.
|
| Opus 4.6 is the last model that's actually useful and can
| "adjust" its perspective to use the newer & better solution.
|
| Where Opus 4.8-5 has over fit training on worse/older but
| "dominant" solutions it refuses to adjust.
|
| Not only does this create an existential threat to adopting
| progress but it also means that if you have a code base that
| has rare but real world tradeoff the newest versions of Opus
| 4.7, 4.8 and 5 are worse than useless and become a major dev
| timesink.
| Fordec wrote:
| Yeah, I've dropped back to 4.8 entirely for the remainder of
| this billing cycle. I'm going to be seriously looking into Qwen
| adoption and harness migration options over the course of
| August.
| thomasfromcdnjs wrote:
| Same.
|
| I could not get Opus 5 to do anything without losing a few
| years of my life from stress.
|
| Fable has been okay but I am doing ML work and not allowed to
| use it which feels insane.
| cesarvarela wrote:
| It is infuriating to interact with, but it is also first in
| many blind test leaderboards on LLMArena
| CuriouslyC wrote:
| Ironically, Opus 5 is the most benchmaxxed model I've seen from
| Anthropic. It is legitimately smart in a lot of ways but it has
| communication issues, both in terms of how it communicates (all
| the autism of GPT class models, without the brevity) and how
| well it catches all the nuance of what you tell it.
| sunaookami wrote:
| Can not confirm, for me it's the complete opposite.
| combyn8tor wrote:
| It works fine for me. Only issue I have is that it has me
| constantly reaching for the dictionary.
| conception wrote:
| https://claude.com/blog/the-new-rules-of-context-engineering...
|
| I don't see anyone talking about how you have to completely
| change your prompting strategies with Op. 5 versus 4.8 to get
| the most success.
| petercooper wrote:
| Hopefully this boils down to the smaller versions they've teased.
| In my experience, Qwen models are the closest to the "less
| knowledge, more intelligence" (yes, the two are hugely
| correlated!) ideal some tool-dependent tasks need. Even the 3.5
| 2B can be easily prompted to always lean on tools and not jump to
| false conclusions (although its actual coding skills are abysmal,
| as you'd expect).
| quotemstr wrote:
| > less knowledge, more intelligence
|
| People produce such models by over-RL-ing smaller models on
| math and coding tasks. I've found the results capable of
| neither innovative work nor thinking outside the box. They're
| straight-A students raised by tiger moments who never let them
| play freely for hours in the dirt.
|
| Perhaps you could say such models are _skilled_ --- but
| intelligent? Not by my measure.
|
| People and AIs alike need diversity of experience and a broad
| liberal arts education to see hidden connections between fields
| and make real advances.
| DC-3 wrote:
| It's amusing to me that AI has become sophisticated enough
| that people have started being racist to it.
| petercooper wrote:
| I agree with you to an extent, but you have certainly given
| me food for thought.
|
| Sticking to LLMs, they seemingly get their intelligence
| (whatever that really means) _from_ building models rich with
| knowledge, so you could have a point. But Qwen models seem to
| be _particularly_ good, even at small model sizes, at
| maintaining both their own knowledge while acquiescing to and
| integrating external information in the moment.
| looksjjhg wrote:
| That took what 2 years? I love how the chip ban made them more
| efficient
| Footprint0521 wrote:
| Facts lol, now all the Chinese models are 1/40th of the cost
| for the same intelligence
| ben8bit wrote:
| Haven't tried this yet, but going to soon! I have to wonder what
| happened at Anthropic. We've cancelled our subscription in favor
| of OpenCode & Codex. Sol is just so good & OC goes so far for
| every $ spent. Claude's become a pain to work with - average
| output with an annoying personality. Who knew this would be an
| issue even a year ago? In any case, loving the stuff from the
| Chinese models!
| tomComb wrote:
| > an annoying personality
|
| I was with you until there. Qwen and the OpenAI models are
| great, aggressive agents, but they're not as good as the
| anthropic models for human interaction. They just don't have
| the subtlety, understanding, or attention to detail.
| ben8bit wrote:
| Really? I've heard so many other people complain about this
| recently. And maybe it's possible that it's the prompt style
| even. But interesting that it's not across the board.
| colingauvin wrote:
| Claude 4.5/4.6 - absolutely agree. Fable 5? From my (limited)
| testing, also reasonable to interact with.
|
| Opus 4.7/4.8/5? Absolutely smug and antagonistic and preachy.
| I'm constantly fighting with it to stop fighting me and
| accept that I occasionally know better. It's really
| frustrating to spend so many tokens of such an expensive
| model arguing with it.
| theropost wrote:
| Anthropic is a bit nuts, I had $260 of credits on my max account
| for the extra usage the other night. It was expiring, so I
| figured I'll fire up an agentic swarm to deep dive and make some
| deep changes to some old cold bases.. literally 25 minutes or
| less, $260 burnt, it didn't get get into the implementation, just
| wrote a ton of useless plans for the most part. It really opened
| my eyes to what they expect to charge people.. wayyyy overpriced.
| polishdude20 wrote:
| You should just spend those towards a cursor subscription.
| cortesoft wrote:
| It's crazy how different the credit cost and subscription cost
| are.
|
| With the $200 subscription, I can have Fable on ultracode
| working for hours and not dent the usage limits.
| AlexandrB wrote:
| VCs are footing the bill for that $200 subscription.
| ericd wrote:
| They have something like 80% gross margins, are at a
| $100B/yr ARR, and are growing at 10x per year... If that
| keeps up, they're going to be doing more revenue than
| Google in a year ($400B ARR, 20% per year growth)
| dexwiz wrote:
| How can you sanely project the last 12 months forward? We
| have seen a huge uptick in usage. Last summer AI was a
| toy to most devs, now every enterprise developer I talked
| to uses it every day. Coding agent providers are surely
| going to hit market saturation in the near future.
| ericd wrote:
| Maybe, maybe not. Personally, I hope local AI eats their
| lunch so that the benefits are more decentralized and
| accrue more to society generally.
|
| I don't think you're right about that last prediction, at
| all. And new use cases are opening up as these get
| smarter. I think things are going to get pretty weird.
|
| But the point was that it really doesn't look like
| they're losing money on users, on average.
| senordevnyc wrote:
| Where did that $100B figure come from? I thought they
| were at ~10B at the end of 2025, so they're either not at
| 100B yet, or they're growing way faster than 10x / year.
| ux266478 wrote:
| At last, a valid usecase for VCs.
| dionian wrote:
| i'll take it, just hope they dont rugpull us soon. im
| sure its coming
| riknos314 wrote:
| The $200 sub is customer acquisition cost to hook devs that
| then become the marketing team trying to get their company
| to bring in Claude (at the highly profitable API price).
| notatoad wrote:
| yeah, i tried out GLM-5.2 when the news was all full of hype
| for that, and it's fine... definitely better value that API
| rates for claude. but comparing the value i got from that to
| the value i get from a claude max subscription... claude is
| way cheaper.
| aenis wrote:
| I managed to lose around $300 in credits I had saved for some
| emergency /fast sessions the following way: switch to Fable.
| Work on the design. Downgrade to Opus for the build. If _any_
| of other parallel Opus session has /fast enabled it seems to
| enable it for the newly spawned session by default. Before I
| knew it, the $300 was gone. I think the bug is now solved, but
| it was rather unpleasant. I dont ever remember bugs that would
| drain my wallet - with claude code its just another Tuesday.
| Still love it.
| tempest_ wrote:
| I dont love it.
|
| Opus 5 is just a token burner.
|
| I use fable plan and spawn opus 4.8 workflows which seems to
| work alright.
| robbru wrote:
| Opus 5 loves to stop working "for safety reasons" and shuts
| down the session! I avoid it at all costs now. Opus 4.8 has
| been my default as well.
| aenis wrote:
| I suspect it must depend on how one manages their codebase
| - wrt to docs, ADRs, and general guardrails.
|
| For me it is not great for design work - Fable is way
| better, and 4.8 was conservative and thus better (Opus 5
| seems to jump to conclusions far more eagerly). But for
| overnight builds, where I give it 8hrs worth of work on
| LLDs created by Fable - its great. Where Opus 4.8 would
| often lose the plot and stop for questions clearly answered
| in the LLD - Opus 5 does manage to complete. Since it
| launched, I don't remember it ever disappointing me with
| builds. But designs? Boy, is this thing explosively stupid
| sometimes.
| ethin wrote:
| For me it's the opposite: I don't have $200 plus to throw
| at Anthropic every Month, and when I do get to use Fable it
| rips through my usage credits like there's absolutely no
| tomorrow.
|
| Of course, the hilarious thing to me is that Anthropic
| likes to claim that the usage limits are because of
| resource allocation problems or something like that.
| Obviously no such issue exists, otherwise they wouldn't
| allow you to bypass it by just paying a bit more and it
| would be a hard limit. So usage credits are entirely their
| way of just screwing you out of more money.
| gnull wrote:
| Claude code is just pool quality. They don't make how this
| thing will behave clear to the user, or give control. They
| fail at anything that needs an abstraction or model, not just
| APIs and shell scripts glued together. And "just ask AI"
| seems to be the default fix.
|
| That vibe coding they brag about as if it was a good thing,
| it shows.
|
| Take their notation for describing permissions. The docs are
| not comprehensive, and in practice it doesn't quite work how
| they describe it.
|
| Or their management of sub-agents. I once lost a sub-agent,
| it finished and disappeared from UI. Apparently, you can't
| bring it back yourself: you have to ask the parent agent to
| do it for you. But the parent was Fable, and I ran out of
| credits, so I was locked out of using my opus sub-agent
| because of it.
|
| Or an even more grotesque example: when you paste your claude
| API token to authorize, it covers characters with *. But it
| seems like an LLM has hallucinated a limit of API key length
| and the tail of your key stays visible.
| thejosh wrote:
| so many ridiculous "how the fuck did this get through basic
| QA?" issues with Claude Code.
|
| I can't believe how many critical bugs fall through.
|
| My favourite one is the bug where Plan mode can execute
| destructive commands inadvertently.
|
| Then all these get closed with `Closing for now -- inactive
| for too long. Please open a new issue if this is still
| relevant.`. Awesome.
| formerly_proven wrote:
| > I can't believe how many critical bugs fall through.
|
| Almost like CC is 100% vibe coded.
| ethin wrote:
| I mean. This is what happens with vibe-coded projects.
| When there's no actual software engineering going on, I
| wouldn't expect anything better than this.
| hungryhobbit wrote:
| What amazes me is how, for a vibe coded product where all
| they have to do is use their AI to fix things ... NOTHING
| EVER GETS FIXED!
|
| I've probably gone to file 20 bugs. In all 20 cases there
| wasn't just one issue already filed for it: there were
| several, each which had a bunch of upvotes. And in all 20
| cases ... every. last. one. ... Anthropic closed the ticket
| with no comment.
|
| IF YOU ARE GOING TO HAVE A SHITTY VIBE CODED PRODUCT, AT
| LEAST USE YOUR SHITTY AI TO FIX THE SHITTY PROBLEMS!
| mikae1 wrote:
| And at that cost they're still not profitable. It's going to be
| a bumpy road ahead...
| arrowleaf wrote:
| I thought they are making a profit on API pricing? A quick
| Google shows somewhere between 50-70% margins on API
| inference.
| bakugo wrote:
| API pricing is almost definitely profitable, but at this
| point I assume it's a small minority of their inference
| traffic compared to subscription usage, and unlikely to
| make up for the rest of their expenses on its own.
| enedil wrote:
| Why would you assume so when companies 150+ people can
| only use API pricing? My assumption is that more people
| use Claude at work than personally.
| tyre wrote:
| People keep saying this but from what we've seen, Anthropic
| models are marginally profitable and earn back their costs
| over their lifetime. The company is burning money building
| the next versions and other ventures (e.g. verticals), but
| the models themselves have been profitable.
| gamblor956 wrote:
| They're EBITDA profitable, not GAAP profitable.
| an0malous wrote:
| What's the blast radius of this bubble popping? It's all
| private investment still right?
| bhewes wrote:
| Two thirds of most of the DC builds are not compute. So
| it's a CRE play the last leg holding up that mess.
| arikrahman wrote:
| Meanwhile I can do all that and more with reasonix harness
| for Deepseek with a cache hit rate of 99%. And that's with
| unsubsidized American providers like cloudflare or Digital
| Ocean
| swalsh wrote:
| I think profitability is a matter of accounting. Inference is
| where money is made, but training is where money is spent. We
| keep getting new models every few months, but frankly the old
| models are still quite usable. I suspect labs will soon start
| specializing in expert models per use case so they can
| increase the lifespan of individual models, and change the
| profitability per model.
| CuriouslyC wrote:
| That's not the only reason to go to expert models. The more
| different domains you try to stuff in there, the more
| parameters the model needs to keep things coherent and not
| overload tokens in a way that induces errors. For example,
| if a model trained only on biology text sees "sonic
| hedgehog" there's no ambiguity, and this compounds for all
| the things that are "overloaded," in the training corpus,
| which turns out to be quite a bit.
| criddell wrote:
| > wayyy overpriced
|
| Maybe they consider that hiring a person to do it would have
| cost at least as much and taken much more time, so paying them
| is a bargain.
| echelon wrote:
| Yeah, but now we can hire the Chinese instead for 1/100th the
| cost. It's an even better deal.
|
| Plus we get to own, keep, run, do whatever with the model. We
| don't feel trapped. Moreover, it's something we can truly
| build on top of and own our own destiny.
|
| Anthropic and OpenAI are the new Oracle (Oracle pre-AI;
| Oracle is even worse now). Expensive, feels like dealing with
| a lawyer, and not at all open. They just became infinitely
| less cool than they were a month ago.
|
| The whole of our industry is going to migrate to open
| weights. We're smart enough to know this is the better deal
| and technical enough to be able to pull it off.
|
| The only thing that might save these OpenAI and Anthropic in
| the near-term is an abundance of enterprise contracts
| negotiated with non-tech companies. They'll soak consulting
| firms and F500 companies for "AI" integrations.
| criddell wrote:
| > the new Oracle
|
| I think that's exactly what they are going for - enterprise
| and government customers.
| pvtmert wrote:
| Anthropic is the new AWS.
|
| Amazon's first principle is the Customer Obsession. Making
| customers happy.
|
| Fun bit is that the human psychology rates personal looking
| fixes better than having no issues at all.
|
| For example, AWS overcharges you, you contact support, and more
| or less hassle free they refund or issue credits. The customer
| feels appreciated, or at least got something "extra" or
| "special treatment".
|
| Meanwhile, any other (small) cloud. Simple, no weird charges.
| Even _most_ of network egress is free. But, no reason to call
| support or feel "extraordinary". Comes out as "meh" against
| Amazon's "top tier" support model...
| john01dav wrote:
| Anthropic's constant changing of its mind leads to
| instability which leads to unhappy customers
| axpy906 wrote:
| I've never gotten a refund from Athropic.
| riknos314 wrote:
| Aws is an infrastructure company that builds services on top
| of that infra to sell more of it at a higher margin.
|
| Anthropic trains models on AWS's (and GCPs, and Microslop's)
| infrastructure, then skims margin off of selling inference
| also on the infrastructure owned by the other companies.
|
| These are extremely different businesses.
| swalsh wrote:
| Its tough to go from max account at home and pay per usage
| enterprise account at work with heavy usage limits... but the
| limits are there because pricing is insane. Feel like I'm in
| the $5 Uber rides phase at home.
| hahahaa wrote:
| The Chinese models are the public transport in the uber
| analogy. Once the price the goes up catch the bus!
| swalsh wrote:
| Lol perfect analogy. I'm still paying for claude because
| the quality is unmatched.
| hahahaa wrote:
| You plugged in a space heater on a roofless house.
|
| There is some element of responsibility on the user to guide
| and monitor the model/harness and not let it rip to burn
| tokens.
| tarnith wrote:
| Hint: The new models are really good at burning tokens.
|
| I've had to use it a bit for work, and it's been remarkable
| watching the degradation in performance with the default
| suggested current models (Opus 5 as a prime example) vs the
| models that got them huge attention a year ago (Opus 4.6)
|
| If you give 4.6 a spec, or existing code to implement a feature
| in, it will ask some pointed questions if there's something
| unclear in the spec, and then produce a plan and move to
| implement it.
|
| 5 will freak out at even a basic task, ask itself if it's own
| assumptions or your instructions are correct, proceed to re-
| assess it's own plan, and it's instructions 3-4 times, and then
| maybe produce code after burning several hundred thousand
| tokens (and quite a bit of time) analyzing existing code and
| thoroughly sweeping it for irrelevant problems both to the task
| it was given and the spec it came up with.
|
| It's quite bizarre to me how well advertised the benchmarks and
| anecdotes from people one shotting MVP browser games are,
| compared to the experience of everyone I know that's had to
| actually use it to accomplish even a relatively basic task.
| brynnbee wrote:
| I had same experience with OpenAI. I have the $200/month plan
| and use 5.6 Sol all the time. What would normally use about 2%
| of my _weekly_ allowance burned through $100 of credits in 40
| minutes.
| d2p wrote:
| I clicked through and it showed Qwen at the top at 55.4 compared
| to 55.3 for Opus Max. I have a screenshot.
|
| Then I clicked away and back, and now it goes Qwen second, with
| 58.4, to Opus Max at top with 59.2.
|
| I have screenshots of both. The description above the chart is
| the same in boh cases:
|
| > Artificial Analysis Agentic Index > Represents the weighted
| average of agentic capabilities benchmarks in the Artificial
| Analysis Intelligence Index (GDPval-AA v2, 3-Banking)
|
| What happened? How can the scores change so much in a few
| seconds?
| h14h wrote:
| They JUST updated their methodology:
|
| https://artificialanalysis.ai/methodology/intelligence-bench...
|
| Edit to provide AA's article explaining it:
|
| https://artificialanalysis.ai/articles/artificial-analysis-i...
| ahartmetz wrote:
| Fixed the result, eh? In both senses of the word.
| gpt5 wrote:
| What was the change?
| johnnyApplePRNG wrote:
| I have been suspicious of these AI leaderboard sites for some
| time now, and this only increases that suspicion.
| torginus wrote:
| In that case they should clearly label that this is a new
| benchmark.
| splatzone wrote:
| Can someone please explain what changed, when it happened,
| and whether it was surreptitious?
| kmeh wrote:
| > HLE, AA-LCR and AA-Omniscience are now graded by GPT-5.6
| Luna (medium), replacing GPT-4o, Qwen3 235B A22B 2507, and
| Gemini 3 Flash Preview respectively. These checks are now
| unified under a more capable modern model, selected for
| strong agreement with human judgment in our grader validation
|
| Interesting that they chose a nano-sized model from OpenAI to
| be a grader for benchmarks involving knowledge and
| hallucination.
| nolok wrote:
| What's interesting is that if you ask 5.6 Sol or Opus 5
| they will tell you it's a bad idea to have the reviewer be
| the dumber of the set as it can't judge them properly to
| decide who is right, and thus if one is better because it
| found an answer that's better but contradict the obvious it
| would be biased against. I know because I just had a
| consensus conversation with them this afternoon about a
| design that was similar (though about something completly
| different than judging agentic quality or whatever).
| WD-42 wrote:
| Same, they just updated it. Hacker news effect?
| apitman wrote:
| Welp. That didn't last long
| personjerry wrote:
| They should probably freeze the results before publishing.
| Gcam wrote:
| Hey! George from the Artificial Analysis team here. We
| published an update today that does result in a change of the
| order, Qwen3.8 Max to second rather than first. The methodology
| change was an already planned upgrade to our equality
| checking/grader models, and brings the latest 3-Banking version
| to Artificial Analysis. Regular updates are normal for us to
| keep our benchmarks up to date.
|
| The order changes but I think the story discussed in this
| thread holds - this is a very impressive release and Qwen3.8
| Max is a huge step up in agentic capabilities.
|
| Relevant blog post (also linked to by others):
| https://artificialanalysis.ai/articles/artificial-analysis-i...
| saretup wrote:
| You gotta admit the timing looks very suspicious.
| TacticalCoder wrote:
| > You gotta admit the timing looks very suspicious.
|
| Do you mean the timing looks like: _" We're SV tech-bros.
| Our benchmarks showed a chinese model above what's
| considered the best model at the moment. So we quickly
| modified the benchmark so that our SV tech-bros don't look
| like they're losing to a chinese model"_?
|
| That's indeed a bit fishy.
| Maxious wrote:
| Luna pricing was just cut by 80%
| https://www.eesel.ai/blog/gpt-5-6-pricing and as the blog
| post states is a more accurate judge than the previous
| methodology.
| atemerev wrote:
| Well, that's the bad index then. It is barely usable in my
| opinion compared to other Chinese frontier models.
| ramon156 wrote:
| which one of the other chinese frontier models is better?
| h14h wrote:
| This has me hopeful for Qwen3.8-27B!
| camnora wrote:
| Qwen is just crushing it overall. I regularly use 3.7-flash for
| everyday coding needs and it gets the job done.
| seizethecheese wrote:
| Opus is still first in Intelligence Index followed by Fable, GPT
| 5.6, Kimi K3 then Qwen 3.8 max.
| https://artificialanalysis.ai/#intelligence
|
| Our leaderboard combines Arena ELO, AA Intelligence index,
| latency and speed and goes: #1 Opus 5 #2 Kimi K3 #3 Qwen3.8 Max
| #4 GPT 5.6 Sol
|
| Source: http://pellmell.ai/leaderboard.
|
| This jumps around a lot based on the top throughput and latency
| of whatever provider happens to be best at the moment.
| d4rkp4ttern wrote:
| All these "intelligence" benchmarks miss something extremely
| important when using an LLM in a code-agent harness: How it
| communicates with you about what it did.
|
| Opus-5 is practically unusable (for complex tasks) in this
| sense - its updates are voluminous, and dense with cryptic
| language (there are numerous reddit threads complaining about
| this, so it's not just me). I often have to ask it to re-state
| concisely in plain terms.
|
| For a fairly gnarly task, after fighting with with Claude-Code
| + Opus-5, I ported my session to Codex + GPT-5.6-sol, and it
| was like a breath of fresh air.
|
| Arguably a key aspect of intelligence is concise, clear
| communication, and current benchmarks miss that, at least as
| far as I'm aware. I would think some arena-type benchmarks
| where humans rate responses would measure this, though I'm not
| sure which those are.
| moffkalast wrote:
| Damn I thought it was my extra instructions, I swear
| everything it writes is in some shorthand with direct
| references to variables that literally nobody could figure
| out unless you literally just wrote that code 5 minutes ago.
| I had it stop writing comments altogether cause it was always
| four lines of complete and utter nonsense, and it doesn't
| even obey that rule half the time. Despite doing an extensive
| back and forth to make a complete plan, 5 seconds into the
| implementation it changes its mind and makes another
| assumption, adding some extra thing that tends to break the
| entire approach and needs follow-ups to repair or cleanup.
| Instruction following is basically non-existent compared to
| Fable, it just does whatever the fuck it wants.
| fearmerchant wrote:
| Everything is load-bearing with 3 measured blockers.
| pixelready wrote:
| Don't forget the smoking guns! I think these new models
| have been reading too many Agatha Christie novels.
| jiggawatts wrote:
| GPT 5.6 has similar language quirks that makes its comments
| nearly unusable.
|
| I wonder if this is a side effect of MoE models -- they can
| write excellent prose, but not simultaneously with writing
| code.
| notfromhere wrote:
| I think it's just where they focused RLHF resources. The
| models have generally only gotten worse at writing.
|
| And writing doesn't have validators like code so you can't
| really scale it in the same way
| chpatrick wrote:
| Is that what it feels like when the models get smarter than
| us?
| micw wrote:
| Guess that's the exact point of the "intelligence"
| benchmarks
| gpt5 wrote:
| A smarter model would know how to communicate with you
| correctly, and not just throw jargon it has just invented
| at you without explaining it.
| IanCal wrote:
| s/model/engineer
| chpatrick wrote:
| But if you have two experts in a field talking to each
| other you wouldn't expect them to dumb down their
| communication.
| riknos314 wrote:
| Effective jargon usage is understood by the target
| audience.
|
| If the AI is communicating to me and can't select the
| appropriate jargon level, it's failing at communicating
| effectively.
| chpatrick wrote:
| Or you're below its level.
| ranguna wrote:
| That's a loadbearing, heavy shaped, second take worthy
| statement
| cloverich wrote:
| Concise, jargon free or limited explanation is the
| opposite of dumbed down. It requires to most skill and
| understanding to do well. Opus 4.8/5, for whatever
| reason, are getting worse at this crucial skill.
| HappyPanacea wrote:
| No, a smart model should also give a concise executive
| summary, "brevity is the soul of wit".
| a2ff6eeb0 wrote:
| So, in short, once models get smart enough they stop
| bothering telling us what they did.
|
| Yeah, makes sense. A parent wouldn't bother explaining
| the details of their job to a toddler.
| FridgeSeal wrote:
| Which would be fine, but the parents also just smeared
| tomato sauce over the walls too, so let's not get too
| ahead of ourselves.
| computably wrote:
| "There is a view in some philosophical circles that
| anything that can be understood by people who have not
| studied philosophy is not profound enough to be worth
| saying. To the contrary, I suspect that whatever cannot be
| said clearly is probably not being thought clearly either."
| msp26 wrote:
| No the models are just ass at communication without being
| directed.
|
| Try asking them to make useful diagrams for some stuff in a
| codebase, out of the box without excessive hand holding
| they don't make good choices about what's worth
| communicating and how to do it.
|
| You see this in their pointless frontend copy all the time
| too.
| embedding-shape wrote:
| Like any time you make them do any UI without strict
| directions they'll almost always add a label describing
| the feature somewhere. Ask for a calculator, and
| instructions for what the different buttons do might
| appear in the bottom out of nowhere for example.
|
| Same concept of "over-sharing" seems to prevalent in a
| bunch of domains when it comes to LLMs, sometimes more
| visible, sometimes less.
| satvikpendem wrote:
| Eh I don't know, I care whether it gets the job done and I
| can see the difference when I review the code, not how well
| it needs to explain the code to me, I can just read it
| myself.
| veber-alex wrote:
| It takes 2 minutes to fix Opus 5
|
| https://code.claude.com/docs/en/output-styles
| veber-alex wrote:
| lol. I am being downvoted for trying to help people out.
|
| This community is pure trash.
| HDBaseT wrote:
| Instead, you should say something like:
|
| "You can adjust the output style in your
| '.claude/settings.local.json' file".
|
| OR
|
| "You can decrease verbosity by doing x, y and z."
| octoberfranklin wrote:
| > and dense with cryptic language
|
| Yeah if you ask it a medical question, it answers in that
| impenetrable jargony style that clinical journals use... full
| of unnecessarily custom adjectives ("orthopedic" instead of
| "of the bone") and discipline-specific terms (anterior,
| distal) even when the user didn't display mastry of this
| terminology (hint to frontier labs: add training cases for
| this; it will improve your model's usability).
|
| My theory is that LLMs perceive the writing styles of various
| fields as being like different (but related) languages, and
| they're inclined to answer a question in the language of its
| source material unless specifically asked otherwise. If you
| add "ELI5" the model treats it as a question plus a
| translation task.
|
| I think this is why programming questions are answered with
| an exaggerated cringey form of HN-speak ("load bearing",
| "gate" as a verb, "dissolves") by some models.
| plaguuuuuu wrote:
| it's like asking a developer to explain something.
|
| I always get Haiku to rephrase anything human-facing.
| hungryhobbit wrote:
| It seems like if latency is having such a big effect that it's
| changing the winners, maybe your tests are awful and shouldn't
| be so latency dependent?
|
| I mean, I get it: how fast a model responds is relevant. But a
| test that changes second by second is far less relevant than a
| test that tells you how smart the model is, and _accounts for_
| latency in some way that isn 't constantly changing the result.
| seizethecheese wrote:
| Latency isn't changing the results for the coding index or
| arena ELO, but neither of those take latency or throughput
| into account, so we added those to our leaderboard as score
| components.
|
| Latency and throughput matter a ton as a user, so I think
| it's actually totally defensible for a leaderboard to bounce
| around a lot as these numbers change. The best model to use
| changes a lot based on these!
| esafak wrote:
| It is also the most expensive open source frontier model, per
| task; cf. _Cost per Intelligence Index Task_. If it is as good as
| the benchmarks indicate it bodes well for Qwen and China. For my
| part, I 'll pass; it is not on the Pareto frontier.
| Fordec wrote:
| Anthropic have a real fight on their hands now. The competition
| is no longer 6 months behind, it's 6 days. If this had come out
| two or three weeks earlier this would be an absolute market
| leader on both quality and timeline.
| brettgo1 wrote:
| Out of curiosity, what's currently the best model I can use
| locally?
| daemonologist wrote:
| With an unlimited budget, Kimi K3 (which is quite comparable to
| this Qwen Max imo). With a normal budget/a PC you might already
| have, probably Qwen 3.6 27B.
| arjie wrote:
| $500k - Kimi K3 (maybe $250k? Haven't done this one)
|
| $25k - DSv4 Flash
|
| $4k - Qwen 3.6 35A3B Q5
|
| $1k - Qwen 3.6 27B Q4
|
| Some people prefer the sense over the MoE YMMV.
| apitman wrote:
| These numbers look about right based on my experiences as
| well. Though for a single user I think 2x DGX Spark (~$10k)
| runs DSv4 Flash fairly well right?
| colingauvin wrote:
| 16 DGX Sparks can run K3 at a reasonable TPS. So that's $64k.
|
| 2 DGX Sparks can run DS4 at 1 million context with 50 TPS so
| that's $8k.
|
| 1 A4500 can run 35A3B. Those are about $1200 new.
|
| 27B actually takes more hardware to run than 35B because
| attention is done differently I believe and therefore KV
| Cache takes a lot of space. It will run on an A4500 but it's
| slow and context will be like 32k.
| proxyscore wrote:
| Does it matter, it's all non deterministic bs ware and deepseek
| is eating the Americans lunch
| bonoboTP wrote:
| I distrust any benchmark where Opus 5 beats Fable 5.
| jjcm wrote:
| China has caught up is the main takeaway here. The SOTA models
| are so close that it's really hard to compare them intelligence
| wise - you have to get a feel for them yourself and what works
| for you.
|
| What I'm really excited for is the 27B model. 3.6 is still the
| king of local, and if 3.8 makes the same improvements it could
| really legitimately make local viable as a default. I'd love to
| run a perpetual agent on 3.8 that's locally driven.
| icedrift wrote:
| I'm still skeptical of the smaller models after the talent
| exodus a few months ago.
| jimbo808 wrote:
| At this point I feel like the only factor differentiating
| SOTA models now is who they're propagandizing you on behalf
| of (not considering agentic tooling/state management, etc).
| Zambyte wrote:
| Qwen 3.6 27b is already a viable default. I'm running it on a
| single 7900 XTX right now for Go development with pi. It's
| great.
| snapplebobapple wrote:
| Works great with room to spare on my lenovo pgx too
| bitexploder wrote:
| I find 35B A3B viable as well, but your harness and runtime
| really matters to get tool calling and such dialed in. In
| fact, I would encourage you to experiment with it some as I
| find I get more reliable output from 35B A3B, though 27B is
| still generally smarter. A3B with a review cycle or two from
| 27B is great for me.
|
| One of the reasons is, with good specs and design, A3B is
| just so fast. It isn't as smart as the 27B model, but it is
| close enough it can usually figure it out with the right
| tools.
| monster_truck wrote:
| Same! The only reason I'm not using it more is because it's
| summertime. I'm not in any hurry.
|
| Setting the memory to "fast timings" is good for 8-12% more
| tokens/second if you haven't tried yet. I miss the slightly
| older days of AMD when powerplay tables were unlocked and we
| could configure the timings and voltages manually, there's
| another 30% being left on the table ez
| MrDrMcCoy wrote:
| What do you mean by 'Setting the memory to "fast timings"'?
| The only runtime I can get working for my GPUs is
| llama.cpp, which I haven't seen anything like that in its
| argument set. My perusal of the options for vllm and sglang
| didn't suggest anything similar either before failing
| miserably.
| doginasuit wrote:
| Another potential takeaway is that the models all gathering
| around the same point supports the idea that there is a ceiling
| to LLM capability.
| AustinDev wrote:
| They always all gather around the same spot then that spot
| moves every 6-9 months. I think the clustering is more likely
| evidence of distillation. I don't personally think
| distillation is a bad thing. If the LLM providers can distill
| all of human output into their models for 'free'. I don't
| think distilling a model from the output of those models is
| morally wrong.
| michelsedgh wrote:
| What an interesting take. One question, do you think
| stealing from a thief is morally okay? I'm just asking no
| judgement on my side.
| AussieWog93 wrote:
| I'd say it's more "Downloading LimeWire Pro from
| LimeWire" than actual theft.
| dullcrisp wrote:
| Why isn't it more like building a hardware store using
| lumber you purchased from a competing hardware store? Or
| founding a school using an education you obtained at a
| different school?
| ygjb wrote:
| Because that doesn't satisfy the narrative of American
| exceptionalism. It's easier to point at something and say
| it was stolen or copied than it is to compete, especially
| with the political climate in the US.
|
| This isn't an anti-American sentiment. It is an anti-
| corporate/regulatory capture/embrace and extinguish
| sentiment (which probably reads the same to many people
| these days).
| Gigachad wrote:
| If the legal system declares the first thief's theft not
| theft then all bets are off.
| BeetleB wrote:
| > If the legal system declares the first thief's theft
| not theft
|
| But they didn't find it. The Big LLM provider accepted
| guilt and paid a fine.
|
| You can argue whether it was a fair amount they paid, but
| there is no legal precedent that was set. It's still
| considered theft.
| kennywinker wrote:
| As i understand it, they accepted guilt for downloading
| stuff illegally. They didn't accept guilt for
| incorporating all of human output into their model
| without consent.
| TheOtherHobbes wrote:
| Copyright law only considers illegal ownership of a work,
| so the crime - or tort - was making/acquiring copies
| without permission or payment.
|
| Training from copies has been ruled fair use because it's
| "transformative" and not simply "derivative."
|
| This is obviously debatable, but that's where the debate
| is at the moment.
| michelsedgh wrote:
| So basically because they just browsed and used the
| information that was mostly public on the internet and
| they didnt copy it, they just learned from it and thats
| fine. Which makes sense. None of the llms let u copy
| someones work exactly anyways... makes total sense
| honestly. So in this case what happens to distilling? Is
| that also learning or ur trying to get to their actual
| weights by kind of reverse engineering it? Where would
| the argument fall there?
| kennywinker wrote:
| > but that's where the debate is at the moment.
|
| Because of the rulings of a couple of judges. Is that
| actually what the majority of _people_ think?
|
| > Copyright law only considers illegal ownership of a
| work
|
| That's definitely not true. File sharing, for example, is
| illegal even if you legally own the original copy you're
| sharing.
|
| Similarly, copyright has something to say if I read a
| legal copy of harry potter and then create a new work in
| that world.
| BeetleB wrote:
| > They didn't accept guilt for incorporating all of human
| output into their model without consent.
|
| Because that use case is actually permitted by law.
| kennywinker wrote:
| I mean... that's one interpretation of the law, sure.
|
| The law was written before the idea of an LLM existed,
| and some judges in some specific cases decided the
| previous law covered this usage.
|
| So, it comes down to if you believe a couple judges
| ruling on a couple cases is the right way to determine a
| world-altering new legal framework.
| nolok wrote:
| > But they didn't find it. The Big LLM provider accepted
| guilt and paid a fine.
|
| That's not how it works. You have to give it back.
|
| Otherwise, the distiller can just pay a fine (no larger
| than the original did) and be okay then, right ?
| throwaway27448 wrote:
| True. It's the courts that failed humanity. Or perhaps
| the shits that invented copyright to start
| mannanj wrote:
| Is it theft if another thief steal's the first thief's
| theft?
| itemize123 wrote:
| question's phrasing made your judgement obvious
| ofjcihen wrote:
| I gathered that the most recent advances haven't been in
| capabilities of the model but more the way that it's able
| to be employed (most recently agents).
| conception wrote:
| If you talk to the Chinese models, even super smart Qwen
| 3.8, you can tell they are distilled just from the verbal
| ticks they have. Gemini, ChatGPT and Claude do not sound
| alike. The Chinese models 100% sound like one of the 3,
| usually Claude. American models are load bearing for this
| LLM generation seam.
| kayson wrote:
| - load bearing -
| AustinDev wrote:
| seams, boundaries, envelopes, etc
| FuckButtons wrote:
| That's definitely my impression of deepseek 0731 after a
| fair bit of use via ds4, it sounds like Claude.
| bossyTeacher wrote:
| > LLM providers can distill all of human output into their
| models for 'free'
|
| Not sure what part of being charged guilty and paying a
| fine you see as "free".
| rlupi wrote:
| I think it's more likely to be the effect of
| synchronization of launches, and the fact that models that
| do not challenge SOTA in some way do not get launched
| (think Gemini Pro delays), launched quietly or do not get
| any attention.
| rllearneratwork wrote:
| the ceiling is to eval's quality
| sscaryterry wrote:
| Don't say that too loud, you may burst the bubble
| prematurely.
| CMay wrote:
| Or are humans more of a bottleneck than before, because to
| improve on the most complex problems that demonstrate
| intelligence you need some way to verify that they are
| correct. If it's hard for humans to even know if something is
| correct, wouldn't that slow everything down and simply put
| limits on the scaling speed of models based on human
| verification?
|
| So instead of relying heavily on human bottlenecks, you focus
| on agentic task verification since that's the low hanging
| fruit and verifiable at scale?
| edg5000 wrote:
| Very interesting point you make! Before LLMs I had a theory
| that we cannot make something more intelligent/complex than
| us.
|
| LLMs are certainly more knowledgable, but maybe not more
| intelligent, arguably. It's possible we're approacing a
| ceiling indeed.
|
| Model capability might be on an asymptote appraching but
| never quite reaching parity with human intelligence.
| miki123211 wrote:
| I don't think there's a ceiling to LLM capability. I do think
| that many software dev tasks are just far below that ceiling,
| and the gains to most dev work won't be that large from now
| on.
|
| Where the new generation of LLMs (Fable, Sol) shines is tasks
| that are much harder than typical soft eng, yet that still
| have a verifiable answer, think mathematical proofs or
| exploits. I think there's still a good amount of low-hanging
| fruit in those (and similar) areas.
|
| The next frontier after that is tasks that don't have
| automatically-verifiable answers, and may not even have
| correct and incorrect ones in the strictest sense of the
| word.
|
| Reasonable lawyers might disagree on the question of "which
| trial strategy do I use given the following set of facts."
| There are answers that are clearly wrong, but being able to
| choose between many plausibly-correct ones requires many
| years of lawyering and seeing many trials play out. I do
| suspect that most lawyers are far below the ceiling that a
| hypothetical immortal lawyer that has practiced for an
| infinite amount of time would have achieved.
| monster_truck wrote:
| Something I don't think many have internalized is that China
| has been as good or better for quite a while now (long before
| anyone was pointing distillation fingers) and enough people
| have finally tried it for themselves that the understanding has
| reached critical mass and the careful narrative of american
| companies is collapsing.
|
| When I finally put $15 into Deepseek and it beat the brakes off
| Codex 5.5 on multiple rather complex projects without any of
| the obnoxious mistakes, I was sick to my stomach with buyers
| remorse. I couldn't believe I ever felt like I was getting my
| moneys worth at $200/mo. I wouldn't even use OAI's models if
| they were free and unlimited at this point, I'll happily pay
| for what I already know works. No reset bingo, no cache errors,
| no annoying shitposters as a primary source of info. Oh, and I
| still had $10 of tokens left
|
| And yes, 3.6 is excellent locally. The rest of this year is
| gonna be awesome
| aliasxneo wrote:
| And I had the opposite experience. It's a really interesting
| phenomenon that I can't really explain. My co-founder swears
| by Deepseek and yet just the other day we were conversing and
| he was telling me about some of the issues with the way the
| AI was behaving and trying to show off the cool workarounds
| he came up with to limit it. I was like, "Interesting, yeah,
| I've literally never had that problem."
|
| I suspect that the models are genuinely close and that
| certain experiences get felt across providers but are
| inconsistent enough to convince people one is superior to the
| other. I for one have tried Deepseek on and off since my co-
| founder is fond of it and I've stopped trying now because I
| never have a good experience.
| wanderlust123 wrote:
| That's not surprising. I have been using Deepseek and it
| consistently produces excellent output given the right
| context howevwr. It depends on the task as it does have
| blindspots.
| surgical_fire wrote:
| That may be true.
|
| I switched to DeepSeek entirely once I decided to put 10
| bucks on it and I realized that it could do whatever I was
| throwing at Claude or ChatGPT prior to that.
|
| I recommended it to one of my friends, and he was surprised
| DeepSeek could solve task that Claude got stuck at. I was
| surprised at it too.
|
| I know others that tried and were less impressed too.
| cyanydeez wrote:
| I think americans assume when they see a chinese or asian
| person working at an american business that they "escaped"
| china as opposed to just being rich enough to go to school
| abroad. and has little to no bearing on the amount of
| intelligent going around.
| pessimizer wrote:
| They've been continuously programmed with insane beliefs
| about China, which is less shocking when you understand
| what insane beliefs that they've had programmed into them
| about their neighbors. The world and your neighborhood are
| full of evil communists and Nazis who are trying to kill
| you all the time.
| swat535 wrote:
| I don't think it's all Americans however there is a
| portion of them who are not able to grasp the world
| outside of their borders.
|
| I think it's mainly due to poor education many receive
| and a very controlled media that suppresses information.
|
| It's shocking considering how much money they spend on
| education compared to other nations.
| TacticalCoder wrote:
| > The world and your neighborhood are full of evil
| communists and Nazis who are trying to kill you all the
| time.
|
| Don't know about that but your neighbors in Iran in early
| january happened to be "nice people" who just followed
| the orders to slaughter 30 000 unarmed civilians.
|
| We could talk about the, what 600 000 deaths, including
| many civilians, in the Ukraine/Russia war.
|
| Or we could talk about the number of nice palestinians
| killed since the beginning of the war in Gaza. Or we
| could go a bit further and talk about the joy and
| celebration in Gaza after their heroes brought back 200
| hostages after having slaughtered 1200 civilians.
|
| You may be living in a place that you think shields you
| from those but I know the ideologies behind these acts.
|
| The fallacy of gray is just that: it's not true that
| there's always a nice middle ground and that there's no
| evil ideology out there.
|
| Something something about the price of liberty being
| eternal vigilance. For there _are_ people abusing your
| blind trust.
| peterashford wrote:
| QED
| 14u2c wrote:
| If Americans see an Asian person working at a US business,
| they will assume that person is an American. They may even
| ask what state you are from. It's honestly one of the nice
| things about the place, you can belong even if you are not
| from there.
| kadushka wrote:
| Opus 5 and 5.6 Sol are definitely not smart enough to do my
| job. They require constant supervision. So why would I want
| to switch to even worse model? Even if it's just slightly
| worse?
| Cookingboy wrote:
| >So why would I want to switch to even worse model?
|
| There would be no reason to if you are in the privileged
| position where cost isn't an issue.
|
| For the rest of us something that's 95% as good for 20% the
| price is a hell of a value proposition.
| esafak wrote:
| That's why every benchmark should show the Pareto
| frontier against cost and latency.
| FuckButtons wrote:
| Pareto optimal dominant vs a human for the same task, not
| an unreasonable framing but that assumes that it can
| actually do the task, which the op was arguing it
| couldn't at all. Which, I suppose you could model as the
| utility of task completion % as being non linear. I have
| heard many people argue that the nature of work is messy
| and complicated and many things they do could not easily
| be emulated or automated. I do wonder how many of those
| activities are actually something that are connected to a
| companies ability to generate revenue or are just the
| messy interactions between people.
| kadushka wrote:
| Cost is absolutely an issue here - my time is worth
| approximately $1000/day, so if a slightly worse model
| wastes one more hour of my time a day than the best
| model, it costs the company >$2k/mo. Fortunately my
| employer understands this well and encourages me to use
| the best models as much as I can.
| surgical_fire wrote:
| This reply must have cost dozens of dollars.
| matheusmoreira wrote:
| > So why would I want to switch to even worse model? Even
| if it's just slightly worse?
|
| Self-hosting is the biggest reason.
| Systemerror7A69 wrote:
| If they already require your constant supervision the
| reason is money.
| surgical_fire wrote:
| > They require constant supervision
|
| I think that may be part of it.
|
| LLMs can be autonomous to an extent. All of them need
| steering - which is why I feel they are more a superpower
| the more I am an expert on the subject matter.
|
| The more you want it to be autonomous, than yeah, you may
| benefit from using the very best the industry has to offer,
| however slightly better it is.
|
| But if you are always in the loop anyway, you may want to
| try DeepSeek. You will get similar results for a fraction
| of the price.
| 4d4m wrote:
| +1 most people are too afraid to try something new. They've
| been roughly on par with their frontier offerings for 6
| months if not more.
| dw_arthur wrote:
| Who is going to break it to the Americans that China is more
| than a slight favorite to win an existential battle over which
| country is better at math?
| londons_explore wrote:
| I just don't think you can combine speed, latency, price and
| intelligence into a single useful metric.
|
| Clearly the weighting of those things depends on the usecase
| dangoodmanUT wrote:
| I'm seeing opus 59.2, Qwen 58.4?
| zmmmmm wrote:
| The fact that the Chinese models have caught up on benchmarks
| suggests to me that its likely we will start to transition now
| into much more of a brand war. It will be subjective qualities
| that drive our decisions more than measures of absolute
| intelligence. Already I am choosing models more because I like
| the personality or style of what they do than because I think
| they have the absolute highest chance of outputting the most
| technically correct answer to any given prompt. It will be very
| interesting to see how things evolve in this direction.
| mindwok wrote:
| For me now it's simply cost and speed. With GPT5.6 and Fable
| (and respective open models since then) we passed a threshold
| where intelligence is sufficient. Now I just need speed of
| iteration and good prices.
| gerdesj wrote:
| My vague equivalent of the pelican riding a bicycle test (for a
| local model without internets) is to ask it: "Where is Yeovil"? I
| don't expect a totally accurate answer for obvious reasons but I
| do enjoy watching the accuracy improve.
|
| Qwen3.6-27B-FP8 currently espouses (see below), which is not too
| bad. The directions are a bit mad but the mileage is about right
| and there is a helicopter manufacturer here and a RNAS (navy not
| airforce) museum nearby at Yeovilton. Cosford is in Shropshire
| which is not a million miles away.
|
| I'm not sure what Pen Di De means but the river Yeo is correct
| ... OK ... "basin like" - again not bad, even if Chinese is not
| the first language here. The model understands that Yeovil is
| named after (or vice versa or at least is associated with) a
| river
|
| Yeovil is the current form of Gifle (Saxon) which I thought meant
| "bend in a river" but WP is currently saying "fork in a river".
| My source is a local museum. There is a fork but was it there
| 2000 odd years ago? My hydrology skills say ... possibly
|
| ---------------------------------------------------- Q: where is
| yeovil:
|
| Yeovil is a town in Somerset, in the South West of England.
|
| It is located roughly: 25 miles (40 km) south-
| west of Exeter 60 miles (100 km) west of Bristol
| 140 miles (225 km) west-south-west of London
|
| Yeovil is known for its historic market town center, RAF Museum
| Cosford (nearby), and as a significant industrial town,
| particularly during World War II for aircraft manufacturing
| (including the Wellington bomber). It sits in thePen Di De
| valley of the River Yeo.
| ngl999 wrote:
| It's censored and it'll spread certain kind of narrative all over
| the world.
| MrDrMcCoy wrote:
| If you're using AI for narratives, you're using it wrong.
| imagetic wrote:
| It's the first model I've used that makes me forget it isn't one
| of the big frontier players after the first prompt. So far I'm
| impressed.
___________________________________________________________________
(page generated 2026-08-07 11:01 UTC)