[HN Gopher] GPT-5.4 Mini and Nano
       ___________________________________________________________________
        
       GPT-5.4 Mini and Nano
        
       Author : meetpateltech
       Score  : 198 points
       Date   : 2026-03-17 17:07 UTC (5 hours ago)
        
 (HTM) web link (openai.com)
 (TXT) w3m dump (openai.com)
        
       | machinecontrol wrote:
       | What's the practical advantage of using a mini or nano model
       | versus the standard GPT model?
        
         | aavci wrote:
         | Cheaper. Every month or so I visit the models used and check
         | whether they can be replaced by the cheapest and smallest model
         | possible for the same task. Some people do fine tuning to
         | achieve this too.
        
       | powera wrote:
       | I've been waiting for this update.
       | 
       | For many "simple" LLM tasks, GPT-5-mini was sufficient 99% of the
       | time. Hopefully these models will do even more and closer to 100%
       | accuracy.
       | 
       | The prices are up 2-4x compared to GPT-5-mini and nano. Were
       | those models just loss leaders, or are these substantially
       | larger/better?
        
         | HugoDias wrote:
         | For us, it was also pretty good, but the performance decreased
         | recently, that forced us to migrate to haiku-4.5. More
         | expensive but much more reliable (when anthropic up, of
         | course).
        
           | throwaway911282 wrote:
           | they dont change the model weights (no frontier lab does). if
           | you have evals and all prompts, tool calls the same, I'm
           | curious how you are saying performance decreased..
        
         | powera wrote:
         | So far on my (simple) benchmarks, GPT-5.4-mini is looking very
         | good. GPT-5.4-mini is about 30% faster than GPT-5-mini.
         | GPT-5.4-mini gets 80% on the "how many Rs in Strawberry" test,
         | and nearly perfect scores on everything else I threw at it.
         | 
         | GPT-5.4-nano is less impressive. I would stick to gpt-5.4-mini
         | where precise data is a requirement. But it is fast, and
         | probably cheaper and better quality than an 8-20B parameter
         | local model would be.
         | 
         | ( https://encyclopedia.foundation/benchmarks/dashboard/ for
         | details - the data is moderately blurry - some outlier (15s)
         | calls are included, a few benchmark questions are ambiguous,
         | and some prices shown are very rough estimates ).
        
       | HugoDias wrote:
       | According to their benchmarks, GPT 5.4 Nano > GPT-5-mini in most
       | areas, but I'm noticing models are getting more expensive and not
       | actually getting cheaper?
       | 
       | GPT 5 mini: Input $0.25 / Output $2.00
       | 
       | GPT 5 nano: Input: $0.05 / Output $0.40
       | 
       | GPT 5.4 mini: Input $0.75 / Output $4.50
       | 
       | GPT 5.4 nano: Input $0.20 / Output $1.25
        
         | simianwords wrote:
         | models are getting costlier but by performance getting cheaper.
         | perhaps they don't see a point supporting really low
         | performance models?
        
           | HugoDias wrote:
           | I would be curious to know if from the enterprise / API
           | consumption perspective, these low-performance models aren't
           | the most used ones. At least it matches our current scenario
           | when it comes to tokens in / tokens out. I'd totally buy the
           | price increase if these are becoming more efficient though,
           | consuming less tokens.
        
         | karmasimida wrote:
         | Those are bigger models. The serving isn't going to be cheaper.
         | 
         | Why expect cheaper then? The performance is also better
        
           | trvz wrote:
           | You seem to have insight into the size of OpenAI's models.
           | 
           | Care to share the parameter counts for them?
        
       | ryao wrote:
       | I will be impressed when they release the weights for these and
       | older models as open source. Until then, this is not that
       | interesting.
        
       | simianwords wrote:
       | why isn't nano available in codex? could be used for ingesting
       | huge amount of logs and other such things
        
         | patates wrote:
         | IMHO the best way is to let a SOTA model have a look at bunch
         | of random samples and write you tools to analyze those.
         | 
         | I think, no model, SOTA or not, has neither the context nor the
         | attention to be able to do anything meaningful with huge amount
         | of logs.
        
       | BoumTAC wrote:
       | To me, mini releases matter much more and better reflect the real
       | progress than SOTA models.
       | 
       | The frontier models have become so good that it's getting almost
       | impossible to notice meaningful differences between them.
       | 
       | Meanwhile, when a smaller / less powerful model releases a new
       | version, the jump in quality is often massive, to the point where
       | we can now use them 100% of the time in many cases.
       | 
       | And since they're also getting dramatically cheaper, it's
       | becoming increasingly compelling to actually run these models in
       | real-life applications.
        
         | pzo wrote:
         | they do are cheaper than SOTA but not getting dramatically
         | cheaper but actually the opposite - GPT 5.4 mini is around ~3x
         | more expensive than GPT 5.0 mini.
         | 
         | Similarly gemini 3.1 flash lite got more expensive than gemini
         | 2.5 flash lite.
        
           | BoumTAC wrote:
           | But they are getting dramatically better.
           | 
           | What's the point of a crazy cheap model if it's shit ?
           | 
           | I code most of the time with haiku 4.5 because it's so good.
           | It's cheaper for me than buying a 23EUR subscription from
           | Anthropic.
        
             | philipkglass wrote:
             | The crazy cheap models may be adequate for a task, and low
             | cost matters with volume. I need to label millions of
             | images to determine if they're sexually suggestive (this
             | includes but is not limited to nudity). The Gemini 2.0
             | Flash Lite model is inexpensive and performs well. Gemini
             | 2.5 Flash Lite is also good, but not noticeably better, and
             | it costs more. When 2.0 gets retired this June my costs are
             | going up.
        
         | brikym wrote:
         | If you're doing something common then maybe there are no
         | differences with SOTA. But I've noticed a few. GPT 5.4 isn't as
         | good at UI work in svelte. Gemini tends to go off and implement
         | stuff even if I prompt it to discuss but it's pretty good at UI
         | code. Claude tends to find out less about my code base than GPT
         | and it abuses the any type in typescript.
        
           | patates wrote:
           | Big part of these differences may be the system prompts
           | and/or the harness.
        
         | sebastiennight wrote:
         | > 100% of the time in many cases
         | 
         | So, every single time, the new model works most of the time?
        
         | zozbot234 wrote:
         | > And since they're also getting dramatically cheaper, it's
         | becoming increasingly compelling to actually run these models
         | in real-life applications.
         | 
         | They're not really cheaper than the SOTA open models on third-
         | party inference platforms, and they're generally dumber. I
         | suppose they're still worth it if you must minimize latency for
         | any given level of smarts, but not really otherwise.
        
         | XCSme wrote:
         | Well, in that case, the difference is quite minimal between 5
         | mini and 5.4 mini
         | 
         | 5.4 mini seems to be a lot more wild/unstable, but with this
         | instability it gets the right answer more often.
         | 
         | https://aibenchy.com/compare/openai-gpt-5-4-mini-medium/open...
        
       | cbg0 wrote:
       | Based on the SWE-Bench it seems like 5.4 mini high is ~= GPT 5.4
       | low in terms of accuracy and price but the latency for mini is
       | considerably higher at 254 seconds vs 171 seconds for GPT5.4.
       | Probably a good option to run at lower effort levels to keep
       | costs down for simpler tasks. Long context performance is also
       | not great.
        
       | beklein wrote:
       | As a big Codex user, with many smaller requests, this one is the
       | highlight: "In Codex, GPT-5.4 mini is available across the Codex
       | app, CLI, IDE extension and web. It uses only 30% of the GPT-5.4
       | quota, letting developers quickly handle simpler coding tasks in
       | Codex for about one-third the cost." + Subagents support will be
       | huge.
        
         | hyperbovine wrote:
         | Having to invoke `/model` according to my perceived complexity
         | of the request is a bit of a deal breaker though.
        
           | serf wrote:
           | you use profiles for that [0], or in the case of a more
           | capable tool (like opencode) they're more confusing referred
           | to as 'agents'[1] , which may or may not coordinate
           | subagents..
           | 
           | So, in opencode you'd make a "PR Meister" and "King of Git
           | Commits" agent that was forced to use 5.4mini or whatever,
           | and whenever it fell down to using that agent it'd do so
           | through the preferred model.
           | 
           | For example, I use the spark models to orchestrate abunch of
           | sub-agents that may or may not use larger models, thus I get
           | sub-agents and concurrency spun up _very fast_ in places
           | where domain depth matter less.
           | 
           | [0]: https://developers.openai.com/codex/config-
           | advanced#profiles [1]: https://opencode.ai/docs/agents/
        
       | miltonlost wrote:
       | Does it still help drive people to psychosis and murder and
       | suicide? Where's the benchmark for that?
        
       | system2 wrote:
       | I am feeling the version fatigue. I cannot deal with their
       | incremental bs versions.
        
       | yomismoaqui wrote:
       | Not comparing with equivalent models from Anthropic or Google,
       | interesting...
        
         | Tiberium wrote:
         | They did actually compare them in the tweet, see
         | https://x.com/OpenAI/status/2033953592424731072
         | 
         | Direct image:
         | https://pbs.twimg.com/media/HDoN4PhasAAinj_?format=png&name=...
        
       | casey2 wrote:
       | I googled all the testimonial names and they are all linked-in
       | mouthpieces.
        
       | Tiberium wrote:
       | I checked the current speed over the API, and so far I'm very
       | impressed. Of course models are usually not as loaded on the
       | release day, but right now:
       | 
       | - Older GPT-5 Mini is about 55-60 tokens/s on API normally,
       | 115-120 t/s when used with service_tier="priority" (2x cost).
       | 
       | - GPT-5.4 Mini averages about 180-190 t/s on API. Priority does
       | nothing for it currently.
       | 
       | - GPT-5.4 Nano is at about 200 t/s.
       | 
       | To put this into perspective, Gemini 3 Flash is about 130 t/s on
       | Gemini API and about 120 t/s on Vertex.
       | 
       | This is raw tokens/s for all models, it doesn't exclude reasoning
       | tokens, but I ran models with none/minimal effort where
       | supported.
       | 
       | And quick price comparisons:
       | 
       | - Claude: Opus 4.6 is $5/$25, Sonnet 4.6 is $3/$15, Haiku 4.5 is
       | $1/$5
       | 
       | - GPT: 5.4 is $2.5/$15 ($5/$22.5 for >200K context), 5.4 Mini is
       | $0.75/$4.5, 5.4 Nano is $0.2/$1.25
       | 
       | - Gemini: 3.1 Pro is $2/$12 ($3/$18 for >200K context), 3 Flash
       | is $0.5/$3, 3.1 Flash Lite is $0.25/$1.5
        
         | coder543 wrote:
         | I wish someone would also thoroughly measure prompt processing
         | speeds across the major providers too. Output speeds are useful
         | too, but more commonly measured.
        
           | JLO64 wrote:
           | In my use case for small models I typically only generate a
           | max of 100 tokens per API call, with the prompt processing
           | taking up the majority of the wait time from the user
           | perspective. I found OAI's models to be quite poor at this
           | and made the switch to Anthropic's API just for this.
           | 
           | I've found Haiku to be a pretty fast at PP, but would be
           | willing to investigate using another provider if they offer
           | faster speeds.
        
           | asselinpaul wrote:
           | OpenRouter has this information
        
             | coder543 wrote:
             | I do not see prompt processing, only some kind of nebulous
             | "throughput" that could be output or input+output, but
             | definitely not input only.
        
         | rattray wrote:
         | Wow. How fast is haiku?
        
         | daniel_iversen wrote:
         | Curious to hear why people pick GPT and Claude over Google
         | (when sometimes you'd think they have a natural advantage on
         | costs, resources and business model etc)?
        
           | coderjames wrote:
           | In my workplace, its availability. We have to use US-only
           | models for government-compliance reasons, so we have access
           | to Opus 4.6 and GPT 5.4, but only Gemini 2.5 which isn't in
           | the same class as the first two.
        
         | rglynn wrote:
         | IME tok/s is only useful with the additional context of ttft
         | and total latency. At this point a given closed-model does not
         | exist in a vaccuum but rather in a wider architecture that
         | affects the actual performance profile for an API consumer.
         | 
         | This isn't usually an issue comparing models within the same
         | provider, but it does mean cross-provider comparison using only
         | tok/s is not apples-to-apples in terms of real-world
         | performance.
        
           | Rapzid wrote:
           | Exactly. Really frustrating they don't advertise TTFT and
           | etc, and that it's really hard to find any info in that
           | regard on newer models.
           | 
           | For voice agents gpt-4.1 and gpt-4.1-mini seem to be the best
           | low latency models when you need to handle bigger data more
           | complex asks.
           | 
           | But they are a year old and trying to figure out if these new
           | models(instant, chat, realtime, mini, nona, wtf) are a good
           | upgrade is very frustrating. AFAICT they aren't; the TTFT
           | latencies are too high.
        
         | msp26 wrote:
         | Man the lowest end pricing has been thoroughly hiked. It was
         | convenient while it lasted.
        
       | 6thbit wrote:
       | Looking at the long context benchmark results for these, sounds
       | like they are best fit for also mini-sized context windows.
       | 
       | Is there any harness with an easy way to pick a model for a
       | subagent based on the required context size the subagent may
       | need?
        
       | bananamogul wrote:
       | They could call them something like "sonnet" and "haiki" maybe.
        
       | reconnecting wrote:
       | All three ChatGPT models (Instant, Thinking, and Pro) have a new
       | knowledge cutoff of _August 2025_.
       | 
       | Seriously?
        
         | zild3d wrote:
         | whats surprising about that? most of the minor version updates
         | from all the labs are post training updates / not changing
         | knowledge cutoff
        
           | reconnecting wrote:
           | Thanks for letting me know, I will be waiting for the major
           | update.
        
             | F7F7F7 wrote:
             | It's been like this since GPT 3.5. This is not a limitation
             | and is generally considered a natural outcome of the
             | process.
             | 
             | So there's no major update in the sense that you might be
             | thinking. Most of the time there's not even an announcement
             | when/if training cut offs are updated. It's just another
             | byline.
             | 
             | A 6 month lag seems to be the standard across the frontier
             | models.
        
               | reconnecting wrote:
               | I've actually started worrying that the amount of false
               | data produced with LLMs on the public internet might
               | provoke a situation where the knowledge cutoff becomes
               | permanently (and silently) frozen. Like we can't trust
               | data after 2025 because it will poison training data at
               | scale, and models will only cover major events without
               | capturing the finer details.
        
               | gwern wrote:
               | I agree. That's why you should write as much as you can
               | now, if you want to get it into the LLMs
               | (https://gwern.net/blog/2024/writing-online). You never
               | know when the window will slam shut and LLM training goes
               | 'hermetic' as they focus on 'civilization in a
               | datacenter' where only extremely vetted whitelisted data
               | gets included in the 'seed' and everything is
               | reconstructed from scratch for the training value &
               | safety.
        
         | dpoloncsak wrote:
         | Do you find the results vary based on whether it uses RAG to
         | hit the internet vs the data being in the weights itself? I'm
         | not sure I've really noticed a difference, but I don't often
         | prompt about current events or anything.
        
           | reconnecting wrote:
           | I noticed that many recent technologies are not familiar to
           | LLMs because of the knowledge cutoff, and thus might not
           | appear in recommendations even if they better match the
           | request.
        
             | dpoloncsak wrote:
             | Oh thats a good point, yeah.
             | 
             | If I told it I'm shopping for a budget-level Mac, it may
             | not recommend the Neo. I'm sure software only moves faster,
             | too. Especially as more code is 'written' blindly, new
             | stacks may never see adoption
        
       | varispeed wrote:
       | I stopped paying attention to GPT-5.x releases, they seem to have
       | been severely dumbed down.
        
       | pscanf wrote:
       | I quite like the GPT models when chatting with them (in fact,
       | they're probably my favorites), but for agentic work I only had
       | bad experiences with them.
       | 
       | They're incredibly slow (via official API or openrouter), but
       | most of all they seem not to understand the instructions that I
       | give them. I'm sure I'm _holding them wrong_, in the sense that
       | I'm not tailoring my prompt for them, but most other models don't
       | have problem with the exact same prompt.
       | 
       | Does anybody else have a similar experience?
        
         | nikanj wrote:
         | Same, and I can't put my finger on the "why" either. Plus I
         | keep hitting guard rails for the strangest reasons, like
         | telling codex "Add code signing to this build pipeline, use the
         | pipeline at ~/myotherproject as reference" and codex tells me
         | "You should not copy other people's code signing keys, I can't
         | help you with this"
        
         | tom1337 wrote:
         | Yea absolutely. I am using GPT 5.2 / 5.2 Codex with OpenCode
         | and it just doesn't get what I am doing or looses context.
         | Claude on the other side (via GitHub Copilot) has no problem
         | and also discovers the repository on it's own in new sessions
         | while I need to basically spoonfeed GPT. I also agree on the
         | speed. Earlier today I tasked GPT 5.2 Codex with a small
         | refactor of a task in our codebase with reasoning to high and
         | it took 20 minutes to move around 20 files.
        
           | furyofantares wrote:
           | I don't know any reason to use 5.2, when 5.3 is quite a bit
           | faster.
        
           | spiderfarmer wrote:
           | If using OpenAI models, use the Codex desktop app, it runs
           | circles around OpenCode.
        
             | qaz_plm wrote:
             | Can you educate me as to what makes Codex app superior
             | using the same GPT model in both? Thx in advance!
        
         | renewiltord wrote:
         | Are you requesting reasoning via param? That was a mistake I
         | was making. However with highest reasoning level I would
         | frequently encounter cyber security violation when using agent
         | that self-modifies.
         | 
         | I prefer Claude models as well or open models for this reason
         | except that Codex subscription gets pretty hefty token space.
        
           | birdsongs wrote:
           | > cyber security violation
           | 
           | Would you mind expanding on this? Do you mean in the
           | resulting code? Or a security problem on your local machine?
           | 
           | I naively use models via our Copilot subscription for small
           | coding tasks, but haven't gone too deep. So this kind of
           | threat model is new to me.
        
             | renewiltord wrote:
             | No, I mean literal API response. They think I'm using it to
             | hack. See related Github issue:
             | https://github.com/anomalyco/opencode/issues/15776
             | 
             | I don't use OpenCode but looks like it also triggered
             | similar use. My message was similar but different.
        
               | birdsongs wrote:
               | Ahhh okay, I see. Thanks!
        
           | pscanf wrote:
           | Yes, I think? But I was talking more specifically about using
           | the models via API in agents I develop, not for agentic
           | coding. Though, thinking about it, I also don't click with
           | the GPT models when I use them for coding (using Codex). They
           | just seem "off" compared to Claude.
        
             | renewiltord wrote:
             | I am also talking about agents I'm developing. They just
             | happen to be self-modifying but they're not _for_ agentic
             | coding. You have to explicitly send the reasoning effort
             | parameter. If you set effort to None (default for gpt-5.4)
             | you get very low intelligence.
        
               | pscanf wrote:
               | Ah OK sorry, I misinterpreted. But yes, I double checked
               | one case and I am indeed setting the parameter explicitly
               | (defaulting to medium effort). But no luck. It feels like
               | the model ignores what I'm telling it.
               | 
               | For example, I pass it a list of database collections and
               | tools to search through them, ask a question that can
               | very obviously be answered with them, and it responds
               | with "I can't tell yet from your current records" (just
               | tested with GPT 5.4-mini).
               | 
               | But I've prodded it a bit more now, and maybe the model
               | doesn't want to answer unless it can be very very
               | confident of the answer it produces. So it's sort of a
               | "soft refusal".
        
             | jorl17 wrote:
             | I like GPT models in Codex, for a fully vibecoded
             | experience (I don't look at code) for my side-projects. In
             | there, they really get the job done: you plan, they say
             | what they'll do, and it shows up done. It's rare I need to
             | push back and point out bugs. I really can't fault them for
             | this very specific use-case.
             | 
             | For anything else, I can't stand them, and it genuinely
             | feels like I am interacting with different models outside
             | of codex:
             | 
             | - They act like terribly arrogant agents. It's just in the
             | way they talk: self-assured, assertive. They don't say they
             | think something, they say it is so. They don't really
             | propose something, they say they're going to do it because
             | it's right.
             | 
             | - If you counter them, their thinking traces are filled
             | with what is virtually identical to: "I must control myself
             | and speak plainly, this human is out of his fucking mind"
             | 
             | - They are slow. Measurably slow. Sonnet is so much faster.
             | With Sonnet models, I can read every token as it comes, but
             | it takes some focusing. With GPT, I can read the whole
             | trace in real-time without any effort. It genuinely gives
             | off this "dumb machine that can't follow me" vibe.
             | 
             | - Paradoxically, even though they are so full of
             | themselves, they insist upon checking things which are
             | obvious. They will say "The fix is to move this bit of code
             | over there [it isn't]" and then immediately start looking
             | at sort of random files to check...what exactly?
             | 
             | - I feel they make perhaps as many mistakes as Sonnet, but
             | they are much less predictable mistakes. The kind that
             | leaves me baffled. This doesn't have to be bad for code
             | quality: Sonnet makes mistakes which _might_ at points even
             | be _harder_ to catch, so might be easier to let slip by.
             | Yet, it just imprints this feeling of distrust in the model
             | which is counter-productive to make me want to come back to
             | it
             | 
             | I didn't compare either with Gemini because Gemini is a
             | joke that "does", and never says what it is "doing", except
             | when it does so by leaving thinking traces in the middle of
             | python code comments. Love my codebase to have "But wait,
             | ..." in the middle of it. A useless model.
             | 
             | I've recently started saying this:
             | 
             | - Anthropic models feel like someone of that level of
             | intelligence thinking through problems and solving them.
             | Sonnet is not Opus -- it is sonnet-level intelligence, and
             | shows it. It approaches problems from a sensible,
             | reasonably predictable way.
             | 
             | - Gemini models feel like a cover for a bunch of inferior
             | developers all cluelessly throwing shit at the wall and
             | seeing what sticks -- yet, ultimately, they only show the
             | final decision. Almost like you're paying a fraudulent
             | agency that doesn't reveal its methods. The thinking is
             | nonsensical and all over the place, and it does eventually
             | achieve some of its goals, but you can't understand what
             | little it shows other than "Running command X" and "Doing
             | Y".
             | 
             | On a final note: when building agentic applications, I used
             | to prefer GPT (a year ago), but I can't stand it now.
             | Robotic, mechanic, constantly mis-using tools. I reach for
             | Sonnet/Opus if I want competence and adherence to prompt,
             | coupled with an impeccable use of tools. I reach for Gemini
             | (mostly flash models) if I want an acceptable experience at
             | a fraction of the price and latency.
        
               | baq wrote:
               | A bit off topic, but reading your post I suddenly
               | realized that if I read it three years ago I'd assume
               | you're either insane or joking. The world moved _fast_
               | looking back.
        
               | pscanf wrote:
               | > They act like terribly arrogant agents
               | 
               | Oh I feel that. I sometimes ask ChatGPT for "a review,
               | pull no punches" of something I'm writing, and my god,
               | the answers _really_ get on my nerves! (They do make some
               | useful points sometimes, though.)
               | 
               | > On a final note: when building agentic applications, I
               | used to prefer GPT (a year ago), but I can't stand it
               | now. Robotic, mechanic, constantly mis-using tools. I
               | reach for Sonnet/Opus if I want competence and adherence
               | to prompt, coupled with an impeccable use of tools. I
               | reach for Gemini (mostly flash models) if I want an
               | acceptable experience at a fraction of the price and
               | latency.
               | 
               | Yeah, this has been almost exactly my experience as well.
        
         | jauntywundrkind wrote:
         | I've had such the opposite experience, but mainly doing agentic
         | coding & little chat.
         | 
         | Codex is an ice man. Every other model will have a thinking
         | output that is meaningful and significant, that is walking
         | through its assumptions. Codex outputs only a very basic idea
         | of what it's thinking about, doesn't verbalize the problem or
         | it's constraints at all.
         | 
         | Codex also is by far the most sycophantic model. I am a capable
         | coder, have my charms, but every single direction change I
         | suggest, codex is all: "that's a great idea, and we should
         | totally go that [very different] direction", try as I might to
         | get it to act like more of a peer.
         | 
         | Opus I think does a better job of working with me to figure out
         | what to build, and understanding the problem more. But I find
         | it still has a propensity for making somewhat weird
         | suggestions. I can watch it talk itself into some weird ideas.
         | Which at least I can stop and alter! But I find its less
         | reliable at kicking out good technical work.
         | 
         | Codex is plenty fast in ChatGPT+. Speed is not the issue. I'm
         | also used to GLM speeds. Having parallel work open, keeping an
         | eye on multiple terminals is just a fact of life now; work
         | needs to optimize itself (organizationally) for parallel
         | workflows if it wants agentic productivity from us.
         | 
         | I have enormous respect for Codex, and think it (by signficiant
         | measure) has the best ability to code. In some ways I think
         | maybe some of the reason it's so good is because it's not
         | trying to convey complex dimensional exploration into a
         | understandable human thought sequence. But I resent how you
         | just have to let it work, before you have a chance to talk with
         | it and intervene. Even when discussing it is extremely
         | extremely terse, and I find I have to ask it again and again
         | and again to expand.
         | 
         | The one caveat i'll add, I've been dabbling elsewhere but
         | mainly i use OpenCode and it's prompt is pretty extensive and
         | may me part of why codex feels like an ice man to me.
         | https://github.com/anomalyco/opencode/blob/dev/packages/open...
        
           | pscanf wrote:
           | > I've had such the opposite experience
           | 
           | Yeah, I've actually heard many other people swear by the GPTs
           | / Codex. I wonder what factors make one "click" with a model
           | and not with another.
           | 
           | > Codex is an ice man.
           | 
           | That might be because OpenAI hides the actual reasoning
           | traces, showing just a summary (if I understood correctly).
        
             | kevinsync wrote:
             | OpenClaw guy (he's Austrian, it's relevant) much prefers
             | Codex over Claude and articulated it as being due to
             | Claude's output feeling very "American" and Codex's output
             | feeling very "German", and I personally really agree with
             | the sentiment.
             | 
             | As an American, Claude feels much more natural to me, with
             | the same overly-optimistic "move fast, break things" ethos
             | that permeates our culture. It takes bigger swings (and
             | misses) at harder-to-quantify concepts than Codex, cuts
             | corners (not intentionally, but it feels like a human who's
             | just moving too fast to see the forest for the trees in the
             | moment), etc. Codex on the other hand feels more grounded,
             | more prone to trying to aggregate blind spots, edge cases,
             | and cover the request more thoroughly than Claude. It's far
             | more pedantic and efficient, almost humorless. The dude
             | also claimed that most of the Codex team is European while
             | Claude team is American, and suggested that as an influence
             | on why this might be.
             | 
             | Anyways, I've found that if I force Claude and Codex to
             | talk to each other, I can get way better results and
             | consistency by using Claude to generate fairly good plans
             | from my detailed requests that it passes to Codex for
             | review and amendment, Claude incorporates the feedback and
             | implements the code, then Codex reviews the commit and
             | patches anything Claude misses. Best of both worlds. YMMV
        
               | pscanf wrote:
               | Oh, interesting perspective. I'm Italian, but from an
               | Alpine valley not far from Austria, so I don't know what
               | I should prefer. :D
               | 
               | But joking aside, putting it like that I'd think I'd
               | prefer the German/Codex way of doing things, yet I'm in
               | camp Claude. But I've always worked better with teammates
               | that balance my fastidiousness, so maybe that's my
               | answer.
        
             | aragonite wrote:
             | Claude Code now hides thinking as well unless you turn on
             | an undocumented setting:
             | 
             | https://github.com/anthropics/claude-
             | code/issues/31326#issue...
             | 
             | https://x.com/nummanali/status/2032451025500528687
        
         | thanhhaimai wrote:
         | Opinions are my own.
         | 
         | For agentic work, both Gemini 3.1 and Opus 4.6 passed the bar
         | for me. I do prefer Opus because my SIs are tuned for that, and
         | I don't want to rewrite them.
         | 
         | But ChatGPT models don't pass the bar. It seems to be trained
         | to be conversational and role-playing. It "acts" like an agent,
         | but it fails to keep the context to really complete the task.
         | It's a bit tiring to always have to double check its work /
         | results.
        
           | kraemahz wrote:
           | I find both Opus 4.6 and GPT-5.4 have weaknesses but tend to
           | support each other. Someone described it to me jokingly as
           | "Claude has ADHD and Codex is autistic." Claude is great at
           | doing something until it gets done and will run for hours on
           | a task without feedback, Codex is often the opposite: it will
           | ask for feedback often and sometimes just stop in the middle
           | of a task saying it's done with step 1 of 5. On the other
           | hand, Codex is a diligent reviewer and will find even subtle
           | bugs that Claude created in its big long-running "until its
           | done" work mode.
        
             | CamperBob2 wrote:
             | Seems like the diagnoses are backwards, in this case.
             | Claude usually stays on task no matter what, but lately
             | Opus 4.6 is showing signs of overuse. I never used to get
             | overload/internal server error messages, but I've seen
             | about a half-dozen of them today alone. And it has been
             | prone to blowing off subtasks that I'd have expected it to
             | resolve.
        
         | ilaksh wrote:
         | These little 5.4 ones are relatively low latency and fast which
         | is what I need for voice applications. But can't quite follow
         | instructions well enough for my task.
         | 
         | That's really the story of my life. Trying to find a smart
         | model with low latency.
         | 
         | Qwen 3.5 9b is almost smart enough and I assume I can run it on
         | a 5090 with very low latency. Almost. So I am thinking I will
         | fine tune it for my application a little.
        
         | hermit_dev wrote:
         | I ran 5.4 Pro on some data analytics (admittedly it was 300+
         | pages). It took forever. Ran the same on Sonnet 4.6, night and
         | day difference. I understand it's like using a V8 engine for a
         | V4 task, but I was curious. These new models look promising
         | though. I'd rather use something like a Haiku most of the time
         | over the best rated. I'm not a rocket scientist or solving the
         | mysteries of the universe. They seem to do a great job 80% of
         | the time.
        
       | dack wrote:
       | i want 5.4 nano to decide whether my prompt needs 5.4 xhigh and
       | route to it automatically
        
         | exitb wrote:
         | Like any work estimation, it will likely disappoint.
        
         | mrtesthah wrote:
         | As per OpenAI themselves, xhigh is only necessary if the agent
         | gets stuck on a long running task. Otherwise it's thinking
         | trades use so many tokens of context that it's less effective
         | than high for a great majority of tasks. This has also been my
         | experience.
        
       | kseniamorph wrote:
       | wow, not bad result on the computer use benchmark for the mini
       | model. for example, Claude Sonnet 4.6 shows 72.5%, almost on par
       | with GPT-5.4 mini (72.1%). but sonnet costs 4x more on input and
       | 3x more on output
        
       | fastpdfai wrote:
       | One thing I really want to find out, is which model and how to
       | process TONS of pdfs very very fast, and very accurate. For
       | prediction of invoice date, accrual accounting and other
       | accounting related purposes. So a decent smart model that is
       | really good at pdf and image reading. While still being very very
       | fast.
        
         | JLO64 wrote:
         | I have a use case somewhat similar to this where I need to
         | convert the content of PDFs in a non standard format to a
         | specific YAML format. I currently use Haiku for this and am
         | pleased with the accuracy/speed (I haven't tried scanned PDFs
         | yet tho) however I've been thinking about fine tuning a small
         | Qwen model for just this task. I can't yet justify the effort
         | to investigate it but I imagine it could work out.
        
       | mikkelam wrote:
       | Why are we treating LLM evaluation like a vibe check rather than
       | an engineering problem?
       | 
       | Most "Model X > Model Y" takes on HN these days (and everywhere)
       | seem based on an hour of unscientific manual prompting. Are we
       | actually running rigorous, version-controlled evals, or just
       | making architectural decisions based on whether a model nailed a
       | regex on the first try this morning?
        
         | tanaros wrote:
         | Whenever somebody makes a benchmark, people complain that the
         | benchmark results are meaningless because they're gamed. I
         | don't know why those same people don't understand that grading
         | on vibes is strictly worse.
        
           | tintor wrote:
           | Depends on benchmark.
           | 
           | If questions are fixed they are trivial to game.
        
         | pizza wrote:
         | There's a Dark Forest problem for evals. As soon as they're
         | made public they start running out of time to be useful. It's
         | also not clear how to predict how the model will perform on a
         | task based on an eval. Or even whether, given two skills that
         | the model can individually do well on in the evals, it still
         | does well on their composition. It might at this point be
         | better to be scientific in unscientific approaches, than to
         | attribute more power to relatively weakly predictive evals than
         | they actually have
        
           | xandrius wrote:
           | Is "Dark Forest problem" an actual name? I just heard of the
           | hypothesis and it has nothing to do with how you used it in
           | this context.
        
             | sebastiennight wrote:
             | I believe the correct term is "Goodhart's Law":
             | https://en.wikipedia.org/wiki/Goodhart%27s_law
        
             | pizza wrote:
             | I meant in the sense of - you have benchmarkers and
             | trainers. If you publicize your evaluation, trainers may
             | likely have their models 'consume' it, even if only
             | indirectly: another person creating their own benchmark
             | from scratch may be influenced by yours, even if the new
             | question sets are clean-room. That, and the rule of thumb
             | that benchmark value dissipates like sqrt(age) [0]
             | 
             | So there is a definite advantage to never publicizing your
             | internal benchmark. But then, no one else can replicate
             | your findings. You should assume that the space of
             | benchmarks that are actually decent at evaluating model
             | performance is much larger and most of the good ones, the
             | ones that were costliest to produce, are hidden, and might
             | not even correspond very well with the public ones. And
             | that the public expensive benchmarks are selective and have
             | a bias towards marketing purposes.
             | 
             | [0] https://www.offconvex.org/2021/04/07/ripvanwinkle/
        
           | H8crilA wrote:
           | Someone else already wrote it, but it's just too funny to not
           | abuse:
           | 
           | Evals are bad because people learn and fit to them. So we do
           | extremely small evals instead.
        
         | Culonavirus wrote:
         | I mean, you vibe check, then you vibe code. Makes perfect
         | sense. (this is a joke)
        
       | beernet wrote:
       | Crazy how OAI is way behind now and the only one to blame is Sam,
       | his ego and lust for influence. Their downwards trajectory of
       | paying accounts since "the move" (DoW deal) is an open secret. If
       | you had placed a new CEO at OAI six months ago and told him to
       | destroy the company, it would have been hard for that CEO to do a
       | better job at that than Sam did. Should have left when he was let
       | go but decided to go full Greg and MAGA instead. Here we are. Go
       | Dario
        
         | beernet wrote:
         | Just to elaborate, as I am getting downvoted by tech bros:
         | 
         | OpenAI restructures after Anthropic captures 70% of new
         | enterprise deals. Claude Code hits $2.5B while Codex lags at
         | $1B ahead of dual IPOs.
         | 
         | Src: https://www.implicator.ai/openai-cuts-its-side-quests-the-
         | en...
        
       | tintor wrote:
       | Several customer testimonials for GPT-5.4 Mini have em dashes in
       | them.
       | 
       | Did GPT write them?
        
       | derefr wrote:
       | OpenAI don't talk about the "size" or "weights" of these models
       | any more. Anyone have any insight into how resource-intensive
       | these Mini/Nano-variant models actually are at this point?
       | 
       | I assume that OpenAI continue to use words like "mini" and "nano"
       | in the names of these model variants, to imply that they reserve
       | the smallest possible resource-units of _their_ inference
       | clusters... but, given OpenAI 's scale, that may well be "one
       | B200" at this point, rather than anything consumers (or even most
       | companies) could afford.
       | 
       | I ask because I'm curious whether the economics of these models'
       | use-cases and call frequency work out (both from the customer
       | perspective, and from OpenAI's perspective) in favor of OpenAI
       | actually hosting inference on these models themselves, vs. it
       | being better if customers (esp. enterprise customers) could
       | instead license these models to run on-prem as black-box software
       | appliances.
       | 
       | But of course, that question is only interesting / only has a
       | non-trivial answer, if these models are small enough that it's
       | actually possible to run them on hardware that costs less to
       | acquire than a year's querying quota for the hosted version.
        
         | technocrat8080 wrote:
         | Have they ever talked about their size or weights?
        
           | derefr wrote:
           | They never put the parameter counts in their model names like
           | other AI companies did, but back in the GPT3 era (i.e. before
           | they had PR people sitting intermediating all their comms
           | channels), OpenAI engineers _would_ disclose this kind of
           | data in their whitepapers  / system cards.
           | 
           | IIRC, GPT-3 itself was admitted to be a 175B model, and its
           | reduced variants were disclosed to have parameter-counts like
           | 1.3B, 6.7B, 13B, etc.
        
             | technocrat8080 wrote:
             | Wow, would love to see a source for this.
        
       | technocrat8080 wrote:
       | 5.4 Mini's OSWorld score is a pleasant surprise. When SOTA scores
       | were still ~30-40 models were too slow and inaccurate for
       | realtime computer use agents (rip Operator/Agent). Curious if
       | anyone's been using these in production.
        
         | Someone1234 wrote:
         | People seem to dismiss OSWorld as "OpenClaw," but I think
         | they're missing how powerful and flexible that type of full-
         | interaction for _safe_ workflows.
         | 
         | We have a legacy Win32 application, and we want to side-by-side
         | compare interactions + responses between it and the web-
         | converted version of the same. Once you've taught the model
         | that "X = Y" between the desktop Vs. web, you've got yourself
         | an automated test suite.
         | 
         | It is possible to do this another way? Sure, but it isn't cost-
         | effective as you scale the workload out to 30+ Win32
         | applications.
        
       | ibrahim_h wrote:
       | The OSWorld numbers are kinda getting lost in the pricing
       | discussion but imo that's the most interesting part. Mini at
       | 72.1% vs 72.4% human baseline is basically noise, so why not just
       | use mini by default unless you're hitting specific failure modes.
       | 
       | Also context bleed into nano subagents in multi-model pipelines
       | -- I've seen orchestrators that just forward the entire message
       | history by default (or something like messages[-N:] without any
       | real budgeting), so your "cheap" extraction step suddenly runs
       | with 30-50K tokens of irrelevant context. And then what's even
       | the point, you've eaten the latency/cost win and added truncation
       | risk on top.
       | 
       | Has anyone actually measured where that cutoff is in practice? At
       | what context size nano stops being meaningfully cheaper/faster in
       | real pipelines, not benchmarks.
        
         | mudkipdev wrote:
         | This is a bot
        
           | ibrahim_h wrote:
           | ironic accusation on a thread about LLMs
        
       | jbellis wrote:
       | Benchmarking these now.
       | 
       | Preregistering my predictions:
       | 
       | Mini: better than Haiku but not as good as Flash 3, especially at
       | reasoning=none.
       | 
       | Nano: worse than Flash 3 Lite. Probably better than Qwen 3.5 27b.
        
       | Rapzid wrote:
       | Oh.. I thought maybe these would be upgrades to gpt-4.1 and
       | gpt-4.1-mini and etc.. But the latency is way too high compared
       | to the 400-600. Yeah, different models and etc but the naming is
       | confusing.
        
       | nicpottier wrote:
       | I've been struggling on finding a reasonably priced model to use
       | with my toy openclaw instance. Opus 4.6 felt kinda magical but
       | that's just too expensive and I'm not risking my max subscription
       | for it.
       | 
       | GPT 5.4 mini is the first alternative that is both affordable and
       | decent. Pretty impressed. On a $20 codex plan I think I'm pretty
       | set and the value is there for me.
        
         | GaggiX wrote:
         | Open source models like MiniMax M2.5, GLM 5, Kimi K2.5 were not
         | decent enough? (via openrouter)
        
           | nicpottier wrote:
           | I will confess that I have not had time to play with those.
           | Will give them a try, thanks for the recommendation.
        
       | simonw wrote:
       | Here's a grid of pelicans for the different models and reasoning
       | levels:
       | https://static.simonwillison.net/static/2026/gpt-5.4-pelican...
        
         | nharada wrote:
         | Surely this task must now be in the training data
        
           | Kye wrote:
           | If it does and works well then it seems like mission
           | accomplished and time for a new benchmark.
        
         | elif wrote:
         | Nano medium must have been run when the servers were on fire
        
       | morpheos137 wrote:
       | i switched to claude when i found chatgpt would argue with just
       | about anything I said even when it was wrong. they have over
       | optimised antisychophancy. i want a model that simulates critical
       | thinking not one that repeats half baked often incomplete dogmas.
       | the chatgpt 5x range is extraordinarily powerful but also extra
       | ordinarily frustrating to try to use for anything creative or
       | productive that is original in my opinion. claude basically is
       | able to think critically while being neither sycophantic or
       | argumentative most of the time in my option with appropriate user
       | prompting. recent chat gpts seem to fight me every step of the
       | way when not doing boiler plate. i don't want to waste my time
       | fighting a tool.
        
       | XCSme wrote:
       | It's odd, that on many benchmarks, including mine[0], Nano does
       | better than Mini.
       | 
       | 5.4 mini seems to struggle with consistency, and even with
       | temperature 0 sometimes gives the correct response, sometimes a
       | wrong one...
       | 
       | [0]: https://aibenchy.com/compare/openai-gpt-5-4-medium/openai-
       | gp...
        
       | michaelgdwn wrote:
       | The Nano tier is the one I'm watching. For agent workflows where
       | you're making dozens of LLM calls per task, the cost per call
       | matters more than peak capability. Would be interesting to see
       | benchmarks on function calling latency specifically -- that's
       | what matters for agents.
        
       ___________________________________________________________________
       (page generated 2026-03-17 23:01 UTC)