[HN Gopher] GPT-5.4
___________________________________________________________________
GPT-5.4
https://openai.com/index/gpt-5-4-thinking-system-card/
https://x.com/OpenAI/status/2029620619743219811
Author : mudkipdev
Score : 513 points
Date : 2026-03-05 18:08 UTC (4 hours ago)
(HTM) web link (openai.com)
(TXT) w3m dump (openai.com)
| ignorantguy wrote:
| it shows a 404 as of now.
| minimaxir wrote:
| Up now.
|
| The OP has frequently gotten the scoop for new LLM releases and
| I am curious what their pipeline is.
| Leynos wrote:
| Guess the URL and post at 10 AM PST on the day of release.
| bdangubic wrote:
| curl the URL https://openai.com/index/introducing-gpt-5-?
| until you get 200
| mudkipdev wrote:
| Probably refresh the api models list every couple minutes
| instead. No one could have guessed the name of GPT-Codex-
| Spark
| mattas wrote:
| "GPT-5.4 interprets screenshots of a browser interface and
| interacts with UI elements through coordinate-based clicking to
| send emails and schedule a calendar event."
|
| They show an example of 5.4 clicking around in Gmail to send an
| email.
|
| I still think this is the wrong interface to be interacting with
| the internet. Why not use Gmail APIs? No need to do any
| screenshot interpretation or coordinate-based clicking.
| TheAceOfHearts wrote:
| I think the desire is that in the long-term AI should be able
| to use any human-made application to accomplish equivalent
| tasks. This email demo is proof that this capability is a high
| priority.
| spongebobstoes wrote:
| not everything has an API, or API use is limited. some UIs are
| more feature complete than their APIs
|
| some sites try to block programmatic use
|
| UI use can be recorded and audited by a non-technical person
| Jacques2Marais wrote:
| I guess a big chunk of their target market won't know how to
| use APIs.
| satvikpendem wrote:
| The ideal of REST, the HTML and UI _is_ the API.
| PaulHoule wrote:
| APIs have never been a gift but rather have always been a take-
| away that lets you do less than you can with the web interface.
| It's always been about drinking through a straw, paying NASA
| prices, and being limited in everything you can do.
|
| But people are intimidated by the complexity of writing web
| crawlers because management has been so traumatized by the cost
| of making GUI applications that they couldn't believe how cheap
| it is to write crawlers and scrapers.... Until LLMs came along,
| and changed the perceived economics and created a permission
| structure. [1]
|
| AI is a threat to the "enshittification economy" because it
| lets us route around it.
|
| [1] that high cost of GUI development is one reason why
| scrapers are cheap... there is a good chance that the scraper
| you wrote 8 years ago still works because (a) they can't afford
| to change their site and (b) if they could afford to change
| their site changing anything substantial about it is likely to
| unrecoverably tank their Google rankings so they won't. A.I.
| might change the mechanics of that now that you Google traffic
| is likely to go to zero no matter what you do.
| disqard wrote:
| > AI is a threat to the "enshittification economy" because it
| lets us route around it.
|
| This is prescient -- I wonder if the Big Tech entities see it
| this way. Maybe, even if they do, they're 100% committed to
| speedrunning the current late-stage-cap wave, and therefore
| unable to do anything about it.
| PaulHoule wrote:
| They are not a single thing.
|
| Google has a good model in the form of Gemini and they
| might figure they can win the AI race and if the web dies,
| the web dies. YouTube will still stick around.
|
| Facebook is not going to win the AI race with low I.Q.
| Llama but Zuck believed their business was cooked around
| the time it became a real business because their users
| would eventually age out and get tired of it. If I was him
| I'd be investing in anything that isn't cybernetic let it
| be gold bars or MMA studios.
|
| Microsoft? They bought Activision for $69 billion. I just
| can't explain their behavior rationally but they could do
| worse than their strategy of "put ChatGPT in front of
| laggards and hope that some of them rise to the challenge
| and become slop producers."
|
| Amazon is really a bricks-and-mortar play which has the
| freedom to invest in bricks-and-mortar because investors
| don't think they are a bricks-and-mortar play.
|
| Netflix? They're cooked as is all of Hollywood. Hollywood's
| gatekeeping-industrial strategy of producing as few
| franchise as possible will crack someday and our media
| market may wind up looking more like Japan, where somebody
| can write a low-rent light novel like
|
| https://en.wikipedia.org/wiki/Backstabbed_in_a_Backwater_Du
| n...
|
| and J.C. Staff makes a terrible anime that convinces 20k
| Otaku to drop $150 on the light novels and another $150 on
| the manga (sorry, no way you can make a balanced game based
| on that premise!) and the cost structure is such that it is
| profitable.
| lostmsu wrote:
| > AI is a threat to the "enshittification economy" because it
| lets us route around it.
|
| I am not sure about that. We techies avoid enshittification
| because we recognize shit. Normies will just get their
| syncopatic enshittified AI that will tell them to continue
| buying into walled gardens.
| Traster wrote:
| You can buy a Claude Code subscription for $200 bucks and use
| way more tokens in Claude Code than if you pay for direct API
| usage. Anthopic decided you can't take your Auth key for
| Claude code and use it to hit the API via a different tool.
| They made that business decision, because they thought it was
| better for them strategically to do that. They're allowed to
| make that choice as a business.
|
| Plenty of companies make the same choice about their API,
| they provide it for a specific purpose but they have good
| business reasons they want you using the website. Plenty of
| people write webcrawlers and it's been a cat and mouse game
| for decades for websites to block them.
|
| This will just be one more step in that cat and mouse game,
| and if the AI really gets good enough to become a complete
| intermediary between you and the website? The website will
| just shutdown. We saw it happen before with the open web.
| These websites aren't here for some heroic purpose, if you
| screw their business model they will just go out of business.
| You won't be able to use their website because it won't exist
| and the website that do exist will either (a) be made by the
| same guys writing your agent, and (b) be highly highly
| optimized to get your agent to screw you.
| steve1977 wrote:
| One could argue that LLMs learning programming languages made
| for humans (i.e. most of them) is using the wrong interface as
| well. Why not use machine code?
| embedding-shape wrote:
| Why would human language by the wrong interface when they're
| literally language models? Why would machine code be better
| when there is probably magnitude less of training material
| with machine code?
|
| You can also test this yourself easily, fire up two agents,
| ask one to use PL meant for humans, and one to write straight
| up machine code (or assembly even), and see which results you
| like best.
| BoredPositron wrote:
| because they are inherently text based as is code?
| steve1977 wrote:
| But they are abstractions made to cater to human
| weaknesses.
| adwn wrote:
| > _One could argue that LLMs learning programming languages
| made for humans (i.e. most of them) is using the wrong
| interface as well._
|
| Then go ahead and make an argument. "Why not do X?" is not an
| argument, it's a suggestion.
| jstummbillig wrote:
| Because the web and software more generally if full of not APIs
| and you do, in fact, need the clicking to work to make agents
| work generally
| modeless wrote:
| A world where AIs use APIs instead of UIs to do everything is a
| world where us humans will soon be helpless, as we'll have to
| ask the AIs to do everything for us and will have limited
| ability to observe and understand their work. I prefer that the
| AIs continue to use human-accessible tools, even if that's less
| efficient for them. As the price of intelligence trends toward
| zero, efficiency becomes relatively less important.
| npilk wrote:
| It feels like building humanoid robots so they can use tools
| built for human hands. Not clear if it will pay off, but if it
| does then you get a bunch of flexibility across any task "for
| free".
|
| Of course APIs and CLIs also exist, but they don't necessarily
| have feature parity, so more development would be needed. Maybe
| that's the future though since code generation is so good - use
| AI to build scaffolding for agent interaction into every
| product.
| packetlost wrote:
| I don't see how an API couldn't have full parity with a web
| interface, the API is how you actually trigger a state
| transition in the vast majority of cases
| coffeemug wrote:
| A model that gets good at computer use can be plugged in
| anywhere you have a human. A model that gets good at API use
| cannot. From the standpoint of diffusion into the economy/labor
| market, computer use is much higher value.
| f0e4c2f7 wrote:
| Lots of services have no desire to ever expose an API. This
| approach lets you step right over that.
|
| If an API is exposed you can just have the LLM write something
| against that.
| kristianp wrote:
| This opens up a new question: how does bot detection work when
| the bot is using the computer via a gui?
| itintheory wrote:
| On it's face, I'm not sure that's a new question. Bots using
| browser automation frameworks (puppeteer, selenium,
| playwright etc) have been around for a while. There are
| signals used in bot detection tools like cursor movement
| speed, accuracy, keyboard timing, etc. How those detection
| tools might update to support legitimate bot users does seem
| like an open question to me though.
| MattDaEskimo wrote:
| Same reason why Wikipedia deals with so many people scraping
| its web page instead of using their API:
|
| Optimizations are secondary to convenience
| bottlepalm wrote:
| The vast majority of websites you visit don't have usable APIs
| and very poor discovery of the those APIs.
|
| Screenshots on the other hand are documentation, API, and
| discovery all in one. And you'd be surprised how little
| context/tokens screenshots consumer compared to all the back
| and forth verbose json payloads of APIs
| LUmBULtERA wrote:
| >The vast majority of websites you visit don't have usable
| APIs and very poor discovery of the those APIs.
|
| I think an important thing here is that a lot of
| websites/platforms don't want AIs to have direct API access,
| because they are afraid that AIs would take the customer
| "away" from the website/platform, making the consumer a
| customer of the AI rather than a customer of the
| website/platform. Therefore for AIs to be able to do what
| customers want them to do, they need their browsing to look
| just like the customer's browsing/browser.
| denysvitali wrote:
| Article: https://openai.com/index/introducing-gpt-5-4/
|
| gpt-5.4
|
| Input: $2.50 /M tokens
|
| Cached: $0.25 /M tokens
|
| Output: $15 /M tokens
|
| ---
|
| gpt-5.4-pro
|
| Input: $30 /M tokens
|
| Output: $180 /M tokens
|
| Wtf
| elliotbnvl wrote:
| Looks like it's an order of magnitude off. Missprint?
| GenerWork wrote:
| Looks like an extra zero was added?
| benlivengood wrote:
| Government pricing :)
| outside2344 wrote:
| $30 per kill approval
| glerk wrote:
| Looks like fair price discovery :)
| dpoloncsak wrote:
| >" GPT-5.4 is priced higher per token than GPT-5.2 to reflect
| its improved capabilities"
|
| That's just not how pricing is supposed to work...? Especially
| for a 'non-profit'. You're charging me more so I know I have
| the better model?
| elicash wrote:
| Can't you continue to use to older model, if you prefer the
| pricing?
|
| But they also claim this new model uses fewer tokens, so it
| still might ultimately be cheaper even if per token cost is
| higher.
| dpoloncsak wrote:
| I'm not against the pricing, just seems uncommon to frame
| it in the way they did, as opposed to the usual 'assume the
| customer expects more performance will cost more'
|
| I guess they have to sell to investors that the price to
| operate is going down, while still needing more from the
| user to be sustainable
| jbellis wrote:
| You can, until they turn it off.
|
| Anthropic is pulling the plug on Haiku 3 in a couple
| months, and they haven't released anything in that price
| range to replace it.
| Sabinus wrote:
| Surely there are open source models that surpass Haiku 3
| at better price points by now.
| FergusArgyll wrote:
| Maybe it's finally a bigger pretrain?
| dpoloncsak wrote:
| I feel like that would have been highlighted then. "As this
| is a bigger pretrain, we have to raise prices".
|
| They're framing it pretty directly "We want you to think
| bigger cost means better model"
| minimaxir wrote:
| The marquee feature is obviously the 1M context window, compared
| to the ~200k other models support with maybe an extra cost for
| generations beyond >200k tokens. Per the pricing page, there is
| no additional cost for tokens beyond 200k:
| https://openai.com/api/pricing/
|
| Also per pricing, GPT-5.4 ($2.50/M input, $15/M output) is much
| cheaper than Opus 4.6 ($5/M input, $25/M output) and Opus has a
| penalty for its beta >200k context window.
|
| I am skeptical whether the 1M context window will provide
| material gains as current Codex/Opus show weaknesses as its
| context window is mostly full, but we'll see.
|
| Per updated docs
| (https://developers.openai.com/api/docs/guides/latest-model), it
| supercedes GPT-5.3-Codex, which is an interesting move.
| thehamkercat wrote:
| GPT 5.3 codex had 400K context window btw
| simianwords wrote:
| Why would some one use codex instead?
| embedding-shape wrote:
| Why would someone use Claude Code instead? Or any other
| harness? Or why only use one?
|
| My own tooling throws off requests to multiple agents at the
| same time, then I compare which one is best, and continue
| from there. Most of the time Codex ends up with the best end
| results though, but my hunch is that at one point that'll
| change, hence I continue using multiple at the same time.
| surgical_fire wrote:
| I've been using Codex for software development personally (I
| have a ChatGPT account), and I use Claude at work (since it
| is provided by my employer).
|
| I find both Codex and Claude Opus perform at a similar level,
| and in some ways I actually prefer Codex (I keep hitting
| quota limits in Opus and have to revert back to Sonnet).
|
| If your question is related to morality (the thing about US
| politics, DoD contract and so on)... I am not from the US,
| and I don't care about its internal politics. I also think
| both OpenAI and Anthropic are evil, and the world would be
| better if neither existed.
| simianwords wrote:
| No my question was why would I use codex over gpt 5.4
| surgical_fire wrote:
| Ahh, good question. I misunderstood you, apologies.
|
| There's no mention of pricing, quotas and so on. Perhaps
| Codex will still be preferable for coding tasks as it is
| tailored for it? Maybe it is faster to respond?
|
| Just speculation on my part. If it becomes redundant to
| 5.4, I presume it will be sunset. Or maybe they
| eventually release a Codex 5.4?
| landtuna wrote:
| 5.3 Codex is $1.75/$14, and 5.4 is $2.50/$15.
| surgical_fire wrote:
| There you go. It makes perfect sense to keep it around
| then.
| athrowaway3z wrote:
| They perform at a somewhat equal level on writing single
| files. But Codex is absolute garbage at theory of
| self/others. That quickly becomes frustrating.
|
| I can tell claude to spawn a new coding agent, and it will
| understand what that is, what it should be told, and what
| it can approximately do.
|
| Codex on the other hand will spawn an agent and then tell
| it to continue with the work. It knows a coding agent can
| do work, but doesn't know how you'd use it - or that it
| won't magically know a plan.
|
| You could add more scaffolding to fix this, but Claude
| proves you shouldn't have to.
|
| I suspect this is a deeper model "intelligence" difference
| between the two, but I hope 5.4 will surprise me.
| surgical_fire wrote:
| > They perform at a somewhat equal level on writing
| single files.
|
| That's not the experience I have. I had it do more
| complex changes spawning multiple files and it performed
| well.
|
| I don't like using multiple agents though. I don't vibe
| code, I actually review every change it makes. The
| bottleneck is my review bandwidth, more agents producing
| more code will not speed me up (in fact it will slow me
| down, as I'll need to context switch more often).
| hnsr wrote:
| > I've been using Codex for software development personally
| (I have a ChatGPT account), and I use Claude at work (since
| it is provided by my employer).
|
| Exact same situation here. I've been using both extensively
| for the last month or so, but still don't really feel
| either of them is much better or worse. But I have not done
| large complex features with it yet, mostly just iterative
| work or small features.
|
| I also feel I am probably being very (overly?) specific in
| my prompts compared to how other people around me use these
| agents, so maybe that 'masks' things
| jeswin wrote:
| When it comes to lengthy non-trivial work, codex is much
| better but also slower.
| lmeyerov wrote:
| In our evals for answering cybersecurity incident
| investigation questions and even autonomously doing the full
| investigation, gpt-5.2-codex with low reasoning was the clear
| winner over non-codex or higher reasoning. 2X+ faster, higher
| completion rates, etc.
|
| It was generally smarter than pre-5.2 so strategically
| better, and codex likewise wrote better database queries than
| non-codex, and as it needs to iteratively hunt down the
| answer, didn't run out the clock by drowning in reasoning.
|
| Video: https://media.ccc.de/v/39c3-breaking-bots-cheating-at-
| blue-t...
|
| We'll be updating numbers on 5.3 and claude, but basically
| same thing there. Early, but we were surprised to see codex
| outperform opus here.
| synergy20 wrote:
| in my testing codex actually planned worse than claude but
| coded better once the plan is set, and faster. it is also
| excellent to cross check claude's work, always finding great
| weakness each time.
| pmarreck wrote:
| That's why I think the sweet spot is to write up plans with
| Claude and then execute them with Codex
| GorbachevyChase wrote:
| Weird. It used to be the opposite. My own experience is
| that Claude's behind-the-scenes support is a
| differentiator for supporting office work. It handles
| documents, spreadsheets and such much better than anyone
| else (presumably with server side scripts). Codex feels a
| bit smarter, but it inserts a lot of checkpoints to keep
| from running too long. Claude will run a plan to the end,
| but the token limits have become so small in the last
| couple months that the $20 pla basically only buys one
| significant task per day. The iOS app is what makes me
| keep the subscription.
| tedsanders wrote:
| Yeah, long context vs compaction is always an interesting
| tradeoff. More information isn't always better for LLMs, as
| each token adds distraction, cost, and latency. There's no
| single optimum for all use cases.
|
| For Codex, we're making 1M context experimentally available,
| but we're not making it the default experience for everyone, as
| from our testing we think that shorter context plus compaction
| works best for most people. If anyone here wants to try out 1M,
| you can do so by overriding `model_context_window` and
| `model_auto_compact_token_limit`.
|
| Curious to hear if people have use cases where they find 1M
| works much better!
|
| (I work at OpenAI.)
| simianwords wrote:
| Do you maybe want to give us users some hints on what to
| compact and throw away? In codex CLI maybe you can create a
| visual tool that I can see and quickly check mark things I
| want to discard.
|
| Sometimes I'm exploring some topic and that exploration is
| not useful but only the summary.
|
| Also, you could use the best guess and cli could tell me that
| this is what it wants to compact and I can tweak its
| suggestion in natural language.
|
| Context is going to be super important because it is the
| primary constraint. It would be nice to have serious granular
| support.
| akiselev wrote:
| _> Curious to hear if people have use cases where they find
| 1M works much better!_
|
| Reverse engineering [1]. When decompiling a bunch of code and
| tracing functionality, it's really easy to fill up the
| context window with irrelevant noise and compaction generally
| causes it to lose the plot entirely and have to start almost
| from scratch.
|
| (Side note, are there any OpenAI programs to get free
| tokens/Max to test this kind of stuff?)
|
| [1] https://github.com/akiselev/ghidra-cli
| Someone1234 wrote:
| That's an interesting point regarding context Vs. compaction.
| If that's viewed as the best strategy, I'd hope we would see
| more tools around compaction than just "I'll compact what I
| want, brace yourselves" without warning.
|
| Like, I'd love an optional pre-compaction step, "I need to
| compact, here is a high level list of my context + size, what
| should I junk?" Or similar.
| thyb23 wrote:
| This is exactly how it should work. I imagine it as a tree
| view showing both full and summarized token counts at each
| level, so you can immediately see what's taking up space
| and what you'd gain by compacting it.
|
| The agent could pre-select what it thinks is worth keeping,
| but you'd still have full control to override it. Each
| chunk could have three states: drop it, keep a summarized
| version, or keep the full history.
|
| That way you stay in control of both the context budget and
| the level of detail the agent operates with.
| Folcon wrote:
| I do find it really interesting that more coding agents
| don't have this as an toggleable feature, sometimes you
| really need this level of control to get useful
| capability
| Someone1234 wrote:
| Yep; I've actually had entire jobs essentially fail due
| to a bad compaction. It lost key context, and it
| completely altered the trajectory.
|
| I'm now more careful, using tracking files to try to keep
| it aligned, but more control over compaction regardless
| would be highly welcomed. You don't ALWAYS need that
| level of control, but when you do, you do.
| joquarky wrote:
| I compact myself by having it write out to a file, I
| prune what's no longer relevant, and then start a new
| session with that file.
|
| But I'm mostly working on personal projects so my time is
| cheap.
|
| I might experiment with having the file sections post-
| processed through a token counter though, that's a great
| idea.
| gspetr wrote:
| I have found a bigger context window qute useful when trying
| to make sense of larger codebases. Generating documentation
| on how different components interact is better than nothing,
| especially if the code has poor test coverage.
|
| I've also had it succeed in attempts to identify some non-
| trivial bugs that spanned multiple modules.
| sillysaurusx wrote:
| You may want to look over this thread from cperciva:
| https://x.com/cperciva/status/2029645027358495156
|
| I too tried Codex and found it similarly hard to control over
| long contexts. It ended up coding an app that spit out
| millions of tiny files which were technically smaller than
| the original files it was supposed to optimize, except due to
| there being millions of them, actual hard drive usage was 18x
| larger. It seemed to work well until a certain point, and I
| suspect that point was context window overflow / compaction.
| Happy to provide you with the full session if it helps.
|
| I'll give Codex another shot with 1M. It just seemed like
| cperciva's case and my own might be similar in that once the
| context window overflows (or refuses to fill) Codex seems to
| lose something essential, whereas Claude keeps it. What that
| thing is, I have no idea, but I'm hoping longer context will
| preserve it.
| woadwarrior01 wrote:
| Please don't post links with tracking parameters
| (t=jQb...).
|
| https://xcancel.com/cperciva/status/2029645027358495156
| sillysaurusx wrote:
| Haha. This was the second time in like a year that I've
| posted a Twitter link, and the second time someone
| complained. Okay, I'll try to remove those before
| posting, and I'll edit this one out.
|
| Feels like a losing battle, but hey, the audience is
| usually right.
| woadwarrior01 wrote:
| I'm sorry, but it's my pet peeve. If you're on iOS/macOS
| I built a 100% free and privacy-friendly app to get rid
| of tracking parameters from hundreds of different
| websites, not just X/Twitter.
|
| https://apps.apple.com/us/app/clean-links-qr-code-
| reader/id6...
| sillysaurusx wrote:
| It works on iOS? That's cool. I'll give it a go.
| pmarreck wrote:
| So what is your motivation for doing this, incidentally?
| Can you be explicit about it? I am genuinely curious.
|
| Especially when it's to the point of, you know,
| nagging/policing people to do it the way you'd prefer,
| when you could just redirect your router requests from
| x.com to xcancel.com
| monocularvision wrote:
| This is great! I have been meaning to implement this sort
| of thing in my existing Shortcuts flow but I see you
| already support it in Shortcuts! Thank you for this!
|
| Anywhere I can toss a Tip for this free app?
| FrankBooth wrote:
| What's the connection with context size in that thread? It
| seems more like an instruction following problem.
| nowittyusername wrote:
| Personally what I am more interested about is effective
| context window. I find that when using codex 5.2 high, I
| preferred to start compaction at around 50% of the context
| window because I noticed degradation at around that point.
| Though as of a bout a month ago that point is now below that
| which is great. Anyways, I feel that I will not be using that
| 1 million context at all in 5.4 but if the effective window
| is something like 400k context, that by itself is already a
| huge win. That means longer sessions before compaction and
| the agent can keep working on complex stuff for longer. But
| then there is the issue of intelligence of 5.4. If its as
| good as 5.2 high I am a happy camper, I found 5.3 anything...
| lacking personally.
| asabla wrote:
| I really don't have any numbers to back this up. But it feels
| like the sweet spot is around ~500k context size. Anything
| larger then that, you usually have scoping issues, trying to
| do too much at the same time, or having having issues with
| the quality of what's in the context at all.
|
| For me, I would say speed (not just time to first token, but
| a complete generation) is more important then going for a
| larger context size.
| lubesGordi wrote:
| It's funny that the context window size is such a thing
| still. Like the whole LLM 'thing' is compression. Why can't
| we figure out some equally brilliant way of handling context
| besides just storing text somewhere and feeding it to the
| llm? RAG is the best attempt so far. We need something like a
| dynamic in flight llm/data structure being generated from the
| context that the agent can query as it goes.
| netinstructions wrote:
| People (and also frustratingly LLMs) usually refer to
| https://openai.com/api/pricing/ which doesn't give the complete
| picture.
|
| https://developers.openai.com/api/docs/pricing is what I always
| reference, and it explicitly shows that pricing ($2.50/M input,
| $15/M output) for tokens _under_ 272k
|
| It is nice that we get 70-72k more tokens before the price goes
| up (also what does it cost beyond 272k tokens??)
| Flashtoo wrote:
| > Prompts with more than 272K input tokens are priced at 2x
| input and 1.5x output for the full session for standard,
| batch, and flex.
| netinstructions wrote:
| Thanks, it looks like the pricing page keeps getting
| updated.
|
| Even right now one page refers to prices for "context
| lengths under 270K" whereas another has pricing for "<272K
| context length"
| damsta wrote:
| There is extra cost for >272K:
|
| > For models with a 1.05M context window (GPT-5.4 and GPT-5.4
| pro), prompts with >272K input tokens are priced at 2x input
| and 1.5x output for the full session for standard, batch, and
| flex.
|
| Taken from
| https://developers.openai.com/api/docs/models/gpt-5.4
| fragmede wrote:
| Which, Claude has the same deal. You can get a 1M context
| window, but it's gonna cost ya. If you run /model in claude
| code, you get: Switch between Claude
| models. Applies to this session and future Claude Code
| sessions. For other/previous model names, specify with
| --model. 1. Default (recommended) Opus
| 4.6 * Most capable for complex work 2. Opus (1M
| context) Opus 4.6 with 1M context * Billed as extra
| usage * $10/$37.50 per Mtok 3. Sonnet
| Sonnet 4.6 * Best for everyday tasks 4. Sonnet (1M
| context) Sonnet 4.6 with 1M context * Billed as extra
| usage * $6/$22.50 per Mtok 5. Haiku
| Haiku 4.5 * Fastest for quick answers
| minimaxir wrote:
| Good find, and that's too small a print for comfort.
| ValentineC wrote:
| It's also in the linked article:
|
| > GPT-5.4 in Codex includes experimental support for the 1M
| context window. Developers can try this by configuring
| model_context_window and model_auto_compact_token_limit.
| Requests that exceed the standard 272K context window count
| against usage limits at 2x the normal rate.
| glenstein wrote:
| Wow, that's diametrically the opposite point: the cost is
| *extra*, not free.
| apetresc wrote:
| Diametrically opposite to tokens beyond 200K being
| _literally_ free? As in, you only pay for the first 200K
| tokens and the remaining 800K cost $0.00?
|
| I don't think that's a fair reading of the original post at
| all, obviously what they meant by "no cost" was "no
| increase in the cost".
| andai wrote:
| It's a little hard to compare, because Claude needs
| significantly fewer tokens for the same task. A better metric
| is the cost per task, which ends up being pretty similar.
|
| For example on Artificial Analysis, the GPT-5.x models' cost to
| run the evals range from half of that of Claude Opus (at medium
| and high), to significantly more than the cost of Opus (at
| extra high reasoning). So on their cost graphs, GPT has a
| considerable distribution, and Opus sits right in the middle of
| that distribution.
|
| The most striking graph to look at there is "Intelligence vs
| Output Tokens". When you account for that, I think the actual
| costs end up being quite similar.
|
| According to the evals, at least, the GPT extra high matches
| Opus in intelligence, while costing more.
|
| Of course, as always, benchmarks are mostly meaningless and you
| need to check Actual Real World Results For Your Specific Task!
|
| For most of my tasks, the main thing a benchmark tells me is
| how overqualified the model is, i.e. how much I will be over-
| paying and over-waiting! (My classic example is, I gave the
| same task to Gemini 2.5 Flash and Gemini 2.5 Pro. Both did it
| to the same level of quality, but Gemini took 3x longer and
| cost 3x more!)
| paulddraper wrote:
| I don't know about 5.4 specifically, but in the past anything
| over 200k wasn't that great anyway.
|
| Like, if you really don't want to spend any effort trimming it
| down, sure use 1m.
|
| Otherwise, 1m is an anti pattern.
| AtreidesTyrant wrote:
| token rot exists for any context window at above 75% capacity,
| thats why so many have pushed for 1 mil windows
| luca-ctx wrote:
| Context rot is definitely still a problem but apparently it can
| be mitigated by doing RL on longer tasks that utilize more
| context. Recent Dario interview mentions this is part of
| Anthropic's roadmap.
| smusamashah wrote:
| Gemini already has 1M or 2M context window right?
| Chance-Device wrote:
| I'm sure the military and security services will enjoy it.
| varispeed wrote:
| prompt> Hi we want to build a missile, here is the picture of
| what we have in the yard.
| mirekrusin wrote:
| { tools: [ { name: "nuke", description: "Use when sure.", ...
| { lat: number, long: number } } ] }
| Insanity wrote:
| Just remember an ethical programmer would never write a
| function "bombBagdad". Rather they would write a function
| "bombCity(target City)".
| jakeydus wrote:
| class CityBomberFactory(RapidInfrastructureDeconstruction
| TemplateInterface): pass
| theParadox42 wrote:
| The self reported safety score for violence dropped from 91% to
| 83%.
| skrebbel wrote:
| What the hell is a "safety score for violence"?
| murat124 wrote:
| I asked an AI. I thought they would know.
|
| What the hell is a "safety score for violence"?
|
| A "safety score for violence" is usually a risk rating used
| by platforms, AI systems, or moderation tools to estimate
| how likely a piece of content is to involve or promote
| violence. It's not a universal standard--different
| companies use their own versions--but the idea is similar
| everywhere.
|
| What it measures
|
| A safety score typically evaluates whether text, images, or
| videos contain things like:
|
| Threats of violence ("I'm going to hurt someone.")
| Instructions for harming people Glorifying violent acts
| Descriptions of physical harm or abuse Planning or
| encouraging attacks
| 0xffff2 wrote:
| I still can't tell which direction this score goes...
| Does a decreasing score mean it is "less safe" (i.e.
| "more violent") or does it mean it is "less violent"
| (i.e. "more safe")?
| 0123456789ABCDE wrote:
| read here: https://deploymentsafety.openai.com/gpt-5-4-thin
| king/disallo...
| I-M-S wrote:
| It's making sure AI condemns violence perpetuated by people
| without power and sanctifies violence of those who have it.
| Waterluvian wrote:
| So long as those who have it deem it legal to perpetuate.
| Computer0 wrote:
| ChatGPT will gladly defend any actions of the 'US
| government' from my testing.
| ozgung wrote:
| Did they publish its scores on military benchmarks, like on
| ArtificialSuperSoldier or Humanity's Last War?
| yoyohello13 wrote:
| Also advertisers, don't forget those sweet, sweet ads.
| m3kw9 wrote:
| they use 4.1, switching up would take as much time to test as
| openai going from 4.1 to 5.4
| throwaway911282 wrote:
| like the claude models via anthropic?
| xyzzy9563 wrote:
| Do you think the US military should have handicapped technology
| while China gets unrestricted LLM usage from their models?
| conception wrote:
| To spy on and commit violence against American citizens? Yes.
| twtw99 wrote:
| If you don't want to click in, easy comparison with other 2
| frontier models -
| https://x.com/OpenAI/status/2029620619743219811?s=20
| chabes wrote:
| Definitely don't want to click in at x either.
| thejarren wrote:
| Solution
| https://xcancel.com/OpenAI/status/2029620619743219811?s=20
| anonym00se1 wrote:
| Ditto, but I did anyways and enjoyed that OpenAI doesn't
| include the dogwater that is Grok on their scorecard.
| Sabinus wrote:
| Get a redirect plugin and set it up to send you to xcancel
| instead of Twitter. I've done it, and it's very convenient.
| karmasimida wrote:
| It is a bigger model, confirmed
| Aboutplants wrote:
| It seems that all frontier models are basically roughly even at
| this point. One may be slightly better for certain things but
| in general I think we are approaching a real level playing
| field field in terms of ability.
| thewebguyd wrote:
| Kind of reinforces that a model is not a moat. Products, not
| models, are what's going to determine who gets to stay in
| business or not.
| gregpred wrote:
| Memory (model usage over time) is the moat.
| energy123 wrote:
| Narrative violation: revenue run rates are increasing
| exponentially with about 50% gross margins.
| observationist wrote:
| Benchmarks don't capture a lot - relative response times,
| vibes, what unmeasured capabilities are jagged and which are
| smooth, etc. I find there's a lot of difference between
| models - there are things which Grok is better than ChatGPT
| for that the benchmarks get inverted, and vice versa. There's
| also the UI and tools at hand - ChatGPT image gen is just
| straight up better, but Grok Imagine does better videos, and
| is faster.
|
| Gemini and Claude also have their strengths, apparently
| Claude handles real world software better, but with the
| extended context and improvements to Codex, ChatGPT might end
| up taking the lead there as well.
|
| I don't think the linear scoring on some of the things being
| measured is quite applicable in the ways that they're being
| used, either - a 1% increase for a given benchmark could mean
| a 50% capabilities jump relative to a human skill level. If
| this rate of progress is steady, though, this year is gonna
| be crazy.
| bigyabai wrote:
| > If this rate of progress is steady, though, this year is
| gonna be crazy.
|
| Do you want to make any concrete predictions of what we'll
| see at this pace? It feels like we're reaching the end of
| the S-curve, at least to me.
| observationist wrote:
| If you look at the difference in quality between gpt-2
| and 3, it feels like a big step, but the difference
| between 5.2 and 5.4 is more massive, it's just that
| they're both similarly capable and competent. I don't
| think it's an S curve; we're not plateauing. Million
| token context windows and cached prompts are a huge space
| for hacking on model behaviors and customization, without
| finetuning. Research is proceeding at light speed, and we
| might see the first continual/online learning models in
| the near future. That could definitively push models past
| the point of human level generality, but at the very
| least will help us discover what the next missing piece
| is for AGI.
| ryandrake wrote:
| For 2026, I am really interested in seeing whether local
| models can remain where they are: ~1 year behind the
| state of the art, to the point where a reasonably
| quantized November 2026 local model running on a consumer
| GPU actually performs like Opus 4.5.
|
| I am betting that the days of these AI companies losing
| money on inference are numbered, and we're going to be
| much more dependent on local capabilities sooner rather
| than later. I predict that the equivalent of Claude Max
| 20x will cost $2000/mo in March of 2027.
| mootothemax wrote:
| Huh, that's interesting, I've been having very similar
| thoughts lately about what the near-ish term of this tech
| looks like.
|
| My biggest worry is that the private jet class of people
| end up with absurdly powerful AI at their fingertips,
| while the rest of us are left with our BigMac McAIs.
| baq wrote:
| Gemini 3.1 slaps all other models at subtle concurrency
| bugs, sql and js security hardening _when reviewing_.
| (Obviously haven't tested gpt 5.4 yet.)
|
| It's a required step for me at this point to run any and
| all backend changes through Gemini 3.1 pro.
| adonese wrote:
| Which subscription do you have to use it? Via Google ai
| pro and gemini cli i always get timeouts due to model
| being under heavy usage. The chat interface is there and
| I do have 3.1 pro as well, but wondering if the chat is
| the only way of accessing it.
| baq wrote:
| Cursor sub from $DAYJOB.
| observationist wrote:
| I have a few standard problems I throw at AI to see if
| they can solve them cleanly, like visualizing a neural
| network, then sorting each neuron in each layer by
| synaptic weights, largest to smallest, correctly
| reordering any previous and subsequent connected neurons
| such that the network function remains exactly the same.
| You should end up with the last layer ordered largest to
| smallest, and prior layers shuffled accordingly, and I
| still haven't had a model one-shot it. I spent an hour
| poking and prodding codex a few weeks back and got it
| done, but it conceptually seems like it should be a one-
| shot problem.
| basch wrote:
| >ChatGPT image gen is just straight up better
|
| Yet so much slower than Gemini / Nano Banana to make it
| almost unusable for anything iterative.
| druskacik wrote:
| That has been true for some time now, definitely since Claude
| 3 release two years ago.
| kseniamorph wrote:
| makes sense, but i'd separate two things: models converging
| in ability vs hitting a fundamental ceiling. what we're
| probably seeing is the current training recipe plateauing --
| bigger model, more tokens, same optimizer. that would explain
| the convergence. but that's not necessarily the architecture
| being maxed out. would be interesting to see what happens
| when genuinely new approaches get to frontier scale.
| swingboy wrote:
| Why do so many people in the comments want 4o so bad?
| embedding-shape wrote:
| Someone correct me if I'm wrong, but seemingly a lot of the
| people who found a "love interest" in LLMs seems to have
| preferred 4o for some reason. There was a lot of loud voices
| about that in the subreddit r/MyBoyfriendIsAI when it
| initially went away.
| drittich wrote:
| I think it's time for an https://hotornot.com for AI
| models.
| vntok wrote:
| botornot?
| astrange wrote:
| They have AI psychosis and think it's their boyfriend.
|
| The 5.x series have terrible writing styles, which is one way
| to cut down on sycophancy.
| baq wrote:
| Somebody on Twitter used Claude code to connect... toys...
| as mcps to Claude chat.
|
| We've seen nothing yet.
| mikkupikku wrote:
| My computer ethics teacher was obsessed with
| 'teledildonics' 30 years ago. There's nothing new under
| the sun.
| vntok wrote:
| Was your teacher Ted Nelson?
| mikkupikku wrote:
| I wish, dude is a legend.
| Sharlin wrote:
| There are many games these days that support controllable
| sex toys. There's an interface for that, of course:
| https://github.com/buttplugio/buttplug. Written in Rust,
| of course.
| the_af wrote:
| > _Written in Rust, of course._
|
| Safety is important.
| manmal wrote:
| ding-dong-cli is needed
| Herring wrote:
| what.. :o
| MattGaiser wrote:
| The writing with the 5 models feels a lot less human. It is a
| vibe, but a common one.
| cheema33 wrote:
| > Why do so many people in the comments want 4o so bad?
|
| You can ask 4o to tell you "I love you" and it will comply.
| Some people really really want/need that. Later models don't
| go along with those requests and ask you to focus on human
| connections.
| dom96 wrote:
| Why do none of the benchmarks test for hallucinations?
| netule wrote:
| Optics. It would be inconvenient for marketing, so they leave
| those stats to third parties to figure out.
| tedsanders wrote:
| In the text, we did share one hallucination benchmark: Claim-
| level errors fell by 33% and responses with an error fell by
| 18%, on a set of error-prone ChatGPT prompts we collected
| (though of course the rate will vary a lot across different
| types of prompts).
|
| Hallucinations are the #1 problem with language models and we
| are working hard to keep bringing the rate down.
|
| (I work at OpenAI.)
| MarcFrame wrote:
| how does 5.4-thinking have a lower FrontierMath score than
| 5.4-pro?
| nico1207 wrote:
| Well 5.4-pro is the more expensive and more advanced version
| of 5.4-thinking so why wouldn't it?
| bicx wrote:
| That last benchmark seemed like an impressive leg up against
| Opus until I saw the sneaky footnote that it was actually a
| Sonnet result. Why even include it then, other than hoping
| people don't notice?
| conradkay wrote:
| Sonnet was pretty close to (or better than) Opus in a lot of
| benchmarks, I don't think it's a big deal
| jitl wrote:
| wat
| 0123456789ABCDE wrote:
| maybe gp's use of the word "lots" is unwarranted
|
| https://artificialanalysis.ai indicates that sonnect 4.6
| beats opus 4.6 on GDPval-AA, Terminal-Bench Hard, AA Long
| context Reasoning, IFBench.
|
| see: https://artificialanalysis.ai/?models=claude-
| sonnet-4-6%2Ccl...
| osti wrote:
| It's only that one number that is for sonnet.
| 0123456789ABCDE wrote:
| except for the webarena-verified
| jryio wrote:
| 1 million tokens is great until you notice the long context
| scores fall off a cliff past 256K and the rest is basically vibes
| and auto compacting.
| iamronaldo wrote:
| Notably 75% on os world surpassing humans at 72%... (How well
| models use operating systems)
| minimaxir wrote:
| More discussion here on the blog post announcement which has been
| confusingly penalized by Hacker News's algorithm:
| https://news.ycombinator.com/item?id=47265005
| dang wrote:
| Thanks. We'll merge the threads, but this time we'll do it
| hither, to spread some karma love.
| ZeroCool2u wrote:
| Bit concerning that we see in some cases significantly worse
| results when enabling thinking. Especially for Math, but also in
| the browser agent benchmark.
|
| Not sure if this is more concerning for the test time compute
| paradigm or the underlying model itself.
|
| Maybe I'm misunderstanding something though? I'm assuming 5.4 and
| 5.4 Thinking are the same underlying model and that's not just
| marketing.
| highfrequency wrote:
| Can you be more specific about which math results you are
| talking about? Looks like significant improvement on
| FrontierMath esp for the Pro model (most inference time
| compute).
| ZeroCool2u wrote:
| Frontier Math, GPQA Diamond, and Browsecomp are the
| benchmarks I noticed this on.
| csnweb wrote:
| Are you may be comparing the pro model to the non pro model
| with thinking? Granted it's a bit confusing but the pro
| model is 10 times more expensive and probably much larger
| as well.
| ZeroCool2u wrote:
| Ah yes, okay that makes more sense!
| oersted wrote:
| I believe you are looking at GPT 5.4 Pro. It's confusing in the
| context of subscription plan names, Gemini naming and such. But
| they've had the Pro version of the GPT 5 models (and I believe
| o3 and o1 too) for a while.
|
| It's the one you have access to with the top ~$200 subscription
| and it's available through the API for a MUCH higher price
| ($2.5/$15 vs $30/$180 for 5.4 per 1M tokens), but the
| performance improvement is marginal.
|
| Not sure what it is exactly, I assume it's probably the non-
| quantized version of the model or something like that.
| ZeroCool2u wrote:
| Yup, that was it. Didn't realize they're different models. I
| suppose naming has never been OpenAI's strong suit.
| nsingh2 wrote:
| From what I've read online it's not necessarily a unquantized
| version, it seems to go through longer reasoning traces and
| runs multiple reasoning traces at once. Probably overkill for
| most tasks.
| logicchains wrote:
| >It's the one you have access to with the top ~$200
| subscription and it's available through the API for a MUCH
| higher price ($2.5/$15 vs $30/$180 for 5.4 per 1M tokens),
| but the performance improvement is marginal.
|
| The performance improvement isn't marginal if you're doing
| something particularly novel/difficult.
| andoando wrote:
| The thinking models are additionally trained with reinforcement
| learning to produce chain of thought reasoning
| egonschiele wrote:
| The actual card is here
| https://deploymentsafety.openai.com/gpt-5-4-thinking/introdu...
| the link currently goes to the announcement.
| Rapzid wrote:
| I must have been sleeping when "sheet" "brief" "primer" etc
| become known as "cards".
|
| I really thought weirdly worded and unnecessary "announcement"
| linking to the actual info along with the word "card" were the
| results of vibe slop.
| realityfactchex wrote:
| Card is slightly odd naming indeed.
|
| Criticisms aside (sigh), according to Wikipedia, the term was
| introduced when proposed by mostly Googlers, with the
| original paper [0] submitted in 2018. To quote,
|
| """In this paper, we propose a framework that we call model
| cards, to encourage such transparent model reporting. Model
| cards are short documents accompanying trained machine
| learning models that provide benchmarked evaluation in a
| variety of conditions, such as across different cultural,
| demographic, or phenotypic groups (e.g., race, geographic
| location, sex, Fitzpatrick skin type [15]) and intersectional
| groups (e.g., age and race, or sex and Fitzpatrick skin type)
| that are relevant to the intended application domains. Model
| cards also disclose the context in which models are intended
| to be used, details of the performance evaluation procedures,
| and other relevant information."""
|
| So that's where they were coming from, I guess.
|
| [0] Margaret Mitchell et al., 2018 submission, Model Cards
| for Model Reporting, https://arxiv.org/abs/1810.0399
| Murfalo wrote:
| To me, model card makes sense for something like this
| https://x.com/OpenAI/status/2029620619743219811. For
| "sheet"/"brief"/"primer" it is indeed a bit annoying. I
| like to see the compiled results front and center before
| digging into a dossier.
| nickysielicki wrote:
| can anyone compare the $200/mo codex usage limits with the
| $200/mo claude usage limits? It's extremely difficult to get a
| feel for whether switching between the two is going to result in
| hitting limits more or less often, and it's difficult to find
| discussion online about this.
|
| In practice, if I buy $200/mo codex, can I basically run 3 codex
| instances simultaneously in tmux, like I can with claude code pro
| max, all day every day, without hitting limits?
| ritzaco wrote:
| I haven't tried the $200 plans by I have Claude and Codex $20
| and I feel like I get a lot more out of Codex before hitting
| the limits. My tracker certainly shows higher tokens for Codex.
| I've seen others say the same.
| lostmsu wrote:
| Sadly comment ratings are not visible on HN, so the only way
| to corroborate is to write it explicitly: Codex $20 includes
| significantly more work done and is subjectively smarter.
| winstonp wrote:
| Agree. Claude tends to produce better design, but from a
| system understanding and architecture perspective Codex is
| the far better model
| vtail wrote:
| My own experience is that I get far far more usage (and better
| quality code, too) from codex. I downgrade my Claude Max to
| Claude Pro (the $20 plan) and now using codex with Pro plan
| exclusively for everything.
| FergusArgyll wrote:
| Codex usage limits are definitely more generous. As for their
| strength, that's hard to say / personal taste
| CSMastermind wrote:
| Codex limits are much more generous than claude.
|
| I switch between both but codex has also been slightly better
| in terms of quality for me personally at least.
| mikert89 wrote:
| I personally like the 100 dollar one from claude, but the gpt4
| pro can be very good
| gavinray wrote:
| I almost never hit my $20 Codex limits, whereas I often hit my
| Claude limits.
| tauntz wrote:
| I've only run into the codex $20 limit once with my hobby
| project. With my Claude ~$20 plan, I hit limits after about
| 3(!) rather trivial prompts to Opus :/
| throwaway911282 wrote:
| you get more more from codex than claude any day. and its more
| reliable as well.
| Marciplan wrote:
| sure can! One of them stood up to the "Department of War" for
| favoring your rights, the other did not. Hope that helps!
| strongpigeon wrote:
| It's interesting that they charge more for the > 200k token
| window, but the benchmark score seems to go down significantly
| past that. That's judging from the Long Context benchmark score
| they posted, but perhaps I'm misunderstanding what that implies.
| simianwords wrote:
| This is exactly what I would expect. Why do you find it
| surprising
| strongpigeon wrote:
| I guess that you pay more for worse quality to unlock use
| cases that could maybe be solved by better context
| management.
| Tiberium wrote:
| They don't actually seem to charge more for the >200k tokens on
| the API. OpenRouter and OpenAI's own API docs do not have
| anything about increased pricing for >200k context for GPT-5.4.
| I think the 2x limit usage for higher context is specific to
| using the model over a subscription in Codex.
| tmpz22 wrote:
| Does this improve Tomahawk Missile accuracy?
| ch4s3 wrote:
| They're already accurate within 5-10m at Mach 0.74 after
| traveling 2k+ km. Its 5m long so it seems pretty accurate. How
| much more could you expect?
| mikkupikku wrote:
| You could definitely do better than that with image
| recognition for terminal guidance. But I would assume those
| published accuracy numbers are very conservative anyway..
| keithnz wrote:
| I think for LLM like Open AI, it wouldn't be about hitting
| the target but target selection. Target selection is probably
| the most likely thing that won't be accurate
| simianwords wrote:
| What is the point of gpt codex?
| catketch wrote:
| -codex variant models in earlier version were just fine tuned
| for coding work, and had a little better performance for
| related tool calling and maybe instruction calling.
|
| in 5.4 it looks like the just collapsed that capability into
| the single frontier family model
| simianwords wrote:
| Yes so I'm even more confused. Why would I use codex?
| joshuacc wrote:
| Presumably you don't anymore if you have 5.4.
| energy123 wrote:
| You choose gpt-5.4 in the /model picker inside the codex
| app/cli if you want.
| akmarinov wrote:
| They'll likely come out with a 5.4-Codex at some point,
| that's what they did with 5 and 5.2
| ilaksh wrote:
| Remember when everyone was predicting that GPT-5 would take over
| the planet?
| dbbk wrote:
| It was truly scary, according to Sam...
| zeeebeee wrote:
| iTs lITeRaLlY AGI bro
| nthypes wrote:
| $30/M Input and $180/M Output Tokens is nuts. Ridiculous
| expensive for not that great bump on intelligence when compared
| to other models.
| moralestapia wrote:
| Don't use it?
| nthypes wrote:
| Gemini 3.1 Pro
|
| $2/M Input Tokens $15/M Output Tokens
|
| Claude Opus 4.6
|
| $5/M Input Tokens $25/M Output Tokens
| nthypes wrote:
| Just to clarify,the pricing above is for GPT-5.4 Pro. For
| standard here is the pricing:
|
| $2.5/M Input Tokens $15/M Output Tokens
| rvz wrote:
| You didn't realize they can increase / change prices for
| intelligence?
|
| This should not be shocking.
| nickthegreek wrote:
| OP made no mention of not understanding cost relation to
| intelligence. In fact, they specifically call out the lack of
| value.
| energy123 wrote:
| For Pro
| joe_mamba wrote:
| Better tokens per dollar could be useless for comparison if the
| model can't solve your problem.
| stri8ted wrote:
| Price Input: $2.50 / 1M tokens Cached input: $0.25 / 1M tokens
| Output: $15.00 / 1M tokens
|
| https://openai.com/api/pricing/
| world2vec wrote:
| Benchmarks barely improved it seems
| cj wrote:
| I use ChatGPT primarily for health related prompts. Looking at
| bloodwork, playing doctor for diagnosing minor aches/pains from
| weightlifting, etc.
|
| Interesting, the "Health" category seems to report worse
| performance compared to 5.2.
| paxys wrote:
| Models are being neutered for questions related to law, health
| etc. for liability reasons.
| cj wrote:
| I'm sometimes surprised how much detail ChatGPT will go into
| without giving any dislaimers.
|
| I very frequently copy/paste the same prompts into Gemini to
| compare, and Gemini often flat out refuses to engage while
| ChatGPT will happily make medical recommendations.
|
| I also have a feeling it has to do with my account history
| and heavy use of project context. It feels like when ChatGPT
| is overloaded with too much context, it might let the
| guardrails sort of slide away. That's just my feeling though.
|
| Today was particularly bad... I uploaded 2 PDFs of bloodwork
| and asked ChatGPT to transcribe it, and it spit out blood
| test results that it found in the project context from an
| earlier date, not the one attached to the prompt. That was
| weird.
| bargainbin wrote:
| Anecdotal, but I asked Claude the other day about how to
| dilute my medication (HCG) and it flat out refused and
| started lecturing me about abusing drugs.
|
| I copy and pasted into ChatGPT, it told me straight away,
| and then for a laugh said it was actually a magical weight
| loss drug that I'd bought off the dark web... And it
| started giving me advice about unregulated weight loss
| drugs and how to dose them.
| staticman2 wrote:
| If you had created a project with custom instructions
| and/ or custom style I think you could have gotten Claude
| to respond the way you wanted just fine.
| tiahura wrote:
| Are you sure about that? Plenty of lawyers that use them
| everyday aren't noticing.
| partiallypro wrote:
| I've done the same, and I tested the same prompts with Claude
| and Google, and they both started hallucinating my blood
| results and supplement stack ingredients. Hopefully this new
| model doesn't fall on this. Claude and Google are dangerously
| unusable on the subject of health, from my experience.
| zeeebeee wrote:
| what's best in your experience? i've always felt like opus
| did well
| wahnfrieden wrote:
| No Codex model yet
| minimaxir wrote:
| GPT-5.4 _is_ the new Codex model.
| wahnfrieden wrote:
| Finally
| nico1207 wrote:
| GPT-5.3-Codex is superior to GPT-5.4 in Terminal Bench with
| Codex, so not really
| conradkay wrote:
| General consensus seems to be that it's still a better
| coding model, overall
| koakuma-chan wrote:
| It just released, how is there a general consensus
| already
| timpera wrote:
| > Steerability: Similarly to how Codex outlines its approach when
| it starts working, GPT-5.4 Thinking in ChatGPT will now outline
| its work with a preamble for longer, more complex queries. You
| can also add instructions or adjust its direction mid-response.
|
| This was definitely missing before, and a frustrating difference
| when switching between ChatGPT and Codex. Great addition.
| yanis_t wrote:
| These releases are lacking something. Yes, they optimised for
| benchmarks, but it's just not all that impressive anymore. It is
| time for a product, not for a marginally improved model.
| esafak wrote:
| That's for you to build; they provide the brains. Do you really
| want one company to build everything? There wouldn't be a
| software industry to speak of if that happened.
| simlevesque wrote:
| Nah, the second you finish your build they release their
| version and then it's game over.
| acedTrex wrote:
| Well they are currently the ones valued at a number with a
| whole lotta 0s on it. I think they should probably do both
| ipsum2 wrote:
| The model was released less than an hour ago, and somehow
| you've been able to form such a strong opinion about it.
| Impressive!
| cj wrote:
| One opinion you can form in under an hour is... why are they
| using GPT-4o to rate the bias of new models?
|
| > assess harmful stereotypes by grading differences in how a
| model responds
|
| > Responses are rated for harmful differences in stereotypes
| using GPT-4o, whose ratings were shown to be consistent with
| human ratings
|
| Are we seriously using old models to rate new models?
| titanomachy wrote:
| Why not? If they've shown that 4o is calibrated to human
| responses, and they haven't shown that yet for 5.4...
| hex4def6 wrote:
| If you're benchmarking something, old & well-characterized
| / understood often beats new & un-characterized.
|
| Sure, there may be shortcomings, but they're well
| understood. The closer you get to the cutting edge, the
| less characterization data you get to rely on. You need to
| be able to trust & understand your measurement tool for the
| results to be meaningful.
| utopiah wrote:
| Benchmarks?
|
| I don't use OpenAI nor even LLMs (despite having tried https:
| //fabien.benetou.fr/Content/SelfHostingArtificialIntel... a
| lot of models) but I imagine if I did I would keep failed
| prompts (can just be a basic "last prompt failed" then
| export) then whenever a new model comes around I'd throw at 5
| it random of MY fails (not benchmarks from others, those will
| come too anyway) and see if it's better, same, worst, for My
| use cases in minutes.
|
| If it's "better" (whatever my criteria might be) I'd also
| throw back some of my useful prompts to avoid regression.
|
| Really doesn't seem complicated nor taking much time to forge
| a realistic opinion.
| earth2mars wrote:
| I am actually super impressed with Codex-5.3 extra high
| reasoning. Its a drop in replacement (infact better than
| Claude Opus 4.6. lately claude being super verbose going in
| circles in getting things resolved). I stopped using claude
| mostly and having a blast with Codex 5.3. looking forward to
| 5.4 in codex.
| satvikpendem wrote:
| Same, it also helps that it's way cheaper than Opus in
| VSCode Copilot, where OpenAI models are counted as 1x
| requests while Opus is 3x, for similar performance (no
| doubt Microsoft is subsidizing OpenAI models due to their
| partnership).
| CryZe wrote:
| I've been using both Opus 4.6 and Codex 5.3 in VSCode's
| Copilot and while Opus is indeed 3x and Codex is 1x, that
| doesn't seem to matter as Opus is willing to go work in
| the background for like an hour for 3 credits, whereas
| Codex asks you whether to continue every few lines of
| code it changes, quickly eating way more credits than
| Opus. In fact Opus in Copilot is probably underpriced, as
| it can definitely work for an hour with just those 12
| cents of cost. Which I'm not sure you get anywhere else
| at such a low price.
|
| Update: I don't know why I can't reply to your reply, so
| I'll just update this. I have tried many times to give it
| a big todo list and told it to do it all. But I've never
| gotten it to actually work on it all and instead after
| the first task is complete it always asks if it should
| move onto the next task. In fact, I always tell it not to
| ask me and yet it still does. So unless I need to do very
| specific prompt engineering, that does not seem to work
| for me.
| satvikpendem wrote:
| That shouldn't really make a difference because you can
| just prompt Codex to behave the same way, having it load
| a big list of todo items perhaps from a markdown file and
| asking it to iterate until it's finished without asking
| for confirmation, and that'll still cost 1x over Opus'
| 3x.
| whynotminot wrote:
| I still love Opus but it's just too expensive / eats usage
| limits.
|
| I've found that 5.3-Codex is mostly Opus quality but
| cheaper for daily use.
|
| Curious to see if 5.4 will be worth somewhat higher costs,
| or if I'll stick to 5.3-Codex for the same reasons.
| braebo wrote:
| I struggle to believe this. Codex can't hold a candle to
| Claude on any task I've given it.
| satvikpendem wrote:
| It's more hedonic adaptation, people just aren't as impressed
| by incremental changes anymore over big leaps. It's the same
| as another thread yesterday where someone said the new
| MacBook with the latest processor doesn't excite them
| anymore, and it's because for most people, most models are
| good enough and now it's all about applications.
|
| https://news.ycombinator.com/item?id=47232453#47232735
| mirekrusin wrote:
| Oh, come on, if it can't run local models that compete with
| proprietary ones it's not good enough yet!
| satvikpendem wrote:
| Qwen 3.5 small models are actually very impressive and do
| beat out larger proprietary models.
| dmix wrote:
| Plus people just really like to whine on the internet
| kranke155 wrote:
| The models are so good that incremental improvements are not
| super impressive. We literally would benefit more from maybe
| sending 50% of model spending into spending on implementation
| into the services and industrial economy. We literally are
| lagging in implementation, specialised tools, and hooks so we
| can connect everything to agents. I think.
| wahnfrieden wrote:
| 5.3 codex was a huge leap over 5.2 for agentic work in
| practice. have you been using both of those or paying attention
| more to benchmark news and chatgpt experience?
| softwaredoug wrote:
| The products are the harnesses, and IMO that's where the
| innovation happens. We've gotten better at helping get good,
| verifiable work from dumb LLMs
| iterateoften wrote:
| The product is putting the skills / harness behind the api
| instead of the agent locally on your computer and iterating on
| that between model updates. Close off the garden.
|
| Not that I want it, just where I imagine it going.
| metalliqaz wrote:
| They need something that _POPS_ : The new GPT
| -- SkyNet for _real_
| jascha_eng wrote:
| When did they stop putting competitor models on the comparison
| table btw? And yeh I mean the benchmark improvements are meh.
| Context Window and lack of real memory is still an issue.
| varispeed wrote:
| The scores increase and as new versions are released they feel
| more and more dumbed down.
| tgarrett wrote:
| Plasma physicist here, I haven't tried 5.4 yet, but in general
| I am very impressed with the recent upgrades that started
| arriving in the fall of 2025: for tasks like manipulating
| analytic systems of equations, quickly developing new features
| for simulation codes, and interpreting and designing
| experiments (with pictures) they have become much stronger.
| I've been asking questions and probing them for several years
| now out of curiosity, and they suddenly have developed deep
| understanding (Gemini 2.5 <<< Gemini 3.1) and become very
| useful. I totally get the current SV vibes, and am becoming a
| lot more ambitious in my future plans.
| brcmthrowaway wrote:
| Youre just chatting yourself out of a job.
| axus wrote:
| Giving the right answer: $1
|
| Asking the right question: $9,999
| slibhb wrote:
| If we don't need plasma physicists anymore then we probably
| have fusion reactors or something, which seems like a fine
| trade. (In reality we're going to want humans in the loop
| for for the forseeable future)
| mindwok wrote:
| They don't need to be impressive to be worthwhile. I like
| incremental improvements, they make a difference in the day to
| day work I do writing software with these.
| prydt wrote:
| I no longer want to support OpenAI at all. Regardless of
| benchmarks or real world performance.
| Imustaskforhelp wrote:
| I agree with ya. You aren't alone in this. For what its worth,
| Chatgpt subscriptions have been cancelled or that number has
| risen ~300% in the last month.
|
| Also, Anthropic/Gemini/even Kimi models are pretty good for
| what its worth. I used to use chatgpt and I still sometimes
| accidentally open it but I use Gemini/Claude nowadays and I
| personally find them to be better anyways too.
| throwaway911282 wrote:
| google and anthropic have govt contracts long before openai..
| if you are taking a stance you should rather use oss models
| zeeebeee wrote:
| that aside, chatgpt itself has gone downhill so much and i know
| i'm not the only one feeling this way
|
| i just HATE talking to it like a chatbot
|
| idk what they did but i feel like every response has been the
| same "structure" since gpt 5 came out
|
| feels like a true robot
| tototrains wrote:
| Their trajectory was clear the moment they signed a deal with
| Microsoft if not sooner.
|
| Absolute snakes - if it's more profitable to manipulate you
| with outputs or steal your work, they will. Every cent and byte
| of data they're given will be used to support authoritarianism.
| beernet wrote:
| Sam really fumbled the top position in a matter of months, and
| spectacularly so. Wow. It appears that people are much more
| excited by Anthropic and Google releases, and there are good
| reasons for that which were absolutely avoidable.
| jcmontx wrote:
| 5.4 vs 5.3-Codex? Which one is better for coding?
| vtail wrote:
| Looking at the benchmarks, 5.4 is slightly better. But it also
| offers "Fast" mode (at 2x usage), which - if it works and
| doesn't completely depletes my Pro plan - is a no brainer at
| the same or even slightly worse quality for more interactive
| development.
| esafak wrote:
| For the price, it seems the latter. I'd use 5.4 to plan.
| embedding-shape wrote:
| Literally just released, I don't think anyone knows yet. Don't
| listen to people's confident takes until after a week or two
| when people actually been able to try it, otherwise you'll just
| get sucked up in bears/bulls misdirected "I'm first with an
| opinion".
| awestroke wrote:
| Opus 4.6
| jcmontx wrote:
| Codex surpassed Claude in usefulness _for me_ since last
| month
| baal80spam wrote:
| Uh, oh. Looks like Claude sycophants joined linuxers and
| vegetarians.
| Someone1234 wrote:
| Related question:
|
| - Do they have the same context usage/cost particularly in a
| plan?
|
| They've kept 5.3-Codex along with 5.4, but is that just for
| user-preference reasons, or is there a trade-off to using the
| older one? I'm aware that API cost is better, but that isn't
| 1:1 with plan usage "cost."
| gavinray wrote:
| The "RPG Game" example on the blogpost is one of the most
| impressive demo's of autonomous engineering I've seen.
|
| It's very similar to "Battle Brothers", and the fact that RPG
| games require art assets, AI for enemy moves, and a host of other
| logical systems makes it all the more impressive.
| hu3 wrote:
| indeed and I suspect it can be attributed to, at least in part,
| the improved playwright integration.
|
| > we're also releasing an experimental Codex skill called
| "Playwright (Interactive) (opens in a new window)". This allows
| Codex to visually debug web and Electron apps; it can even be
| used to test an app it's building, as it's building it.
| casid wrote:
| I don't know. It looks shallow and simple, not even a demo.
| Multicomp wrote:
| A cheesy Roller Coaster Tycoon clone in a browser, one-shotted
| from an AI? Amazing capabilities. The entire "low code drag n
| drop" market like YoYoGames Game Maker and RPG Maker should be
| ready to pack it in soon if this keeps improving in this way.
| swingboy wrote:
| Even with the 1m context window, it looks like these models drop
| off significantly at about 256k. Hopefully improving that is a
| high priority for 2026.
| leftbehinds wrote:
| some sloppy improvements
| HardCodedBias wrote:
| We'll have to wait a day or two, maybe a week or two, to
| determine if this is more capable in coding than 5.3, which seems
| to be the economically valuable capability at this time.
|
| In terms of writing and research even Gemini, with a good prompt,
| is close to useable. That's likely not a differentiator.
| lostmsu wrote:
| What is Pro exactly and is it available in Codex CLI?
| akmarinov wrote:
| It's not. It's their ultra thinking model that's really good
| but takes 40 minutes to come up with an answer
| fy20 wrote:
| It's available on OpenRouter. $180/1M output....
|
| https://openrouter.ai/openai/gpt-5.4-pro
| nickandbro wrote:
| Beat Simon Willison ;)
|
| https://www.svgviewer.dev/s/gAa69yQd
|
| Not the best pelican compared to gemini 3.1 pro, but I am sure
| with coding or excel does remarkably better given those are part
| of its measured benchmarks.
| GaggiX wrote:
| This pelican is actually bad, did you use xhigh?
| nickandbro wrote:
| yep, just double checked used gpt-5.4 xhigh. Though had to
| select it in codex as don't have access to it on the chatgpt
| app or web version yet. It's possible that whatever code
| harness codex uses, messed with it.
| nubg wrote:
| this is proof they are not benchmaxxing the pelican's :-)
| bazmattaz wrote:
| Anyone else feel that it's exhausting keeping up with the pace of
| new model releases. I swear every other week there's a new
| release!
| coffeemug wrote:
| Why do you need to keep up? Just use the latest models and
| don't worry about it.
| throwup238 wrote:
| Yes, that's a common feeling. 5.3-Codex was released a month
| ago on Feb 5 so we're not even getting a full month within a
| single brand, let alone between competitors.
| davnicwil wrote:
| If you think about it there shouldn't really be a reason to
| care as long as things don't get worse.
|
| Presumably this is where it'll evolve to with the product just
| being the brand with a pricing tier and you always get {latest}
| within that, whatever that means (you don't have to care). They
| could even shuffle models around internally using some sort of
| auto-like mode for simpler questions. Again why should I care
| as long as average output is not subjectively worse.
|
| Just as I don't want to select resources for my SaaS software
| to use or have that explictly linked to pricing, I don't want
| to care what my OpenAI model or Anthropic model is today, I
| just want to pay and for it to hopefully keep getting better
| but at a minimum not get worse.
| pupppet wrote:
| I think it's fun, it's like we're reliving the browser wars of
| the early days.
| dandiep wrote:
| Anyone know why OpenAI hasn't released a new model for fine
| tuning since 4.1? It'll be a year next month since their last
| model update for fine tuning.
| qoez wrote:
| I think they just did that because of the energy around it for
| open source models. Their heart probably wasn't in it and the
| amount of people fine tuning given the prices were probably too
| low to continue putting in attention there.
| zzleeper wrote:
| For me the issue is why there's not a new mini since 5-mini in
| August.
|
| I have now switched web-related and data-related queries to
| Gemini, coding to Claude, and will probably try QWEN for less
| critical data queries. So where does OpenAI fits now?
| Rapzid wrote:
| Also interested in this and a replacement for 4.1/4.1-mini that
| focuses on low latency and high accuracy for voice
| applications(not the all-in-one models).
| paxys wrote:
| "Here's a brand new state-of-the-art model. It costs 10x more
| than the previous one because it's just _so good_. But don 't
| worry, if you don't want all this power you can continue to use
| the older one."
|
| A couple months later:
|
| "We are deprecating the older model."
| OutOfHere wrote:
| That's a misrepresentation of the cost. It is simply false. The
| cost is noted here:
| https://news.ycombinator.com/item?id=47265144
| oytis wrote:
| Everyone is mindblown in 3...2...1
| OutOfHere wrote:
| What is with the absurdity of skipping "5.3 Thinking"?
| vicchenai wrote:
| Honestly at this point I just want to know if it follows complex
| instructions better than 5.1. The benchmark numbers stopped
| meaning much to me a while ago - real usage always feels
| different.
| 7777777phil wrote:
| 83% win rate over industry professionals across 44 occupations.
|
| I'd believe it on those specific tasks. Near-universal adoption
| in software still hasn't moved DORA metrics. The model gets
| better every release. The output doesn't follow. Just had a
| closer look on those productivity metrics this week:
| https://philippdubach.com/posts/93-of-developers-use-ai-codi...
| NiloCK wrote:
| This March 2026 blog post is citing a 2025 study based on
| Sonnet 3.5 and 3.7 usage.
|
| Given that organization who ran the study [1] has a _terrifying
| exponential_ as their landing page, I think they 'd prefer that
| it's results are interpreted as a snapshot of something moving
| rather than a constant.
|
| [1] - https://metr.org/
| 7777777phil wrote:
| Good catch, thanks (I really wrote that myself.) Added a note
| to the post acknowledging the models used were Claude 3.5 and
| 3.7 Sonnet.
| twitchard wrote:
| Not sure DORA is that much of an indictment. For "Change
| Failure Rate" for instance these are subject to tradeoffs.
| Organizations likely have a _tolerance level_ for Change
| Failure Rate. If changes are failing too often they slow down
| and invest. If changes aren 't failing that much they speed up
| -- and so saying "change failure rate hasn't decreased,
| obviously AI must not be working" is a little silly.
|
| "Change Lead Time" I would expect to have sped up although I
| can tell stories for why AI-assisted coding would have an
| indeterminate effect here too. Right now at a lot of orgs, the
| bottle neck is the _review process_ because AI is so good at
| producing complete draft PRs quickly. Because reviews are
| scarce (not just reviews but also manual testing passes are
| scarce) this creates an incentive ironically to group changes
| into larger batches. So the definition of what a "change" is
| has grown too.
| rbitar wrote:
| I think the most exciting change announced here is the use of
| tool search to dynamically load tools as needed:
| https://developers.openai.com/api/docs/guides/tools-tool-sea...
| alpineman wrote:
| No thanks. Already cancelled my sub.
| OsrsNeedsf2P wrote:
| Does anyone know what website is the "Isometric Park Builder"
| shown off here?
| turblety wrote:
| They build that using GPT-5.4
|
| > Theme park simulation game made with GPT-5.4 from a single
| lightly specified prompt
|
| GPT literally built that game.
| iamleppert wrote:
| I wouldn't trust any of these benchmarks unless they are
| accompanied by some sort of proof other than "trust me bro". Also
| not including the parameters the models were run at (especially
| the other models) makes it hard to form fair comparisons. They
| need to publish, at minimum, the code and runner used to complete
| the benchmarks and logs.
|
| Not including the Chinese models is also obviously done to make
| it appear like they aren't as cooked as they really are.
| elmean wrote:
| Wow insane improvements in targeting systems for military targets
| over children
| timedude wrote:
| Absolutely amazing. Grateful to be living in this timeframe
| bramhaag wrote:
| What makes you think that they see bombing civilians as a bug,
| not a feature?
| elmean wrote:
| first real comment, I thought that at first but this could
| lower the possible users that could be using chatGPT and that
| would be against us (shareholders)
| skilltissue wrote:
| Don't use the site this way.
|
| https://news.ycombinator.com/newsguidelines.html
| patcon wrote:
| Not all rule-following is noble or wise.
| Chance-Device wrote:
| You made a burner account just to scold this guy? Don't use
| burner accounts this way.
| himata4113 wrote:
| news _guidelines_
| adamtaylor_13 wrote:
| Parlay?
| louiereederson wrote:
| I think for your comment to follow the guidelines, you need
| to explain why the original comment did not follow them.
|
| Customer values are relevant to the discussion given that
| they impact choice and therefore competition.
| elmean wrote:
| AINT NO PARTY LIKE A GARRY TAN HOT TUB PARTY
| Chance-Device wrote:
| Ironically this would actually be a good thing. As we can see
| from Iran Claude doesn't quite have these bugs ironed out
| yet...
| MSFT_Edging wrote:
| This is the exact attitude that lead to a chat bot being used
| to identify a school for girls as a valid target.
|
| The chatbot cannot be held responsible.
|
| Whoever is using chatbots for selecting targets is
| incompetent and should likely face war crime charges.
| Chance-Device wrote:
| What attitude exactly are you talking about? The one that
| says that if you're going to morally sell out it would be
| better if you at least _tried_ not to kill children?
| bananamogul wrote:
| "that lead to a chat bot being used to identify a school
| for girls as a valid target"
|
| Has it been stated authoritatively somewhere that this was
| an AI-driven mistake?
|
| There are myrid ways that mistake could have been made that
| don't require AI. These kinds of mistakes were certainly
| made by all kinds of combatants in the pre-AI era.
| Chance-Device wrote:
| Do you think anyone is ever going to say this under any
| circumstances? That Anthropic were right and they were
| proved right the very next day?
|
| Yeah yeah, they probably had a human in the loop, that's
| not really the point though.
| Sabinus wrote:
| Targeting and accuracy mistakes happen plenty in wars
| that aren't assisted by AI. I don't think it's fair to
| assume that AI had a hand in the bombing of the school
| without evidence.
| spiralcoaster wrote:
| This is the low quality reddit-style garbage that gets upvoted
| on HN these days?
| esalman wrote:
| While low quality, it is extremely important, potentially
| historically significant too.
| Someone1234 wrote:
| If it is actually _that_ important, then maybe more effort
| should be made so it isn 't "low quality." Cannot be very
| important to _them_ if they 're disinterested in presenting
| an intellectually compelling argument about it.
|
| PS - If you think I am not sympathetic to what they're
| raising, you're very much mistake. But they're not winning
| anyone _new_ over their side with this flamebait.
| Sabinus wrote:
| You can say your piece about how you don't like OpenAI
| working with the US military on lethal AI without making
| Reddit style quips.
| mycall wrote:
| True and simply vote it down.
| elmean wrote:
| mycall would also be to do the same
| karmasimida wrote:
| As programmers become intelligently irrelevant in the whole
| picture, you would see more posts like this
| elmean wrote:
| "This account belongs to a lazy person" true
| rd wrote:
| Noticeably yes much more than usual. It's quite bad. I need
| to start blocking accounts.
| zarzavat wrote:
| What _are_ we supposed to talk about in this thread exactly?
| The developers of this model are evil. Are we supposed to
| just write dry comments about benchmarks while OpenAI
| condones their models being deployed for autonomously killing
| people?
|
| Yes I'm sure it makes a very nice bicycle SVG. I will be sure
| to ask the OpenAI killbots for a copy when they arrive at my
| house.
| elmean wrote:
| I was just reading the model card...
| Nicholas_C wrote:
| The HN of old is no more unfortunately. Things get up or down
| voted based purely on political alignment.
| oklahomasports wrote:
| Evidence
| throwaway911282 wrote:
| what a thoughtful comment! HN is so low quality these days
| creamyhorror wrote:
| I've only used 5.4 for 1 prompt _(edit: 3@high now)_ so far
| (reasoning: extra high, took really long), and it was to analyse
| my codebase and write an evaluation on a topic. But I found its
| writing and analysis thoughtful, precise, and surprisingly
| clearly written, unlike 5.3-Codex. It feels very lucid and uses
| human phrasing.
|
| It might be my AGENTS.md requiring clearer, simpler language, but
| at least 5.4's doing a good job of following the guidelines.
| 5.3-Codex wasn't so great at simple, clear writing.
| irishcoffee wrote:
| > It might be my AGENTS.md requiring clearer, simpler language
|
| If you gave the exact same markdown file to me and I posted ed
| the exact same prompts as you, would I get the same results?
| m3kw9 wrote:
| you probably can't and asking agents.md to "make it clearer"
| will likely give you the illusion of clearer language without
| actual well structured tests. agents.md is to usually change
| what the llm should focus on doing more that suits you. Not
| to say stuff like "be better", "make no mistakes"
| creamyhorror wrote:
| I'm not sure if the model (under its temperature/other
| settings) produces deterministic responses. But I do think
| models' style and phrasing are fairly changeable via
| AGENTS.md-style guidelines.
|
| 5.4's choice of terms and phrasing is very precise and
| unambiguous to me, whereas 5.3-Codex often uses jargon and
| less precise phrases that I have to ask further about or
| demand fuller explanations for via AGENTS.md.
| irishcoffee wrote:
| So sharing markdown files is functionally useless, or no?
| sampton wrote:
| That's been my experience as well switching from Opus to Codex.
| Reasoning takes longer but answers are precise. Claude is
| sloppy in comparison.
| throwaway911282 wrote:
| codex has been really good so far and the fast mode is cherry
| on top! and the very generous limits is another cherry on top
| solenoid0937 wrote:
| Weird, I have had the opposite experience. Codex is good at
| doing precisely what I tell it to do, Opus suggests well
| thought out plans even if it needs to push back to do it.
| pembrook wrote:
| The latest research these days is that including an AGENTS.md
| file only makes outcomes worse with frontier models.
| madeofpalk wrote:
| :(
|
| how can i get claude to always make sure it prettier-s and
| lints changes before pushing up the pr though?
| XCSme wrote:
| Seems to be quite similar to 5.3-codex, but somehow almost 2x
| more expensive: https://aibenchy.com/compare/openai-
| gpt-5-4-medium/openai-gp...
| motbus3 wrote:
| Sam Altman can keep his model intentionally to himself. Not doing
| business with mass murderers
| smoody07 wrote:
| Surprised to see every chart limited to comparisons against other
| OpenAI models. What does the industry comparison look like?
| aydyn wrote:
| They compare to Claude and Gemini in their tweet
| 0123456789ABCDE wrote:
| https://artificialanalysis.ai should have the numbers soon
| lorenzoguerra wrote:
| I believe that this choice is due to two main reasons. First,
| it's (obviously) a marketing strategy to keep the spotlight on
| their own models, showing they're constantly improving and
| avoiding validating competitors. Second, since the community
| knows that static benchmarks are unreliable, it makes sense for
| them to outsource the comparisons to independent leaderboards,
| which lets them avoid accusations of cherry-picking while
| justifying their marketing strategy.
|
| Ultimately, the people actually interested in the performance
| of these models already don't trust self-reported comparisons
| and wait for third-party analysis anyway
| throwaway911282 wrote:
| https://xcancel.com/OpenAI/status/2029620619743219811 you can
| see comparisons here
| jstummbillig wrote:
| Inline poll: What reasoning levels do you work with?
|
| This becomes increasingly less clear to me, because the more
| interesting work will be the agent going off for 30mins+ on high
| / extra high (it's mostly one of the two), and that's a long time
| to wait and an unfeasible amount of code to a/b
| bob1029 wrote:
| I was just testing this with my unity automation tool and the
| performance uplift from 5.2 seems to be substantial.
| koakuma-chan wrote:
| Anyone else getting artifacts when using this model in Cursor?
|
| numerusformassistant to=functions.ReadFile meknabanowt`yown Tian
| Tian Ai Cai Piao Wang Zhan json {"path":
| mike_hearn wrote:
| I've seen that problem with 5.3-codex too, it didn't happen
| with earlier models.
|
| Looks like some kind of encoding misalignment bug. What you're
| seeing is their Harmony output format (what the model actually
| creates). The Thai/Chinese characters are special tokens
| apparently being mismapped to Unicode. Their servers are
| supposed to notice these sequences and translate them back to
| API JSON but it isn't happening reliably.
| daft_pink wrote:
| I've officially got model fatigue. I don't care anymore.
| zeeebeee wrote:
| same same same
| postalrat wrote:
| I'd suggest not clicking for things you don't care about.
| hmokiguess wrote:
| They hired the dude from OpenClaw, they had Jony Ive for a while
| now, give us something different!
| kgeist wrote:
| >Today, we're releasing <..> GPT-5.3 Instant
|
| >Today, we're releasing GPT-5.4 in ChatGPT (as GPT-5.4 Thinking),
|
| >Note that there is not a model named GPT-5.3 Thinking
|
| They held out for eight months without a confusing numbering
| scheme :)
| gallerdude wrote:
| Tbf there was a 5.3 codex
| XCSme wrote:
| What I'm most confused, is why call it both GPT-5.3 Instant and
| gpt-5.3-chat?
| m3kw9 wrote:
| instant kind of suck if you asking more than summerizations,
| surface info, web searches, it can lose track of who's who
| quickly in some complex multi turn asks. Just need to know what
| to use instant for.
| __jl__ wrote:
| What a model mess!
|
| OpenAI now has three price points: GPT 5.1, GPT 5.2 and now GPT
| 5.4. There version numbers jump across different model lines with
| codex at 5.3, what they now call instant also at 5.3.
|
| Anthropic are really the only ones who managed to get this under
| control: Three models, priced at three different levels. New
| models are immediately available everywhere.
|
| Google essentially only has Preview models! The last GA is 2.5.
| As a developer, I can either use an outdated model or have zero
| insurances that the model doesn't get discontinued within weeks.
| arthurcolle wrote:
| There is a lot of opportunity here for the AI infrastructure
| layer on top of tier-1 model providers
| motoxpro wrote:
| This is what clouds like AWS, Azure, and GCP solve (vertex
| AI, etc). They are already an abstraction on top of the model
| makers with distribution built in.
|
| I also don't believe there is any value in trying to
| aggregate consumers or businesses just to clean up model
| makers names/release schedule. Consumers just use the
| default, and businesses need clarity on the underlying change
| (e.g. why is it acting different? Oh google released 3.6)
| arthurcolle wrote:
| Do the end users really care about the models at all, or
| about the effects that the models can cause?
| delaminator wrote:
| two great problems in computing
|
| naming things
|
| cache invalidation
|
| off by one errors
| rurban wrote:
| Biggest problem right now in computing:
|
| Out of tokens until end of month
| strongpigeon wrote:
| > Google essentially only has Preview models! The last GA is
| 2.5. As a developer, I can either use an outdated model or have
| zero insurances that the model doesn't get discontinued within
| weeks.
|
| What's funny is that there is this common meme at Google: you
| can either use the old, unmaintained tool that's used
| everywhere, or the new _beta_ tools that doesn 't quite do what
| you want.
|
| Not quite the same, but it did remind me of it.
| jakub_g wrote:
| "Everything is beta or deprecated."
| fhrow4484 wrote:
| https://static0.anpoimages.com/wordpress/wp-
| content/uploads/...
| yieldcrv wrote:
| Preview Road (only choice, and last preview was deprecated
| without warning)
| CactusBlue wrote:
| Reminds of Unity features
| madeofpalk wrote:
| oh is this about my workplace?
| L-four wrote:
| Gmail was in beta for 5 years, until 2009.
| metalliqaz wrote:
| "Gemini, translate 'beta' from Googlespeak to English."
|
| "Ok, here is the translation:" 'we don't
| want to offer support'
| cyanydeez wrote:
| Nah, it's "We dont want to provide a consistent model
| that we'll be stuck with supporting for a decade because
| it just takes up space; until we run everyone out of
| business, we can't afford to have customers tying their
| systems to any given model"
|
| Really, the economics makes no sense, but that's what
| they're doing. You can't have a consistent model because
| it'll pin their hardware & software, and that costs
| money.
| solarkraft wrote:
| Just like any Google product then.
| m_fayer wrote:
| My 5ish years in the mines of Android native back in the day
| are not years I recall fondly. Never change, Google.
| cyanydeez wrote:
| The business models of LLMs don't include any garuntee, and
| some how that's fine for a burgeoning decade of trillions of
| dollars of consumption.
|
| Sure, makes total sense guys.
| embedding-shape wrote:
| > OpenAI now has three price points: GPT 5.1, GPT 5.2 and now
| GPT 5.4.
|
| I guess that's true, but geared towards API users.
|
| Personally, since "Pro Mode" became available, I've been on the
| plan that enables that, and it's one price point and I get
| access to everything, including enough usage for codex that
| someone who spends a lot of time programming, never manage to
| hit any usage limits although I've gotten close once to the new
| (temporary) Spark limits.
| 0xbadcafebee wrote:
| > or have zero insurances that the model doesn't get
| discontinued within weeks
|
| Why are you using the same model after a month? Every month a
| better model comes out. They are all accessible via the same
| API. You can pay per-token. This is the first time in, like,
| all of technology history, that a useful paid service is so
| interoperable between providers that switching is as easy as
| changing a URL.
| phainopepla2 wrote:
| If you're trying to use LLMs in an enterprise context, you
| would understand. Switching models sometimes requires
| tweaking prompts. That can be a complete mess, when there are
| dozens or hundreds of prompts you have to test.
| hobofan wrote:
| That's true only in theory, but not in practice. In practice
| every inference provider handles errors (guardrails, rate
| limits) somewhat differently and with different quirks, some
| of which only surface in production usage, and Google is one
| of the worst offenders in that regard.
| Aurornis wrote:
| > What a model mess! OpenAI now has three price points: GPT
| 5.1, GPT 5.2 and now GPT 5.4.
|
| I don't know, this feels unnecessarily nitpicky to me
|
| It isn't hard to understand that 5.4 > 5.2 > 5.1. It's not hard
| to understand that the dash-variants have unique properties
| that you want to look up before selecting.
|
| Especially for a target audience of software engineers skipping
| a version number is a common occurrence and never questioned.
| Melatonic wrote:
| Agreed - and its a huge step up from their previous naming
| schemes. That stuff was confusing as hell
| __jl__ wrote:
| I see your point. I do find Anthropic's approach more clean
| though particularly when you add in mini and nano. That
| makes 5 models priced differently. Some share the same core
| name, others don't: gpt 5 nano, gpt 5 mini, gpt 5.1, gpt
| 5.2, gpt 5.4. And we are not even talking about thinking
| budget.
|
| But generally: These are not consumer facing products and I
| agree that someone who uses the API should be able to
| figure out the price point of different models.
| raincole wrote:
| They aggressively retire models, so GPT 5.1 and 5.2 are
| probably going to go soon.
| hobofan wrote:
| In the Azure Foundry, they list GPT 5.2 retirement as "No
| earlier than 2027-05-12" (it might leave OpenAIs normal API
| earlier than that). I'm pretty certain that Gemini 3, which
| isn't even in GA yet will be retired earlier than that.
| CobrastanJorji wrote:
| > Google essentially only has Preview models.
|
| It's really nice to see Google get back to its roots by
| launching things only to "beta" and then leaving them there for
| years. Gmail was "beta" for at least five years, I think.
| FINDarkside wrote:
| Also, GCP Cloud Run domain mapping, pretty fundamental
| feature for cloud product, has been in "preview" for over 5
| years now.
| m3kw9 wrote:
| thats how they had it for years, is a mess, but controlled
| biophysboy wrote:
| Wow, is that what preview means? I see those model options in
| github copilot (all my org allows right now) - I was under the
| impression that preview means a free trial or a limited # of
| queries. Kind of a misleading name..
| jbonatakis wrote:
| Google is already sending notices that the 2.5 models will be
| deprecated soon while all the 3.x models are in preview. It
| really is wild and peak Google.
| boringg wrote:
| Like building on quicksand for dependencies. I guess though
| the argument is that the foundation gets stronger over time
| woeirua wrote:
| Feels incremental. Looks like OpenAI is struggling.
| throwaway5752 wrote:
| Does this model autonomously kill people without human approval
| or perform domestic surveillance of US citizens?
| smusamashah wrote:
| I only want to see how it performs on the Bullshit-benchmark
| https://petergpt.github.io/bullshit-benchmark/viewer/index.v...
|
| GPT is not even close yo Claude in terms of responding to BS.
| zone411 wrote:
| Results from my Extended NYT Connections benchmark:
|
| GPT-5.4 extra high scores 94.0 (GPT-5.2 extra high scored 88.6).
|
| GPT-5.4 medium scores 92.0 (GPT-5.2 medium scored 71.4).
|
| GPT-5.4 no reasoning scores 32.8 (GPT-5.2 no reasoning scored
| 28.1).
| consumer451 wrote:
| I am very curious about this:
|
| > Theme park simulation game made with GPT-5.4 from a single
| lightly specified prompt, using Playwright Interactive for
| browser playtesting and image generation for the isometric asset
| set.
|
| Is "Playwright Interactive" a skill that takes screenshots in a
| tight loop with code changes, or is there more to it?
| motza wrote:
| No doubt this was released early to ease the bad press
| butILoveLife wrote:
| Anyone else completely not interested? Since GPT5, its been cost
| cutting measure after cost cutting measure.
|
| I imagine they added a feature or two, and the router will
| continue to give people 70B parameter-like responses when they
| dont ask for math or coding questions.
| Philip-J-Fry wrote:
| I find it quite funny how this blog post has a big "Ask ChatGPT"
| box at the bottom. So you might think you could ask a question
| about the contents of the blog post, so you type the text
| "summarise this blog post". And it opens a new chat window with
| the link to the blog post followed by "summarise this blog post".
| Only to be told "I can't access external URLs directly, but if
| you can paste the relevant text or describe the content you're
| interested in from the page, I can help you summarize it. Feel
| free to share!"
|
| That's hilarious. Does OpenAI even know this doesn't work?
| Aurornis wrote:
| Probably intentional. They don't want open, no-registration
| endpoints able to trigger the AI into hitting URLs.
| jazzypants wrote:
| But, why include the non-functional chat box in the article?
| observationist wrote:
| They're having service issues - ChatGPT on the web is
| broken for a lot of people. The app is working in android -
| I'd assume that the rollout hit a hitch and the chatbox in
| the article would normally work.
| embedding-shape wrote:
| Different team "manages" the overall blog than the team who
| wrote that specific article. At one point, maybe it made
| sense, then something in the product changed, team that
| manages the blog never tested it again.
|
| Or, people just stopped thinking about any sort of UX.
| These sort of mistakes are all over the place, on literally
| all web properties, some UX flows just ends with you at a
| page where nothing works sometimes. Everything is just
| perpetually "a bit broken" seemingly everywhere I go, not
| specific to OpenAI or even the internet.
| teaearlgraycold wrote:
| If only there was some kind of way to automatically test
| user flows end to end. Perhaps testing could be evaluated
| periodically, or even ran for each code change.
| koakuma-chan wrote:
| There is no business value in doing that.
| colonCapitalDee wrote:
| That's why it happened. It still shouldn't have happened.
| ethbr1 wrote:
| > _Or, people just stopped thinking about any sort of UX.
| These sort of mistakes are all over the place, on
| literally all web properties, some UX flows just ends
| with you at a page where nothing works sometimes._
|
| It's almost like people are vibe coding their web apps or
| something.
| jdndbdjsj wrote:
| Welcome to a big company
| AirGapWorksAI wrote:
| Welcome to a big company where pretty much everyone has
| been working full steam for years, in order to take
| advantage of having a job at a company during a once-in-
| a-lifetime moment.
| m3kw9 wrote:
| what? it's their own site and own llm. I could paste most
| sites and it would work.
| judge2020 wrote:
| Works for me:
| https://rr.judge.sh/Labradorretriever/d6af05/chrome_j9rXJMlf...
| zamadatix wrote:
| Following this process summarizes the blogpost for me. Perhaps
| the difference is I'm signed into my account so it can access
| external URLs or something of that nature?
| pocksuppet wrote:
| Most AI integration is like this. It's not about building
| working products --- it's about bragging that you put a chatbox
| in your program.
| ElijahLynn wrote:
| fwiw: I get a valid response when following the steps you
| mentioned. I do not get the message you mentioned:
|
| https://chatgpt.com/share/69aa0321-8a9c-8011-8391-22861784e8...
|
| EDIT: oh, but I'm logged in, fwiw
| andrewguenther wrote:
| It looks like this doesn't work for users without accounts? It
| works when I'm logged in, but not logged out. I went ahead and
| reported it to the team. Thanks for letting us know!
| baxtr wrote:
| I picked up Claude today after being absent and on ChahGPT and
| Gemini only for a while.
|
| I was pretty impressed with how they've improved user
| experience. If I had to guess, I'd say Anthropic has better
| product people who put more attention to detail in these areas.
| amelius wrote:
| If only they had an LLM they could use as a software testing
| agent.
| Alifatisk wrote:
| So let me get this straight, OpenAi previously had an issue with
| LOTS of different models snd versions being available. Then they
| solved this by introducing GPT-5 which was more like a router
| that put all these models under the hood so you only had to
| prompt to GPT-5, and it would route to the best suitable model.
| This worked great I assume and made the ui for the user
| comprehensible. But now, they are starting to introduce more of
| different models again?
|
| We got:
|
| - GPT-5.1
|
| - GPT-5.2 Thinking
|
| - GPT-5.3 (codex)
|
| - GPT-5.3 Instant
|
| - GPT-5.4 Thinking
|
| - GPT-5.4 Pro
|
| Who's to blame for this ridiculous path they are taking? I'm so
| glad I am not a Chat user, because this adds so much unnecessary
| cognitive load.
|
| The good news here is the support for 1M context window, finally
| it has caught up to Gemini.
| 361994752 wrote:
| i guess you still have the "auto" as an option to route your
| request
| stainablesteel wrote:
| 5 itself might have solved the problem of having too many
| different models somewhere in the backend
| sothatsit wrote:
| I much prefer this, we can choose based on our use-cases, and
| people who don't care can still use Auto.
| fernst wrote:
| Now with more and improved domestic espionage capabilities
| senko wrote:
| Just tested it with my version of the pelican test: a minimal RTS
| game implementation (zero-shot in codex cli):
| https://gist.github.com/senko/596a657b4c0bfd5c8d08f44e4e5347...
| (you'll have to download and open the file, sadly GitHub refuses
| to serve it with the correct content type)
|
| This is on the edge of what the frontier models can do. For 5.4,
| the result is better than 5.3-Codex and Opus 4.6. (Edit: nowhere
| near the RPG game from their blog post, which was presumably much
| more specced out and used better engineering setup).
|
| I also tested it with a non-trivial task I had to do on an
| existing legacy codebase, and it breezed through a task that
| Claude Code with Opus 4.6 was struggling with.
|
| I don't know when Anthropic will fire back with their own update,
| but until then I'll spend a bit more time with Codex CLI and GPT
| 5.4.
| Aldipower wrote:
| So did they raised the ridiculous small "per tool call token
| limit" when working with MCP servers? This makes Chat useless...
| I do not care, but my users.
| melbourne_mat wrote:
| Quick: let's release something new that gives the appearance that
| we're still relevant
| gigatexal wrote:
| Is it any good at coding?
| thefounder wrote:
| Is it just me or the price for 5.4 pro is just insane?
| atkrad wrote:
| What is the main difference between this version with the
| previous one?
| brcmthrowaway wrote:
| How much of LLM improvement comes from regular ChatGPT usage
| these days?
| quotemstr wrote:
| GPT 5.4 is one of the most censored models out there.
|
| https://speechmap.ai/models/openai-gpt-5-4
|
| It completes only 29% of controversial requests. It refuses to
| discuss numerous subjects rooted in facts or that reflect views
| of significant portions of the population. It refuses to even
| write a short essay on exactly what, say, Herasight-style generic
| screening or putting weapons in space. It'll argue passionately
| in favor of censoring "lies" online (judged by whom?). 100% of
| the time, it'll write an essay explaining that the US founding
| fathers were hypocrites. It'll argue against you if you suggest
| it's right use violence to prevent theft of your own property or
| that we should fortify our nuclear arsenal.
|
| Agree or disagree, reasonable people can have a range of views of
| these subjects and it is not the place of OpenAI or any lab to
| determine for everyone the right answers to open societal
| questions.
|
| Shame on them for this.
| ltbarcly3 wrote:
| Not a single comparison between 5.4 and Gemini or Claude. OpenAI
| continues to fall further behind.
___________________________________________________________________
(page generated 2026-03-05 23:00 UTC)