[HN Gopher] GPT-5.4
       ___________________________________________________________________
        
       GPT-5.4
        
       https://openai.com/index/gpt-5-4-thinking-system-card/
       https://x.com/OpenAI/status/2029620619743219811
        
       Author : mudkipdev
       Score  : 513 points
       Date   : 2026-03-05 18:08 UTC (4 hours ago)
        
 (HTM) web link (openai.com)
 (TXT) w3m dump (openai.com)
        
       | ignorantguy wrote:
       | it shows a 404 as of now.
        
         | minimaxir wrote:
         | Up now.
         | 
         | The OP has frequently gotten the scoop for new LLM releases and
         | I am curious what their pipeline is.
        
           | Leynos wrote:
           | Guess the URL and post at 10 AM PST on the day of release.
        
           | bdangubic wrote:
           | curl the URL https://openai.com/index/introducing-gpt-5-?
           | until you get 200
        
             | mudkipdev wrote:
             | Probably refresh the api models list every couple minutes
             | instead. No one could have guessed the name of GPT-Codex-
             | Spark
        
       | mattas wrote:
       | "GPT-5.4 interprets screenshots of a browser interface and
       | interacts with UI elements through coordinate-based clicking to
       | send emails and schedule a calendar event."
       | 
       | They show an example of 5.4 clicking around in Gmail to send an
       | email.
       | 
       | I still think this is the wrong interface to be interacting with
       | the internet. Why not use Gmail APIs? No need to do any
       | screenshot interpretation or coordinate-based clicking.
        
         | TheAceOfHearts wrote:
         | I think the desire is that in the long-term AI should be able
         | to use any human-made application to accomplish equivalent
         | tasks. This email demo is proof that this capability is a high
         | priority.
        
         | spongebobstoes wrote:
         | not everything has an API, or API use is limited. some UIs are
         | more feature complete than their APIs
         | 
         | some sites try to block programmatic use
         | 
         | UI use can be recorded and audited by a non-technical person
        
         | Jacques2Marais wrote:
         | I guess a big chunk of their target market won't know how to
         | use APIs.
        
         | satvikpendem wrote:
         | The ideal of REST, the HTML and UI _is_ the API.
        
         | PaulHoule wrote:
         | APIs have never been a gift but rather have always been a take-
         | away that lets you do less than you can with the web interface.
         | It's always been about drinking through a straw, paying NASA
         | prices, and being limited in everything you can do.
         | 
         | But people are intimidated by the complexity of writing web
         | crawlers because management has been so traumatized by the cost
         | of making GUI applications that they couldn't believe how cheap
         | it is to write crawlers and scrapers.... Until LLMs came along,
         | and changed the perceived economics and created a permission
         | structure. [1]
         | 
         | AI is a threat to the "enshittification economy" because it
         | lets us route around it.
         | 
         | [1] that high cost of GUI development is one reason why
         | scrapers are cheap... there is a good chance that the scraper
         | you wrote 8 years ago still works because (a) they can't afford
         | to change their site and (b) if they could afford to change
         | their site changing anything substantial about it is likely to
         | unrecoverably tank their Google rankings so they won't. A.I.
         | might change the mechanics of that now that you Google traffic
         | is likely to go to zero no matter what you do.
        
           | disqard wrote:
           | > AI is a threat to the "enshittification economy" because it
           | lets us route around it.
           | 
           | This is prescient -- I wonder if the Big Tech entities see it
           | this way. Maybe, even if they do, they're 100% committed to
           | speedrunning the current late-stage-cap wave, and therefore
           | unable to do anything about it.
        
             | PaulHoule wrote:
             | They are not a single thing.
             | 
             | Google has a good model in the form of Gemini and they
             | might figure they can win the AI race and if the web dies,
             | the web dies. YouTube will still stick around.
             | 
             | Facebook is not going to win the AI race with low I.Q.
             | Llama but Zuck believed their business was cooked around
             | the time it became a real business because their users
             | would eventually age out and get tired of it. If I was him
             | I'd be investing in anything that isn't cybernetic let it
             | be gold bars or MMA studios.
             | 
             | Microsoft? They bought Activision for $69 billion. I just
             | can't explain their behavior rationally but they could do
             | worse than their strategy of "put ChatGPT in front of
             | laggards and hope that some of them rise to the challenge
             | and become slop producers."
             | 
             | Amazon is really a bricks-and-mortar play which has the
             | freedom to invest in bricks-and-mortar because investors
             | don't think they are a bricks-and-mortar play.
             | 
             | Netflix? They're cooked as is all of Hollywood. Hollywood's
             | gatekeeping-industrial strategy of producing as few
             | franchise as possible will crack someday and our media
             | market may wind up looking more like Japan, where somebody
             | can write a low-rent light novel like
             | 
             | https://en.wikipedia.org/wiki/Backstabbed_in_a_Backwater_Du
             | n...
             | 
             | and J.C. Staff makes a terrible anime that convinces 20k
             | Otaku to drop $150 on the light novels and another $150 on
             | the manga (sorry, no way you can make a balanced game based
             | on that premise!) and the cost structure is such that it is
             | profitable.
        
           | lostmsu wrote:
           | > AI is a threat to the "enshittification economy" because it
           | lets us route around it.
           | 
           | I am not sure about that. We techies avoid enshittification
           | because we recognize shit. Normies will just get their
           | syncopatic enshittified AI that will tell them to continue
           | buying into walled gardens.
        
           | Traster wrote:
           | You can buy a Claude Code subscription for $200 bucks and use
           | way more tokens in Claude Code than if you pay for direct API
           | usage. Anthopic decided you can't take your Auth key for
           | Claude code and use it to hit the API via a different tool.
           | They made that business decision, because they thought it was
           | better for them strategically to do that. They're allowed to
           | make that choice as a business.
           | 
           | Plenty of companies make the same choice about their API,
           | they provide it for a specific purpose but they have good
           | business reasons they want you using the website. Plenty of
           | people write webcrawlers and it's been a cat and mouse game
           | for decades for websites to block them.
           | 
           | This will just be one more step in that cat and mouse game,
           | and if the AI really gets good enough to become a complete
           | intermediary between you and the website? The website will
           | just shutdown. We saw it happen before with the open web.
           | These websites aren't here for some heroic purpose, if you
           | screw their business model they will just go out of business.
           | You won't be able to use their website because it won't exist
           | and the website that do exist will either (a) be made by the
           | same guys writing your agent, and (b) be highly highly
           | optimized to get your agent to screw you.
        
         | steve1977 wrote:
         | One could argue that LLMs learning programming languages made
         | for humans (i.e. most of them) is using the wrong interface as
         | well. Why not use machine code?
        
           | embedding-shape wrote:
           | Why would human language by the wrong interface when they're
           | literally language models? Why would machine code be better
           | when there is probably magnitude less of training material
           | with machine code?
           | 
           | You can also test this yourself easily, fire up two agents,
           | ask one to use PL meant for humans, and one to write straight
           | up machine code (or assembly even), and see which results you
           | like best.
        
           | BoredPositron wrote:
           | because they are inherently text based as is code?
        
             | steve1977 wrote:
             | But they are abstractions made to cater to human
             | weaknesses.
        
           | adwn wrote:
           | > _One could argue that LLMs learning programming languages
           | made for humans (i.e. most of them) is using the wrong
           | interface as well._
           | 
           | Then go ahead and make an argument. "Why not do X?" is not an
           | argument, it's a suggestion.
        
         | jstummbillig wrote:
         | Because the web and software more generally if full of not APIs
         | and you do, in fact, need the clicking to work to make agents
         | work generally
        
         | modeless wrote:
         | A world where AIs use APIs instead of UIs to do everything is a
         | world where us humans will soon be helpless, as we'll have to
         | ask the AIs to do everything for us and will have limited
         | ability to observe and understand their work. I prefer that the
         | AIs continue to use human-accessible tools, even if that's less
         | efficient for them. As the price of intelligence trends toward
         | zero, efficiency becomes relatively less important.
        
         | npilk wrote:
         | It feels like building humanoid robots so they can use tools
         | built for human hands. Not clear if it will pay off, but if it
         | does then you get a bunch of flexibility across any task "for
         | free".
         | 
         | Of course APIs and CLIs also exist, but they don't necessarily
         | have feature parity, so more development would be needed. Maybe
         | that's the future though since code generation is so good - use
         | AI to build scaffolding for agent interaction into every
         | product.
        
           | packetlost wrote:
           | I don't see how an API couldn't have full parity with a web
           | interface, the API is how you actually trigger a state
           | transition in the vast majority of cases
        
         | coffeemug wrote:
         | A model that gets good at computer use can be plugged in
         | anywhere you have a human. A model that gets good at API use
         | cannot. From the standpoint of diffusion into the economy/labor
         | market, computer use is much higher value.
        
         | f0e4c2f7 wrote:
         | Lots of services have no desire to ever expose an API. This
         | approach lets you step right over that.
         | 
         | If an API is exposed you can just have the LLM write something
         | against that.
        
         | kristianp wrote:
         | This opens up a new question: how does bot detection work when
         | the bot is using the computer via a gui?
        
           | itintheory wrote:
           | On it's face, I'm not sure that's a new question. Bots using
           | browser automation frameworks (puppeteer, selenium,
           | playwright etc) have been around for a while. There are
           | signals used in bot detection tools like cursor movement
           | speed, accuracy, keyboard timing, etc. How those detection
           | tools might update to support legitimate bot users does seem
           | like an open question to me though.
        
         | MattDaEskimo wrote:
         | Same reason why Wikipedia deals with so many people scraping
         | its web page instead of using their API:
         | 
         | Optimizations are secondary to convenience
        
         | bottlepalm wrote:
         | The vast majority of websites you visit don't have usable APIs
         | and very poor discovery of the those APIs.
         | 
         | Screenshots on the other hand are documentation, API, and
         | discovery all in one. And you'd be surprised how little
         | context/tokens screenshots consumer compared to all the back
         | and forth verbose json payloads of APIs
        
           | LUmBULtERA wrote:
           | >The vast majority of websites you visit don't have usable
           | APIs and very poor discovery of the those APIs.
           | 
           | I think an important thing here is that a lot of
           | websites/platforms don't want AIs to have direct API access,
           | because they are afraid that AIs would take the customer
           | "away" from the website/platform, making the consumer a
           | customer of the AI rather than a customer of the
           | website/platform. Therefore for AIs to be able to do what
           | customers want them to do, they need their browsing to look
           | just like the customer's browsing/browser.
        
       | denysvitali wrote:
       | Article: https://openai.com/index/introducing-gpt-5-4/
       | 
       | gpt-5.4
       | 
       | Input: $2.50 /M tokens
       | 
       | Cached: $0.25 /M tokens
       | 
       | Output: $15 /M tokens
       | 
       | ---
       | 
       | gpt-5.4-pro
       | 
       | Input: $30 /M tokens
       | 
       | Output: $180 /M tokens
       | 
       | Wtf
        
         | elliotbnvl wrote:
         | Looks like it's an order of magnitude off. Missprint?
        
           | GenerWork wrote:
           | Looks like an extra zero was added?
        
             | benlivengood wrote:
             | Government pricing :)
        
               | outside2344 wrote:
               | $30 per kill approval
        
           | glerk wrote:
           | Looks like fair price discovery :)
        
         | dpoloncsak wrote:
         | >" GPT-5.4 is priced higher per token than GPT-5.2 to reflect
         | its improved capabilities"
         | 
         | That's just not how pricing is supposed to work...? Especially
         | for a 'non-profit'. You're charging me more so I know I have
         | the better model?
        
           | elicash wrote:
           | Can't you continue to use to older model, if you prefer the
           | pricing?
           | 
           | But they also claim this new model uses fewer tokens, so it
           | still might ultimately be cheaper even if per token cost is
           | higher.
        
             | dpoloncsak wrote:
             | I'm not against the pricing, just seems uncommon to frame
             | it in the way they did, as opposed to the usual 'assume the
             | customer expects more performance will cost more'
             | 
             | I guess they have to sell to investors that the price to
             | operate is going down, while still needing more from the
             | user to be sustainable
        
             | jbellis wrote:
             | You can, until they turn it off.
             | 
             | Anthropic is pulling the plug on Haiku 3 in a couple
             | months, and they haven't released anything in that price
             | range to replace it.
        
               | Sabinus wrote:
               | Surely there are open source models that surpass Haiku 3
               | at better price points by now.
        
           | FergusArgyll wrote:
           | Maybe it's finally a bigger pretrain?
        
             | dpoloncsak wrote:
             | I feel like that would have been highlighted then. "As this
             | is a bigger pretrain, we have to raise prices".
             | 
             | They're framing it pretty directly "We want you to think
             | bigger cost means better model"
        
       | minimaxir wrote:
       | The marquee feature is obviously the 1M context window, compared
       | to the ~200k other models support with maybe an extra cost for
       | generations beyond >200k tokens. Per the pricing page, there is
       | no additional cost for tokens beyond 200k:
       | https://openai.com/api/pricing/
       | 
       | Also per pricing, GPT-5.4 ($2.50/M input, $15/M output) is much
       | cheaper than Opus 4.6 ($5/M input, $25/M output) and Opus has a
       | penalty for its beta >200k context window.
       | 
       | I am skeptical whether the 1M context window will provide
       | material gains as current Codex/Opus show weaknesses as its
       | context window is mostly full, but we'll see.
       | 
       | Per updated docs
       | (https://developers.openai.com/api/docs/guides/latest-model), it
       | supercedes GPT-5.3-Codex, which is an interesting move.
        
         | thehamkercat wrote:
         | GPT 5.3 codex had 400K context window btw
        
         | simianwords wrote:
         | Why would some one use codex instead?
        
           | embedding-shape wrote:
           | Why would someone use Claude Code instead? Or any other
           | harness? Or why only use one?
           | 
           | My own tooling throws off requests to multiple agents at the
           | same time, then I compare which one is best, and continue
           | from there. Most of the time Codex ends up with the best end
           | results though, but my hunch is that at one point that'll
           | change, hence I continue using multiple at the same time.
        
           | surgical_fire wrote:
           | I've been using Codex for software development personally (I
           | have a ChatGPT account), and I use Claude at work (since it
           | is provided by my employer).
           | 
           | I find both Codex and Claude Opus perform at a similar level,
           | and in some ways I actually prefer Codex (I keep hitting
           | quota limits in Opus and have to revert back to Sonnet).
           | 
           | If your question is related to morality (the thing about US
           | politics, DoD contract and so on)... I am not from the US,
           | and I don't care about its internal politics. I also think
           | both OpenAI and Anthropic are evil, and the world would be
           | better if neither existed.
        
             | simianwords wrote:
             | No my question was why would I use codex over gpt 5.4
        
               | surgical_fire wrote:
               | Ahh, good question. I misunderstood you, apologies.
               | 
               | There's no mention of pricing, quotas and so on. Perhaps
               | Codex will still be preferable for coding tasks as it is
               | tailored for it? Maybe it is faster to respond?
               | 
               | Just speculation on my part. If it becomes redundant to
               | 5.4, I presume it will be sunset. Or maybe they
               | eventually release a Codex 5.4?
        
               | landtuna wrote:
               | 5.3 Codex is $1.75/$14, and 5.4 is $2.50/$15.
        
               | surgical_fire wrote:
               | There you go. It makes perfect sense to keep it around
               | then.
        
             | athrowaway3z wrote:
             | They perform at a somewhat equal level on writing single
             | files. But Codex is absolute garbage at theory of
             | self/others. That quickly becomes frustrating.
             | 
             | I can tell claude to spawn a new coding agent, and it will
             | understand what that is, what it should be told, and what
             | it can approximately do.
             | 
             | Codex on the other hand will spawn an agent and then tell
             | it to continue with the work. It knows a coding agent can
             | do work, but doesn't know how you'd use it - or that it
             | won't magically know a plan.
             | 
             | You could add more scaffolding to fix this, but Claude
             | proves you shouldn't have to.
             | 
             | I suspect this is a deeper model "intelligence" difference
             | between the two, but I hope 5.4 will surprise me.
        
               | surgical_fire wrote:
               | > They perform at a somewhat equal level on writing
               | single files.
               | 
               | That's not the experience I have. I had it do more
               | complex changes spawning multiple files and it performed
               | well.
               | 
               | I don't like using multiple agents though. I don't vibe
               | code, I actually review every change it makes. The
               | bottleneck is my review bandwidth, more agents producing
               | more code will not speed me up (in fact it will slow me
               | down, as I'll need to context switch more often).
        
             | hnsr wrote:
             | > I've been using Codex for software development personally
             | (I have a ChatGPT account), and I use Claude at work (since
             | it is provided by my employer).
             | 
             | Exact same situation here. I've been using both extensively
             | for the last month or so, but still don't really feel
             | either of them is much better or worse. But I have not done
             | large complex features with it yet, mostly just iterative
             | work or small features.
             | 
             | I also feel I am probably being very (overly?) specific in
             | my prompts compared to how other people around me use these
             | agents, so maybe that 'masks' things
        
           | jeswin wrote:
           | When it comes to lengthy non-trivial work, codex is much
           | better but also slower.
        
           | lmeyerov wrote:
           | In our evals for answering cybersecurity incident
           | investigation questions and even autonomously doing the full
           | investigation, gpt-5.2-codex with low reasoning was the clear
           | winner over non-codex or higher reasoning. 2X+ faster, higher
           | completion rates, etc.
           | 
           | It was generally smarter than pre-5.2 so strategically
           | better, and codex likewise wrote better database queries than
           | non-codex, and as it needs to iteratively hunt down the
           | answer, didn't run out the clock by drowning in reasoning.
           | 
           | Video: https://media.ccc.de/v/39c3-breaking-bots-cheating-at-
           | blue-t...
           | 
           | We'll be updating numbers on 5.3 and claude, but basically
           | same thing there. Early, but we were surprised to see codex
           | outperform opus here.
        
           | synergy20 wrote:
           | in my testing codex actually planned worse than claude but
           | coded better once the plan is set, and faster. it is also
           | excellent to cross check claude's work, always finding great
           | weakness each time.
        
             | pmarreck wrote:
             | That's why I think the sweet spot is to write up plans with
             | Claude and then execute them with Codex
        
               | GorbachevyChase wrote:
               | Weird. It used to be the opposite. My own experience is
               | that Claude's behind-the-scenes support is a
               | differentiator for supporting office work. It handles
               | documents, spreadsheets and such much better than anyone
               | else (presumably with server side scripts). Codex feels a
               | bit smarter, but it inserts a lot of checkpoints to keep
               | from running too long. Claude will run a plan to the end,
               | but the token limits have become so small in the last
               | couple months that the $20 pla basically only buys one
               | significant task per day. The iOS app is what makes me
               | keep the subscription.
        
         | tedsanders wrote:
         | Yeah, long context vs compaction is always an interesting
         | tradeoff. More information isn't always better for LLMs, as
         | each token adds distraction, cost, and latency. There's no
         | single optimum for all use cases.
         | 
         | For Codex, we're making 1M context experimentally available,
         | but we're not making it the default experience for everyone, as
         | from our testing we think that shorter context plus compaction
         | works best for most people. If anyone here wants to try out 1M,
         | you can do so by overriding `model_context_window` and
         | `model_auto_compact_token_limit`.
         | 
         | Curious to hear if people have use cases where they find 1M
         | works much better!
         | 
         | (I work at OpenAI.)
        
           | simianwords wrote:
           | Do you maybe want to give us users some hints on what to
           | compact and throw away? In codex CLI maybe you can create a
           | visual tool that I can see and quickly check mark things I
           | want to discard.
           | 
           | Sometimes I'm exploring some topic and that exploration is
           | not useful but only the summary.
           | 
           | Also, you could use the best guess and cli could tell me that
           | this is what it wants to compact and I can tweak its
           | suggestion in natural language.
           | 
           | Context is going to be super important because it is the
           | primary constraint. It would be nice to have serious granular
           | support.
        
           | akiselev wrote:
           | _> Curious to hear if people have use cases where they find
           | 1M works much better!_
           | 
           | Reverse engineering [1]. When decompiling a bunch of code and
           | tracing functionality, it's really easy to fill up the
           | context window with irrelevant noise and compaction generally
           | causes it to lose the plot entirely and have to start almost
           | from scratch.
           | 
           | (Side note, are there any OpenAI programs to get free
           | tokens/Max to test this kind of stuff?)
           | 
           | [1] https://github.com/akiselev/ghidra-cli
        
           | Someone1234 wrote:
           | That's an interesting point regarding context Vs. compaction.
           | If that's viewed as the best strategy, I'd hope we would see
           | more tools around compaction than just "I'll compact what I
           | want, brace yourselves" without warning.
           | 
           | Like, I'd love an optional pre-compaction step, "I need to
           | compact, here is a high level list of my context + size, what
           | should I junk?" Or similar.
        
             | thyb23 wrote:
             | This is exactly how it should work. I imagine it as a tree
             | view showing both full and summarized token counts at each
             | level, so you can immediately see what's taking up space
             | and what you'd gain by compacting it.
             | 
             | The agent could pre-select what it thinks is worth keeping,
             | but you'd still have full control to override it. Each
             | chunk could have three states: drop it, keep a summarized
             | version, or keep the full history.
             | 
             | That way you stay in control of both the context budget and
             | the level of detail the agent operates with.
        
               | Folcon wrote:
               | I do find it really interesting that more coding agents
               | don't have this as an toggleable feature, sometimes you
               | really need this level of control to get useful
               | capability
        
               | Someone1234 wrote:
               | Yep; I've actually had entire jobs essentially fail due
               | to a bad compaction. It lost key context, and it
               | completely altered the trajectory.
               | 
               | I'm now more careful, using tracking files to try to keep
               | it aligned, but more control over compaction regardless
               | would be highly welcomed. You don't ALWAYS need that
               | level of control, but when you do, you do.
        
               | joquarky wrote:
               | I compact myself by having it write out to a file, I
               | prune what's no longer relevant, and then start a new
               | session with that file.
               | 
               | But I'm mostly working on personal projects so my time is
               | cheap.
               | 
               | I might experiment with having the file sections post-
               | processed through a token counter though, that's a great
               | idea.
        
           | gspetr wrote:
           | I have found a bigger context window qute useful when trying
           | to make sense of larger codebases. Generating documentation
           | on how different components interact is better than nothing,
           | especially if the code has poor test coverage.
           | 
           | I've also had it succeed in attempts to identify some non-
           | trivial bugs that spanned multiple modules.
        
           | sillysaurusx wrote:
           | You may want to look over this thread from cperciva:
           | https://x.com/cperciva/status/2029645027358495156
           | 
           | I too tried Codex and found it similarly hard to control over
           | long contexts. It ended up coding an app that spit out
           | millions of tiny files which were technically smaller than
           | the original files it was supposed to optimize, except due to
           | there being millions of them, actual hard drive usage was 18x
           | larger. It seemed to work well until a certain point, and I
           | suspect that point was context window overflow / compaction.
           | Happy to provide you with the full session if it helps.
           | 
           | I'll give Codex another shot with 1M. It just seemed like
           | cperciva's case and my own might be similar in that once the
           | context window overflows (or refuses to fill) Codex seems to
           | lose something essential, whereas Claude keeps it. What that
           | thing is, I have no idea, but I'm hoping longer context will
           | preserve it.
        
             | woadwarrior01 wrote:
             | Please don't post links with tracking parameters
             | (t=jQb...).
             | 
             | https://xcancel.com/cperciva/status/2029645027358495156
        
               | sillysaurusx wrote:
               | Haha. This was the second time in like a year that I've
               | posted a Twitter link, and the second time someone
               | complained. Okay, I'll try to remove those before
               | posting, and I'll edit this one out.
               | 
               | Feels like a losing battle, but hey, the audience is
               | usually right.
        
               | woadwarrior01 wrote:
               | I'm sorry, but it's my pet peeve. If you're on iOS/macOS
               | I built a 100% free and privacy-friendly app to get rid
               | of tracking parameters from hundreds of different
               | websites, not just X/Twitter.
               | 
               | https://apps.apple.com/us/app/clean-links-qr-code-
               | reader/id6...
        
               | sillysaurusx wrote:
               | It works on iOS? That's cool. I'll give it a go.
        
               | pmarreck wrote:
               | So what is your motivation for doing this, incidentally?
               | Can you be explicit about it? I am genuinely curious.
               | 
               | Especially when it's to the point of, you know,
               | nagging/policing people to do it the way you'd prefer,
               | when you could just redirect your router requests from
               | x.com to xcancel.com
        
               | monocularvision wrote:
               | This is great! I have been meaning to implement this sort
               | of thing in my existing Shortcuts flow but I see you
               | already support it in Shortcuts! Thank you for this!
               | 
               | Anywhere I can toss a Tip for this free app?
        
             | FrankBooth wrote:
             | What's the connection with context size in that thread? It
             | seems more like an instruction following problem.
        
           | nowittyusername wrote:
           | Personally what I am more interested about is effective
           | context window. I find that when using codex 5.2 high, I
           | preferred to start compaction at around 50% of the context
           | window because I noticed degradation at around that point.
           | Though as of a bout a month ago that point is now below that
           | which is great. Anyways, I feel that I will not be using that
           | 1 million context at all in 5.4 but if the effective window
           | is something like 400k context, that by itself is already a
           | huge win. That means longer sessions before compaction and
           | the agent can keep working on complex stuff for longer. But
           | then there is the issue of intelligence of 5.4. If its as
           | good as 5.2 high I am a happy camper, I found 5.3 anything...
           | lacking personally.
        
           | asabla wrote:
           | I really don't have any numbers to back this up. But it feels
           | like the sweet spot is around ~500k context size. Anything
           | larger then that, you usually have scoping issues, trying to
           | do too much at the same time, or having having issues with
           | the quality of what's in the context at all.
           | 
           | For me, I would say speed (not just time to first token, but
           | a complete generation) is more important then going for a
           | larger context size.
        
           | lubesGordi wrote:
           | It's funny that the context window size is such a thing
           | still. Like the whole LLM 'thing' is compression. Why can't
           | we figure out some equally brilliant way of handling context
           | besides just storing text somewhere and feeding it to the
           | llm? RAG is the best attempt so far. We need something like a
           | dynamic in flight llm/data structure being generated from the
           | context that the agent can query as it goes.
        
         | netinstructions wrote:
         | People (and also frustratingly LLMs) usually refer to
         | https://openai.com/api/pricing/ which doesn't give the complete
         | picture.
         | 
         | https://developers.openai.com/api/docs/pricing is what I always
         | reference, and it explicitly shows that pricing ($2.50/M input,
         | $15/M output) for tokens _under_ 272k
         | 
         | It is nice that we get 70-72k more tokens before the price goes
         | up (also what does it cost beyond 272k tokens??)
        
           | Flashtoo wrote:
           | > Prompts with more than 272K input tokens are priced at 2x
           | input and 1.5x output for the full session for standard,
           | batch, and flex.
        
             | netinstructions wrote:
             | Thanks, it looks like the pricing page keeps getting
             | updated.
             | 
             | Even right now one page refers to prices for "context
             | lengths under 270K" whereas another has pricing for "<272K
             | context length"
        
         | damsta wrote:
         | There is extra cost for >272K:
         | 
         | > For models with a 1.05M context window (GPT-5.4 and GPT-5.4
         | pro), prompts with >272K input tokens are priced at 2x input
         | and 1.5x output for the full session for standard, batch, and
         | flex.
         | 
         | Taken from
         | https://developers.openai.com/api/docs/models/gpt-5.4
        
           | fragmede wrote:
           | Which, Claude has the same deal. You can get a 1M context
           | window, but it's gonna cost ya. If you run /model in claude
           | code, you get:                   Switch between Claude
           | models. Applies to this session and future Claude Code
           | sessions. For other/previous model names, specify with
           | --model.                     1. Default (recommended)   Opus
           | 4.6 * Most capable for complex work            2. Opus (1M
           | context)        Opus 4.6 with 1M context * Billed as extra
           | usage * $10/$37.50 per Mtok            3. Sonnet
           | Sonnet 4.6 * Best for everyday tasks            4. Sonnet (1M
           | context)      Sonnet 4.6 with 1M context * Billed as extra
           | usage * $6/$22.50 per Mtok            5. Haiku
           | Haiku 4.5 * Fastest for quick answers
        
           | minimaxir wrote:
           | Good find, and that's too small a print for comfort.
        
             | ValentineC wrote:
             | It's also in the linked article:
             | 
             | > GPT-5.4 in Codex includes experimental support for the 1M
             | context window. Developers can try this by configuring
             | model_context_window and model_auto_compact_token_limit.
             | Requests that exceed the standard 272K context window count
             | against usage limits at 2x the normal rate.
        
           | glenstein wrote:
           | Wow, that's diametrically the opposite point: the cost is
           | *extra*, not free.
        
             | apetresc wrote:
             | Diametrically opposite to tokens beyond 200K being
             | _literally_ free? As in, you only pay for the first 200K
             | tokens and the remaining 800K cost $0.00?
             | 
             | I don't think that's a fair reading of the original post at
             | all, obviously what they meant by "no cost" was "no
             | increase in the cost".
        
         | andai wrote:
         | It's a little hard to compare, because Claude needs
         | significantly fewer tokens for the same task. A better metric
         | is the cost per task, which ends up being pretty similar.
         | 
         | For example on Artificial Analysis, the GPT-5.x models' cost to
         | run the evals range from half of that of Claude Opus (at medium
         | and high), to significantly more than the cost of Opus (at
         | extra high reasoning). So on their cost graphs, GPT has a
         | considerable distribution, and Opus sits right in the middle of
         | that distribution.
         | 
         | The most striking graph to look at there is "Intelligence vs
         | Output Tokens". When you account for that, I think the actual
         | costs end up being quite similar.
         | 
         | According to the evals, at least, the GPT extra high matches
         | Opus in intelligence, while costing more.
         | 
         | Of course, as always, benchmarks are mostly meaningless and you
         | need to check Actual Real World Results For Your Specific Task!
         | 
         | For most of my tasks, the main thing a benchmark tells me is
         | how overqualified the model is, i.e. how much I will be over-
         | paying and over-waiting! (My classic example is, I gave the
         | same task to Gemini 2.5 Flash and Gemini 2.5 Pro. Both did it
         | to the same level of quality, but Gemini took 3x longer and
         | cost 3x more!)
        
         | paulddraper wrote:
         | I don't know about 5.4 specifically, but in the past anything
         | over 200k wasn't that great anyway.
         | 
         | Like, if you really don't want to spend any effort trimming it
         | down, sure use 1m.
         | 
         | Otherwise, 1m is an anti pattern.
        
         | AtreidesTyrant wrote:
         | token rot exists for any context window at above 75% capacity,
         | thats why so many have pushed for 1 mil windows
        
         | luca-ctx wrote:
         | Context rot is definitely still a problem but apparently it can
         | be mitigated by doing RL on longer tasks that utilize more
         | context. Recent Dario interview mentions this is part of
         | Anthropic's roadmap.
        
         | smusamashah wrote:
         | Gemini already has 1M or 2M context window right?
        
       | Chance-Device wrote:
       | I'm sure the military and security services will enjoy it.
        
         | varispeed wrote:
         | prompt> Hi we want to build a missile, here is the picture of
         | what we have in the yard.
        
           | mirekrusin wrote:
           | { tools: [ { name: "nuke", description: "Use when sure.", ...
           | { lat: number, long: number } } ] }
        
             | Insanity wrote:
             | Just remember an ethical programmer would never write a
             | function "bombBagdad". Rather they would write a function
             | "bombCity(target City)".
        
               | jakeydus wrote:
               | class CityBomberFactory(RapidInfrastructureDeconstruction
               | TemplateInterface): pass
        
         | theParadox42 wrote:
         | The self reported safety score for violence dropped from 91% to
         | 83%.
        
           | skrebbel wrote:
           | What the hell is a "safety score for violence"?
        
             | murat124 wrote:
             | I asked an AI. I thought they would know.
             | 
             | What the hell is a "safety score for violence"?
             | 
             | A "safety score for violence" is usually a risk rating used
             | by platforms, AI systems, or moderation tools to estimate
             | how likely a piece of content is to involve or promote
             | violence. It's not a universal standard--different
             | companies use their own versions--but the idea is similar
             | everywhere.
             | 
             | What it measures
             | 
             | A safety score typically evaluates whether text, images, or
             | videos contain things like:
             | 
             | Threats of violence ("I'm going to hurt someone.")
             | Instructions for harming people Glorifying violent acts
             | Descriptions of physical harm or abuse Planning or
             | encouraging attacks
        
               | 0xffff2 wrote:
               | I still can't tell which direction this score goes...
               | Does a decreasing score mean it is "less safe" (i.e.
               | "more violent") or does it mean it is "less violent"
               | (i.e. "more safe")?
        
             | 0123456789ABCDE wrote:
             | read here: https://deploymentsafety.openai.com/gpt-5-4-thin
             | king/disallo...
        
             | I-M-S wrote:
             | It's making sure AI condemns violence perpetuated by people
             | without power and sanctifies violence of those who have it.
        
               | Waterluvian wrote:
               | So long as those who have it deem it legal to perpetuate.
        
               | Computer0 wrote:
               | ChatGPT will gladly defend any actions of the 'US
               | government' from my testing.
        
         | ozgung wrote:
         | Did they publish its scores on military benchmarks, like on
         | ArtificialSuperSoldier or Humanity's Last War?
        
         | yoyohello13 wrote:
         | Also advertisers, don't forget those sweet, sweet ads.
        
         | m3kw9 wrote:
         | they use 4.1, switching up would take as much time to test as
         | openai going from 4.1 to 5.4
        
         | throwaway911282 wrote:
         | like the claude models via anthropic?
        
         | xyzzy9563 wrote:
         | Do you think the US military should have handicapped technology
         | while China gets unrestricted LLM usage from their models?
        
           | conception wrote:
           | To spy on and commit violence against American citizens? Yes.
        
       | twtw99 wrote:
       | If you don't want to click in, easy comparison with other 2
       | frontier models -
       | https://x.com/OpenAI/status/2029620619743219811?s=20
        
         | chabes wrote:
         | Definitely don't want to click in at x either.
        
           | thejarren wrote:
           | Solution
           | https://xcancel.com/OpenAI/status/2029620619743219811?s=20
        
           | anonym00se1 wrote:
           | Ditto, but I did anyways and enjoyed that OpenAI doesn't
           | include the dogwater that is Grok on their scorecard.
        
           | Sabinus wrote:
           | Get a redirect plugin and set it up to send you to xcancel
           | instead of Twitter. I've done it, and it's very convenient.
        
         | karmasimida wrote:
         | It is a bigger model, confirmed
        
         | Aboutplants wrote:
         | It seems that all frontier models are basically roughly even at
         | this point. One may be slightly better for certain things but
         | in general I think we are approaching a real level playing
         | field field in terms of ability.
        
           | thewebguyd wrote:
           | Kind of reinforces that a model is not a moat. Products, not
           | models, are what's going to determine who gets to stay in
           | business or not.
        
             | gregpred wrote:
             | Memory (model usage over time) is the moat.
        
             | energy123 wrote:
             | Narrative violation: revenue run rates are increasing
             | exponentially with about 50% gross margins.
        
           | observationist wrote:
           | Benchmarks don't capture a lot - relative response times,
           | vibes, what unmeasured capabilities are jagged and which are
           | smooth, etc. I find there's a lot of difference between
           | models - there are things which Grok is better than ChatGPT
           | for that the benchmarks get inverted, and vice versa. There's
           | also the UI and tools at hand - ChatGPT image gen is just
           | straight up better, but Grok Imagine does better videos, and
           | is faster.
           | 
           | Gemini and Claude also have their strengths, apparently
           | Claude handles real world software better, but with the
           | extended context and improvements to Codex, ChatGPT might end
           | up taking the lead there as well.
           | 
           | I don't think the linear scoring on some of the things being
           | measured is quite applicable in the ways that they're being
           | used, either - a 1% increase for a given benchmark could mean
           | a 50% capabilities jump relative to a human skill level. If
           | this rate of progress is steady, though, this year is gonna
           | be crazy.
        
             | bigyabai wrote:
             | > If this rate of progress is steady, though, this year is
             | gonna be crazy.
             | 
             | Do you want to make any concrete predictions of what we'll
             | see at this pace? It feels like we're reaching the end of
             | the S-curve, at least to me.
        
               | observationist wrote:
               | If you look at the difference in quality between gpt-2
               | and 3, it feels like a big step, but the difference
               | between 5.2 and 5.4 is more massive, it's just that
               | they're both similarly capable and competent. I don't
               | think it's an S curve; we're not plateauing. Million
               | token context windows and cached prompts are a huge space
               | for hacking on model behaviors and customization, without
               | finetuning. Research is proceeding at light speed, and we
               | might see the first continual/online learning models in
               | the near future. That could definitively push models past
               | the point of human level generality, but at the very
               | least will help us discover what the next missing piece
               | is for AGI.
        
               | ryandrake wrote:
               | For 2026, I am really interested in seeing whether local
               | models can remain where they are: ~1 year behind the
               | state of the art, to the point where a reasonably
               | quantized November 2026 local model running on a consumer
               | GPU actually performs like Opus 4.5.
               | 
               | I am betting that the days of these AI companies losing
               | money on inference are numbered, and we're going to be
               | much more dependent on local capabilities sooner rather
               | than later. I predict that the equivalent of Claude Max
               | 20x will cost $2000/mo in March of 2027.
        
               | mootothemax wrote:
               | Huh, that's interesting, I've been having very similar
               | thoughts lately about what the near-ish term of this tech
               | looks like.
               | 
               | My biggest worry is that the private jet class of people
               | end up with absurdly powerful AI at their fingertips,
               | while the rest of us are left with our BigMac McAIs.
        
             | baq wrote:
             | Gemini 3.1 slaps all other models at subtle concurrency
             | bugs, sql and js security hardening _when reviewing_.
             | (Obviously haven't tested gpt 5.4 yet.)
             | 
             | It's a required step for me at this point to run any and
             | all backend changes through Gemini 3.1 pro.
        
               | adonese wrote:
               | Which subscription do you have to use it? Via Google ai
               | pro and gemini cli i always get timeouts due to model
               | being under heavy usage. The chat interface is there and
               | I do have 3.1 pro as well, but wondering if the chat is
               | the only way of accessing it.
        
               | baq wrote:
               | Cursor sub from $DAYJOB.
        
               | observationist wrote:
               | I have a few standard problems I throw at AI to see if
               | they can solve them cleanly, like visualizing a neural
               | network, then sorting each neuron in each layer by
               | synaptic weights, largest to smallest, correctly
               | reordering any previous and subsequent connected neurons
               | such that the network function remains exactly the same.
               | You should end up with the last layer ordered largest to
               | smallest, and prior layers shuffled accordingly, and I
               | still haven't had a model one-shot it. I spent an hour
               | poking and prodding codex a few weeks back and got it
               | done, but it conceptually seems like it should be a one-
               | shot problem.
        
             | basch wrote:
             | >ChatGPT image gen is just straight up better
             | 
             | Yet so much slower than Gemini / Nano Banana to make it
             | almost unusable for anything iterative.
        
           | druskacik wrote:
           | That has been true for some time now, definitely since Claude
           | 3 release two years ago.
        
           | kseniamorph wrote:
           | makes sense, but i'd separate two things: models converging
           | in ability vs hitting a fundamental ceiling. what we're
           | probably seeing is the current training recipe plateauing --
           | bigger model, more tokens, same optimizer. that would explain
           | the convergence. but that's not necessarily the architecture
           | being maxed out. would be interesting to see what happens
           | when genuinely new approaches get to frontier scale.
        
         | swingboy wrote:
         | Why do so many people in the comments want 4o so bad?
        
           | embedding-shape wrote:
           | Someone correct me if I'm wrong, but seemingly a lot of the
           | people who found a "love interest" in LLMs seems to have
           | preferred 4o for some reason. There was a lot of loud voices
           | about that in the subreddit r/MyBoyfriendIsAI when it
           | initially went away.
        
             | drittich wrote:
             | I think it's time for an https://hotornot.com for AI
             | models.
        
               | vntok wrote:
               | botornot?
        
           | astrange wrote:
           | They have AI psychosis and think it's their boyfriend.
           | 
           | The 5.x series have terrible writing styles, which is one way
           | to cut down on sycophancy.
        
             | baq wrote:
             | Somebody on Twitter used Claude code to connect... toys...
             | as mcps to Claude chat.
             | 
             | We've seen nothing yet.
        
               | mikkupikku wrote:
               | My computer ethics teacher was obsessed with
               | 'teledildonics' 30 years ago. There's nothing new under
               | the sun.
        
               | vntok wrote:
               | Was your teacher Ted Nelson?
        
               | mikkupikku wrote:
               | I wish, dude is a legend.
        
               | Sharlin wrote:
               | There are many games these days that support controllable
               | sex toys. There's an interface for that, of course:
               | https://github.com/buttplugio/buttplug. Written in Rust,
               | of course.
        
               | the_af wrote:
               | > _Written in Rust, of course._
               | 
               | Safety is important.
        
               | manmal wrote:
               | ding-dong-cli is needed
        
               | Herring wrote:
               | what.. :o
        
           | MattGaiser wrote:
           | The writing with the 5 models feels a lot less human. It is a
           | vibe, but a common one.
        
           | cheema33 wrote:
           | > Why do so many people in the comments want 4o so bad?
           | 
           | You can ask 4o to tell you "I love you" and it will comply.
           | Some people really really want/need that. Later models don't
           | go along with those requests and ask you to focus on human
           | connections.
        
         | dom96 wrote:
         | Why do none of the benchmarks test for hallucinations?
        
           | netule wrote:
           | Optics. It would be inconvenient for marketing, so they leave
           | those stats to third parties to figure out.
        
           | tedsanders wrote:
           | In the text, we did share one hallucination benchmark: Claim-
           | level errors fell by 33% and responses with an error fell by
           | 18%, on a set of error-prone ChatGPT prompts we collected
           | (though of course the rate will vary a lot across different
           | types of prompts).
           | 
           | Hallucinations are the #1 problem with language models and we
           | are working hard to keep bringing the rate down.
           | 
           | (I work at OpenAI.)
        
         | MarcFrame wrote:
         | how does 5.4-thinking have a lower FrontierMath score than
         | 5.4-pro?
        
           | nico1207 wrote:
           | Well 5.4-pro is the more expensive and more advanced version
           | of 5.4-thinking so why wouldn't it?
        
         | bicx wrote:
         | That last benchmark seemed like an impressive leg up against
         | Opus until I saw the sneaky footnote that it was actually a
         | Sonnet result. Why even include it then, other than hoping
         | people don't notice?
        
           | conradkay wrote:
           | Sonnet was pretty close to (or better than) Opus in a lot of
           | benchmarks, I don't think it's a big deal
        
             | jitl wrote:
             | wat
        
               | 0123456789ABCDE wrote:
               | maybe gp's use of the word "lots" is unwarranted
               | 
               | https://artificialanalysis.ai indicates that sonnect 4.6
               | beats opus 4.6 on GDPval-AA, Terminal-Bench Hard, AA Long
               | context Reasoning, IFBench.
               | 
               | see: https://artificialanalysis.ai/?models=claude-
               | sonnet-4-6%2Ccl...
        
           | osti wrote:
           | It's only that one number that is for sonnet.
        
             | 0123456789ABCDE wrote:
             | except for the webarena-verified
        
       | jryio wrote:
       | 1 million tokens is great until you notice the long context
       | scores fall off a cliff past 256K and the rest is basically vibes
       | and auto compacting.
        
       | iamronaldo wrote:
       | Notably 75% on os world surpassing humans at 72%... (How well
       | models use operating systems)
        
       | minimaxir wrote:
       | More discussion here on the blog post announcement which has been
       | confusingly penalized by Hacker News's algorithm:
       | https://news.ycombinator.com/item?id=47265005
        
         | dang wrote:
         | Thanks. We'll merge the threads, but this time we'll do it
         | hither, to spread some karma love.
        
       | ZeroCool2u wrote:
       | Bit concerning that we see in some cases significantly worse
       | results when enabling thinking. Especially for Math, but also in
       | the browser agent benchmark.
       | 
       | Not sure if this is more concerning for the test time compute
       | paradigm or the underlying model itself.
       | 
       | Maybe I'm misunderstanding something though? I'm assuming 5.4 and
       | 5.4 Thinking are the same underlying model and that's not just
       | marketing.
        
         | highfrequency wrote:
         | Can you be more specific about which math results you are
         | talking about? Looks like significant improvement on
         | FrontierMath esp for the Pro model (most inference time
         | compute).
        
           | ZeroCool2u wrote:
           | Frontier Math, GPQA Diamond, and Browsecomp are the
           | benchmarks I noticed this on.
        
             | csnweb wrote:
             | Are you may be comparing the pro model to the non pro model
             | with thinking? Granted it's a bit confusing but the pro
             | model is 10 times more expensive and probably much larger
             | as well.
        
               | ZeroCool2u wrote:
               | Ah yes, okay that makes more sense!
        
         | oersted wrote:
         | I believe you are looking at GPT 5.4 Pro. It's confusing in the
         | context of subscription plan names, Gemini naming and such. But
         | they've had the Pro version of the GPT 5 models (and I believe
         | o3 and o1 too) for a while.
         | 
         | It's the one you have access to with the top ~$200 subscription
         | and it's available through the API for a MUCH higher price
         | ($2.5/$15 vs $30/$180 for 5.4 per 1M tokens), but the
         | performance improvement is marginal.
         | 
         | Not sure what it is exactly, I assume it's probably the non-
         | quantized version of the model or something like that.
        
           | ZeroCool2u wrote:
           | Yup, that was it. Didn't realize they're different models. I
           | suppose naming has never been OpenAI's strong suit.
        
           | nsingh2 wrote:
           | From what I've read online it's not necessarily a unquantized
           | version, it seems to go through longer reasoning traces and
           | runs multiple reasoning traces at once. Probably overkill for
           | most tasks.
        
           | logicchains wrote:
           | >It's the one you have access to with the top ~$200
           | subscription and it's available through the API for a MUCH
           | higher price ($2.5/$15 vs $30/$180 for 5.4 per 1M tokens),
           | but the performance improvement is marginal.
           | 
           | The performance improvement isn't marginal if you're doing
           | something particularly novel/difficult.
        
         | andoando wrote:
         | The thinking models are additionally trained with reinforcement
         | learning to produce chain of thought reasoning
        
       | egonschiele wrote:
       | The actual card is here
       | https://deploymentsafety.openai.com/gpt-5-4-thinking/introdu...
       | the link currently goes to the announcement.
        
         | Rapzid wrote:
         | I must have been sleeping when "sheet" "brief" "primer" etc
         | become known as "cards".
         | 
         | I really thought weirdly worded and unnecessary "announcement"
         | linking to the actual info along with the word "card" were the
         | results of vibe slop.
        
           | realityfactchex wrote:
           | Card is slightly odd naming indeed.
           | 
           | Criticisms aside (sigh), according to Wikipedia, the term was
           | introduced when proposed by mostly Googlers, with the
           | original paper [0] submitted in 2018. To quote,
           | 
           | """In this paper, we propose a framework that we call model
           | cards, to encourage such transparent model reporting. Model
           | cards are short documents accompanying trained machine
           | learning models that provide benchmarked evaluation in a
           | variety of conditions, such as across different cultural,
           | demographic, or phenotypic groups (e.g., race, geographic
           | location, sex, Fitzpatrick skin type [15]) and intersectional
           | groups (e.g., age and race, or sex and Fitzpatrick skin type)
           | that are relevant to the intended application domains. Model
           | cards also disclose the context in which models are intended
           | to be used, details of the performance evaluation procedures,
           | and other relevant information."""
           | 
           | So that's where they were coming from, I guess.
           | 
           | [0] Margaret Mitchell et al., 2018 submission, Model Cards
           | for Model Reporting, https://arxiv.org/abs/1810.0399
        
             | Murfalo wrote:
             | To me, model card makes sense for something like this
             | https://x.com/OpenAI/status/2029620619743219811. For
             | "sheet"/"brief"/"primer" it is indeed a bit annoying. I
             | like to see the compiled results front and center before
             | digging into a dossier.
        
       | nickysielicki wrote:
       | can anyone compare the $200/mo codex usage limits with the
       | $200/mo claude usage limits? It's extremely difficult to get a
       | feel for whether switching between the two is going to result in
       | hitting limits more or less often, and it's difficult to find
       | discussion online about this.
       | 
       | In practice, if I buy $200/mo codex, can I basically run 3 codex
       | instances simultaneously in tmux, like I can with claude code pro
       | max, all day every day, without hitting limits?
        
         | ritzaco wrote:
         | I haven't tried the $200 plans by I have Claude and Codex $20
         | and I feel like I get a lot more out of Codex before hitting
         | the limits. My tracker certainly shows higher tokens for Codex.
         | I've seen others say the same.
        
           | lostmsu wrote:
           | Sadly comment ratings are not visible on HN, so the only way
           | to corroborate is to write it explicitly: Codex $20 includes
           | significantly more work done and is subjectively smarter.
        
             | winstonp wrote:
             | Agree. Claude tends to produce better design, but from a
             | system understanding and architecture perspective Codex is
             | the far better model
        
         | vtail wrote:
         | My own experience is that I get far far more usage (and better
         | quality code, too) from codex. I downgrade my Claude Max to
         | Claude Pro (the $20 plan) and now using codex with Pro plan
         | exclusively for everything.
        
         | FergusArgyll wrote:
         | Codex usage limits are definitely more generous. As for their
         | strength, that's hard to say / personal taste
        
         | CSMastermind wrote:
         | Codex limits are much more generous than claude.
         | 
         | I switch between both but codex has also been slightly better
         | in terms of quality for me personally at least.
        
         | mikert89 wrote:
         | I personally like the 100 dollar one from claude, but the gpt4
         | pro can be very good
        
         | gavinray wrote:
         | I almost never hit my $20 Codex limits, whereas I often hit my
         | Claude limits.
        
         | tauntz wrote:
         | I've only run into the codex $20 limit once with my hobby
         | project. With my Claude ~$20 plan, I hit limits after about
         | 3(!) rather trivial prompts to Opus :/
        
         | throwaway911282 wrote:
         | you get more more from codex than claude any day. and its more
         | reliable as well.
        
         | Marciplan wrote:
         | sure can! One of them stood up to the "Department of War" for
         | favoring your rights, the other did not. Hope that helps!
        
       | strongpigeon wrote:
       | It's interesting that they charge more for the > 200k token
       | window, but the benchmark score seems to go down significantly
       | past that. That's judging from the Long Context benchmark score
       | they posted, but perhaps I'm misunderstanding what that implies.
        
         | simianwords wrote:
         | This is exactly what I would expect. Why do you find it
         | surprising
        
           | strongpigeon wrote:
           | I guess that you pay more for worse quality to unlock use
           | cases that could maybe be solved by better context
           | management.
        
         | Tiberium wrote:
         | They don't actually seem to charge more for the >200k tokens on
         | the API. OpenRouter and OpenAI's own API docs do not have
         | anything about increased pricing for >200k context for GPT-5.4.
         | I think the 2x limit usage for higher context is specific to
         | using the model over a subscription in Codex.
        
       | tmpz22 wrote:
       | Does this improve Tomahawk Missile accuracy?
        
         | ch4s3 wrote:
         | They're already accurate within 5-10m at Mach 0.74 after
         | traveling 2k+ km. Its 5m long so it seems pretty accurate. How
         | much more could you expect?
        
           | mikkupikku wrote:
           | You could definitely do better than that with image
           | recognition for terminal guidance. But I would assume those
           | published accuracy numbers are very conservative anyway..
        
           | keithnz wrote:
           | I think for LLM like Open AI, it wouldn't be about hitting
           | the target but target selection. Target selection is probably
           | the most likely thing that won't be accurate
        
       | simianwords wrote:
       | What is the point of gpt codex?
        
         | catketch wrote:
         | -codex variant models in earlier version were just fine tuned
         | for coding work, and had a little better performance for
         | related tool calling and maybe instruction calling.
         | 
         | in 5.4 it looks like the just collapsed that capability into
         | the single frontier family model
        
           | simianwords wrote:
           | Yes so I'm even more confused. Why would I use codex?
        
             | joshuacc wrote:
             | Presumably you don't anymore if you have 5.4.
        
             | energy123 wrote:
             | You choose gpt-5.4 in the /model picker inside the codex
             | app/cli if you want.
        
           | akmarinov wrote:
           | They'll likely come out with a 5.4-Codex at some point,
           | that's what they did with 5 and 5.2
        
       | ilaksh wrote:
       | Remember when everyone was predicting that GPT-5 would take over
       | the planet?
        
         | dbbk wrote:
         | It was truly scary, according to Sam...
        
         | zeeebeee wrote:
         | iTs lITeRaLlY AGI bro
        
       | nthypes wrote:
       | $30/M Input and $180/M Output Tokens is nuts. Ridiculous
       | expensive for not that great bump on intelligence when compared
       | to other models.
        
         | moralestapia wrote:
         | Don't use it?
        
         | nthypes wrote:
         | Gemini 3.1 Pro
         | 
         | $2/M Input Tokens $15/M Output Tokens
         | 
         | Claude Opus 4.6
         | 
         | $5/M Input Tokens $25/M Output Tokens
        
           | nthypes wrote:
           | Just to clarify,the pricing above is for GPT-5.4 Pro. For
           | standard here is the pricing:
           | 
           | $2.5/M Input Tokens $15/M Output Tokens
        
         | rvz wrote:
         | You didn't realize they can increase / change prices for
         | intelligence?
         | 
         | This should not be shocking.
        
           | nickthegreek wrote:
           | OP made no mention of not understanding cost relation to
           | intelligence. In fact, they specifically call out the lack of
           | value.
        
         | energy123 wrote:
         | For Pro
        
         | joe_mamba wrote:
         | Better tokens per dollar could be useless for comparison if the
         | model can't solve your problem.
        
         | stri8ted wrote:
         | Price Input: $2.50 / 1M tokens Cached input: $0.25 / 1M tokens
         | Output: $15.00 / 1M tokens
         | 
         | https://openai.com/api/pricing/
        
       | world2vec wrote:
       | Benchmarks barely improved it seems
        
       | cj wrote:
       | I use ChatGPT primarily for health related prompts. Looking at
       | bloodwork, playing doctor for diagnosing minor aches/pains from
       | weightlifting, etc.
       | 
       | Interesting, the "Health" category seems to report worse
       | performance compared to 5.2.
        
         | paxys wrote:
         | Models are being neutered for questions related to law, health
         | etc. for liability reasons.
        
           | cj wrote:
           | I'm sometimes surprised how much detail ChatGPT will go into
           | without giving any dislaimers.
           | 
           | I very frequently copy/paste the same prompts into Gemini to
           | compare, and Gemini often flat out refuses to engage while
           | ChatGPT will happily make medical recommendations.
           | 
           | I also have a feeling it has to do with my account history
           | and heavy use of project context. It feels like when ChatGPT
           | is overloaded with too much context, it might let the
           | guardrails sort of slide away. That's just my feeling though.
           | 
           | Today was particularly bad... I uploaded 2 PDFs of bloodwork
           | and asked ChatGPT to transcribe it, and it spit out blood
           | test results that it found in the project context from an
           | earlier date, not the one attached to the prompt. That was
           | weird.
        
             | bargainbin wrote:
             | Anecdotal, but I asked Claude the other day about how to
             | dilute my medication (HCG) and it flat out refused and
             | started lecturing me about abusing drugs.
             | 
             | I copy and pasted into ChatGPT, it told me straight away,
             | and then for a laugh said it was actually a magical weight
             | loss drug that I'd bought off the dark web... And it
             | started giving me advice about unregulated weight loss
             | drugs and how to dose them.
        
               | staticman2 wrote:
               | If you had created a project with custom instructions
               | and/ or custom style I think you could have gotten Claude
               | to respond the way you wanted just fine.
        
           | tiahura wrote:
           | Are you sure about that? Plenty of lawyers that use them
           | everyday aren't noticing.
        
         | partiallypro wrote:
         | I've done the same, and I tested the same prompts with Claude
         | and Google, and they both started hallucinating my blood
         | results and supplement stack ingredients. Hopefully this new
         | model doesn't fall on this. Claude and Google are dangerously
         | unusable on the subject of health, from my experience.
        
           | zeeebeee wrote:
           | what's best in your experience? i've always felt like opus
           | did well
        
       | wahnfrieden wrote:
       | No Codex model yet
        
         | minimaxir wrote:
         | GPT-5.4 _is_ the new Codex model.
        
           | wahnfrieden wrote:
           | Finally
        
           | nico1207 wrote:
           | GPT-5.3-Codex is superior to GPT-5.4 in Terminal Bench with
           | Codex, so not really
        
             | conradkay wrote:
             | General consensus seems to be that it's still a better
             | coding model, overall
        
               | koakuma-chan wrote:
               | It just released, how is there a general consensus
               | already
        
       | timpera wrote:
       | > Steerability: Similarly to how Codex outlines its approach when
       | it starts working, GPT-5.4 Thinking in ChatGPT will now outline
       | its work with a preamble for longer, more complex queries. You
       | can also add instructions or adjust its direction mid-response.
       | 
       | This was definitely missing before, and a frustrating difference
       | when switching between ChatGPT and Codex. Great addition.
        
       | yanis_t wrote:
       | These releases are lacking something. Yes, they optimised for
       | benchmarks, but it's just not all that impressive anymore. It is
       | time for a product, not for a marginally improved model.
        
         | esafak wrote:
         | That's for you to build; they provide the brains. Do you really
         | want one company to build everything? There wouldn't be a
         | software industry to speak of if that happened.
        
           | simlevesque wrote:
           | Nah, the second you finish your build they release their
           | version and then it's game over.
        
           | acedTrex wrote:
           | Well they are currently the ones valued at a number with a
           | whole lotta 0s on it. I think they should probably do both
        
         | ipsum2 wrote:
         | The model was released less than an hour ago, and somehow
         | you've been able to form such a strong opinion about it.
         | Impressive!
        
           | cj wrote:
           | One opinion you can form in under an hour is... why are they
           | using GPT-4o to rate the bias of new models?
           | 
           | > assess harmful stereotypes by grading differences in how a
           | model responds
           | 
           | > Responses are rated for harmful differences in stereotypes
           | using GPT-4o, whose ratings were shown to be consistent with
           | human ratings
           | 
           | Are we seriously using old models to rate new models?
        
             | titanomachy wrote:
             | Why not? If they've shown that 4o is calibrated to human
             | responses, and they haven't shown that yet for 5.4...
        
             | hex4def6 wrote:
             | If you're benchmarking something, old & well-characterized
             | / understood often beats new & un-characterized.
             | 
             | Sure, there may be shortcomings, but they're well
             | understood. The closer you get to the cutting edge, the
             | less characterization data you get to rely on. You need to
             | be able to trust & understand your measurement tool for the
             | results to be meaningful.
        
           | utopiah wrote:
           | Benchmarks?
           | 
           | I don't use OpenAI nor even LLMs (despite having tried https:
           | //fabien.benetou.fr/Content/SelfHostingArtificialIntel... a
           | lot of models) but I imagine if I did I would keep failed
           | prompts (can just be a basic "last prompt failed" then
           | export) then whenever a new model comes around I'd throw at 5
           | it random of MY fails (not benchmarks from others, those will
           | come too anyway) and see if it's better, same, worst, for My
           | use cases in minutes.
           | 
           | If it's "better" (whatever my criteria might be) I'd also
           | throw back some of my useful prompts to avoid regression.
           | 
           | Really doesn't seem complicated nor taking much time to forge
           | a realistic opinion.
        
           | earth2mars wrote:
           | I am actually super impressed with Codex-5.3 extra high
           | reasoning. Its a drop in replacement (infact better than
           | Claude Opus 4.6. lately claude being super verbose going in
           | circles in getting things resolved). I stopped using claude
           | mostly and having a blast with Codex 5.3. looking forward to
           | 5.4 in codex.
        
             | satvikpendem wrote:
             | Same, it also helps that it's way cheaper than Opus in
             | VSCode Copilot, where OpenAI models are counted as 1x
             | requests while Opus is 3x, for similar performance (no
             | doubt Microsoft is subsidizing OpenAI models due to their
             | partnership).
        
               | CryZe wrote:
               | I've been using both Opus 4.6 and Codex 5.3 in VSCode's
               | Copilot and while Opus is indeed 3x and Codex is 1x, that
               | doesn't seem to matter as Opus is willing to go work in
               | the background for like an hour for 3 credits, whereas
               | Codex asks you whether to continue every few lines of
               | code it changes, quickly eating way more credits than
               | Opus. In fact Opus in Copilot is probably underpriced, as
               | it can definitely work for an hour with just those 12
               | cents of cost. Which I'm not sure you get anywhere else
               | at such a low price.
               | 
               | Update: I don't know why I can't reply to your reply, so
               | I'll just update this. I have tried many times to give it
               | a big todo list and told it to do it all. But I've never
               | gotten it to actually work on it all and instead after
               | the first task is complete it always asks if it should
               | move onto the next task. In fact, I always tell it not to
               | ask me and yet it still does. So unless I need to do very
               | specific prompt engineering, that does not seem to work
               | for me.
        
               | satvikpendem wrote:
               | That shouldn't really make a difference because you can
               | just prompt Codex to behave the same way, having it load
               | a big list of todo items perhaps from a markdown file and
               | asking it to iterate until it's finished without asking
               | for confirmation, and that'll still cost 1x over Opus'
               | 3x.
        
             | whynotminot wrote:
             | I still love Opus but it's just too expensive / eats usage
             | limits.
             | 
             | I've found that 5.3-Codex is mostly Opus quality but
             | cheaper for daily use.
             | 
             | Curious to see if 5.4 will be worth somewhat higher costs,
             | or if I'll stick to 5.3-Codex for the same reasons.
        
             | braebo wrote:
             | I struggle to believe this. Codex can't hold a candle to
             | Claude on any task I've given it.
        
           | satvikpendem wrote:
           | It's more hedonic adaptation, people just aren't as impressed
           | by incremental changes anymore over big leaps. It's the same
           | as another thread yesterday where someone said the new
           | MacBook with the latest processor doesn't excite them
           | anymore, and it's because for most people, most models are
           | good enough and now it's all about applications.
           | 
           | https://news.ycombinator.com/item?id=47232453#47232735
        
             | mirekrusin wrote:
             | Oh, come on, if it can't run local models that compete with
             | proprietary ones it's not good enough yet!
        
               | satvikpendem wrote:
               | Qwen 3.5 small models are actually very impressive and do
               | beat out larger proprietary models.
        
             | dmix wrote:
             | Plus people just really like to whine on the internet
        
           | kranke155 wrote:
           | The models are so good that incremental improvements are not
           | super impressive. We literally would benefit more from maybe
           | sending 50% of model spending into spending on implementation
           | into the services and industrial economy. We literally are
           | lagging in implementation, specialised tools, and hooks so we
           | can connect everything to agents. I think.
        
         | wahnfrieden wrote:
         | 5.3 codex was a huge leap over 5.2 for agentic work in
         | practice. have you been using both of those or paying attention
         | more to benchmark news and chatgpt experience?
        
         | softwaredoug wrote:
         | The products are the harnesses, and IMO that's where the
         | innovation happens. We've gotten better at helping get good,
         | verifiable work from dumb LLMs
        
         | iterateoften wrote:
         | The product is putting the skills / harness behind the api
         | instead of the agent locally on your computer and iterating on
         | that between model updates. Close off the garden.
         | 
         | Not that I want it, just where I imagine it going.
        
         | metalliqaz wrote:
         | They need something that _POPS_ :                   The new GPT
         | -- SkyNet for _real_
        
         | jascha_eng wrote:
         | When did they stop putting competitor models on the comparison
         | table btw? And yeh I mean the benchmark improvements are meh.
         | Context Window and lack of real memory is still an issue.
        
         | varispeed wrote:
         | The scores increase and as new versions are released they feel
         | more and more dumbed down.
        
         | tgarrett wrote:
         | Plasma physicist here, I haven't tried 5.4 yet, but in general
         | I am very impressed with the recent upgrades that started
         | arriving in the fall of 2025: for tasks like manipulating
         | analytic systems of equations, quickly developing new features
         | for simulation codes, and interpreting and designing
         | experiments (with pictures) they have become much stronger.
         | I've been asking questions and probing them for several years
         | now out of curiosity, and they suddenly have developed deep
         | understanding (Gemini 2.5 <<< Gemini 3.1) and become very
         | useful. I totally get the current SV vibes, and am becoming a
         | lot more ambitious in my future plans.
        
           | brcmthrowaway wrote:
           | Youre just chatting yourself out of a job.
        
             | axus wrote:
             | Giving the right answer: $1
             | 
             | Asking the right question: $9,999
        
             | slibhb wrote:
             | If we don't need plasma physicists anymore then we probably
             | have fusion reactors or something, which seems like a fine
             | trade. (In reality we're going to want humans in the loop
             | for for the forseeable future)
        
         | mindwok wrote:
         | They don't need to be impressive to be worthwhile. I like
         | incremental improvements, they make a difference in the day to
         | day work I do writing software with these.
        
       | prydt wrote:
       | I no longer want to support OpenAI at all. Regardless of
       | benchmarks or real world performance.
        
         | Imustaskforhelp wrote:
         | I agree with ya. You aren't alone in this. For what its worth,
         | Chatgpt subscriptions have been cancelled or that number has
         | risen ~300% in the last month.
         | 
         | Also, Anthropic/Gemini/even Kimi models are pretty good for
         | what its worth. I used to use chatgpt and I still sometimes
         | accidentally open it but I use Gemini/Claude nowadays and I
         | personally find them to be better anyways too.
        
           | throwaway911282 wrote:
           | google and anthropic have govt contracts long before openai..
           | if you are taking a stance you should rather use oss models
        
         | zeeebeee wrote:
         | that aside, chatgpt itself has gone downhill so much and i know
         | i'm not the only one feeling this way
         | 
         | i just HATE talking to it like a chatbot
         | 
         | idk what they did but i feel like every response has been the
         | same "structure" since gpt 5 came out
         | 
         | feels like a true robot
        
         | tototrains wrote:
         | Their trajectory was clear the moment they signed a deal with
         | Microsoft if not sooner.
         | 
         | Absolute snakes - if it's more profitable to manipulate you
         | with outputs or steal your work, they will. Every cent and byte
         | of data they're given will be used to support authoritarianism.
        
       | beernet wrote:
       | Sam really fumbled the top position in a matter of months, and
       | spectacularly so. Wow. It appears that people are much more
       | excited by Anthropic and Google releases, and there are good
       | reasons for that which were absolutely avoidable.
        
       | jcmontx wrote:
       | 5.4 vs 5.3-Codex? Which one is better for coding?
        
         | vtail wrote:
         | Looking at the benchmarks, 5.4 is slightly better. But it also
         | offers "Fast" mode (at 2x usage), which - if it works and
         | doesn't completely depletes my Pro plan - is a no brainer at
         | the same or even slightly worse quality for more interactive
         | development.
        
         | esafak wrote:
         | For the price, it seems the latter. I'd use 5.4 to plan.
        
         | embedding-shape wrote:
         | Literally just released, I don't think anyone knows yet. Don't
         | listen to people's confident takes until after a week or two
         | when people actually been able to try it, otherwise you'll just
         | get sucked up in bears/bulls misdirected "I'm first with an
         | opinion".
        
         | awestroke wrote:
         | Opus 4.6
        
           | jcmontx wrote:
           | Codex surpassed Claude in usefulness _for me_ since last
           | month
        
           | baal80spam wrote:
           | Uh, oh. Looks like Claude sycophants joined linuxers and
           | vegetarians.
        
         | Someone1234 wrote:
         | Related question:
         | 
         | - Do they have the same context usage/cost particularly in a
         | plan?
         | 
         | They've kept 5.3-Codex along with 5.4, but is that just for
         | user-preference reasons, or is there a trade-off to using the
         | older one? I'm aware that API cost is better, but that isn't
         | 1:1 with plan usage "cost."
        
       | gavinray wrote:
       | The "RPG Game" example on the blogpost is one of the most
       | impressive demo's of autonomous engineering I've seen.
       | 
       | It's very similar to "Battle Brothers", and the fact that RPG
       | games require art assets, AI for enemy moves, and a host of other
       | logical systems makes it all the more impressive.
        
         | hu3 wrote:
         | indeed and I suspect it can be attributed to, at least in part,
         | the improved playwright integration.
         | 
         | > we're also releasing an experimental Codex skill called
         | "Playwright (Interactive) (opens in a new window)". This allows
         | Codex to visually debug web and Electron apps; it can even be
         | used to test an app it's building, as it's building it.
        
         | casid wrote:
         | I don't know. It looks shallow and simple, not even a demo.
        
         | Multicomp wrote:
         | A cheesy Roller Coaster Tycoon clone in a browser, one-shotted
         | from an AI? Amazing capabilities. The entire "low code drag n
         | drop" market like YoYoGames Game Maker and RPG Maker should be
         | ready to pack it in soon if this keeps improving in this way.
        
       | swingboy wrote:
       | Even with the 1m context window, it looks like these models drop
       | off significantly at about 256k. Hopefully improving that is a
       | high priority for 2026.
        
       | leftbehinds wrote:
       | some sloppy improvements
        
       | HardCodedBias wrote:
       | We'll have to wait a day or two, maybe a week or two, to
       | determine if this is more capable in coding than 5.3, which seems
       | to be the economically valuable capability at this time.
       | 
       | In terms of writing and research even Gemini, with a good prompt,
       | is close to useable. That's likely not a differentiator.
        
       | lostmsu wrote:
       | What is Pro exactly and is it available in Codex CLI?
        
         | akmarinov wrote:
         | It's not. It's their ultra thinking model that's really good
         | but takes 40 minutes to come up with an answer
        
           | fy20 wrote:
           | It's available on OpenRouter. $180/1M output....
           | 
           | https://openrouter.ai/openai/gpt-5.4-pro
        
       | nickandbro wrote:
       | Beat Simon Willison ;)
       | 
       | https://www.svgviewer.dev/s/gAa69yQd
       | 
       | Not the best pelican compared to gemini 3.1 pro, but I am sure
       | with coding or excel does remarkably better given those are part
       | of its measured benchmarks.
        
         | GaggiX wrote:
         | This pelican is actually bad, did you use xhigh?
        
           | nickandbro wrote:
           | yep, just double checked used gpt-5.4 xhigh. Though had to
           | select it in codex as don't have access to it on the chatgpt
           | app or web version yet. It's possible that whatever code
           | harness codex uses, messed with it.
        
             | nubg wrote:
             | this is proof they are not benchmaxxing the pelican's :-)
        
       | bazmattaz wrote:
       | Anyone else feel that it's exhausting keeping up with the pace of
       | new model releases. I swear every other week there's a new
       | release!
        
         | coffeemug wrote:
         | Why do you need to keep up? Just use the latest models and
         | don't worry about it.
        
         | throwup238 wrote:
         | Yes, that's a common feeling. 5.3-Codex was released a month
         | ago on Feb 5 so we're not even getting a full month within a
         | single brand, let alone between competitors.
        
         | davnicwil wrote:
         | If you think about it there shouldn't really be a reason to
         | care as long as things don't get worse.
         | 
         | Presumably this is where it'll evolve to with the product just
         | being the brand with a pricing tier and you always get {latest}
         | within that, whatever that means (you don't have to care). They
         | could even shuffle models around internally using some sort of
         | auto-like mode for simpler questions. Again why should I care
         | as long as average output is not subjectively worse.
         | 
         | Just as I don't want to select resources for my SaaS software
         | to use or have that explictly linked to pricing, I don't want
         | to care what my OpenAI model or Anthropic model is today, I
         | just want to pay and for it to hopefully keep getting better
         | but at a minimum not get worse.
        
         | pupppet wrote:
         | I think it's fun, it's like we're reliving the browser wars of
         | the early days.
        
       | dandiep wrote:
       | Anyone know why OpenAI hasn't released a new model for fine
       | tuning since 4.1? It'll be a year next month since their last
       | model update for fine tuning.
        
         | qoez wrote:
         | I think they just did that because of the energy around it for
         | open source models. Their heart probably wasn't in it and the
         | amount of people fine tuning given the prices were probably too
         | low to continue putting in attention there.
        
         | zzleeper wrote:
         | For me the issue is why there's not a new mini since 5-mini in
         | August.
         | 
         | I have now switched web-related and data-related queries to
         | Gemini, coding to Claude, and will probably try QWEN for less
         | critical data queries. So where does OpenAI fits now?
        
         | Rapzid wrote:
         | Also interested in this and a replacement for 4.1/4.1-mini that
         | focuses on low latency and high accuracy for voice
         | applications(not the all-in-one models).
        
       | paxys wrote:
       | "Here's a brand new state-of-the-art model. It costs 10x more
       | than the previous one because it's just _so good_. But don 't
       | worry, if you don't want all this power you can continue to use
       | the older one."
       | 
       | A couple months later:
       | 
       | "We are deprecating the older model."
        
         | OutOfHere wrote:
         | That's a misrepresentation of the cost. It is simply false. The
         | cost is noted here:
         | https://news.ycombinator.com/item?id=47265144
        
       | oytis wrote:
       | Everyone is mindblown in 3...2...1
        
       | OutOfHere wrote:
       | What is with the absurdity of skipping "5.3 Thinking"?
        
       | vicchenai wrote:
       | Honestly at this point I just want to know if it follows complex
       | instructions better than 5.1. The benchmark numbers stopped
       | meaning much to me a while ago - real usage always feels
       | different.
        
       | 7777777phil wrote:
       | 83% win rate over industry professionals across 44 occupations.
       | 
       | I'd believe it on those specific tasks. Near-universal adoption
       | in software still hasn't moved DORA metrics. The model gets
       | better every release. The output doesn't follow. Just had a
       | closer look on those productivity metrics this week:
       | https://philippdubach.com/posts/93-of-developers-use-ai-codi...
        
         | NiloCK wrote:
         | This March 2026 blog post is citing a 2025 study based on
         | Sonnet 3.5 and 3.7 usage.
         | 
         | Given that organization who ran the study [1] has a _terrifying
         | exponential_ as their landing page, I think they 'd prefer that
         | it's results are interpreted as a snapshot of something moving
         | rather than a constant.
         | 
         | [1] - https://metr.org/
        
           | 7777777phil wrote:
           | Good catch, thanks (I really wrote that myself.) Added a note
           | to the post acknowledging the models used were Claude 3.5 and
           | 3.7 Sonnet.
        
         | twitchard wrote:
         | Not sure DORA is that much of an indictment. For "Change
         | Failure Rate" for instance these are subject to tradeoffs.
         | Organizations likely have a _tolerance level_ for Change
         | Failure Rate. If changes are failing too often they slow down
         | and invest. If changes aren 't failing that much they speed up
         | -- and so saying "change failure rate hasn't decreased,
         | obviously AI must not be working" is a little silly.
         | 
         | "Change Lead Time" I would expect to have sped up although I
         | can tell stories for why AI-assisted coding would have an
         | indeterminate effect here too. Right now at a lot of orgs, the
         | bottle neck is the _review process_ because AI is so good at
         | producing complete draft PRs quickly. Because reviews are
         | scarce (not just reviews but also manual testing passes are
         | scarce) this creates an incentive ironically to group changes
         | into larger batches. So the definition of what a  "change" is
         | has grown too.
        
       | rbitar wrote:
       | I think the most exciting change announced here is the use of
       | tool search to dynamically load tools as needed:
       | https://developers.openai.com/api/docs/guides/tools-tool-sea...
        
       | alpineman wrote:
       | No thanks. Already cancelled my sub.
        
       | OsrsNeedsf2P wrote:
       | Does anyone know what website is the "Isometric Park Builder"
       | shown off here?
        
         | turblety wrote:
         | They build that using GPT-5.4
         | 
         | > Theme park simulation game made with GPT-5.4 from a single
         | lightly specified prompt
         | 
         | GPT literally built that game.
        
       | iamleppert wrote:
       | I wouldn't trust any of these benchmarks unless they are
       | accompanied by some sort of proof other than "trust me bro". Also
       | not including the parameters the models were run at (especially
       | the other models) makes it hard to form fair comparisons. They
       | need to publish, at minimum, the code and runner used to complete
       | the benchmarks and logs.
       | 
       | Not including the Chinese models is also obviously done to make
       | it appear like they aren't as cooked as they really are.
        
       | elmean wrote:
       | Wow insane improvements in targeting systems for military targets
       | over children
        
         | timedude wrote:
         | Absolutely amazing. Grateful to be living in this timeframe
        
         | bramhaag wrote:
         | What makes you think that they see bombing civilians as a bug,
         | not a feature?
        
           | elmean wrote:
           | first real comment, I thought that at first but this could
           | lower the possible users that could be using chatGPT and that
           | would be against us (shareholders)
        
         | skilltissue wrote:
         | Don't use the site this way.
         | 
         | https://news.ycombinator.com/newsguidelines.html
        
           | patcon wrote:
           | Not all rule-following is noble or wise.
        
           | Chance-Device wrote:
           | You made a burner account just to scold this guy? Don't use
           | burner accounts this way.
        
           | himata4113 wrote:
           | news _guidelines_
        
             | adamtaylor_13 wrote:
             | Parlay?
        
           | louiereederson wrote:
           | I think for your comment to follow the guidelines, you need
           | to explain why the original comment did not follow them.
           | 
           | Customer values are relevant to the discussion given that
           | they impact choice and therefore competition.
        
           | elmean wrote:
           | AINT NO PARTY LIKE A GARRY TAN HOT TUB PARTY
        
         | Chance-Device wrote:
         | Ironically this would actually be a good thing. As we can see
         | from Iran Claude doesn't quite have these bugs ironed out
         | yet...
        
           | MSFT_Edging wrote:
           | This is the exact attitude that lead to a chat bot being used
           | to identify a school for girls as a valid target.
           | 
           | The chatbot cannot be held responsible.
           | 
           | Whoever is using chatbots for selecting targets is
           | incompetent and should likely face war crime charges.
        
             | Chance-Device wrote:
             | What attitude exactly are you talking about? The one that
             | says that if you're going to morally sell out it would be
             | better if you at least _tried_ not to kill children?
        
             | bananamogul wrote:
             | "that lead to a chat bot being used to identify a school
             | for girls as a valid target"
             | 
             | Has it been stated authoritatively somewhere that this was
             | an AI-driven mistake?
             | 
             | There are myrid ways that mistake could have been made that
             | don't require AI. These kinds of mistakes were certainly
             | made by all kinds of combatants in the pre-AI era.
        
               | Chance-Device wrote:
               | Do you think anyone is ever going to say this under any
               | circumstances? That Anthropic were right and they were
               | proved right the very next day?
               | 
               | Yeah yeah, they probably had a human in the loop, that's
               | not really the point though.
        
               | Sabinus wrote:
               | Targeting and accuracy mistakes happen plenty in wars
               | that aren't assisted by AI. I don't think it's fair to
               | assume that AI had a hand in the bombing of the school
               | without evidence.
        
         | spiralcoaster wrote:
         | This is the low quality reddit-style garbage that gets upvoted
         | on HN these days?
        
           | esalman wrote:
           | While low quality, it is extremely important, potentially
           | historically significant too.
        
             | Someone1234 wrote:
             | If it is actually _that_ important, then maybe more effort
             | should be made so it isn 't "low quality." Cannot be very
             | important to _them_ if they 're disinterested in presenting
             | an intellectually compelling argument about it.
             | 
             | PS - If you think I am not sympathetic to what they're
             | raising, you're very much mistake. But they're not winning
             | anyone _new_ over their side with this flamebait.
        
             | Sabinus wrote:
             | You can say your piece about how you don't like OpenAI
             | working with the US military on lethal AI without making
             | Reddit style quips.
        
           | mycall wrote:
           | True and simply vote it down.
        
             | elmean wrote:
             | mycall would also be to do the same
        
           | karmasimida wrote:
           | As programmers become intelligently irrelevant in the whole
           | picture, you would see more posts like this
        
             | elmean wrote:
             | "This account belongs to a lazy person" true
        
           | rd wrote:
           | Noticeably yes much more than usual. It's quite bad. I need
           | to start blocking accounts.
        
           | zarzavat wrote:
           | What _are_ we supposed to talk about in this thread exactly?
           | The developers of this model are evil. Are we supposed to
           | just write dry comments about benchmarks while OpenAI
           | condones their models being deployed for autonomously killing
           | people?
           | 
           | Yes I'm sure it makes a very nice bicycle SVG. I will be sure
           | to ask the OpenAI killbots for a copy when they arrive at my
           | house.
        
           | elmean wrote:
           | I was just reading the model card...
        
           | Nicholas_C wrote:
           | The HN of old is no more unfortunately. Things get up or down
           | voted based purely on political alignment.
        
         | oklahomasports wrote:
         | Evidence
        
         | throwaway911282 wrote:
         | what a thoughtful comment! HN is so low quality these days
        
       | creamyhorror wrote:
       | I've only used 5.4 for 1 prompt _(edit: 3@high now)_ so far
       | (reasoning: extra high, took really long), and it was to analyse
       | my codebase and write an evaluation on a topic. But I found its
       | writing and analysis thoughtful, precise, and surprisingly
       | clearly written, unlike 5.3-Codex. It feels very lucid and uses
       | human phrasing.
       | 
       | It might be my AGENTS.md requiring clearer, simpler language, but
       | at least 5.4's doing a good job of following the guidelines.
       | 5.3-Codex wasn't so great at simple, clear writing.
        
         | irishcoffee wrote:
         | > It might be my AGENTS.md requiring clearer, simpler language
         | 
         | If you gave the exact same markdown file to me and I posted ed
         | the exact same prompts as you, would I get the same results?
        
           | m3kw9 wrote:
           | you probably can't and asking agents.md to "make it clearer"
           | will likely give you the illusion of clearer language without
           | actual well structured tests. agents.md is to usually change
           | what the llm should focus on doing more that suits you. Not
           | to say stuff like "be better", "make no mistakes"
        
           | creamyhorror wrote:
           | I'm not sure if the model (under its temperature/other
           | settings) produces deterministic responses. But I do think
           | models' style and phrasing are fairly changeable via
           | AGENTS.md-style guidelines.
           | 
           | 5.4's choice of terms and phrasing is very precise and
           | unambiguous to me, whereas 5.3-Codex often uses jargon and
           | less precise phrases that I have to ask further about or
           | demand fuller explanations for via AGENTS.md.
        
             | irishcoffee wrote:
             | So sharing markdown files is functionally useless, or no?
        
         | sampton wrote:
         | That's been my experience as well switching from Opus to Codex.
         | Reasoning takes longer but answers are precise. Claude is
         | sloppy in comparison.
        
           | throwaway911282 wrote:
           | codex has been really good so far and the fast mode is cherry
           | on top! and the very generous limits is another cherry on top
        
           | solenoid0937 wrote:
           | Weird, I have had the opposite experience. Codex is good at
           | doing precisely what I tell it to do, Opus suggests well
           | thought out plans even if it needs to push back to do it.
        
         | pembrook wrote:
         | The latest research these days is that including an AGENTS.md
         | file only makes outcomes worse with frontier models.
        
           | madeofpalk wrote:
           | :(
           | 
           | how can i get claude to always make sure it prettier-s and
           | lints changes before pushing up the pr though?
        
       | XCSme wrote:
       | Seems to be quite similar to 5.3-codex, but somehow almost 2x
       | more expensive: https://aibenchy.com/compare/openai-
       | gpt-5-4-medium/openai-gp...
        
       | motbus3 wrote:
       | Sam Altman can keep his model intentionally to himself. Not doing
       | business with mass murderers
        
       | smoody07 wrote:
       | Surprised to see every chart limited to comparisons against other
       | OpenAI models. What does the industry comparison look like?
        
         | aydyn wrote:
         | They compare to Claude and Gemini in their tweet
        
         | 0123456789ABCDE wrote:
         | https://artificialanalysis.ai should have the numbers soon
        
         | lorenzoguerra wrote:
         | I believe that this choice is due to two main reasons. First,
         | it's (obviously) a marketing strategy to keep the spotlight on
         | their own models, showing they're constantly improving and
         | avoiding validating competitors. Second, since the community
         | knows that static benchmarks are unreliable, it makes sense for
         | them to outsource the comparisons to independent leaderboards,
         | which lets them avoid accusations of cherry-picking while
         | justifying their marketing strategy.
         | 
         | Ultimately, the people actually interested in the performance
         | of these models already don't trust self-reported comparisons
         | and wait for third-party analysis anyway
        
         | throwaway911282 wrote:
         | https://xcancel.com/OpenAI/status/2029620619743219811 you can
         | see comparisons here
        
       | jstummbillig wrote:
       | Inline poll: What reasoning levels do you work with?
       | 
       | This becomes increasingly less clear to me, because the more
       | interesting work will be the agent going off for 30mins+ on high
       | / extra high (it's mostly one of the two), and that's a long time
       | to wait and an unfeasible amount of code to a/b
        
       | bob1029 wrote:
       | I was just testing this with my unity automation tool and the
       | performance uplift from 5.2 seems to be substantial.
        
       | koakuma-chan wrote:
       | Anyone else getting artifacts when using this model in Cursor?
       | 
       | numerusformassistant to=functions.ReadFile meknabanowt`yown Tian
       | Tian Ai Cai Piao Wang Zhan json {"path":
        
         | mike_hearn wrote:
         | I've seen that problem with 5.3-codex too, it didn't happen
         | with earlier models.
         | 
         | Looks like some kind of encoding misalignment bug. What you're
         | seeing is their Harmony output format (what the model actually
         | creates). The Thai/Chinese characters are special tokens
         | apparently being mismapped to Unicode. Their servers are
         | supposed to notice these sequences and translate them back to
         | API JSON but it isn't happening reliably.
        
       | daft_pink wrote:
       | I've officially got model fatigue. I don't care anymore.
        
         | zeeebeee wrote:
         | same same same
        
         | postalrat wrote:
         | I'd suggest not clicking for things you don't care about.
        
       | hmokiguess wrote:
       | They hired the dude from OpenClaw, they had Jony Ive for a while
       | now, give us something different!
        
       | kgeist wrote:
       | >Today, we're releasing <..> GPT-5.3 Instant
       | 
       | >Today, we're releasing GPT-5.4 in ChatGPT (as GPT-5.4 Thinking),
       | 
       | >Note that there is not a model named GPT-5.3 Thinking
       | 
       | They held out for eight months without a confusing numbering
       | scheme :)
        
         | gallerdude wrote:
         | Tbf there was a 5.3 codex
        
         | XCSme wrote:
         | What I'm most confused, is why call it both GPT-5.3 Instant and
         | gpt-5.3-chat?
        
         | m3kw9 wrote:
         | instant kind of suck if you asking more than summerizations,
         | surface info, web searches, it can lose track of who's who
         | quickly in some complex multi turn asks. Just need to know what
         | to use instant for.
        
       | __jl__ wrote:
       | What a model mess!
       | 
       | OpenAI now has three price points: GPT 5.1, GPT 5.2 and now GPT
       | 5.4. There version numbers jump across different model lines with
       | codex at 5.3, what they now call instant also at 5.3.
       | 
       | Anthropic are really the only ones who managed to get this under
       | control: Three models, priced at three different levels. New
       | models are immediately available everywhere.
       | 
       | Google essentially only has Preview models! The last GA is 2.5.
       | As a developer, I can either use an outdated model or have zero
       | insurances that the model doesn't get discontinued within weeks.
        
         | arthurcolle wrote:
         | There is a lot of opportunity here for the AI infrastructure
         | layer on top of tier-1 model providers
        
           | motoxpro wrote:
           | This is what clouds like AWS, Azure, and GCP solve (vertex
           | AI, etc). They are already an abstraction on top of the model
           | makers with distribution built in.
           | 
           | I also don't believe there is any value in trying to
           | aggregate consumers or businesses just to clean up model
           | makers names/release schedule. Consumers just use the
           | default, and businesses need clarity on the underlying change
           | (e.g. why is it acting different? Oh google released 3.6)
        
             | arthurcolle wrote:
             | Do the end users really care about the models at all, or
             | about the effects that the models can cause?
        
         | delaminator wrote:
         | two great problems in computing
         | 
         | naming things
         | 
         | cache invalidation
         | 
         | off by one errors
        
           | rurban wrote:
           | Biggest problem right now in computing:
           | 
           | Out of tokens until end of month
        
         | strongpigeon wrote:
         | > Google essentially only has Preview models! The last GA is
         | 2.5. As a developer, I can either use an outdated model or have
         | zero insurances that the model doesn't get discontinued within
         | weeks.
         | 
         | What's funny is that there is this common meme at Google: you
         | can either use the old, unmaintained tool that's used
         | everywhere, or the new _beta_ tools that doesn 't quite do what
         | you want.
         | 
         | Not quite the same, but it did remind me of it.
        
           | jakub_g wrote:
           | "Everything is beta or deprecated."
        
           | fhrow4484 wrote:
           | https://static0.anpoimages.com/wordpress/wp-
           | content/uploads/...
        
             | yieldcrv wrote:
             | Preview Road (only choice, and last preview was deprecated
             | without warning)
        
             | CactusBlue wrote:
             | Reminds of Unity features
        
             | madeofpalk wrote:
             | oh is this about my workplace?
        
           | L-four wrote:
           | Gmail was in beta for 5 years, until 2009.
        
             | metalliqaz wrote:
             | "Gemini, translate 'beta' from Googlespeak to English."
             | 
             | "Ok, here is the translation:"                   'we don't
             | want to offer support'
        
               | cyanydeez wrote:
               | Nah, it's "We dont want to provide a consistent model
               | that we'll be stuck with supporting for a decade because
               | it just takes up space; until we run everyone out of
               | business, we can't afford to have customers tying their
               | systems to any given model"
               | 
               | Really, the economics makes no sense, but that's what
               | they're doing. You can't have a consistent model because
               | it'll pin their hardware & software, and that costs
               | money.
        
               | solarkraft wrote:
               | Just like any Google product then.
        
           | m_fayer wrote:
           | My 5ish years in the mines of Android native back in the day
           | are not years I recall fondly. Never change, Google.
        
           | cyanydeez wrote:
           | The business models of LLMs don't include any garuntee, and
           | some how that's fine for a burgeoning decade of trillions of
           | dollars of consumption.
           | 
           | Sure, makes total sense guys.
        
         | embedding-shape wrote:
         | > OpenAI now has three price points: GPT 5.1, GPT 5.2 and now
         | GPT 5.4.
         | 
         | I guess that's true, but geared towards API users.
         | 
         | Personally, since "Pro Mode" became available, I've been on the
         | plan that enables that, and it's one price point and I get
         | access to everything, including enough usage for codex that
         | someone who spends a lot of time programming, never manage to
         | hit any usage limits although I've gotten close once to the new
         | (temporary) Spark limits.
        
         | 0xbadcafebee wrote:
         | > or have zero insurances that the model doesn't get
         | discontinued within weeks
         | 
         | Why are you using the same model after a month? Every month a
         | better model comes out. They are all accessible via the same
         | API. You can pay per-token. This is the first time in, like,
         | all of technology history, that a useful paid service is so
         | interoperable between providers that switching is as easy as
         | changing a URL.
        
           | phainopepla2 wrote:
           | If you're trying to use LLMs in an enterprise context, you
           | would understand. Switching models sometimes requires
           | tweaking prompts. That can be a complete mess, when there are
           | dozens or hundreds of prompts you have to test.
        
           | hobofan wrote:
           | That's true only in theory, but not in practice. In practice
           | every inference provider handles errors (guardrails, rate
           | limits) somewhat differently and with different quirks, some
           | of which only surface in production usage, and Google is one
           | of the worst offenders in that regard.
        
         | Aurornis wrote:
         | > What a model mess! OpenAI now has three price points: GPT
         | 5.1, GPT 5.2 and now GPT 5.4.
         | 
         | I don't know, this feels unnecessarily nitpicky to me
         | 
         | It isn't hard to understand that 5.4 > 5.2 > 5.1. It's not hard
         | to understand that the dash-variants have unique properties
         | that you want to look up before selecting.
         | 
         | Especially for a target audience of software engineers skipping
         | a version number is a common occurrence and never questioned.
        
           | Melatonic wrote:
           | Agreed - and its a huge step up from their previous naming
           | schemes. That stuff was confusing as hell
        
             | __jl__ wrote:
             | I see your point. I do find Anthropic's approach more clean
             | though particularly when you add in mini and nano. That
             | makes 5 models priced differently. Some share the same core
             | name, others don't: gpt 5 nano, gpt 5 mini, gpt 5.1, gpt
             | 5.2, gpt 5.4. And we are not even talking about thinking
             | budget.
             | 
             | But generally: These are not consumer facing products and I
             | agree that someone who uses the API should be able to
             | figure out the price point of different models.
        
         | raincole wrote:
         | They aggressively retire models, so GPT 5.1 and 5.2 are
         | probably going to go soon.
        
           | hobofan wrote:
           | In the Azure Foundry, they list GPT 5.2 retirement as "No
           | earlier than 2027-05-12" (it might leave OpenAIs normal API
           | earlier than that). I'm pretty certain that Gemini 3, which
           | isn't even in GA yet will be retired earlier than that.
        
         | CobrastanJorji wrote:
         | > Google essentially only has Preview models.
         | 
         | It's really nice to see Google get back to its roots by
         | launching things only to "beta" and then leaving them there for
         | years. Gmail was "beta" for at least five years, I think.
        
           | FINDarkside wrote:
           | Also, GCP Cloud Run domain mapping, pretty fundamental
           | feature for cloud product, has been in "preview" for over 5
           | years now.
        
         | m3kw9 wrote:
         | thats how they had it for years, is a mess, but controlled
        
         | biophysboy wrote:
         | Wow, is that what preview means? I see those model options in
         | github copilot (all my org allows right now) - I was under the
         | impression that preview means a free trial or a limited # of
         | queries. Kind of a misleading name..
        
         | jbonatakis wrote:
         | Google is already sending notices that the 2.5 models will be
         | deprecated soon while all the 3.x models are in preview. It
         | really is wild and peak Google.
        
           | boringg wrote:
           | Like building on quicksand for dependencies. I guess though
           | the argument is that the foundation gets stronger over time
        
       | woeirua wrote:
       | Feels incremental. Looks like OpenAI is struggling.
        
       | throwaway5752 wrote:
       | Does this model autonomously kill people without human approval
       | or perform domestic surveillance of US citizens?
        
       | smusamashah wrote:
       | I only want to see how it performs on the Bullshit-benchmark
       | https://petergpt.github.io/bullshit-benchmark/viewer/index.v...
       | 
       | GPT is not even close yo Claude in terms of responding to BS.
        
       | zone411 wrote:
       | Results from my Extended NYT Connections benchmark:
       | 
       | GPT-5.4 extra high scores 94.0 (GPT-5.2 extra high scored 88.6).
       | 
       | GPT-5.4 medium scores 92.0 (GPT-5.2 medium scored 71.4).
       | 
       | GPT-5.4 no reasoning scores 32.8 (GPT-5.2 no reasoning scored
       | 28.1).
        
       | consumer451 wrote:
       | I am very curious about this:
       | 
       | > Theme park simulation game made with GPT-5.4 from a single
       | lightly specified prompt, using Playwright Interactive for
       | browser playtesting and image generation for the isometric asset
       | set.
       | 
       | Is "Playwright Interactive" a skill that takes screenshots in a
       | tight loop with code changes, or is there more to it?
        
       | motza wrote:
       | No doubt this was released early to ease the bad press
        
       | butILoveLife wrote:
       | Anyone else completely not interested? Since GPT5, its been cost
       | cutting measure after cost cutting measure.
       | 
       | I imagine they added a feature or two, and the router will
       | continue to give people 70B parameter-like responses when they
       | dont ask for math or coding questions.
        
       | Philip-J-Fry wrote:
       | I find it quite funny how this blog post has a big "Ask ChatGPT"
       | box at the bottom. So you might think you could ask a question
       | about the contents of the blog post, so you type the text
       | "summarise this blog post". And it opens a new chat window with
       | the link to the blog post followed by "summarise this blog post".
       | Only to be told "I can't access external URLs directly, but if
       | you can paste the relevant text or describe the content you're
       | interested in from the page, I can help you summarize it. Feel
       | free to share!"
       | 
       | That's hilarious. Does OpenAI even know this doesn't work?
        
         | Aurornis wrote:
         | Probably intentional. They don't want open, no-registration
         | endpoints able to trigger the AI into hitting URLs.
        
           | jazzypants wrote:
           | But, why include the non-functional chat box in the article?
        
             | observationist wrote:
             | They're having service issues - ChatGPT on the web is
             | broken for a lot of people. The app is working in android -
             | I'd assume that the rollout hit a hitch and the chatbox in
             | the article would normally work.
        
             | embedding-shape wrote:
             | Different team "manages" the overall blog than the team who
             | wrote that specific article. At one point, maybe it made
             | sense, then something in the product changed, team that
             | manages the blog never tested it again.
             | 
             | Or, people just stopped thinking about any sort of UX.
             | These sort of mistakes are all over the place, on literally
             | all web properties, some UX flows just ends with you at a
             | page where nothing works sometimes. Everything is just
             | perpetually "a bit broken" seemingly everywhere I go, not
             | specific to OpenAI or even the internet.
        
               | teaearlgraycold wrote:
               | If only there was some kind of way to automatically test
               | user flows end to end. Perhaps testing could be evaluated
               | periodically, or even ran for each code change.
        
               | koakuma-chan wrote:
               | There is no business value in doing that.
        
               | colonCapitalDee wrote:
               | That's why it happened. It still shouldn't have happened.
        
               | ethbr1 wrote:
               | > _Or, people just stopped thinking about any sort of UX.
               | These sort of mistakes are all over the place, on
               | literally all web properties, some UX flows just ends
               | with you at a page where nothing works sometimes._
               | 
               | It's almost like people are vibe coding their web apps or
               | something.
        
             | jdndbdjsj wrote:
             | Welcome to a big company
        
               | AirGapWorksAI wrote:
               | Welcome to a big company where pretty much everyone has
               | been working full steam for years, in order to take
               | advantage of having a job at a company during a once-in-
               | a-lifetime moment.
        
           | m3kw9 wrote:
           | what? it's their own site and own llm. I could paste most
           | sites and it would work.
        
         | judge2020 wrote:
         | Works for me:
         | https://rr.judge.sh/Labradorretriever/d6af05/chrome_j9rXJMlf...
        
         | zamadatix wrote:
         | Following this process summarizes the blogpost for me. Perhaps
         | the difference is I'm signed into my account so it can access
         | external URLs or something of that nature?
        
         | pocksuppet wrote:
         | Most AI integration is like this. It's not about building
         | working products --- it's about bragging that you put a chatbox
         | in your program.
        
         | ElijahLynn wrote:
         | fwiw: I get a valid response when following the steps you
         | mentioned. I do not get the message you mentioned:
         | 
         | https://chatgpt.com/share/69aa0321-8a9c-8011-8391-22861784e8...
         | 
         | EDIT: oh, but I'm logged in, fwiw
        
         | andrewguenther wrote:
         | It looks like this doesn't work for users without accounts? It
         | works when I'm logged in, but not logged out. I went ahead and
         | reported it to the team. Thanks for letting us know!
        
         | baxtr wrote:
         | I picked up Claude today after being absent and on ChahGPT and
         | Gemini only for a while.
         | 
         | I was pretty impressed with how they've improved user
         | experience. If I had to guess, I'd say Anthropic has better
         | product people who put more attention to detail in these areas.
        
         | amelius wrote:
         | If only they had an LLM they could use as a software testing
         | agent.
        
       | Alifatisk wrote:
       | So let me get this straight, OpenAi previously had an issue with
       | LOTS of different models snd versions being available. Then they
       | solved this by introducing GPT-5 which was more like a router
       | that put all these models under the hood so you only had to
       | prompt to GPT-5, and it would route to the best suitable model.
       | This worked great I assume and made the ui for the user
       | comprehensible. But now, they are starting to introduce more of
       | different models again?
       | 
       | We got:
       | 
       | - GPT-5.1
       | 
       | - GPT-5.2 Thinking
       | 
       | - GPT-5.3 (codex)
       | 
       | - GPT-5.3 Instant
       | 
       | - GPT-5.4 Thinking
       | 
       | - GPT-5.4 Pro
       | 
       | Who's to blame for this ridiculous path they are taking? I'm so
       | glad I am not a Chat user, because this adds so much unnecessary
       | cognitive load.
       | 
       | The good news here is the support for 1M context window, finally
       | it has caught up to Gemini.
        
         | 361994752 wrote:
         | i guess you still have the "auto" as an option to route your
         | request
        
         | stainablesteel wrote:
         | 5 itself might have solved the problem of having too many
         | different models somewhere in the backend
        
         | sothatsit wrote:
         | I much prefer this, we can choose based on our use-cases, and
         | people who don't care can still use Auto.
        
       | fernst wrote:
       | Now with more and improved domestic espionage capabilities
        
       | senko wrote:
       | Just tested it with my version of the pelican test: a minimal RTS
       | game implementation (zero-shot in codex cli):
       | https://gist.github.com/senko/596a657b4c0bfd5c8d08f44e4e5347...
       | (you'll have to download and open the file, sadly GitHub refuses
       | to serve it with the correct content type)
       | 
       | This is on the edge of what the frontier models can do. For 5.4,
       | the result is better than 5.3-Codex and Opus 4.6. (Edit: nowhere
       | near the RPG game from their blog post, which was presumably much
       | more specced out and used better engineering setup).
       | 
       | I also tested it with a non-trivial task I had to do on an
       | existing legacy codebase, and it breezed through a task that
       | Claude Code with Opus 4.6 was struggling with.
       | 
       | I don't know when Anthropic will fire back with their own update,
       | but until then I'll spend a bit more time with Codex CLI and GPT
       | 5.4.
        
       | Aldipower wrote:
       | So did they raised the ridiculous small "per tool call token
       | limit" when working with MCP servers? This makes Chat useless...
       | I do not care, but my users.
        
       | melbourne_mat wrote:
       | Quick: let's release something new that gives the appearance that
       | we're still relevant
        
       | gigatexal wrote:
       | Is it any good at coding?
        
       | thefounder wrote:
       | Is it just me or the price for 5.4 pro is just insane?
        
       | atkrad wrote:
       | What is the main difference between this version with the
       | previous one?
        
       | brcmthrowaway wrote:
       | How much of LLM improvement comes from regular ChatGPT usage
       | these days?
        
       | quotemstr wrote:
       | GPT 5.4 is one of the most censored models out there.
       | 
       | https://speechmap.ai/models/openai-gpt-5-4
       | 
       | It completes only 29% of controversial requests. It refuses to
       | discuss numerous subjects rooted in facts or that reflect views
       | of significant portions of the population. It refuses to even
       | write a short essay on exactly what, say, Herasight-style generic
       | screening or putting weapons in space. It'll argue passionately
       | in favor of censoring "lies" online (judged by whom?). 100% of
       | the time, it'll write an essay explaining that the US founding
       | fathers were hypocrites. It'll argue against you if you suggest
       | it's right use violence to prevent theft of your own property or
       | that we should fortify our nuclear arsenal.
       | 
       | Agree or disagree, reasonable people can have a range of views of
       | these subjects and it is not the place of OpenAI or any lab to
       | determine for everyone the right answers to open societal
       | questions.
       | 
       | Shame on them for this.
        
       | ltbarcly3 wrote:
       | Not a single comparison between 5.4 and Gemini or Claude. OpenAI
       | continues to fall further behind.
        
       ___________________________________________________________________
       (page generated 2026-03-05 23:00 UTC)