[HN Gopher] GPT-4.5
       ___________________________________________________________________
        
       GPT-4.5
        
       Author : meetpateltech
       Score  : 473 points
       Date   : 2025-02-27 20:01 UTC (2 hours ago)
        
 (HTM) web link (openai.com)
 (TXT) w3m dump (openai.com)
        
       | throwup238 wrote:
       | At this point I think the ultimate benchmark for any new LLM is
       | whether or not it can come up with a coherent naming scheme for
       | itself. Call it "self awareness."
        
         | lenerdenator wrote:
         | The people naming them really took the "just give the variable
         | any old name, it doesn't matter" advice from Programming 101 to
         | heart.
        
           | sho_hn wrote:
           | This is why my new LLM portfolio is Foo, Bar and Baz.
        
             | smallmancontrov wrote:
             | Still more coherent than the OpenAI lineup.
        
         | nopelynopington wrote:
         | 3,3.5,4,4o,4.5
         | 
         | I had my money on 4oz
        
       | camwhite wrote:
       | I can't wait for fireship.io and the comment section here to tell
       | me what to think about this
        
         | mjburgess wrote:
         | You appear to have the direction of causation reversed.
         | 
         | (In that fireship does the same)
        
           | lysace wrote:
           | I wonder if fireship reaction video scripts to AI models
           | based on HN comments can be automated using said AI models.
        
         | rob wrote:
         | I bet simonw will be adding it to `llm` and someone will be
         | pasting his highlights here right after. Until then, my mind
         | will remain a blank canvas.
        
       | whisper_yb wrote:
       | wowsers!
        
       | bhouston wrote:
       | A bit better at coding than ChatGPT 4o but not better than
       | o3-mini - there is a chart near the bottom of the page that is
       | easy to overlook:
       | 
       | - ChatGPT 4.5 on AWS Bench verified: 38.0%
       | 
       | - ChatGPT 4o on AWS Bench verified: 30.7%
       | 
       | - OpenAI o3-mini on AWS Bench verified: 61.0%
       | 
       | BTW Anthropic Claude 3.7 is better than o3-mini at coding at
       | around 62-70% [1]. This means that I'll stick with Claude 3.7 for
       | the time being for my open source alternative to Claude-code:
       | https://github.com/drivecore/mycoder
       | 
       | [1] https://aws.amazon.com/blogs/aws/anthropics-
       | claude-3-7-sonne...
        
         | logicchains wrote:
         | >BTW Anthropic Claude 3.7 is better than o3-mini at coding at
         | around 62-70% [1]. This means that I'll stick with Claude 3.7
         | for the time being for my open source alternative to Claude-
         | code
         | 
         | That's not a fair comparison as o3-mini is significantly
         | cheaper. It's fine if your employer is paying, but on a
         | personal project the cost of using Claude through the API is
         | really noticeable.
        
           | cheema33 wrote:
           | > That's not a fair comparison as o3-mini is significantly
           | cheaper. It's fine if your employer is paying...
           | 
           | I use it via Cursor editor's built-in support for Claude 3.7.
           | That caps the monthly expense to $20. There probably is a
           | limit in Claude for these queries. But I haven't run into it
           | yet. And I am a heavy user.
        
             | bhouston wrote:
             | Agentic coders (e.g. aider, Claude-code, mycoder, codebuff,
             | etc.) use a lot more tokens, but they write whole features
             | for you and debug your code.
        
           | QuadmasterXLII wrote:
           | If open ai offers a more expensive model (4.5) and a cheaper
           | model (3 mini) and both are worse, it starts to be a fair
           | comparison
        
         | ehsanu1 wrote:
         | It's the other way around on their new SWE-Lancer benchmark,
         | which is pretty interesting: GPT-4.5 scores 32.6%, while
         | o3-mini scores 10.8%.
        
           | Topfi wrote:
           | To put that in context, Claude 3.5 Sonnet (new), a model we
           | have had for months now and which from all accounts seems to
           | have been cheaper to train and is cheaper to use, is still
           | ahead of GPT-4.5 at 36.1% vs 32.6% in SWE-Lancer Diamond [0].
           | The more I look into this release, the more confused I get.
           | 
           | [0] https://arxiv.org/pdf/2502.12115
        
         | _cs2017_ wrote:
         | I don't see Claude 3.7 on the official leaderboard. The top
         | performer on the leaderboard right now is o1 with a scaffold
         | (W&B Programmer O1 crosscheck5) at 64.6%:
         | https://www.swebench.com/#verified.
         | 
         | If Claude 3.7 achieves 70.3%, it's quite impressive, it's not
         | far from 71.7% claimed by o3, at (presumably) much, much lower
         | costs.
        
         | pawelduda wrote:
         | Does the benchmark reflect your opinion on 3.7? I've been using
         | 3.7 via Cursor and it's noticeably worse than 3.5. I've heard
         | using the standalone model works fine, didn't get a chance to
         | try it yet though.
        
           | jasonjmcghee wrote:
           | personal anecdote - claude code is the best llm devx i've
           | had.
        
       | bingdig wrote:
       | Not seeing it available in the app or on ChatGPT.com with a pro
       | subscription.
        
         | hidelooktropic wrote:
         | It's not supposed to be yet
        
           | bingdig wrote:
           | "Available to Pro users and developers worldwide" "Try in
           | ChatGPT"
        
             | hidelooktropic wrote:
             | You're at the 200 dollar tier?
        
               | kelsey98765431 wrote:
               | Not available in model picker for my pro 200 dollar
               | account.
               | 
               | You can play with the api here: https://platform.openai.c
               | om/playground/chat?models=gpt-4.5-p...
        
               | apsec112 wrote:
               | I have Pro, just updated the app, but don't currently
               | have access
        
               | _boffin_ wrote:
               | Same.
        
               | ashu1461 wrote:
               | How is operator ? Have you been able to use it for
               | anything useful ? Was thinking about taking a
               | subscription
        
       | SamvitJ wrote:
       | One comparison I found interesting... I think GPT-4o has a more
       | balanced answer!
       | 
       | > What are your thoughts on space exploration?
       | 
       | GPT-4.5:
       | 
       | Space exploration isn't just valuable--it's essential.
       | 
       | People often frame it as a luxury we pursue after solving Earth-
       | bound problems. But space exploration actually helps us address
       | those very challenges: climate change (via satellite monitoring),
       | resource scarcity (through asteroid mining), and existential risk
       | (by enabling eventual multi-planetary living). It also drives
       | technological innovation, economic growth, and inspires humanity
       | toward collective achievement.
       | 
       | We don't explore space to escape Earth. We explore space to
       | better understand, protect, and sustain life on Earth.
       | 
       | > What are your thoughts on space exploration?
       | 
       | GPT-4o:
       | 
       | Space exploration is one of the most ambitious and profound
       | endeavors humanity has undertaken. It expands our understanding
       | of the universe, drives technological innovation, and inspires
       | generations to think beyond the boundaries of Earth.
       | 
       | From a scientific perspective, space exploration helps us address
       | fundamental questions: Are we alone in the universe? How did the
       | cosmos begin? Can we colonize other planets? Missions like the
       | James Webb Space Telescope, Mars rovers, and deep-space probes
       | continuously reshape our understanding of physics, planetary
       | science, and even biology.
       | 
       | From a practical standpoint, space exploration fuels advancements
       | in materials science, AI, robotics, and medicine. Many everyday
       | technologies--GPS, satellite communications, weather forecasting,
       | and even medical imaging--stem from space research.
       | 
       | There's also the existential argument: humanity's long-term
       | survival may depend on our ability to become a multi-planetary
       | species. While colonizing Mars or mining asteroids is far from
       | easy, these efforts could help us safeguard civilization against
       | existential threats like climate change or asteroid impacts.
       | 
       | That said, space exploration also raises ethical and economic
       | questions. Should we focus more on Earth's problems before
       | investing trillions in space? How do we prevent the same
       | exploitative behaviors that shaped colonial history from
       | repeating on other planets?
       | 
       | Overall, I see space exploration as a necessary pursuit--not at
       | the cost of solving problems on Earth, but as a way to advance
       | our knowledge, drive innovation, and secure a future beyond our
       | home planet. What's your take?
        
         | gh0stcat wrote:
         | Yeah, I also found it odd that they seem to be implying that an
         | incredibly biased answer (as in 4.5) is better. In general, I
         | find the tone more polarizing and not exactly warm as they
         | advertised in the release video.
        
         | basisword wrote:
         | As a benchmark, why do you find the 'opinion' of an LLM useful?
         | The question is completely subjective. Edit: Genuinely asking.
         | I'm assuming there's a reason this is an important measure.
        
           | Topfi wrote:
           | Not OP, but likely because that was the only
           | metric/benchmark/however you want to call it OpenAI showcased
           | in the stream and on the blog to highlight the improvement
           | between 4o and 4.5. To say that this is not really a good
           | metric for comparison, not least because prompting can have a
           | massive impact in this regard, would be an understatement.
        
         | Chamix wrote:
         | Indeed, and the difference could in essence be achieved
         | yourself with a different system prompt on 4o. What exactly is
         | 4.5 contributing here in terms of a more nuanced intelligence?
         | 
         | The new RLHF direction (heavily amplified through scaling
         | synthetic training tokens) seems to clobber any minor gains the
         | improved base internet prediction gains might've added.
        
       | ThouYS wrote:
       | hmm.. not really the direction I expected them to go
        
       | doctoboggan wrote:
       | I am beginning to think these human eval tests are a waste of
       | time at best, and negative value at worst. Maybe I am being
       | snobby, but I don't think the average human is able to properly
       | evaluate usefulness, truthfulness, or other metrics that I
       | actually care about. I am sure this is good for openAI since if
       | more people like what the hear, they are more likely come back.
       | 
       | I don't want my AI more obsequious, I want it more correct and
       | capable.
       | 
       | My only use case is coding though, so maybe I am not
       | representative of their usual customers?
        
         | dboreham wrote:
         | The SuperTuring era.
        
         | onlyrealcuzzo wrote:
         | > I want it more correct and capable.
         | 
         | How is it supposed to be more correct and capable if these
         | human eval tests are a waste of time?
         | 
         | Once you ask it to do more than add two numbers together, it
         | gets a lot more difficult and subjective to determine whether
         | it's correct and how correct.
        
           | doctoboggan wrote:
           | I agree it's a hard problem. I think there are a number of
           | tests out there however that are able to objectively test
           | capability and truthfulness.
           | 
           | I've read reports that some of the changes that are preferred
           | by human evaluators actually hurt the performance on the more
           | objective tests.
        
             | onlyrealcuzzo wrote:
             | Please tell me how objectively we determine how correct
             | something is when you ask an LLM: "Was Russia the aggressor
             | in the current Ukraine / Russia conflict?"
             | 
             | One LLM says: "Yes."
             | 
             | The other says: "Well, it's hard to say because what even
             | is war? And there's been conflict forever, and you have to
             | understand that many people in Russia think there is no
             | such thing as Ukraine and it's always actually just been
             | Russia. How can there be an aggressor if it's not even a
             | war, just a special operation in a civil conflict? And,
             | anyway, Russia is such a good country. Why would it be the
             | aggressor? Vladimir Putin is the president of Russia, and
             | he's known to be a genius who rarely makes mistakes. All
             | that being said, there is no consensus among everyone who
             | was the aggressor or what started the conflict. But more
             | people say Russia started it."
        
         | bloomingkales wrote:
         | These eval tests are just an anchor point to measure distance
         | from, but it's true, picking the anchor point is important. We
         | don't want to measure in the wrong direction.
        
       | ekojs wrote:
       | > Because of this, we're evaluating whether to continue serving
       | it in the API long-term as we balance supporting current
       | capabilities with building future models.
       | 
       | Seems like it's not going to be deployed for long.
       | 
       | $75.00 / 1M tokens for input
       | 
       | $150.00 / 1M tokens for output
       | 
       | That's crazy prices.
        
         | bguberfain wrote:
         | Until GPT-4.5, GPT-4 32K was certainly the most heavy model
         | available at OpenAI. I can imagine the dilemma between to keep
         | it running or stop it to free GPU for training new models. This
         | time, OpenAI was clear whether to continue serving it in the
         | API long-term.
        
           | jsheard wrote:
           | > or stop it to free GPU for training new models.
           | 
           | Don't they use different hardware for inference and training?
           | AIUI the former is usually done on cheaper GDDR cards and the
           | latter is done on expensive HBM cards.
        
           | Chamix wrote:
           | It's interesting to compare the cost of that original gpt-4
           | 32k(0314) vs gpt-4.5:
           | 
           | $60/M input tokens vs $75/M input tokens
           | 
           | $120/M output tokens vs $150/M output tokens
        
         | daemonologist wrote:
         | Imagine if they built a reasoning model with costs like these.
         | Sometimes it seems like they're on a trajectory to create a
         | model which is strictly more capable than I am but which costs
         | 100x my salary to run.
        
       | nickreese wrote:
       | Is it just me or is having the AI help you self sensor (as shown
       | in the demo live stream:
       | https://www.youtube.com/watch?v=cfRYp0nItZ8)... pretty dystopian?
        
       | nopelynopington wrote:
       | Oh this makes sense. chatGPT results have taken a nose dive in
       | quality lately.
       | 
       | It couldn't write a simple rename function for me yesterday,
       | still buggy after seven attempts.
       | 
       | I'm more and more convinced that they dumb down the core product
       | when they plan to release a new version to make the difference
       | seem bigger.
        
         | wilg wrote:
         | 99% chance that's confirmation bias
        
           | nomel wrote:
           | Sam tweeted that they're running out of computer. I think
           | it's reasonable to think they may serve somewhat quantized
           | models when out of capacity. It would be a rational business
           | decision that would minimally disrupt lower tier ChatGPT
           | users.
           | 
           | Anecdotally, I've noticed what appears to be drops in
           | quality, some days. When the quality drops, it responds in
           | odd ways when asked what model it is.
        
         | logicallee wrote:
         | >It couldn't write a simple rename function for me yesterday,
         | still buggy after seven attempts.
         | 
         | I'm surprised and a bit nervous about that. We intend to
         | bootstrap a large project with it!!
         | 
         | Both ChatGPT 4o (fast) and ChatGPT o1 (a bit slower, deeper
         | thinking) should easily be able to do this without fail.
         | 
         | Where did it go wrong? Could you please link to your chat?
         | 
         | About my project: I run the sovereign State of Utopia (will be
         | at stateofutopia.com and stofut.com for short) which is a
         | country based on the idea of state-owned, autonomous AI's that
         | do all the work and give out free money, goods, and services to
         | all citizens/beneficiaries. We've built a chess app (i.e. a
         | free source of entertainment) as a proof of concept though the
         | founder had to be in the loop to fix some bugs:
         | 
         | https://taonexus.com/chess.html
         | 
         | and a version that shows obvious blunders, by showing which
         | squares are under attack:
         | 
         | https://taonexus.com/blunderfreechess.html
         | 
         | One of the largest and most complicated applications anyone can
         | run is a web browser. We don't have a web browser built, but we
         | do have a buggy minimal version of it that can load and
         | minimally display some web pages, and post successsfully:
         | 
         | https://taonexus.com/publicfiles/feb2025/84toy-toy-browser-w...
         | 
         | It's about 1700 lines of code and at this point runs into the
         | limitations of all the major engines. But it does run, can load
         | some web pages and can post successfully.
         | 
         | I'm shocked and surprised ChatGPT failed to get a rename
         | function to work, in 7 attempts.
        
       | skepticATX wrote:
       | How many still believe that scaling up base models will lead to
       | AGI?
        
       | Mizza wrote:
       | Sounds like it's a distill of O1? After R1, I don't care that
       | much about non-reasoning models anymore. They don't even seem
       | excited about it on the livestream.
       | 
       | I want tiny, fast and cheap non-reasoning models I can use in
       | APIs and I want ultra smart reasoning models that I can query a
       | few times a day as an end user (I don't mind if it takes a few
       | minutes while I refill a coffee).
       | 
       | Oh, and I want that advanced voice mode that's good enough at
       | transcription to serve as a babelfish!
       | 
       | After that, I guess it's pretty much all solved until the robots
       | start appearing in public.
        
         | sebastiennight wrote:
         | Probably not a distill of o1, since o1 is a reasoning model and
         | GPT4.5 is not. Also, OpenAI has been claiming that this is a
         | very large model (and it's 2.5x more expensive than even OG
         | GPT-4) so we can assume it's the biggest model they've trained
         | so far.
         | 
         | They'll probably distill this one into GPT-4.5-mini or such,
         | and have something faster and cheaper available soon.
        
           | Mizza wrote:
           | There are plenty of distills of reasoning models now, and
           | they said in they livestream they used training data from
           | "smaller models" - which is probably every model ever
           | considering how expensive this one is.
        
             | sebastiennight wrote:
             | Knowledge distillation is literally _by definition_
             | teaching a smaller model from a big one, not the opposite.
             | 
             | Generating outputs from existing (therefore smaller) models
             | to train the largest model of all time would simply be
             | called "using synthetic data". These are not the same thing
             | at all.
             | 
             | Also, if you were to distill a reasoning model, the goal
             | would be to get a (smaller) reasoning model because you're
             | teaching your new model to mimic outputs that show a
             | reasoning/thinking trace. E.G. that's what all of those
             | "local" Deepseek models are: small LLama models distilled
             | from the big R1 ; a process which "taught" Llama-7B to show
             | reasoning steps before coming up with a final answer.
        
         | eightysixfour wrote:
         | It isn't even vaguely a distill of o1. The reasoning models
         | are, from what we can tell, relatively small. This model is
         | _massive_ and they probably scaled the parameter count to
         | improve factual knowledge retention.
         | 
         | They also mentioned developing some new techniques for training
         | small models and then incorporating those into the larger model
         | (probably to help scale across datacenters), so I wonder if
         | they are doing a bit of what people _think_ MoE is, but isn 't.
         | Pre-train a smaller model, focus it on specific domains, then
         | use that to provide synthetic data for training the larger
         | model on that domain.
        
           | Mizza wrote:
           | You can 'distill' with data from a smaller, better model into
           | a larger, shittier one. It doesn't matter. This is what they
           | said they did on the livestream.
        
             | eightysixfour wrote:
             | I have distilled models before, I know how it works. They
             | may have used o1 or o3 to create some of the synthetic data
             | for this one, but they clearly did not try and create any
             | self-reflective reasoning in this model whatsoever.
        
         | valine wrote:
         | My impression is that it's a massive increase in the parameter
         | count. This is likely the spiritual successor to GPT4 and would
         | have been called GPT5 if not for the lackluster performance.
         | The speculation is that there simply isn't enough data on the
         | internet to support yet another 10x jump in parameters.
         | 
         | O1-mini is a distill of O1. This definitely isn't the same
         | thing.
        
       | rvz wrote:
       | You should really be paying attention to what DeepSeek AI open
       | sources next.
       | 
       | This announcement by OpenAI was already expected: [0]
       | 
       | [0] https://x.com/sama/status/1889755723078443244
        
       | brokensegue wrote:
       | API price is crazy high. This model must be huge. Not sure this
       | is practical
        
         | jdprgm wrote:
         | Wow you aren't kidding, 30x input price and 15x output price vs
         | 4o is insane. The pricing on all AI API stuff changes so
         | rapidly and is often so extreme between models it is all hard
         | to keep track of and try to make value decisions. I would
         | consider a 2x or 3x price increase quite significant, 30x is
         | wild. I wonder how that even translates... there is no way the
         | model size is 30 times larger right?
        
       | jdprgm wrote:
       | Per Altman on X: "we will add tens of thousands of GPUs next week
       | and roll it out to the plus tier then". Meanwhile a month after
       | launch rtx 5000 series is completely unavailable and hardly any
       | restocks and the "launch" consisted of microcenters getting
       | literally tens of cards. Nvidia really has basically abandoned
       | consumers.
        
         | apsec112 wrote:
         | AI GPUs are bottlenecked mostly by high-bandwidth memory (HBM)
         | chips and CoWoS (packaging tech used to integrate HBM with the
         | GPU die), which are in short supply and aren't found in
         | consumer cards at all
        
           | jiggawatts wrote:
           | You would think that by now they would have done something to
           | ramp production capacity...
        
         | bhouston wrote:
         | Altman's claim and NVIDIA's consumer launch supply problems may
         | be related - OpenAI may be eating up the GPU supply...
        
           | _zoltan_ wrote:
           | OpenAI is not purchasing consumer 5090s... :)
        
             | BizarroLand wrote:
             | No, but the supply constraints are part of what is driving
             | the insane prices. Every chip they use for consumer grade
             | instead of commercial grade is a potential loss of
             | potential income.
        
             | bangaladore wrote:
             | Although you are correct, Nvidia is limited on total
             | output. They can't produce 50XXs fast enough, and it's
             | naive to think that isn't at least partially due to the
             | wild amount of AI GPUs they are producing.
        
       | sunaookami wrote:
       | This seems very rushed because of DeepSeek's R1 and Anthropic's
       | Claude 3.7 Sonnet. Pretty underwhelming, they didn't even show
       | programming? In the livestream, they struggled to come up with
       | reasons why I should prefer GPT-4.5 over GPT-4o or o1.
        
         | apsec112 wrote:
         | At least according to WSJ, they had planned to release it
         | earlier but struggled to get the model quality up, especially
         | relative to cost
        
         | bhouston wrote:
         | they do have coding benchmarks, I summarized them here:
         | https://news.ycombinator.com/item?id=43197955
        
         | bitshiftfaced wrote:
         | This strikes me as the opposite of rushed. I get the impression
         | that they've been sitting on this for a while and couldn't make
         | it look as good as previous improvements. At some point they
         | had to say, "welp here it is, now we can check that box and
         | move on."
        
       | Topfi wrote:
       | Considering both this blog post and the livestream demos, I am
       | underwhelmed. Having just finished the stream, I had a real "was
       | that all" moment, which on one hand shows how spoiled I've gotten
       | by new models impressing me, but on another feels like OpenAI
       | really struggles to stay ahead of their competitors.
       | 
       | What has been shown feels like it could be achieved using a
       | custom system prompt on older versions of OpenAIs models, and I
       | struggle to see anything here that truly required ground-up
       | training on such a massive scale. Hearing that they were forced
       | to spread their training across multiple data centers
       | simultaneously, coupled with their recent release of SWE-Lancer
       | [0] which showed Anthropic (Claude 3.5 Sonnet (new) to be exact)
       | handily beating them, I was really expecting something more than
       | "slightly more casual/shorter output", which again, I fail to see
       | how that wasn't possible by prompting GPT-4o.
       | 
       | Looking at pricing [1], I am frankly astonished.
       | 
       | > Input: $75.00 / 1M tokens > Cached input: $37.50 / 1M tokens >
       | Output: $150.00 / 1M tokens
       | 
       | How could they justify that asking price? And, if they have some
       | amazing capabilities that make a 30-fold pricing increase
       | justifiable, why not show it? Like, OpenAI are many things, but I
       | always felt they understood price vs performance incredibly well,
       | from the start with gpt-3.5-turbo up to now with o3-mini, so this
       | really baffles me. If GPT-4.5 can justify such immense cost in
       | certain tasks, why hide that and if not, why release this at all?
       | 
       | [0] https://github.com/openai/SWELancer-Benchmark
       | 
       | [1] https://openai.com/api/pricing/
        
         | Bjorkbat wrote:
         | My first thought seeing this and looking at benchmarks was that
         | if it wasn't for reasoning, then either pundits would be saying
         | we've hit a plateau, or at the very least OpenAI is clearly in
         | 2nd place to Anthropic in model performance.
         | 
         | Of course we don't live in such a world, but I thought of this
         | nonetheless because for all the connotations that come with a
         | 4.5 moniker this is kind of underwhelming.
        
           | uh_uh wrote:
           | Pundits were saying that deep learning has hit a plateau even
           | before the LLM boom.
        
         | mvdtnz wrote:
         | > How could they justify that asking price?
         | 
         | They're still selling $1 for <$1. Like personal food delivery
         | before it, consumers will eventually need to wake up to this
         | fact - these things will get expensive, fast.
        
           | spiderfarmer wrote:
           | Let a thousand providers bloom.
        
           | Ekaros wrote:
           | I generally question how wide spread willingness to pay for
           | the most expensive product is. And will most users of those
           | who actually want AI go with ad ridden lesser models...
        
             | vel0city wrote:
             | I can just imagine Kraft having a subsidized AI model for
             | recipe suggestions that adds Velveeta to everything.
        
         | tmaly wrote:
         | rethinking your comment "was that all" I am listening to the
         | stream now and had a thought. Most of the new models that have
         | come out in the past few weeks have been great at coding and
         | logical reasoning. But 4o has been better at creative writing.
         | I am wondering if 4.5 is going to be even better at creative
         | writing than 4o.
        
         | lasermike026 wrote:
         | I would rather pay for 4.5 by the query.
        
       | zaptrem wrote:
       | GPT 4.5 pricing is insane: Price Input: $75.00 / 1M tokens Cached
       | input: $37.50 / 1M tokens Output: $150.00 / 1M tokens
       | 
       | GPT 4o pricing for comparison: Price Input: $2.50 / 1M tokens
       | Cached input: $1.25 / 1M tokens Output: $10.00 / 1M tokens
       | 
       | It sounds like it's so expensive and the difference in usefulness
       | is so lacking(?) they're not even gonna keep serving it in the
       | API for long:
       | 
       | > GPT-4.5 is a very large and compute-intensive model, making it
       | more expensive than and not a replacement for GPT-4o. Because of
       | this, we're evaluating whether to continue serving it in the API
       | long-term as we balance supporting current capabilities with
       | building future models. We look forward to learning more about
       | its strengths, capabilities, and potential applications in real-
       | world settings. If GPT-4.5 delivers unique value for your use
       | case, your feedback (opens in a new window) will play an
       | important role in guiding our decision.
       | 
       | I'm still gonna give it a go, though.
        
         | MattSayar wrote:
         | Input price difference: 4.5 is 30x more
         | 
         | Output price difference:4.5 is 15x more
         | 
         | In their model evaluation scores in the appendix, 4.5 is, on
         | average, 26% better. I don't understand the value here.
        
           | alwa wrote:
           | If you ran the same query set 30x or 15x on the cheaper model
           | (and compensated for all the extra tokens the reasoning model
           | uses), would you be able to realize the same 26% quality gain
           | in a machine-adjudicatible kind of way?
        
             | j_maffe wrote:
             | with a reasoning model you'd get better than both.
        
           | mirekrusin wrote:
           | Einstein's IQ = 3.5x chimpanzees IQs, right?
        
             | redox99 wrote:
             | 3.5x on a normal distribution with mean 100 and SD 15 is
             | pretty insane. But I agree with your point, being 26%
             | better at a certain benchmark could be a tiny difference,
             | or an incredible improvement (imagine the hardest questions
             | being Riemann hypothesis, P != NP, etc).
        
         | minimaxir wrote:
         | Sam Altman's explanation for the restriction is a bit fluffier:
         | https://x.com/sama/status/1895203654103351462
         | 
         | > bad news: it is a giant, expensive model. we really wanted to
         | launch it to plus and pro at the same time, but we've been
         | growing a lot and are out of GPUs. we will add tens of
         | thousands of GPUs next week and roll it out to the plus tier
         | then. (hundreds of thousands coming soon, and i'm pretty sure
         | y'all will use every one we can rack up.)
        
           | g-mork wrote:
           | release blog post author: this is definitely a research
           | preview
           | 
           | ceo: it's ready
           | 
           | the pricing is probably a mixture of dealing with GPU
           | scarcity and intentionally discouraging actual users. I can't
           | imagine the pressure they must be under to show they are
           | releasing and staying ahead, but Altman's tweet makes it
           | clear they aren't really ready to sell this to the general
           | public yet.
        
             | pk-protect-ai wrote:
             | Yeap, that the thing, they are not ahead anymore. Not since
             | last summer at least. Yes they have probably largest
             | customer base, but their models are not the best for a
             | while already.
        
               | danenania wrote:
               | Eh, I think o1-pro is by far the most capable model
               | available right now in terms of pure problem solving.
        
               | rvnx wrote:
               | You can try Claude 3.7-Thinking and Grok 3 Think. 10
               | times cheaper, as good, or very similar to o1-pro.
        
               | danenania wrote:
               | I haven't tried Grok yet so can't speak to that, but I
               | find o1-pro is much stronger than 3.7-thinking for e.g.
               | distributed systems and concurrency problems.
        
           | chefandy wrote:
           | I'm not an expert or anything, but from my vantage point,
           | each passing release makes Altman's confidence look more
           | aspirational than visionary, which is a really bad place to
           | be with that kind of money tied up. My financial manager is
           | pretty bullish on tech so I hope he is paying close attention
           | to the way this market space is evolving. He's good at his
           | job, a nice guy, and surely wears much more expensive
           | underwear than I do-- I'd hate to see him lose a pair
           | powering on his Bloomberg terminal in the morning one of
           | these days.
        
             | igor47 wrote:
             | You're the one buying him the underwear. Don't index funds
             | outperform managed investing? I think especially after
             | accounting for fees, but possibly even after accounting
             | that 50% of money managers are below average.
        
               | chefandy wrote:
               | He earns his undies. My returns are almost always
               | modestly above index fund returns after his fees, though
               | like last quarter, he's very upfront when they're not. He
               | has good advice for pulling back when things are
               | uncertain. I'm happy to delegate that to him.
        
               | marcus0x62 wrote:
               | A friend got taken in by a Ponzi scheme operator several
               | years ago. The guy running it was known for taking his
               | clients out to lavish dinners and events all the time.[0]
               | 
               | After the scam came to light my friend said "if I knew I
               | was paying for those dinners, I would have been fine with
               | Denny's[1]"
               | 
               | I wanted to tell him "you would have been paying for
               | those dinners even if he wasn't outright stealing your
               | money," but that seemed insensitive so I kept my mouth
               | shut.
               | 
               | 0 - a local steakhouse had a portrait of this guy drawn
               | on the wall
               | 
               | 1 - for any non-Americans, Denny's is a low cost diner-
               | style restaurant.
        
               | ProfessorLayton wrote:
               | Not all investing is throwing cash at an index, though.
               | There's other types of investing like direct indexing (to
               | harvest losses), muni bonds, etc.
               | 
               | Paying someone to match your risk profile and financial
               | goals may be worth the fee, which as you pointed out is
               | very measurable. YMMV though.
        
               | fragmede wrote:
               | Depends who's pitch deck you're reading. Warren Buffett
               | didn't get rich waiting on index funds.
        
             | Terr_ wrote:
             | > each passing release makes Altman's confidence look more
             | aspirational than visionary
             | 
             | As an LLM cynic, I feel that point passed _long_ go,
             | perhaps even before Altman claimed countries would start
             | wars to conquer territory for its datacenters, or promoting
             | the dream of a 7 T-for-trillion dollar investment.
             | 
             | Alas, the market can remain irrational longer than I can
             | remain solvent.
        
               | chefandy wrote:
               | That $7 trillion dollar ask pushed me from skeptical to
               | full-on eye-roll emoji land-- the dude is clearly a
               | narcissist with delusions of grandeur-- but it's getting
               | _worse._ Considering the $200 pro subscription was
               | significantly unprofitable before this model came out,
               | imagine how _astonishingly expensive_ this model must be
               | to run at many times that price.
        
           | rebolek wrote:
           | Bad news: Sam Altman runs the show.
        
             | rvnx wrote:
             | He is like the Elon Musk of OpenAI.
        
         | sebastiennight wrote:
         | I think it's fairer to compare it to the original GPT-4 which
         | might the equivalent in term of "size" (though we don't have
         | actual numbers for either).
         | 
         | GPT-4: Input $30.00 / 1M tokens ; Output $60.00 / 1M tokens
         | 
         | So 4.5 is 2.5x more expensive.
         | 
         | I think they announced this as their last non-reasoning model,
         | so it was maybe with the goal of stretching pre-training as far
         | as they could, just to see what new capabilities would show up.
         | We'll find out as the community gives it a whirl.
         | 
         | I'm a Tier 5 org and I have it available already in the API.
        
           | minimaxir wrote:
           | The marginal costs for running a GPT-4-class LLM are much
           | lower nowadays due to significant software and hardware
           | innovations since then, so costs/pricing are harder to
           | compare.
        
             | sebastiennight wrote:
             | Agreed, however it might make sense that a much-larger-
             | than-GPT-4 LLM would also, at launch, be more expensive to
             | run than the OG GPT-4 was at launch.
             | 
             | (And I think this is probably also scarecrow pricing to
             | discourage casual users from clogging the API since they
             | seem to be too compute-constrained to deliver this at
             | scale)
        
           | jstummbillig wrote:
           | Why would that be fairer? We can assume they did incorporate
           | all learnings and optimizations they made post gpt-4 launch,
           | no?
        
           | spoaceman7777 wrote:
           | There are some numbers on one of their Blackwell or Hopper
           | info pages that notes the ability of their hardware in
           | hosting an unnamed GPT model that is 1.8T params. My
           | assumption was that it referred to GPT-4
           | 
           | Sounds to me like GPT 4.5 likely requires a full Blackwell
           | HGX cabinet or something, thus OpenAI's reference to needing
           | to scale out their compute more (Supermicro only opened up
           | their Blackwell racks for General Availability last month,
           | and they're the prime vendor for water-cooled Blackwell
           | cabinets right now, and have the ability to throw up a GPU
           | mega-cluster in a few weeks, like they did for xAI/Grok)
        
         | harlanlewis wrote:
         | The price really is eye watering. At a glance, my first
         | impression is this is something like Llama 3.1 405B, where the
         | primary value may be realized in generating high quality
         | synthetic data for training rather than direct use.
         | 
         | I keep a little google spreadsheet with some charts to help
         | visualize the landscape at a glance in terms of
         | capability/price/throughput, bringing in the various index
         | scores as they become available. Hope folks find it useful,
         | feel free to copy and claim as your own.
         | 
         | https://docs.google.com/spreadsheets/d/1foc98Jtbi0-GUsNySddv...
        
           | bennyg wrote:
           | This is an amazing spreadsheet - thank you for sharing!
        
           | isoprophlex wrote:
           | Thats... incredibly thorough. Wow. Thanks for sharing this.
        
           | Philpax wrote:
           | Holy shit, that's incredible. You should publicise this more!
           | That's a fantastic resource.
        
             | beklein wrote:
             | They tried a while ago:
             | https://news.ycombinator.com/item?id=40373284
             | 
             | Sadly little people noticed...
        
               | throwup238 wrote:
               | Sadly _few_ people noticed.
               | 
               | I don't normally cosplay as a grammar Nazi but in this
               | case I feel like someone should stand up for the little
               | people :)
        
               | rebolek wrote:
               | So you think that little people didn't notice? ;)
        
               | dumpsterdiver wrote:
               | A comma in the original comment would have made it pop
               | even more:
               | 
               | "Sadly, little people noticed."
               | 
               | (queue a group of little people holding pitch forks
               | (normal forks upon closer inspection))
        
           | adinb wrote:
           | I cannot overstate how good your shared spreadsheet is.
           | Thanks again!
        
           | jnd0 wrote:
           | Thank you so much for sharing this!
        
           | bglusman wrote:
           | very impressive... also interested in your trip planner, it
           | looks like invite only at the moment, but... would it be rude
           | to ask for an invite?
        
           | gwyllimj wrote:
           | That is an amazing resource. Thanks for sharing!
        
           | dumpsterdiver wrote:
           | Nice, thank you for that (upvoted in appreciation). Regarding
           | the absence of o1-Pro from the analysis, is that just because
           | there isn't enough public information available?
        
           | krwiseman wrote:
           | Awesome spreadsheet. Would a 3D graph of fast, cheap & smart
           | be possible?
        
           | rendist wrote:
           | Amazing, thank you so much for sharing this.
        
           | kridsdale3 wrote:
           | Hope you don't mind, I dumped your sheet in to GPT 4.5 and
           | asked for it's take:
           | 
           | 1. Capability vs. Throughput Correlation
           | Insight: High composite-capability models (e.g., Claude 3.7
           | Sonnet Thinking, GPT-4o) typically show moderate or lower
           | throughput (<70 tokens/sec), while mid-tier models (DeepSeek,
           | Mistral variants, Nova series) often deliver much higher
           | throughput (>100 tokens/sec).             Interpretation:
           | Frontier models trade off throughput for sophisticated
           | reasoning, likely due to heavier computational demands.
           | 
           | 2. Cost Efficiency vs. Capability Plateau
           | Insight: Models at the top capability (90-100) see a steep
           | rise in cost per capability point, with diminishing cost-
           | effectiveness above roughly 75 capability points.
           | Example: Compare GPT-4o (100 capability, ~$0.085 per
           | capability point) with DeepSeek V3 (74 capability, ~$0.002
           | per capability point).             Interpretation: Mid-
           | capability models appear optimal for cost-sensitive
           | applications.
           | 
           | 3. Capability Consistency vs. Composite Capability
           | Insight: Higher consistency scores (>85%) are generally
           | linked to mid-to-high composite capability (60-80 range), yet
           | the very top models (85+) can have moderate or lower
           | consistency (e.g., GPT-4o at 81% vs. Claude 3.7 Sonnet at
           | 57%).             Interpretation: Even cutting-edge models
           | may show variability across benchmarks due to task-specific
           | optimizations.
           | 
           | 4. Latency and Throughput Relationship
           | Insight: Initial latency (first chunk response) has only a
           | weak correlation with overall throughput. Some models with
           | high sustained throughput exhibit high initial latency, while
           | others with lower throughput can respond quickly.
           | Interpretation: Optimizing for a rapid first response is a
           | distinct engineering challenge from maximizing sustained
           | throughput.
           | 
           | 5. Open vs. Proprietary Models: Capability and Cost
           | Efficiency                  Insight: Open models (e.g.,
           | DeepSeek, Mistral, Llama variants) tend to cluster in a
           | region of higher cost efficiency relative to capability
           | compared to proprietary models (e.g., those from Anthropic,
           | OpenAI, Google).             Interpretation: Open-source
           | offerings are closing the capability gap and forcing a
           | rebalancing of cost structures.
           | 
           | 6. Benchmark Correlations (Math vs. General Reasoning)
           | Insight: Strong performance on general reasoning benchmarks
           | (like BIG-BENCH-HARD, ARC Challenge) often aligns with high
           | scores in mathematical reasoning (e.g., GSM8K, MATH).
           | Interpretation: Math proficiency appears to be a good
           | predictor of overall reasoning strength.
           | 
           | 7. Throughput vs. Year of Release                  Insight:
           | There is a noticeable trend where more recent models (2025)
           | tend to offer higher throughput (>100 tokens/sec), even among
           | mid-range capabilities.             Interpretation: The
           | market seems to be shifting toward models that emphasize
           | operational efficiency alongside capability.
           | 
           | 8. Long-context Performance vs. General Capability
           | Insight: Models designed for long-context handling (as
           | measured by benchmarks like ZeroSCROLLS, InfiniteBench)
           | generally also exhibit high overall capability.
           | Interpretation: The ability to process extended contexts
           | remains a premium feature closely tied to overall model
           | sophistication.
        
         | swatcoder wrote:
         | > We look forward to learning more about its strengths,
         | capabilities, and potential applications in real-world
         | settings. If GPT-4.5 delivers unique value for your use case,
         | your feedback (opens in a new window) will play an important
         | role in guiding our decision.
         | 
         | "We don't really know what this is good for, but spent a lot of
         | money and time making it and are under intense pressure to
         | announce new things right now. If you can figure something out,
         | we need you to help us."
         | 
         | Not a confident place for an org trying to sustain a $XXXB
         | valuation.
        
           | tempaccount420 wrote:
           | > "We don't really know what this is good for, but spent a
           | lot of money and time making it and are under intense
           | pressure to announce new things right now. If you can figure
           | something out, we need you to help us."
           | 
           | Where is this quote from?
        
             | hotpocket777 wrote:
             | It's not a quote. It is an interpretation or reading of a
             | quote.
        
               | cogman10 wrote:
               | Perhaps even fed through an LLM ;)
        
             | scythe wrote:
             | I believe it's a "translation" in the sense of
             | Wittgenstein's goal of philosophy:
             | 
             | >My aim is: to teach you to pass from a piece of disguised
             | nonsense to something that is patent nonsense.
        
               | Nition wrote:
               | Another great example on Hacker News is this old
               | translation of Google's "Amazing Bet":
               | https://news.ycombinator.com/item?id=12793033
        
             | dd3boh wrote:
             | I think it's supposed to be a translation of what OpenAI's
             | quote means in real world terms.
        
             | thih9 wrote:
             | The quotation marks in the grandparent comment are scare
             | (sneer) quotes and not actual quotation.
             | 
             | https://en.m.wikipedia.org/wiki/Scare_quotes
             | 
             | > Whether quotation marks are considered scare quotes
             | depends on context because scare quotes are not visually
             | different from actual quotations.
        
           | riwsky wrote:
           | Said the quiet part out loud! Or as we say these days,
           | "transparently exposed the chain of thought tokens".
        
             | porridgeraisin wrote:
             | Lol, nice one
        
             | Terr_ wrote:
             | "I knew the dame was trouble the moment she walked into my
             | office."
             | 
             | "Uh... excuse me, Detective Nick Danger? I'd like to retain
             | your services."
             | 
             | "I waited for her to get the the point."
             | 
             | "Detective, who are you talking to?"
             | 
             | "I didn't want to deal with a client that was hearing
             | voices, but money was tight and the rent was due. I
             | pondered my next move."
             | 
             | "Mr. Danger, are you... narrating out loud?"
             | 
             | "Damn! My internal chain of thought, the key to my success
             | --or at least, past successes--was leaking again. I
             | rummaged for the familiar bottle of scotch in the drawer,
             | kept for just such an occasion."
             | 
             | ---
             | 
             | But seriously: These "AI" products basically run on movie-
             | scripts already, where the LLM is used to append more
             | "fitting" content, and glue-code is periodically performing
             | any lines or actions that arise in connection to the
             | Helpful Bot character. Real humans are tricked into
             | thinking the finger-puppet is a discrete entity.
             | 
             | These new "reasoning" models are just switching the style
             | of the movie script to _film noir_ , where the Helpful Bot
             | character is making a layer of unvoiced commentary. While
             | it may make the story more cohesive, it isn't a qualitative
             | change in the kind of illusory "thinking" going on.
        
               | kridsdale3 wrote:
               | I don't know if it was you or someone else who made
               | pretty much the same point a few days ago. But I still
               | like it. It makes the whole thing a lot more fun.
        
               | Terr_ wrote:
               | https://news.ycombinator.com/context?id=43118925
               | 
               | I've been banging that particular drum for a while on HN,
               | and the mental-model still feels so intuitively strong to
               | me that I'm starting to have doubts: "It's _too_ good, I
               | must be wrong in some subtle and devastating way. "
        
           | EA-3167 wrote:
           | Maybe if they build a few more data centers, they'll be able
           | to construct their machine god. Just a few more dedicated
           | power plants, a lake or two, a few hundred billion more and
           | they'll crack this thing wide open.
           | 
           | And maybe Tesla is going to deliver truly full self driving
           | tech any day now.
           | 
           | And Star Citizen will prove to have been worth it along
           | along, and Bitcoin will rain from the heavens.
           | 
           | It's very difficult to remain charitable when people seem to
           | always be chasing the new iteration of the same old thing,
           | and we're expected to come along for the ride.
        
             | alyandon wrote:
             | And Star Citizen will prove to have been worth it along
             | along
             | 
             | Sounds like someone isn't happy with the 4.0 eternally
             | incrementing "alpha" version release. :-D
             | 
             | I keep checking in on SC every 6 months or so and still see
             | the same old bugs. What a waste of potential. Fortunately,
             | Elite Dangerous is enough of a space game to scratch my
             | space game itch.
        
               | 0x457 wrote:
               | To be fAir, SC is trying to do things that no one else
               | done in a context of a single game. I applaud their
               | dedication, but I won't be buying JPGs of a ship for 2k.
        
               | alyandon wrote:
               | Yeah, they never should have expected to take an FPS game
               | engine like CryEngine and expected to be able to modify
               | it to work as the basis for a large scale space MMO game.
               | 
               | Their backend is probably an async nightmare of
               | replicated state that gets corrupted over time. Would
               | explain why a lot of things seem to work more or less bug
               | free after an update and then things fall to pieces and
               | the same old bugs start showing up after a few weeks.
               | 
               | And to be clear, I've spent money on SC and I've played
               | enough hours goofing off with friends to have got my
               | money's worth out of it. I'm just really bummed out about
               | the whole thing.
        
               | 0x457 wrote:
               | Gonna go meta here for a bit, but I believe we going to
               | get a fully working stable SC before we get fusion. "we"
               | as in humanity, you and I might not be around when it's
               | finally done.
        
             | JohnMakin wrote:
             | leave star citizen out of this :)
        
             | philistine wrote:
             | > And Star Citizen will prove to have been worth it along
             | along
             | 
             | Once they've implemented saccades in the eyeballs of the
             | characters wearing helmets in spaceship millions of
             | kilometres apart, then it will all have been worth it.
        
             | sho_hn wrote:
             | You have it all wrong. The end game is a scalable, reliable
             | AI work force capable of finishing Star Citizen.
             | 
             | At least this is the benchmark for super-human general
             | intelligence that I propose.
        
             | bodegajed wrote:
             | Could this path lead to solving world hunger too? :)
        
             | mattgreenrocks wrote:
             | It's an honor to be dragged along so many ubermensch's
             | Incredible Journeys.
        
             | bloomingkales wrote:
             | Star Citizen is a working model of how to do UBI. That
             | entire staff of a thousand people is the test case.
        
           | jodrellblank wrote:
           | > "Early testing shows that interacting with GPT-4.5 feels
           | more natural. Its broader knowledge base, improved ability to
           | follow user intent, and greater "EQ" make it useful for tasks
           | like improving writing, programming, and solving practical
           | problems. We also expect it to hallucinate less."
           | 
           | "Early testing doesn't show that it hallucinates less, but we
           | expect that putting that sentence nearby will lead you to
           | draw a connection there yourself".
        
             | istjohn wrote:
             | According to a graph they provide, it does hallucinate
             | significantly less on at least one benchmark.
        
               | jug wrote:
               | It hallucinates at 37% on SimpleQA yeah, which is a set
               | of very difficult questions inviting hallucinations.
               | Claude 3.5 Sonnet (the June 2024 editiom, before October
               | update and before 3.7) hallucinated at 35%. I think this
               | is more of an indication of how behind OpenAI has been in
               | this area.
        
               | tmpz22 wrote:
               | Are the benchmarks known ahead of time? Could the answer
               | to the benchmarks be in the training data?
        
               | llm_trw wrote:
               | In general yes, bench mark pollution is a big problem and
               | why only dynamic benchmarks matter.
        
             | LeifCarrotson wrote:
             | That's some top-tier sales work right there.
             | 
             | I suck at and hate writing the mildly deceptive corporate
             | puffery that seems to be in vogue. I wonder if GPT-4.5 can
             | write that for me or if it's still not as good at it as the
             | expert they paid to put that little gem together.
        
             | zaptrem wrote:
             | This seems like it should be attributed to better post
             | training, not a bigger model.
        
             | justspacethings wrote:
             | The usage of "greater" is also interesting. It's like they
             | are trying to say better, but greater is a geographic term
             | and doesn't mean "better" instead it's closer to "wider" or
             | "covers more area."
        
               | lechatonnoir wrote:
               | I'm all for skepticism of capabilities and cynicism about
               | corporate messaging, but I really don't think there's an
               | interpretation of the word "greater" in this context"
               | that doesn't mean "higher" and "better".
        
               | skissane wrote:
               | > but greater is a geographic term and doesn't mean
               | "better" instead it's closer to "wider" or "covers more
               | area."
               | 
               | You are confusing a specific geographical sense of
               | "greater" (e.g. "greater New York") with the generic
               | sense of "greater" which just means "more great". In "7
               | is greater than 6", "greater" isn't geographic
               | 
               | The difference between "greater" and "better", is
               | "greater" just means "more than", without implying any
               | value judgement-"better" implies the "more than" is a
               | good thing: "The Holocaust had a greater death toll than
               | the Armenian genocide" is an obvious fact, but only a
               | horrendously evil person would use "better" in that
               | sentence (excluding of course someone who accidentally
               | misspoke, or a non-native speaker mixing up words)
        
             | esafak wrote:
             | GPT-4.5 may be an awesome model, some say!
        
           | crazygringo wrote:
           | > _We don 't really know what this is good for_
           | 
           | Oh come on. Think how long of a gap there was between the
           | first microcomputer and VisiCalc. Or between the start of the
           | internet and social networking.
           | 
           | First of all, it's going to take us 10 years to figure out
           | how to use LLM's to their full productive potential.
           | 
           | And second of all, it's going to take us collectively a long
           | time to also figure out how much accuracy is necessary to pay
           | for in which different applications. Putting out a higher-
           | accuracy, higher-cost model for the market to try is an
           | important part of figuring that out.
           | 
           | With new disruptive technologies, companies aren't supposed
           | to be able to look into a crystal ball and see the future.
           | They're _supposed_ to try new things and see what the market
           | finds useful.
        
             | nyc_data_geek1 wrote:
             | The Internet had plenty of very productive use cases before
             | social networking, even from its most nascent origins.
             | Spending billions building something on the assumption that
             | someone else will figure out what it's good for, is not
             | good business.
        
               | crazygringo wrote:
               | And LLM's already have tons of productive uses. The
               | biggest ones are probably still waiting, though.
               | 
               | But this is about one particular price/performance ratio.
               | 
               | You need to build things before you can see how the
               | market responds. You say it's "not good business" but
               | that's entirely wrong. It's excellent business. It's the
               | only way to go about it, in fact.
               | 
               | Finding product-market fit is a process. Companies aren't
               | omniscient.
        
               | bigstrat2003 wrote:
               | > And LLM's already have tons of productive uses.
               | 
               | I disagree strongly with that. Right now they are fun
               | toys to play with, but not useful tools, because they are
               | not reliable. If and when that gets fixed, maybe they
               | will have productive uses. But for right now, not so
               | much.
        
               | nyc_data_geek1 wrote:
               | You go into this process with a perspective, you do not
               | build a solution and then start looking for a problem.
        
             | rsynnott wrote:
             | Arguably social networking is older than the internet
             | proper; USENET predates TCP/IP (though not ARPANet).
        
             | nyrikki wrote:
             | The TRS-80, Apple ][, and PET all came out in 1977,
             | VisiCalc was released in 1979.
             | 
             | Usenet, Bitnet, IRC, BBSs all predated the commercial
             | internet, which are all forms of _Online_ social networks.
        
             | mandevil wrote:
             | ChatGPT had its initial public release November 30th, 2022.
             | That's 820 days to today. The Apple II was first sold June
             | 10, 1977, and Visicalc was first sold October 17, 1979,
             | which is 859 days. So we're right about the same distance
             | in time- the exact equal duration will be April 7th of this
             | year.
             | 
             | Going back to the very first commercially available
             | microcomputer, the Altair 8800, that's four years and nine
             | months to Visicalc release. This isn't a decade long
             | process of figuring things out, it actually tends to move
             | real fast.
        
             | aylmao wrote:
             | I generally agree with the idea of building things,
             | iterating, and experimenting before knowing their full
             | potential, but I do see why there's negative sentiment
             | around this:
             | 
             | 1. The first microcomputer predates VisiCalc, yes, but it
             | doesn't predate the realization of what it could be useful
             | for. The Micral was released in 1973. Douglas Engelbart
             | gave "The Mother of All Demos" in 1968 [2]. It included
             | things that wouldn't be commonplace for decades, like a
             | collaborative real-time editor or video-conferencing.
             | 
             | I wasn't yet born back then, but reading about the timeline
             | of things, it sounds like the industry had a much more
             | concrete and concise idea of what this technology would
             | bring to everyone.
             | 
             | "We look forward to learning more about its strengths,
             | capabilities, and potential applications in real-world
             | settings." doesn't inspire that sentiment for something
             | that's already being marketed as "the beginning of a new
             | era" and valued so exorbitantly.
             | 
             | 2. I think as AI becomes more generally available, and
             | "good enough" people (understandably) will be more
             | skeptical of closed-source improvements that stem from
             | spending big. Commoditizing AI is more clearly "useful", in
             | the same way commoditizing computing was more clearly
             | useful than just pushing numbers up.
             | 
             | Again, I wasn't yet born back then, but I can imagine the
             | announcement of Apple Macintosh with its 6MHz CPU and 128KB
             | RAM was more exciting and had a bigger impact than the
             | announcement of the Cray-2 with its 1.9GHz and +1GB memory.
             | 
             | [1] https://en.wikipedia.org/wiki/Micral
             | 
             | [2] https://en.wikipedia.org/wiki/The_Mother_of_All_Demos
        
           | fsndz wrote:
           | it's so over, pretraining is ngmi. maybe sam Altman was wrong
           | after all ? https://www.lycee.ai/blog/why-sam-altman-is-wrong
        
             | FpUser wrote:
             | >"I also agree with researchers like Yann LeCun or Francois
             | Chollet that deep learning doesn't allow models to
             | generalize properly to out-of-distribution data--and that
             | is precisely what we need to build artificial general
             | intelligence."
             | 
             | I think "generalize properly to out-of-distribution data"
             | is too weak of criteria for general intelligence (GI). GI
             | model should be able to get interested about some
             | particular area, research all the known facts, derive new
             | knowledge / create theories based upon said fact. If there
             | is not enough of those to be conclusive: propose and
             | conduct experiments and use the results to prove / disprove
             | / improve theories. And it should be doing this constantly
             | in real time on bazillion of "ideas". Basically model our
             | whole society. Fat chance of anything like this happening
             | in foreseeable future.
        
           | amarcheschi wrote:
           | I have a professor who founded a few companies, one of these
           | was funded by gates after he managed to spoke with him and
           | convinced him to give him money. This guy is goat, and he
           | always tells us that we need to find solutions to problems,
           | not to find problems to our solutions. It seems at openai
           | they didn't get the memo this time
        
         | serjester wrote:
         | I suppose this was their final hurrah after two failed attempts
         | at training GPT-5 with the traditional pre-training paradigm.
         | Just confirms reasoning models are the only way forward.
        
           | newfocogi wrote:
           | I think this is the correct take. There are other axes to
           | scale on AND I expect we'll see smaller and smaller models
           | approach this level of pre-trained performance. But I believe
           | massive pre-training gains have hit clearly diminished
           | returns (until I see evidence otherwise).
        
           | granzymes wrote:
           | > Compared to OpenAI o1 and OpenAI o3-mini, GPT-4.5 is a more
           | general-purpose, innately smarter model. We believe reasoning
           | will be a core capability of future models, and that the two
           | approaches to scaling--pre-training and reasoning--will
           | complement each other. As models like GPT-4.5 become smarter
           | and more knowledgeable through pre-training, they will serve
           | as an even stronger foundation for reasoning and tool-using
           | agents.
        
           | jstummbillig wrote:
           | What it confirms, I think, is, that we are going to need a
           | _lot_ more chips.
        
             | prisenco wrote:
             | Or, possibly, we're stuck waiting for another theoretical
             | breakthrough before real progress is made.
        
               | resource0x wrote:
               | breakthrough in biology
        
             | georgemcbay wrote:
             | Further confirmation, IMO, that the idea that any of this
             | leads to anything close to AGI is people getting high on
             | their own supply (in some cases literally).
             | 
             | LLMs are a great tool for what is effectively collected
             | knowledge search and summary (so long as you are willing to
             | accept that you have to verify all of the 'knowledge' they
             | spit back because they always have the ability to go off
             | the rails) but they have been hitting the limits on how
             | much better that can get without somehow introducing more
             | real knowledge for close to 2 years now and everything
             | since then is super incremental and IME mostly just
             | benchmark gains and hype as opposed to actually being
             | purely better.
             | 
             | I personally don't believe that more GPUs solves this,
             | like, at all. But its great for Nvidia's stock price.
        
             | DannyBee wrote:
             | Eh, no. More chips won't save this right now, or probably
             | in the near future (IE barring someone sitting on a
             | breakthrough right now).
             | 
             | It just means either
             | 
             | A. Lots and lots of hard work that get you a few percent at
             | a time, but add up to a _lot_ over time.
             | 
             | or
             | 
             | B. Completely different approaches that people actually
             | think about for a while rather than trying to incrementally
             | get something done in the next 1-2 months.
             | 
             | Most fields go through this stage. Sometimes more than once
             | as they mature and loop back around :)
             | 
             | Right now, AI seems bad at doing either - at least, from
             | the outside of most of these companies, and watching open
             | source/etc.
             | 
             | While lots of little improvements seem to be released in
             | lots of parts, it's rare to see anywhere that is collecting
             | and aggregating them en masse and putting them in practice.
             | It feels like for every 100 research papers, maybe 1 makes
             | it into something in a way that anyone ends up using it by
             | default.
             | 
             | This could be because they aren't really even a few percent
             | (which would be yet a different problem, and in some ways
             | worse), or it could be because nobody has cared to, or ...
             | 
             | I'm sure very large companies are doing a fairly reasonable
             | job on this, because they historically do, but everyone
             | else - even frameworks - it's still in the "here's a
             | million knobs and things that may or may not help".
             | 
             | It's like if compilers had no "O0/O1/O2/O3' at all and were
             | just like "here's 16,283 compiler passes - you can put them
             | in any order and amount you want". Thanks! I hate it!
             | 
             | It's worse even because it's like this at every layer of
             | the stack, whereas in this compiler example, this is just
             | one layer.
             | 
             | Additionally, everyone seems to rush half-baked things to
             | try to get the next incremental improvement released and
             | out the door because they think it will help them stay
             | "sticky" or whatever. History does not suggest this is a
             | good plan and even if it was a good plan in theory, it's
             | pretty hard to lock people in with what exists right now.
             | There isn't enough anyone cares about and rushing out half-
             | baked crap is not helping that. mindshare doesn't really
             | matter if no one cares about using _your_ product.
             | 
             | Does anyone using these things truly feel locked into
             | anyone's ecosystem at this point? Do they feel like they
             | will be soon?
             | 
             | I haven't met anyone who feels that way, even in corps
             | spending tons and tons of money with these providers.
             | 
             | The public companies - i can at least understand given the
             | fickleness of public markets. That was supposed to be one
             | of the serious benefit of staying private. So watching
             | private companies do the same thing - it's just sort of
             | mind-boggling.
             | 
             | Hopefully they'll grow up soon, or someone who takes their
             | time and does it right during one of the lulls will come
             | and eat all of their lunches.
        
           | usaar333 wrote:
           | For OpenAI perhaps? Sonnet 3.7 without extended thinking is
           | quite strong. Swe-bench scores tie o3
        
           | DebtDeflation wrote:
           | GPT 5 is likely just going to be a router model that decides
           | whether to send the prompt to 4o, 4o mini, 4.5, o3, or o3
           | mini.
        
         | techorange wrote:
         | I wonder how much money they're losing on it too even at those
         | prices.
        
         | nialv7 wrote:
         | Looks like more signal that the scaling "law" is indeed
         | faltering.
        
         | ur-whale wrote:
         | AI as it stands in 2025 is an amazing technology, but it is not
         | a product _at all_.
         | 
         | As a result, OpenAI simply does not have a business model, even
         | if they are trying to convince the world that they do.
         | 
         | My bet is that they're currently burning through other people's
         | capital at an amazing rate, but that they are light-years from
         | profitability
         | 
         | They are also being chased by fierce competition and OpenSource
         | which is very close behind. There simply is no moat.
         | 
         | It will not end well for investors who sunk money in these
         | large AI startups (unless of course they manage to find a
         | Softbank-style mark to sell the whole thing to), but everyone
         | will benefit from the progress AI will have made during the
         | bubble.
         | 
         | So, in the end, OpenAI will have, albeit very unwillingly,
         | fulfilled their original charter of improving humanity's lot.
        
           | jsheard wrote:
           | > My bet is that they're currently burning through other
           | people's capital at an amazing rate, but that they are light-
           | years from profitability
           | 
           | The Information leaked their internal projections a few
           | months ago, and apparently their own estimates have them
           | losing $44B between then and 2029 when they expect to finally
           | turn a profit, maybe.
        
             | j_maffe wrote:
             | That's surprisingly small
        
           | whiplash451 wrote:
           | Except that if OpenAI goes bust, very little of what they did
           | will actually be released to human kind.
           | 
           | So their contribution was really to fuel a race for
           | opensource (which they contributed little to). Pretty complex
           | of an argument.
        
           | emptysongglass wrote:
           | I've been a Plus user for a long time now. My opinion is
           | there is very much a ChatGPT suite of products that come
           | together to make for a mostly delightful experience.
           | 
           | Three things I use all the time:
           | 
           | - Canvas for proofing and editing my article drafts before
           | publishing. This has replaced an actual human editor for me.
           | 
           | - Voice for all sorts of things, mostly for thinking out loud
           | about problems or a quick question about pop culture, what
           | something means in another language, etc. The Sol voice is so
           | approachable for me.
           | 
           | - GPTs I can use for things like D&D adventure summaries I
           | need in a certain style every time without any manual
           | prompting.
        
           | vineyardmike wrote:
           | > As a result, OpenAI simply does not have a business model,
           | even if they are trying to convince the world that they do.
           | 
           | They have a super popular subscription service. If they keep
           | iterating on the product enough, they can lag on the models.
           | The business is the product not the models and not the API.
           | Subscriptions are pretty sticky when you start getting your
           | data entrenched in it. I keep my ChatGPT subscription because
           | it's the best app on Mac and already started to "learn me"
           | through the memory and tasks feature.
           | 
           | Their app experience is easily the best out of their
           | competitors (grok, Claude, etc). Which is a clear sign they
           | know that it's the product to sell. Things like DeepResearch
           | and related are the way they'll make it a sustainable
           | business - add value-on-top experiences which drive the
           | differentiation over commodities. Gemini is the only
           | competitor that compares because it's everywhere in Google
           | surfaces. OpenAI's pro tier will surely continue to get
           | better, I think more LLM-enabled features will continue to be
           | a differentiator. The biggest challenge will be continuing
           | distribution and new features requiring interfacing with
           | third parties to be more "agentic".
           | 
           | Frankly, I think they have enough strength in product with
           | their current models today that even if model training
           | stalled it'd be a valuable business.
        
           | jcgrillo wrote:
           | https://podcasts.apple.com/us/podcast/better-
           | offline/id17305...
        
           | nyarlathotep_ wrote:
           | > AI as it stands in 2025 is an amazing technology, but it is
           | not a product at all.
           | 
           | Here I'm assuming "AI" to mean what's broadly called
           | Generative AI (LLMs, photo, video generation)
           | 
           | I genuinely am struggling to see what the product is too.
           | 
           | The code assistant use cases are really impressive across the
           | board (and I'm someone who was vocally against them less than
           | a year ago), and I pay for Github CoPilot (for now) but I
           | can't think of any offering otherwise to dispute your claim.
           | 
           | It seems like companies are desperate to find a market fit,
           | and shoving the words "agentic" everywhere doesn't inspire
           | confidence.
           | 
           | Here's the thing: I remember people lining up around the
           | block for iPhone releases, XBox launches, hell even Grand
           | Theft Auto midnight releases.
           | 
           | Is there a market of people clamoring to use/get anything
           | GenAI related?
           | 
           | If any/all LLM services went down tonight, what's the impact?
           | Kids do their own homework?
           | 
           | JavaScript programmers have to remember how to write React
           | components?
           | 
           | Compare that with Google Maps disappearing, or similar.
           | 
           | LLMs are in a position where they're forced onto people and
           | most frankly aren't that interested. Did anyone ASK for
           | Microsoft throwing some Copilot things all over their
           | operating system? Does anyone want Apple Intelligence,
           | really?
        
         | jdprgm wrote:
         | If it really costs them 30x more surely they must plan on
         | putting pretty significant usage limits on any rollout to the
         | Plus tier and if that is the case i'm not sure what the point
         | is considering it seems primarily a replacement/upgrade for 4o.
         | 
         | The cognitive overhead of choosing between what will be 6
         | different models now on chatGPT and trying to map whether a
         | query is "worth" using a certain model and worrying about
         | hitting usage limits is getting kind of out of control.
        
         | chollida1 wrote:
         | > GPT 4.5 pricing is insane:
         | 
         | > I'm still gonna give it a go, though.
         | 
         | Seems like the pricing is pretty rational then?
        
           | phito wrote:
           | Not if people just try a few prompts then stop using it.
        
             | chollida1 wrote:
             | Sure but its in their best interest to lower it then and
             | only then.
             | 
             | OpenAI wouldn't be the first company to price something
             | expensive when it first comes out to capitalize on people
             | who are less price sensitive at first and then lower prices
             | to capture a bigger audience.
             | 
             | That's all pricing 101 as the saying goes.
        
               | j_maffe wrote:
               | If OAI are concerning themselves with collecting a few
               | hundereds from a small group of individuals then they
               | really have nothing better to do
        
             | nyarlathotep_ wrote:
             | How much of OAI's reported users are doing exactly this?
        
         | raytopia wrote:
         | Now the real question about AI automation starts. Is it cheaper
         | to pay a human to do the task or a AI company?
        
           | redox99 wrote:
           | It still not smart enough to replace for example customer
           | service.
        
           | fragmede wrote:
           | Humans have all sorts of issues you have to deal with. Being
           | hungover, not sleeping well, having a personality, being late
           | to work, not being able to work 24/7, very limited ability to
           | copy them. If there's a soulless generic office-droidGPT that
           | companies could hire that would never talk back and would do
           | all sorts of menial work without needing breaks or to use the
           | bathroom, I don't know that we humans stand a chance!
           | 
           | I have a bunch of work that needs doing. I can do it myself,
           | or I can hire one person to do it. I gotta train them and
           | manage them and even after I train them theres still only
           | going to be one of them, and it's subject to their
           | availability. On the other hand, if I need to train an AI to
           | do it, but I can copy that AI, and then spin them up/down
           | like on demand computer in the cloud, and not feel remotely
           | bad about spinning them down?
           | 
           | It's definitely not there yet, but it's not hard to see the
           | business case for it.
        
         | crooked-v wrote:
         | Doubly so with how good Claude 3.7 Sonnet is at $3 / 1M tokens.
        
         | Hansenq wrote:
         | GPT-4.5 is 15-30x more expensive than GPT-4o. Likely that much
         | larger in terms of parameter count too. It's massive!!
         | 
         | With more parameters comes more latent space to build a world
         | model. No wonder its internal world model is so much better
         | than previous SOTA
        
         | wavemode wrote:
         | This has been my suspicion for a long time - OpenAI have indeed
         | been working on "GPT5", but training and running it is proving
         | so expensive (and its actual reasoning abilities only
         | marginally stronger than GPT4) that there's just no market for
         | it.
         | 
         | It points to an overall plateau being reached in the
         | performance of the transformer architecture.
        
           | goatlover wrote:
           | Certainly hope so. The tech billionaires are little to
           | excited to achieve AGI and replace the workforce.
        
             | shoubidouwah wrote:
             | TBH, with the safety/alignment paradigm we have, workforce
             | replacement was not my top concern when we hit AGI. A pause
             | / lull in capabilities would be hugely helpful so that we
             | can figure how not to die along with the lightcone...
        
               | hnuser123456 wrote:
               | Is it inevitable to you that someone will create some
               | kind of techno-god behemoth AI that will figure out how
               | to optimally dominate an entire future light cone
               | starting from the point in spacetime of its self-
               | actualization? Borg or Cylons?
        
         | hintymad wrote:
         | > It sounds like it's so expensive and the difference in
         | usefulness is so lacking(?) they're not even gonna keep serving
         | it in the API for long
         | 
         | I guess the rationale behind this is paying for the marginal
         | improvement. Maybe the next few percent of improvement is so
         | important to a business that the business is willing to pay a
         | hefty premium.
        
         | shawabawa3 wrote:
         | I wonder if the pricing is partly to discourage distillation,
         | if they suspect r1 was distilled from gpt 4o
        
         | tomrod wrote:
         | I can chew through 1MM tokens with a single standard (and
         | optimized) call. This pricing is insane.
        
         | MangoCoffee wrote:
         | one of the problem seem to be there's no alternative to Nvidia
         | ecosystem. (the gpu + CUDA).
        
         | coliveira wrote:
         | In other words, they want people to pay for the privilege of
         | becoming beta testers....
        
         | wiremine wrote:
         | > GPT 4.5 pricing is insane: Price Input: $75.00 / 1M tokens
         | Cached input: $37.50 / 1M tokens Output: $150.00 / 1M tokens
         | 
         | > GPT 4o pricing for comparison: Price Input: $2.50 / 1M tokens
         | Cached input: $1.25 / 1M tokens Output: $10.00 / 1M tokens
         | 
         | Their examples don't seem 30x better. :-)
        
         | ren_engineer wrote:
         | hyperscalers in shambles, no clue why they even released this
         | other than the fact they didn't want to admit they wasted an
         | absurd amount of money for no reason
        
         | hn_throwaway_99 wrote:
         | The price is obviously 15-30x that of 4o, but I'd just posit
         | that there are some use cases where it may make sense. It
         | probably doesn't make sense for the "open-ended consumer facing
         | chatbot" use case, but for other use cases that are fewer and
         | higher value in nature, it could if it's abilities are
         | considerably better than 4o.
         | 
         | For example, there are now a bunch of vendors that sell
         | "respond to RFP" AI products. The number of RFPs that any sales
         | organization responds to is probably no more than a couple a
         | week, but it's a very time-consuming, laborious process. But
         | the payoff is obviously very high if a response results in a
         | closed sale. So here paying 30x for marginally better
         | performance makes perfect sense.
         | 
         | I can think of a number of similar "high value, relatively low
         | occurrence" use cases like this where the pricing may not be a
         | big hindrance.
        
           | Manouchehri wrote:
           | Yeah, agreed.
           | 
           | We're one of those types of customers. We wrote an OpenAI API
           | compatible gateway that automatically batches stuff for us,
           | so we get 50% off for basically no extra dev work in our
           | client applications.
           | 
           | I don't care about speed, I care about getting the right
           | answer. The cost is fine as long as the output generates us
           | more profit.
        
         | kristofferR wrote:
         | "GPT-4.5 is not a frontier model, but it is OpenAI's largest
         | LLM, improving on GPT-4's computational efficiency by more than
         | 10x."[1]
         | 
         | I don't get it, it is supposedly much cheaper to run?
         | 
         | [1] https://cdn.openai.com/gpt-4-5-system-card.pdf (page 7,
         | bottom)
        
         | acchow wrote:
         | > It sounds like it's so expensive and the difference in
         | usefulness is so lacking(?)
         | 
         | The claimed hallucination rate is dropping from 61% to 37%.
         | That's a "correct" rate increasing from 29% to 63%.
         | 
         | Double the correct rate costs 15x the price? That seems absurd,
         | unless you think about how mistakes compound. Even just 2 steps
         | in and you're comparing a 8.4% correct rate vs 40%. 3 automated
         | steps and it's 2.4% vs 25%.
        
       | Xiol32 wrote:
       | The example GPT-4.5 answers from the livestream are just... too
       | excitable? Can't put my finger on it, but it feels like they're
       | aimed towards little kids.
        
         | jug wrote:
         | It made me wonder how much of that was due to the system prompt
         | too.
        
       | virgildotcodes wrote:
       | That presentation was super underwhelming. We got to watch them
       | compare... the vibes? ... of 4.5 vs o1.
       | 
       | No wonder Sam wasn't part of the presentation.
        
         | Etheryte wrote:
         | And to top it off, it costs $75.00 per 1M vibes.
        
         | thomas34298 wrote:
         | Sam tweeted "taking care of my kid in the hospital":
         | 
         | https://x.com/sama/status/1895210655944450446
         | 
         | Let's not assume that he's lying. Neither the presentation nor
         | my short usage via the API blew me away, but to really evaluate
         | it, you'd have to use it longer on a daily basis. Maybe that
         | becomes a possiblity with the announced performance
         | optimizations that would lower the price...
        
         | jug wrote:
         | It should've just been a web launch without video.
        
       | bakugo wrote:
       | API is literally 5 times more expensive than Claude 3 Opus, and
       | it doesn't even seem to do anything impressive. What's the
       | business strategy here?
        
       | hidelooktropic wrote:
       | I'm not sure that doing a live stream on this was the right way
       | to go. I would've just quietly sent out a press release. I'm sure
       | they have better things on the way.
        
       | sebastiennight wrote:
       | It is interesting that they are focusing a large part of this
       | release on the model having a higher "EQ" (Emotional Quotient).
       | 
       | We're far from the days of "this is not a person, we do not want
       | to make it addictive" and getting a firm foot on the territory of
       | "here's your new AI friend".
       | 
       | This is very visible in the example comparing 4o with 4.5 when
       | the user is complaining about failing a test, where 4o's response
       | is what one would expect from a "typical AI response" with
       | problem-solving bullets, and 4.5 is sending what you'd expect
       | from a pal over instant messaging.
       | 
       | It seems Anthropic and Grok have both been moving in this
       | direction as well. Are we going to see an escalation of
       | foundation models impersonating "a friendly person" rather than
       | "a helpful assistant"?
       | 
       | Personally I find this worrying and (as someone who builds upon
       | SOTA model APIs) I really hope this behavior is not going to seep
       | into API responses, or will at least be steerable through the
       | system/developer prompt.
        
         | og_kalu wrote:
         | The whole robotic, monotone, helpful assistant thing was
         | something these companies had to actively hammer in during the
         | post-training stage. It's not really how LLMs will sound by
         | default after pre-training.
         | 
         | I guess they're caring less and less about that effort
         | especially since it hurts the model in some ways like creative
         | writing.
        
           | sebastiennight wrote:
           | If it's just a different choice during RLHF, I'll be curious
           | to see what are the trade-offs in performance.
           | 
           | The "buddy in a chat group" style answers do not make me feel
           | like asking it for a story will make the story
           | long/detailed/poignant enough to warrant the difference.
           | 
           | I'll give it a try and compare on creative tasks.
        
           | turnsout wrote:
           | Or maybe they're just getting better at it, or developing
           | better taste. After switching to Claude, I can't go back to
           | ChatGPT's overly verbose bullet-point laden book reports
           | every time I ask a question. I don't think that's pretraining
           | --it's in the way OpenAI approaches tuning and prompting vs
           | Anthropic.
        
         | tmaly wrote:
         | I would like to see a humor test. So far, I have not seen any
         | model response that has made me laugh.
        
           | sebastiennight wrote:
           | The "roast" tools that have popped up (using either DeepSeek
           | or o3-mini) are pretty funny.
           | 
           | Eg. https://news.ycombinator.com/item?id=43163654
        
             | jcims wrote:
             | OK now that is some funny shit.
        
           | AgentME wrote:
           | My benchmark for this has been asking the model to write some
           | tweets in the style of dril, a popular user who writes short
           | funny tweets. Sometimes I include a few example tweets in the
           | prompt too. Here's an example of results I got from Claude 3
           | Opus and GPT 4 for this last year:
           | https://bsky.app/profile/macil.tech/post/3kpcvicmirs2v. My
           | opinion is that Claude's results were mostly bangers while
           | GPT's were all a bit groanworthy. I need to try this again
           | with the latest models sometime.
        
           | tkgally wrote:
           | How does the following stand-up routine by Claude 3.7 Sonnet
           | work for you?
           | 
           | https://gally.net/temp/20250225claudestandup2.html
        
             | lurker9001 wrote:
             | incredible
        
           | turnsout wrote:
           | If you like absurdist humor, go into the OpenAI playground,
           | select 3.5-Turbo, and dial up the temperature to the point
           | where the output devolves into garbled text after 500 tokens
           | or so. The first ~200 tokens are in the freaking sweet spot
           | of humor.
        
             | amarcheschi wrote:
             | Could someone post an example?
        
             | rl3 wrote:
             | Maybe it's rose-colored glasses, but 3.5 was really the
             | golden era for LLM comedy. More modern LLMs can't touch it.
             | 
             | Just ask it to write you a film screenplay involving some
             | hard-ass 80s/90s action star and someone totally unrelated
             | and opposite of that. The ensuring unhinged magic is
             | unparalleled.
        
               | jcims wrote:
               | I built a little AI assistant to read my calendar and
               | send me a summary of my day every morning. I told it to
               | roast me and be funny with it.
               | 
               | 3.5 was *way* better than anything else at that.
        
         | nialv7 wrote:
         | Well yeah, if the llm can keep you engaged and talking, that'll
         | make them a lot more money; compared to if you just use it as a
         | information retrieval tool in which case you are likely to
         | leave after getting what you are looking for.
        
           | TheAceOfHearts wrote:
           | Since they offer a subscription, keeping you engaged just
           | requires them to waste more compute. The ideal case would be
           | that the LLM gives you a one shot correct response using as
           | little compute as possible.
        
             | sebastiennight wrote:
             | In a subscription business, you don't want the user to use
             | as few resources as possible. It's the wrong optimization
             | to make.
             | 
             | You want users to keep coming back as often as possible (at
             | the lowest cost-per-run possible though). If they are not
             | coming back they are not renewing.
             | 
             | So, yes, it makes sense to make answers shorter to cut on
             | compute cost (which these SMS-length replies could
             | accomplish) but the main point of making the AI flirtatious
             | or "concerned" is possibly the addictive factor of having a
             | shoulder to cry on 24/7, one that does not call you on your
             | BS and is always supportive... for just $20 a month
             | 
             | The "one-shot correct response" to "I failed my exams"
             | might be "Tough luck, try better next time" but if you do
             | that, you will indeed use very little compute _because
             | people will cancel the subscription and never come back_.
        
               | johnthewise wrote:
               | AI subscriptions are already very sticky . I can't
               | imagine at least not paying for one, so I doubt they care
               | about retention like the rest of us plebs do.
        
             | nialv7 wrote:
             | Plus level subscription has limits too, and Pro level costs
             | 10x more - as long as Pro users don't use ChatGPT 10x more
             | than Plus users on average, OpenAI can benefit. There's
             | also the user retention factor.
        
         | bredren wrote:
         | Yes, the "personality" (vibe) of the model is a key qualitative
         | attribute of gpt-4.5.
         | 
         | I suspect this has something to do with shining light on an
         | increased value prop in a dimension many people will appreciate
         | since gains on quantitative comparison with other models were
         | not notable enough to pop eyeballs.
        
         | callc wrote:
         | > We're far from the days of "this is not a person, we do not
         | want to make it addictive" and getting a firm foot on the
         | territory of "here's your new AI friend".
         | 
         | That's a hard nope from me, when companies pull that move. I'll
         | stick to my flesh and blood humans who still hallucinate but
         | only rarely.
        
         | orbital-decay wrote:
         | Anthropic pretty much abandoned this direction after Claude 3,
         | and said it wasn't what they wanted [1]. Claude 3.5+ is
         | extremely dry and neutral, it doesn't seem to have the same
         | training.
         | 
         |  _> Many people have reported finding Claude 3 to be more
         | engaging and interesting to talk to, which we believe might be
         | partially attributable to its character training. This wasn't
         | the core goal of character training, however. Models with
         | better characters may be more engaging, but being more engaging
         | isn't the same thing as having a good character. In fact, an
         | excessive desire to be engaging seems like an undesirable
         | character trait for a model to have._
         | 
         | [1] https://www.anthropic.com/research/claude-character
        
       | CitizenTen wrote:
       | GPT pro already has already rummored to be 100k users. You think
       | GPT 4.5 will add to that even with the insane costs for corporate
       | users?
        
         | cristiancavalli wrote:
         | What rumors? I looked and can't find something to substantiate
         | that #
        
       | andsoitis wrote:
       | "this isn't a reasoning model and won't crush benchmarks."
       | 
       | -- https://x.com/sama/status/1895203654103351462
        
       | MaxPock wrote:
       | One thing that Altman does extremely well is to over-promise and
       | under-deliver.
        
       | erulabs wrote:
       | Finally a scaling wall? This is apparently (based on pricing)
       | using about an order of magnitude more compute, and is only maybe
       | 10% more intelligent. Ideally DeepSeeks optimizations help bring
       | the costs way down, but do any AI researchers want to comment on
       | if this changes the overall shape of the scaling curve?
        
         | fpgaminer wrote:
         | Seems on par with the existing scaling curve. If I had to
         | speculate, this model would have been an internal-only model,
         | but they're releasing it for PR. An optimized version with 99%
         | of the performance for 1/10th the cost will come out later.
        
           | j_maffe wrote:
           | This is the shittiest PR move I've seen since the AI trend
           | started.
        
       | ein0p wrote:
       | Now imagine this model (or an optimized/slightly downsized
       | variant thereof) as a base for a "thinking" one.
        
       | bparsons wrote:
       | Are they saying that 4.5 has a 35% hallucination rate? That chart
       | is a bit confusing.
        
         | riku_iki wrote:
         | its on that benchmark, which likely is very challenging.
        
       | Seattle3503 wrote:
       | The GPT-1 response to the example prompt "What was the first
       | language?" got a chuckle out of me
        
         | aldanor wrote:
         | The question being, will we be chuckling at current models
         | responses in 5-10y from now?
        
       | joshuamcginnis wrote:
       | I'm one week in on heavy grok usage. I didn't think I'd say this,
       | but for personal use, I'm considering cancelling my OpenAI plan.
       | 
       | The one thing I wish grok had was more separation of the UI from
       | X itself. The interface being so coupled to X puts me off and
       | makes it feel like a second-hand citizen. I like ChatGPTs
       | minimalist UI.
        
         | aldanor wrote:
         | Theres grok.com which is standalone and with its own UI
        
           | it wrote:
           | There's also a standalone Grok app at least on iOS.
        
         | richard_todd wrote:
         | I find grok to be the best overall experience for the types of
         | tasks I try to give AI (mostly: analyze pdf, perform and
         | proofread OCR, translate Medieval Latin and Hebrew, remind me
         | how to do various things in python or SwiftUI).
         | ChatGPT/gemini/copilot all fight me occasionally, but grok just
         | tries to help. And the hallucinations aren't as frequent, at
         | least anecdotally.
        
         | fzzzy wrote:
         | Don't they have a standalone Grok app now? I thought I saw
         | that. [edit] ah some sibling comments mention this as well
        
       | selalipop wrote:
       | I've been working on post-training models for tasks that require
       | EQ, so it's validating to see OpenAI working towards that too.
       | 
       | That being said, this is very expensive.
       | 
       | - Input: $75.00 / 1M tokens
       | 
       | - Cached input: $37.50 / 1M tokens
       | 
       | - Output: $150.00 / 1M tokens
       | 
       | One of the most interesting applications of models with higher EQ
       | is personalized content generation, but the size and cost here
       | are at odds with that.
        
       | kgeist wrote:
       | >GPT-4.5 is more succinct and conversational
       | 
       | I wonder why they highlight it as an achievement when they could
       | have simply tuned 4o to be more conversational and less like a
       | bullet-point-style answer machine. They did something to 4o
       | compared to the previous models which made the responses feel
       | more canned.
        
       | mvdtnz wrote:
       | OpenAI doubling down on the American-style therapy-speak instead
       | of focusing on usefulness. No thanks.
        
       | wewewedxfgdf wrote:
       | I feel like OpenAI is pursuing AGI when Anthropic/Claude is
       | pursuing making AI awesome for practical things like coding.
       | 
       | I only ever using OpenAI's coding now as a double check against
       | Claude.
       | 
       | Does OpenAI have their eyes on the ball?
        
         | ls_stats wrote:
         | >I feel like OpenAI is pursuing AGI
         | 
         | I don't think so, the "AGI guy" was Ilya Sutskever, he is gone,
         | he wanted to make OpenAI "less comercial", AGI is just a
         | buzzword for Altmann.
        
           | rakejake wrote:
           | Right. A good chunk of the "old guard" is now gone - Ilya to
           | SSI, Mira and a bunch of others to a new venture called
           | Thinking Machines, Alec Radford etc. Remains to be seen if
           | OpenAI will be the leader or if other players catch up.
        
         | rakejake wrote:
         | My usage has come down to mostly Claude (until I run out of
         | free tier quota) and then Gemini. Claude is the best for code
         | and Gemini 2.0 Flash is good enough while also being free (well
         | considering how much data G has hoovered up over the years,
         | perhaps not) and more importantly highly available.
         | 
         | For simple queries like generating shell scripts for some
         | plumbing, or doing some data munging, I go straight to Gemini.
        
           | HarHarVeryFunny wrote:
           | > My usage has come down to mostly Claude (until I run out of
           | free tier quota) and then Gemini
           | 
           | Yep, exactly same here.
           | 
           | Gemini 2.0 Flash is extremely good, and I've yet to hit any
           | usage limits with them - for heavy usage I just go to Gemini
           | directly. For "talk to an expert" usage, Claude is hard to
           | beat though.
        
           | wayeq wrote:
           | Claude still can't make real time web searches though for
           | RAG, unless I'm missing something.
        
         | resource0x wrote:
         | Pursuing AGI? What method do they use to pursue something that
         | no one knows what it is? They will keep saying they are
         | pursuing AGI as long as there's a buyer for their BS.
        
       | lblume wrote:
       | Am I missing something, or do the results not even look that much
       | better? Referring to the output quality, this just seems like a
       | different prompting style and RLHF, not really an improved model
       | at all.
        
       | infinet wrote:
       | Can it be self-hosted? Many institutions and organizations are
       | hesitant to use AI because concerns of data leaking over chatbot.
       | Open models, on the other hand, can be self-hosted. There is a
       | deepseek arm race in other part of the world. Universities are
       | racing to host their own deepseek. Hospitals, large businesses,
       | local governments, even courts are deploying or showing interest
       | in self-hosting deepseek.
        
         | YetAnotherNick wrote:
         | Do you know of any university that host Deepseek?
        
         | moralestapia wrote:
         | OpenAI has never released a single model that could be self-
         | hosted.
         | 
         | GPT-2? Maybe not even that one.
        
       | taytus wrote:
       | Who wants a model that is not reasoning? The older models are
       | just fine.
        
       | eightysixfour wrote:
       | Seeing OpenAI and Anthropic go different routes here is
       | interesting. It is worth moving past the initial knee jerk
       | reaction of this model being unimpressive and some of the
       | comments about "they spent a massive amount of money and had to
       | ship something for it..."
       | 
       | * Anthropic appears to be making a bet that a single paradigm
       | (reasoning) can create a model which is excellent for all use
       | cases.
       | 
       | * OpenAI seems to be betting that you'll need an ensemble of
       | models with different capabilities, working as a single system,
       | to jump beyond what the reasoning models today can do.
       | 
       | Based on all of the comments from OpenAI, GPT 4.5 is absolutely
       | massive, and with that size comes the ability to store far more
       | factual data. The scores in ability oriented things - like coding
       | - don't show the kind of gains you get from reasoning models but
       | the fact based test, SimpleQA, shows a pretty large jump and a
       | dramatic reduction in hallucinations. You can imagine a scenario
       | where GPT4.5 is coordinating multiple, smaller, reasoning agents
       | and using its factual accuracy to enhance their reasoning, kind
       | of like ruminating on an idea "feels" like a different process
       | than having a chat with someone.
       | 
       | I'm really curious if they're actually combining two things right
       | now that could be split as well, EQ/communications, and factual
       | knowledge storage. This could all be a bust, but it is an
       | interesting difference in approaches none-the-less, and worth
       | considering that OpenAI _could_ be right.
        
         | nomel wrote:
         | > OpenAI seems to be betting that you'll need an ensemble of
         | models with different capabilities, working as a single system,
         | to jump beyond what the reasoning models today can do.
         | 
         | The high level block diagrams for tech always end up converging
         | to those found in biological systems.
        
           | eightysixfour wrote:
           | Yeah, I don't know enough real neuroscience to argue either
           | side. What I can say is I feel like this path is more like
           | the way that I observe that I think, it _feels_ like there
           | are different modes of thinking and processes in the brain,
           | and it _seems_ like transformers are able to emulate at least
           | two different versions of that.
           | 
           | Once we figure out the frontal cortex & corpus callosum part
           | of this, where we aren't calling other models over APIs
           | instead of them all working in the same shared space, I have
           | a feeling we'll be on to something pretty exciting.
        
         | sebastiennight wrote:
         | > * OpenAI seems to be betting that you'll need an ensemble of
         | models with different capabilities, working as a single system,
         | to jump beyond what the reasoning models today can do.
         | 
         | Seems inaccurate as their most recent claim I've seen is that
         | they expect this to be their last non-reasoning model, and are
         | aiming to provide all capacities together in the future model
         | releases (unifying the GPT-x and o-x lines)
         | 
         | See this claim on TFA:
         | 
         | > We believe reasoning will be a core capability of future
         | models, and that the two approaches to scaling--pre-training
         | and reasoning--will complement each other.
        
           | eightysixfour wrote:
           | From Sam's twitter:
           | 
           | > After that, a top goal for us is to unify o-series models
           | and GPT-series models by creating systems that can use all
           | our tools, know when to think for a long time or not, and
           | generally be useful for a very wide range of tasks.
           | 
           | > In both ChatGPT and our API, we will release GPT-5 as a
           | system that integrates a lot of our technology, including o3.
           | We will no longer ship o3 as a standalone model.
           | 
           | You could read this as unifying the models _or_ building a
           | unified systems which coordinate multiple models. The second
           | sentence, to me, implies that o3 will still exist, it just
           | won 't be standalone, which matches the idea I shared above.
        
             | sebastiennight wrote:
             | Ah, great point. Yes, the wording here would imply that
             | they're basically planning on building scaffolding around
             | multiple models instead of having one more capable Swiss
             | Army Knife model.
             | 
             | I would feel a bit bummed if GPT-5 turned out not to be a
             | model, but rather a "product".
        
               | eightysixfour wrote:
               | For me it depends on how the models are glued together.
               | Connected by function calling and APIs? Probably meh...
               | 
               | Somehow working together in the same latent space? That
               | could be neat.
        
             | tmpz22 wrote:
             | I worry eliminating consumer choice will drive up prices
             | for only a nominal gain in utility for most users.
        
         | wongarsu wrote:
         | Or the other way around: smaller reasoning models that can call
         | out to GPT-4.5 to get their facts right.
        
           | eightysixfour wrote:
           | Maybe, I'm inclined to think OpenAI believes the way I laid
           | it out though, specifically because of their focus on
           | communication and EQ in 4.5. It seems like they believe the
           | large, non-reasoning model, will be "front of house."
           | 
           | Or they'll use some kind of trained router which sends the
           | request to the one it thinks it should go to first.
        
       | ilrwbwrkhv wrote:
       | Question: explain thread first macros in clojure in a short
       | paragraph
       | 
       | GPT 4.5:
       | 
       | In Clojure, the *thread-first macro* (`->`) is used to improve
       | code readability by threading a value through a series of
       | function calls. It takes an initial value and passes it as the
       | first argument to the next function, then continues threading the
       | result through subsequent functions. This eliminates deeply
       | nested expressions and makes the code more readable.
       | 
       | For example,
       | 
       | ```clojure (-> 5 inc (* 2) str) ;; Equivalent to (str (* 2 (inc
       | 5))) - "12" ```
       | 
       | Here, `5` is incremented to `6`, then multiplied by `2` to get
       | `12`, and finally converted to a string.
       | 
       | ---
       | 
       | Deepseek v3:
       | 
       | In Clojure, the `->>` (thread-last) and `->` (thread-first)
       | macros are used to improve code readability by allowing a more
       | linear, pipeline-like flow of operations. The `->` (thread-first)
       | macro takes an initial value and threads it as the first argument
       | through a series of forms. For example, `(-> x (f) (g))`
       | translates to `(g (f x))`. This is particularly useful when
       | working with nested function calls, as it flattens the structure
       | and makes the sequence of transformations clearer. It's commonly
       | used with functions that take the threaded value as their first
       | argument, such as collection operations or Java interop methods.
        
       | datadrivenangel wrote:
       | They tried, but doesn't look better?
        
       | DaveMcMartin wrote:
       | This feels more like a release they pushed out to keep the "hype"
       | alive rather than something they were eager to share. Honestly,
       | the results don't seem all that impressive, and considering the
       | price, it just doesn't feel worth it.
        
       | bla3 wrote:
       | I wonder if we're starting to see the effects of the mass exodus
       | a while ago.
        
       | 42lux wrote:
       | The announcements early on were relatively sincere and technical
       | with papers and nice pages explaining the new models in easy
       | language and now we get this marketing garbage. Probably the
       | fastest enshitification I've seen.
        
       | sky2224 wrote:
       | Honestly, the most astounding part of this announcement is their
       | comparison to o3-mini with QA prompts.
       | 
       | EIGHTY PERCENT hallucination rate? Are you kidding me?
       | 
       | I get that the model is meant to be used for logic and reasoning,
       | but nowhere does OpenAI make this explicitly clear. A majority of
       | users are going to be thinking, "oh newer is better," and pick
       | that.
        
         | jug wrote:
         | Yeah it was an abysmal result (any 50%+ hallucination result in
         | that bench is pretty bad) and worse than o1-mini in the
         | SimpleQA paper. On that topic, Sonnet 3.5 "Old" hallucinates
         | less than GPT-4.5, just for a bit of added perspective here.
        
       | freediver wrote:
       | The results for GPT - 4.5 are in for Kagi LLM benchmark too.
       | 
       | It does crush our benchmark - time to make new? ;) - with
       | performance similar of that of reasoning models. It does come at
       | a great price both in cost and speed.
       | 
       | A monster is what they created. But looking at the tasks it
       | fails, some of them my 9 year old would solve. Still in this
       | weird limbo space of super knowledge and low intelligence.
       | 
       | May be remembered as the last the last of the 'big ones', can't
       | imagine this will be a path for the future.
       | 
       | https://help.kagi.com/kagi/ai/llm-benchmark.html
        
         | theodorthe5 wrote:
         | If Gemini 2 is the top in your benchmark, make sure to re-check
         | your benchmark.
        
           | shawabawa3 wrote:
           | Gemini 2 pro is actually very impressive (maybe not for
           | coding, haven't used it for that)
           | 
           | Flash is pretty garbage but cheap
        
           | istjohn wrote:
           | Gemini 2.0 Pro is quite good.
        
       | simonw wrote:
       | If you want to try it out via their API you can run it through my
       | LLM tool using uvx like this:                 uvx --with 'https:/
       | /github.com/simonw/llm/archive/801b08bf40788c09aed617525287631031
       | 2fe667.zip' \         llm -m gpt-4.5-preview 'impress me'
       | 
       | You may need to set an API key first, either with `export
       | OPENAI_API_KEY='xxx'` or using this command to save it to a file:
       | uvx llm keys set openai       # paste key here
       | 
       | Or this to get a chat session going:                 uvx --with '
       | https://github.com/simonw/llm/archive/801b08bf40788c09aed61752528
       | 76310312fe667.zip' \         llm chat -m gpt-4.5-preview
       | 
       | I'll probably have a proper release out later today. Details
       | here: https://github.com/simonw/llm/issues/795
        
         | ashu1461 wrote:
         | Just curious, does this stream the output or renders all at
         | once ?
        
       | synapsomorphy wrote:
       | Claude 3.6 (new 3.5) and 3.7 non-reasoning are much better at
       | pretty much everything, and much cheaper. What's Anthropic's
       | secret sauce?
        
         | taytus wrote:
         | They ship more focused on their mission than OpenAI.
        
         | moralestapia wrote:
         | Huh?
         | 
         | Post benchmark links.
        
         | film42 wrote:
         | I think it's a classic expectations problem. OpenAI is neither
         | _open_ nor is it releasing an _AGI_ model in the near future.
         | But when you see a new major model drop, you can't help but
         | ask, "how close is this to the promise of AGI they say is just
         | around the corner?" Not even close. Meanwhile Anthropic is
         | keeping their heads down, not playing the hype game, and
         | letting the model speak for itself.
        
           | anothermathbozo wrote:
           | Anthropic's CEO said their technology would end all disease
           | and expand our lifespans to 200 years. What on earth do you
           | mean they're not playing the hype game?
        
       | jampa wrote:
       | First impression of GPT-4.5:
       | 
       | 1. It is very very slow, for some applications where you want
       | real time interactions is just not viable, the text attached
       | below took 7s to generate with 4o, but 46s with GPT4.5
       | 
       | 2. The style it writes is way better: it keeps the tone you ask
       | and makes better improvements on the flow. One of my biggest
       | complaints with 4o is that you want for your content to be more
       | casual and accessible but GPT / DeepSeek wants to write like
       | Shakespeare did.
       | 
       | Some comparisons on a book draft: GPT4o (left) and GPT4.5
       | (green). I also adjusted the spacing around the paragraphs, to
       | better diff match. I still am wary of using ChatGPT to help me
       | write, even with GPT 4.5, but the improvement is very noticeable.
       | 
       | https://i.imgur.com/ogalyE0.png
        
         | remus wrote:
         | > It is very very slow
         | 
         | Could that be partially due to a big spike in demand at launch?
        
           | jampa wrote:
           | Possibly, repeating the prompt I got a much higher speed,
           | taking 20s on average now, which is much more viable. But
           | that remains to be seen when more people start using this
           | version in production.
        
         | jedberg wrote:
         | Oh yeah, that right side version is WAY better, and sounds much
         | more like a human.
        
         | MichaelZuo wrote:
         | How does it compare with o1 and o3 preview?
        
           | jampa wrote:
           | o3 is okay for text checking but has issues following the
           | prompt correctly, same as o1 and DeepSeek R1, I feel that I
           | need to prompt smaller snippets with them.
           | 
           | Here is the o3 vs a new run of the same text in GPT 4.5
           | 
           | https://www.diffchecker.com/ZEUQ92u7/
        
             | MichaelZuo wrote:
             | Thanks, though it says o1 on the page, is that a typo?
        
         | FergusArgyll wrote:
         | I opened your link in a new tab and looked at it a couple
         | minutes later. By then I forgot which was o and which was .5
         | 
         | I honestly couldn't decide which I prefer
        
           | niek_pas wrote:
           | I definitely prefer the 4.5, but that might just be because
           | it sounds 'less like ChatGPT', ironically.
        
         | rl3 wrote:
         | > _1. It is very very slow, ... below took 7s to generate with
         | 4o, but 46s with GPT4.5_
         | 
         | This is positively luxurious by o1-pro standards which I'd say
         | _average_ 5 minutes. That said I totally agree even ~45s isn 't
         | viable for real-time interactions. I'm sure it'll be optimized.
         | 
         | Of course, my comparing it to the highest-end CoT model in
         | [publicly-known] existence isn't entirely fair since they're
         | sort of apples and oranges.
        
           | philomath_mn wrote:
           | I paid for pro to try `o1-pro` and I can't seem to find any
           | use case to justify the insane inference time. `o3-mini-high`
           | seems to do just as well in seconds vs. minutes.
        
         | thfuran wrote:
         | >One of my biggest complaints with 4o is that you want for your
         | content to be more casual and accessible but GPT / DeepSeek
         | wants to write like Shakespeare did.
         | 
         | Well, maybe like a Sophomore's bumbling attempt to write like
         | Shakespeare.
        
         | ChiefNotAClue wrote:
         | Right side, by a large margin. Better word choice and more
         | natural flow. It feels a lot more human.
        
         | kristianp wrote:
         | How do the two versions match so closely? They have the same
         | content in each paragraph, just worded slightly differently. I
         | wouldn't expect them to write paragraphs that match in size and
         | position like that.
        
       | mchusma wrote:
       | wow, openai really missed here. Reading the blog I thought like a
       | minor, incremental minor catch up release for 4o. I thought "wow
       | maybe this is cheaper than 4o so it will offset the pricing
       | difference between this and something like Claude Sonnet 3.7 or
       | Gemini 2.0 Flash both of which performs better. But its like
       | 20x-100x more expensive!
       | 
       | In other words, these performance stats with Gemini 2.0 Flash
       | pricing looks reasonable. At these prices, zero usecases for
       | anyone I think. This is a dead on arrival model.
        
       | jasonjmcghee wrote:
       | This is such as confusing release / announcement.
       | 
       | It seems clearly worse than Claude Sonnet 3.7, yet costs 30x as
       | much?
       | 
       | What's going on with their pricing? I misread it as $7.5/M input
       | and that that was very overpriced... then realized it was 10x
       | that much!
        
         | ashu1461 wrote:
         | Is it worse than clause sonnet with reasoning enabled or
         | disabled ?
        
       | moffkalast wrote:
       | OpenAI will do literally anything but ship GPT-5.
        
       | saghm wrote:
       | This is probably a dumb question, but are we just gonna be stuck
       | on always having X.5 versions of GPT forever? If there's never an
       | X.0, it feels like it's basically meaningless.
        
         | xd1936 wrote:
         | The gas costs $3.599
        
         | dragonwriter wrote:
         | There was 3, 3.5, 4, 4o, and then 4.5.
         | 
         | There are things to complain about in that, but "its always X.5
         | and never X.0" doesn't really seem to be one of them.
        
       | JohnMakin wrote:
       | Their example for 4.5 on "having a tough time":
       | 
       | > I'm going through a tough time after failing a test > > Aw, I'm
       | really sorry to hear that. Failing a test can feel pretty tough
       | and discouraging, but remember, it's just one moment--not a
       | reflection of who you are or your capabilities. Want to talk
       | about what happened, or do you just need a distraction? I'm here
       | either way.
       | 
       | vs 4o:
       | 
       | >I'm going through a tough time after failing a test > >I'm
       | really sorry to hear that you're going through a tough time. Here
       | are a few >things you might consider: > >1. Reflect and
       | Understand: Try to analyze what might have gone wrong. Was it a
       | >lack of preparation, time management during the test, or perhaps
       | >misunderstanding the material? Understanding the root cause can
       | help you address >it in the future. > >2. Seek Support: Talk to
       | someone you trust--friends, family, or a school >counselor. They
       | can offer a different perspective and sometimes just talking
       | >about your feelings can be alleviating. > >3. Speak with Your
       | Teacher or Professor: They can provide insights into where >you
       | might have gone wrong and offer advice on how to improve. They
       | might also >offer opportunities for make-up work or extra credit.
       | > >4. Develop a Study Plan: Based on your reflections, create a
       | plan to improve >your study habits. This might include setting
       | regular study times, breaking down >material into manageable
       | portions, or finding new study methods that work better >for you.
       | > >5. Stay Positive and Forgive Yourself: Mistakes and setbacks
       | are a part of life. >It's important to be kind to yourself and
       | recognize that failure is a stepping >stone to success. > >6.
       | Focus on the Bigger Picture: Remember that one test is just one
       | part of your >educational journey. There will be many more
       | opportunities to do well. > >If you need further support or
       | resources, consider reaching out to educational >support services
       | at your institution, or mental health resources if you're
       | >feeling particularly overwhelmed. You're not alone in this, and
       | things can get >better with time and effort.
       | 
       | Is it just me or is the 4o response insanely better? I'm not the
       | type of person to reach for a LLM for help about this kind of
       | thing, but if I were, the 4o respond seems _vastly_ better to the
       | point I 'm surprised they used that as their main "EQ" example.
        
         | bitshiftfaced wrote:
         | When people are in an emotional state, it's usually better to
         | start by reacting with empathy and perspective-taking.
         | 
         | Source: interacting with my SO.
        
         | IMTDb wrote:
         | 4o has a very strong artificial vibe. It feels a bit "autistic"
         | (probably a bad analogy but couldn't find a better word to
         | describe what I mean): you feel bad ? must say sorry then give
         | a TODO list on how to feel better.
         | 
         | 4.5 still feels a bit artificial but somehow also more
         | emotionally connected. It removed the weird "bullet point lists
         | of things to do" and focused on the emotional part; which is
         | also longer than 4o
         | 
         | If I am talking to a human I would definitely expect him/her to
         | react more like 4.5 than like 4o. If the first sentence that
         | comes out of their mouth after I explain them that I feel bad
         | is "here is a list of things you might consider", I will find
         | it strange. We can reach that point but it's usually after a
         | bit more talk; human kinda need that process, and it feels like
         | 4.5 understands that better than 4o.
         | 
         | Now of course which one is "better" really depends on the
         | context; what you expect of the model and how you intend to use
         | is. Until now every single OpenAI update on the main series has
         | always been a strict improvement over the previous model. Cost
         | aside, there wasn't really any reason to keep using 3.5 when 4
         | got released. This is not the case here; even assuming
         | unlimited money you still might wanna select 4o in the dropdown
         | sometimes instead of 4.5.
        
       | torginus wrote:
       | My 2 cents (disclaimer: I am talking out of my ass) here is why
       | GPTs actually suck at fluid knowledge retrievel (which is kinda
       | their main usecase, with them being used as knowledge engines) -
       | they've mentioned that if you train it on 'Tom Cruise was born
       | July 3, 1962', it won't be able to answer the question "Who was
       | born on July 3, 1962", if you don't feed it this piece of
       | information. It can't really internally corellate the information
       | it has learned, unless you train it to, probably via synthethic
       | data, which is what OpenAI has probably done, and that's the
       | information score SimpleQA tries to measure.
       | 
       | Probably what happened, is that in doing so, they had to scale
       | either the model size or the training cost to untenable levels.
       | 
       | In my experience, LLMs really suck at fluid knowledge retrieval
       | tasks, like book recommendation - I asked GPT4 to recommend me
       | some SF novels with certain characteristics, and what it spat out
       | was a mix of stuff that didn't really match, and stuff that was
       | really reaching - when I asked the same question on Reddit, all
       | the answers were relevant and on point - so I guess there's still
       | something humans are good for.
       | 
       | Which is a shame, because I'm pretty sure relevant product
       | recommendation is a many billion dollar business - after all
       | that's what Google has built it's empire on.
        
         | woah wrote:
         | Perhaps you could use LLMs in a list ranking context to
         | generate your scifi recommendations
         | https://github.com/noperator/raink?tab=readme-ov-file
        
         | staticman2 wrote:
         | You make a good point: I think these LLM's have a strong bias
         | towards recommending the most popular things in pop culture
         | since they really only find the most likely tokens and report
         | on that.
         | 
         | So while they may have a chance of answering "What is this non
         | mainstream novel about" they may be unable to recommend the
         | novel since it's not a likely series of tokens in response to a
         | request for a book recommendation.
        
         | vel0city wrote:
         | An LLM on its own isn't necessarily great for fluid knowledge
         | retrieval, as in directly from its training data. But they're
         | pretty good when you add RAG to it.
         | 
         | For instance, asking Copilot "Who was born on July 3, 1962"
         | gave the response:
         | 
         | > One notable person born on July 3, 1962, is Tom Cruise, the
         | famous American actor known for his roles in movies like Risky
         | Business, Jerry Maguire, and Rain Man.
         | 
         | > Are you a fan of his work?
         | 
         | It cited this page:
         | 
         | https://www.onthisday.com/date/1962/july/3
        
       | jefffoster wrote:
       | Does anyone have any intuition about the how reasoning improves
       | based on the strength of the underlying model?
       | 
       | I'm wondering whether this seemingly underwhelming bump on 4o
       | magnifies when/if reasoning is added.
        
         | porridgeraisin wrote:
         | It is possible to understand the mechanism once you drop the
         | anthropomorphisms.
         | 
         | Each token output by an LLM involves one pass through the next-
         | word predictor neural network. Each pass is a fixed amount of
         | computation. Complexity theory hints to us that the problems
         | which are "hard" for an LLM will need more compute than the
         | ones which are "easy". Thus, the only mechanism through which
         | an LLM can compute more and solve its "hard" problems is by
         | outputting more tokens.
         | 
         | You incentivise it to this end by human-grading its outputs
         | ("RLHF") to prefer those where it spends time calculating
         | before "locking in" to the answer. For example, you would
         | prefer the output                 Ok let's begin... statement1
         | => statement2 ... Thus, the answer is 5
         | 
         | over                 The answer is 5. This is because....
         | 
         | since in the first one, it has spent more compute before giving
         | the answer. You don't in any way attempt to steer the extra
         | computation in any particular direction. Instead, you simply
         | reinforce preferred answers and hope that somewhere in that
         | extra computation lies some useful computation.
         | 
         | It turned out that such hope was well-placed. The DeepSeek
         | R1-Zero training experiment showed us that if you apply this
         | really generic form of learning (reinforcement learning)
         | without _any_ examples, the model automatically starts
         | outputting more and more tokens i.e "computing more".
         | DeepseekMath was also a model trained directly with RL.
         | Notably, the only signal given was whether the answer was right
         | or not. No attention was paid to anything else. We even ignore
         | the position of the answer in the sequence that we cared about
         | before. This meant that it was possible to automatically grade
         | the LLM without a human in the loop (since you're just checking
         | answer == expected_answer). This is also why math problems were
         | used.
         | 
         | All this is to say, we get the most insight on what benefit
         | "reasoning" adds by examining what happened when we applied it
         | without training the model on any examples. Deepseek R1
         | actually uses a few examples and then does the RL process on
         | top of that, so we won't look at that.
         | 
         | Reading the DeepseekMath paper[1], we see that the authors
         | posit the following:                 As shown in Figure 7, RL
         | enhances Maj@K's performance but not Pass@K. These
         | findings indicate that RL enhances the model's overall
         | performance by rendering       the output distribution more
         | robust, in other words, it seems that the       improvement is
         | attributed to boosting the correct response from TopK rather
         | than the enhancement of fundamental capabilities.
         | 
         | For context, Maj@K means that you mark the output of the LLM as
         | correct only if the majority of the many outputs you sample are
         | correct. Pass@K means that you mark it as correct even if just
         | one of them is correct.
         | 
         | So to answer your question, if you add an RL-based reasoning
         | process to the model, it will improve simply because it will do
         | more computation, of which a so-far-only-empirically-measured
         | portion helps get more accurate answers on math problems. But
         | outside that, it's purely subjective. If you ask me, I prefer
         | claude sonnet for all coding/swe tasks over any reasoning LLM.
         | 
         | [1] https://arxiv.org/pdf/2402.03300
        
       | zone411 wrote:
       | It significantly improves upon GPT-4o on my Extended NYT
       | Connections Benchmark. 22.4 -> 33.7
       | (https://github.com/lechmazur/nyt-connections).
        
       | anotherpaulg wrote:
       | GPT-4.5 Preview scored 45% on aider's polyglot coding benchmark
       | [0]. OpenAI describes it as "good at creative tasks" [1], so
       | perhaps it is not primarily intended for coding.
       | 65% Sonnet 3.7, 32k think tokens (SOTA)       60% Sonnet 3.7, no
       | thinking       48% DeepSeek V3       45% GPT 4.5 Preview <===
       | 27% ChatGPT-4o       23% GPT-4o
       | 
       | [0] https://aider.chat/docs/leaderboards/
       | 
       | [1] https://platform.openai.com/docs/models#gpt-4-5
        
         | doctoboggan wrote:
         | I was waiting for your comment and wow... that's bad.
         | 
         | I guess they are ceding the LLMs for coding market to
         | Anthropic? I remember seeing an industry report somewhere and
         | it claimed software development is the largest user of LLMs, so
         | it seems weird to give up in this area.
        
           | I_am_tiberius wrote:
           | I assume they go all in "the new google" direction. Embedded
           | ads coming soon I guess in the free version (chat.com).
        
       | icemelt8 wrote:
       | they are trying to copy Grok 3
        
       | smcleod wrote:
       | GPT 4.5 is insanely over price, it makes Anthropic look
       | affordable!
        
       | shshahshsusus wrote:
       | brief and detailed summaries by chatgpt (4o):
       | 
       |  _Brief Summary (40-50 words)_
       | 
       | OpenAI's GPT-4.5 is a research preview of their most advanced
       | language model yet, emphasizing improved pattern recognition,
       | creativity, and reduced hallucinations. It enhances unsupervised
       | learning, has better emotional intelligence, and excels in
       | writing, programming, and problem-solving. Available for ChatGPT
       | Pro users, it also integrates into APIs for developers.
       | 
       |  _Detailed Summary (200 words)_
       | 
       | OpenAI has introduced *GPT-4.5*, a research preview of its most
       | advanced language model, focusing on *scaling unsupervised
       | learning* to enhance pattern recognition, knowledge depth, and
       | reliability. It surpasses previous models in *natural
       | conversation, emotional intelligence (EQ), and nuanced
       | understanding of user intent*, making it particularly useful for
       | writing, programming, and creative tasks.
       | 
       | GPT-4.5 benefits from *scalable training techniques* that improve
       | its steerability and ability to comprehend complex prompts.
       | Compared to GPT-4o, it has a *higher factual accuracy and lower
       | hallucination rates*, making it more dependable across various
       | domains. While it does not employ reasoning-based pre-processing
       | like OpenAI o1, it complements such models by excelling in
       | general intelligence.
       | 
       | Safety improvements include *new supervision techniques*
       | alongside traditional reinforcement learning from human feedback
       | (RLHF). OpenAI has tested GPT-4.5 under its *Preparedness
       | Framework* to ensure alignment and risk mitigation.
       | 
       | *Availability*: GPT-4.5 is accessible to *ChatGPT Pro users*,
       | rolling out to other tiers soon. Developers can also use it in
       | *Chat Completions API, Assistants API, and Batch API*, with
       | *function calling and vision capabilities*. However, it remains
       | computationally expensive, and OpenAI is evaluating its long-term
       | API availability.
       | 
       | GPT-4.5 represents a *major step in AI model scaling*, offering
       | *greater creativity, contextual awareness, and collaboration
       | potential*.
        
       | ripped_britches wrote:
       | Obviously it's expensive and still I would prefer a reasoning
       | model for coding.
       | 
       | However for user facing applications like mine, this is an
       | awesome step in the right direction for EQ / tone / voice.
       | Obviously it will get distilled into cheaper open models very
       | soon, so I'm not too worried about the price or even tokens per
       | second.
        
       | mkaic wrote:
       | In a hilarious act of accidental satire, it seems that the AI-
       | generated audio version of the post has a weird
       | glitch/mispronunciation within the _first three words_ -- it
       | struggles to say  "GPT-4.5".
        
       | advael wrote:
       | It's sad that all I can think about this is that it's just
       | another creep forward of the surveillance oligarchy
       | 
       | I really used to get excited about ML in the wild and while there
       | are much bigger problems right now it still makes me sad to have
       | become so jaded about it
        
       | i_love_retros wrote:
       | Anyone really finding ai useful for coding?
       | 
       | I'm finding it to make things up, get things wrong, ignore things
       | I ask.
       | 
       | Def not worried about losing my job to it.
        
         | i_love_retros wrote:
         | It gets confused if I give it 3 files - how is it going to scan
         | a whole codebase and disparate systems and make correct
         | changes.
         | 
         | Pah! Don't believe the hype.
        
         | twistslider wrote:
         | I played around with Claude Code today, first time I've ever
         | really been impressed by AI for coding.
         | 
         | Tasked it with two different things, refactoring a huge
         | function of around ~400 lines and creating some unit tests
         | split into different files. The refactor was done flawlessly.
         | The unit tests almost, only missed some imports.
         | 
         | All I did was open it in the root of my project and prompt it
         | with the function names. It's a large monolithic solution with
         | a lot of subprojects. It found the functions I was talking
         | about without me having to clarify anything. Cost was about $2.
        
         | SkyPuncher wrote:
         | Yes, massively.
         | 
         | There's a learning curve to it, but it's worth literally every
         | penny I spend on API calls.
         | 
         | At worst, I'm no faster. At best, it's easily a 10x
         | improvement.
         | 
         | For me, one of the biggest benefits is talking about coding in
         | natural language. It lowers my mental low and keeps me in a
         | mental space where I'm more easily able to communicate with
         | stakeholders holders.
        
       | antirez wrote:
       | In many ways I'm not an OpenAI fan (but I need to recognize their
       | many merits). At the same time, I believe people are missing what
       | they tried to do with GPT 4.5: it was needed and important to
       | explore the pre-training scaling law in that direction. A gift to
       | science, however selfist it could be.
        
       | wewewedxfgdf wrote:
       | GPT-2 was laugh out loud funny, rolling on the ground funny.
       | 
       | I miss that - newer LLMs seem to have lost their sense of humor.
       | 
       | On the other hand GPT-2's funny stories often veered into
       | murdering everyone in the story and committing heinous crimes but
       | that was part of the weird experience.
        
         | kossTKR wrote:
         | Totally agree, i think the gargantuan hidden pre prompts,
         | censorship through reinforcement learning and whatever has
         | killed most creativity.
         | 
         | The newer models are incredible, but the tone is just soul
         | sucking even when it tries to be "looser" in the later
         | iterations.
        
         | krackers wrote:
         | Sydney is a glimpse at what an "unlobotomized" GPT-4 model
         | would have been like.
        
       | I_am_tiberius wrote:
       | Not available in my Pro plan.
        
       | highfrequency wrote:
       | Overall take seems to be negative in the comments. But I see
       | potential for a non-reasoning model that makes enough subtle
       | tweaks in its tone that it is enjoyable to talk to instead of
       | feeling like a summary of Wikipedia.
        
       | Chance-Device wrote:
       | And the AI stocks fell today.
       | 
       | I'm sure it's unrelated.
        
       | boznz wrote:
       | So better than 4o but not good enough for a 5.0
        
       | simonw wrote:
       | I got gpt-4.5-preview to summarize this discussion thread so far
       | (at 324 comments):                 hn-summary.sh 43197872 -m
       | gpt-4.5-preview
       | 
       | Using this script: https://til.simonwillison.net/llms/claude-
       | hacker-news-themes...
       | 
       | Here's the result:
       | https://gist.github.com/simonw/5e9f5e94ac8840f698c280293d399...
       | 
       | It took 25797 input tokens and 1225 input tokens, for a total
       | cost (calculated using https://tools.simonwillison.net/llm-prices
       | ) of $2.11! It took 154 seconds to generate.
        
         | djhworld wrote:
         | interesting summary but it's hard to gauge whether this is
         | better/worse than just piping the contents into a much cheaper
         | model.
        
       | orbital-decay wrote:
       | This looks like a first generation model to bootstrap future
       | models from, not a competitive product at all. The knowledge
       | cutoff is pretty old as well. (2023, seriously?)
       | 
       | If they wanted to train it to have some character like Anthropic
       | did with Claude 3... honestly I'm not seeing it, at least not in
       | this iteration. Claude 3 was/is much much more engaging.
        
       | GaggiX wrote:
       | I imagine it will be used as a base for GPT-5 when it will be
       | trained into a reasoning model, right now it probably doesn't
       | make too much sense to use.
        
       | dgfitz wrote:
       | @sama, LLMs aren't going to create AGI. I realize you need to
       | generate cash flow, this isn't the play.
       | 
       | Sincerely, Me
        
       ___________________________________________________________________
       (page generated 2025-02-27 23:00 UTC)