[HN Gopher] GPT-4.5
___________________________________________________________________
GPT-4.5
Author : meetpateltech
Score : 473 points
Date : 2025-02-27 20:01 UTC (2 hours ago)
(HTM) web link (openai.com)
(TXT) w3m dump (openai.com)
| throwup238 wrote:
| At this point I think the ultimate benchmark for any new LLM is
| whether or not it can come up with a coherent naming scheme for
| itself. Call it "self awareness."
| lenerdenator wrote:
| The people naming them really took the "just give the variable
| any old name, it doesn't matter" advice from Programming 101 to
| heart.
| sho_hn wrote:
| This is why my new LLM portfolio is Foo, Bar and Baz.
| smallmancontrov wrote:
| Still more coherent than the OpenAI lineup.
| nopelynopington wrote:
| 3,3.5,4,4o,4.5
|
| I had my money on 4oz
| camwhite wrote:
| I can't wait for fireship.io and the comment section here to tell
| me what to think about this
| mjburgess wrote:
| You appear to have the direction of causation reversed.
|
| (In that fireship does the same)
| lysace wrote:
| I wonder if fireship reaction video scripts to AI models
| based on HN comments can be automated using said AI models.
| rob wrote:
| I bet simonw will be adding it to `llm` and someone will be
| pasting his highlights here right after. Until then, my mind
| will remain a blank canvas.
| whisper_yb wrote:
| wowsers!
| bhouston wrote:
| A bit better at coding than ChatGPT 4o but not better than
| o3-mini - there is a chart near the bottom of the page that is
| easy to overlook:
|
| - ChatGPT 4.5 on AWS Bench verified: 38.0%
|
| - ChatGPT 4o on AWS Bench verified: 30.7%
|
| - OpenAI o3-mini on AWS Bench verified: 61.0%
|
| BTW Anthropic Claude 3.7 is better than o3-mini at coding at
| around 62-70% [1]. This means that I'll stick with Claude 3.7 for
| the time being for my open source alternative to Claude-code:
| https://github.com/drivecore/mycoder
|
| [1] https://aws.amazon.com/blogs/aws/anthropics-
| claude-3-7-sonne...
| logicchains wrote:
| >BTW Anthropic Claude 3.7 is better than o3-mini at coding at
| around 62-70% [1]. This means that I'll stick with Claude 3.7
| for the time being for my open source alternative to Claude-
| code
|
| That's not a fair comparison as o3-mini is significantly
| cheaper. It's fine if your employer is paying, but on a
| personal project the cost of using Claude through the API is
| really noticeable.
| cheema33 wrote:
| > That's not a fair comparison as o3-mini is significantly
| cheaper. It's fine if your employer is paying...
|
| I use it via Cursor editor's built-in support for Claude 3.7.
| That caps the monthly expense to $20. There probably is a
| limit in Claude for these queries. But I haven't run into it
| yet. And I am a heavy user.
| bhouston wrote:
| Agentic coders (e.g. aider, Claude-code, mycoder, codebuff,
| etc.) use a lot more tokens, but they write whole features
| for you and debug your code.
| QuadmasterXLII wrote:
| If open ai offers a more expensive model (4.5) and a cheaper
| model (3 mini) and both are worse, it starts to be a fair
| comparison
| ehsanu1 wrote:
| It's the other way around on their new SWE-Lancer benchmark,
| which is pretty interesting: GPT-4.5 scores 32.6%, while
| o3-mini scores 10.8%.
| Topfi wrote:
| To put that in context, Claude 3.5 Sonnet (new), a model we
| have had for months now and which from all accounts seems to
| have been cheaper to train and is cheaper to use, is still
| ahead of GPT-4.5 at 36.1% vs 32.6% in SWE-Lancer Diamond [0].
| The more I look into this release, the more confused I get.
|
| [0] https://arxiv.org/pdf/2502.12115
| _cs2017_ wrote:
| I don't see Claude 3.7 on the official leaderboard. The top
| performer on the leaderboard right now is o1 with a scaffold
| (W&B Programmer O1 crosscheck5) at 64.6%:
| https://www.swebench.com/#verified.
|
| If Claude 3.7 achieves 70.3%, it's quite impressive, it's not
| far from 71.7% claimed by o3, at (presumably) much, much lower
| costs.
| pawelduda wrote:
| Does the benchmark reflect your opinion on 3.7? I've been using
| 3.7 via Cursor and it's noticeably worse than 3.5. I've heard
| using the standalone model works fine, didn't get a chance to
| try it yet though.
| jasonjmcghee wrote:
| personal anecdote - claude code is the best llm devx i've
| had.
| bingdig wrote:
| Not seeing it available in the app or on ChatGPT.com with a pro
| subscription.
| hidelooktropic wrote:
| It's not supposed to be yet
| bingdig wrote:
| "Available to Pro users and developers worldwide" "Try in
| ChatGPT"
| hidelooktropic wrote:
| You're at the 200 dollar tier?
| kelsey98765431 wrote:
| Not available in model picker for my pro 200 dollar
| account.
|
| You can play with the api here: https://platform.openai.c
| om/playground/chat?models=gpt-4.5-p...
| apsec112 wrote:
| I have Pro, just updated the app, but don't currently
| have access
| _boffin_ wrote:
| Same.
| ashu1461 wrote:
| How is operator ? Have you been able to use it for
| anything useful ? Was thinking about taking a
| subscription
| SamvitJ wrote:
| One comparison I found interesting... I think GPT-4o has a more
| balanced answer!
|
| > What are your thoughts on space exploration?
|
| GPT-4.5:
|
| Space exploration isn't just valuable--it's essential.
|
| People often frame it as a luxury we pursue after solving Earth-
| bound problems. But space exploration actually helps us address
| those very challenges: climate change (via satellite monitoring),
| resource scarcity (through asteroid mining), and existential risk
| (by enabling eventual multi-planetary living). It also drives
| technological innovation, economic growth, and inspires humanity
| toward collective achievement.
|
| We don't explore space to escape Earth. We explore space to
| better understand, protect, and sustain life on Earth.
|
| > What are your thoughts on space exploration?
|
| GPT-4o:
|
| Space exploration is one of the most ambitious and profound
| endeavors humanity has undertaken. It expands our understanding
| of the universe, drives technological innovation, and inspires
| generations to think beyond the boundaries of Earth.
|
| From a scientific perspective, space exploration helps us address
| fundamental questions: Are we alone in the universe? How did the
| cosmos begin? Can we colonize other planets? Missions like the
| James Webb Space Telescope, Mars rovers, and deep-space probes
| continuously reshape our understanding of physics, planetary
| science, and even biology.
|
| From a practical standpoint, space exploration fuels advancements
| in materials science, AI, robotics, and medicine. Many everyday
| technologies--GPS, satellite communications, weather forecasting,
| and even medical imaging--stem from space research.
|
| There's also the existential argument: humanity's long-term
| survival may depend on our ability to become a multi-planetary
| species. While colonizing Mars or mining asteroids is far from
| easy, these efforts could help us safeguard civilization against
| existential threats like climate change or asteroid impacts.
|
| That said, space exploration also raises ethical and economic
| questions. Should we focus more on Earth's problems before
| investing trillions in space? How do we prevent the same
| exploitative behaviors that shaped colonial history from
| repeating on other planets?
|
| Overall, I see space exploration as a necessary pursuit--not at
| the cost of solving problems on Earth, but as a way to advance
| our knowledge, drive innovation, and secure a future beyond our
| home planet. What's your take?
| gh0stcat wrote:
| Yeah, I also found it odd that they seem to be implying that an
| incredibly biased answer (as in 4.5) is better. In general, I
| find the tone more polarizing and not exactly warm as they
| advertised in the release video.
| basisword wrote:
| As a benchmark, why do you find the 'opinion' of an LLM useful?
| The question is completely subjective. Edit: Genuinely asking.
| I'm assuming there's a reason this is an important measure.
| Topfi wrote:
| Not OP, but likely because that was the only
| metric/benchmark/however you want to call it OpenAI showcased
| in the stream and on the blog to highlight the improvement
| between 4o and 4.5. To say that this is not really a good
| metric for comparison, not least because prompting can have a
| massive impact in this regard, would be an understatement.
| Chamix wrote:
| Indeed, and the difference could in essence be achieved
| yourself with a different system prompt on 4o. What exactly is
| 4.5 contributing here in terms of a more nuanced intelligence?
|
| The new RLHF direction (heavily amplified through scaling
| synthetic training tokens) seems to clobber any minor gains the
| improved base internet prediction gains might've added.
| ThouYS wrote:
| hmm.. not really the direction I expected them to go
| doctoboggan wrote:
| I am beginning to think these human eval tests are a waste of
| time at best, and negative value at worst. Maybe I am being
| snobby, but I don't think the average human is able to properly
| evaluate usefulness, truthfulness, or other metrics that I
| actually care about. I am sure this is good for openAI since if
| more people like what the hear, they are more likely come back.
|
| I don't want my AI more obsequious, I want it more correct and
| capable.
|
| My only use case is coding though, so maybe I am not
| representative of their usual customers?
| dboreham wrote:
| The SuperTuring era.
| onlyrealcuzzo wrote:
| > I want it more correct and capable.
|
| How is it supposed to be more correct and capable if these
| human eval tests are a waste of time?
|
| Once you ask it to do more than add two numbers together, it
| gets a lot more difficult and subjective to determine whether
| it's correct and how correct.
| doctoboggan wrote:
| I agree it's a hard problem. I think there are a number of
| tests out there however that are able to objectively test
| capability and truthfulness.
|
| I've read reports that some of the changes that are preferred
| by human evaluators actually hurt the performance on the more
| objective tests.
| onlyrealcuzzo wrote:
| Please tell me how objectively we determine how correct
| something is when you ask an LLM: "Was Russia the aggressor
| in the current Ukraine / Russia conflict?"
|
| One LLM says: "Yes."
|
| The other says: "Well, it's hard to say because what even
| is war? And there's been conflict forever, and you have to
| understand that many people in Russia think there is no
| such thing as Ukraine and it's always actually just been
| Russia. How can there be an aggressor if it's not even a
| war, just a special operation in a civil conflict? And,
| anyway, Russia is such a good country. Why would it be the
| aggressor? Vladimir Putin is the president of Russia, and
| he's known to be a genius who rarely makes mistakes. All
| that being said, there is no consensus among everyone who
| was the aggressor or what started the conflict. But more
| people say Russia started it."
| bloomingkales wrote:
| These eval tests are just an anchor point to measure distance
| from, but it's true, picking the anchor point is important. We
| don't want to measure in the wrong direction.
| ekojs wrote:
| > Because of this, we're evaluating whether to continue serving
| it in the API long-term as we balance supporting current
| capabilities with building future models.
|
| Seems like it's not going to be deployed for long.
|
| $75.00 / 1M tokens for input
|
| $150.00 / 1M tokens for output
|
| That's crazy prices.
| bguberfain wrote:
| Until GPT-4.5, GPT-4 32K was certainly the most heavy model
| available at OpenAI. I can imagine the dilemma between to keep
| it running or stop it to free GPU for training new models. This
| time, OpenAI was clear whether to continue serving it in the
| API long-term.
| jsheard wrote:
| > or stop it to free GPU for training new models.
|
| Don't they use different hardware for inference and training?
| AIUI the former is usually done on cheaper GDDR cards and the
| latter is done on expensive HBM cards.
| Chamix wrote:
| It's interesting to compare the cost of that original gpt-4
| 32k(0314) vs gpt-4.5:
|
| $60/M input tokens vs $75/M input tokens
|
| $120/M output tokens vs $150/M output tokens
| daemonologist wrote:
| Imagine if they built a reasoning model with costs like these.
| Sometimes it seems like they're on a trajectory to create a
| model which is strictly more capable than I am but which costs
| 100x my salary to run.
| nickreese wrote:
| Is it just me or is having the AI help you self sensor (as shown
| in the demo live stream:
| https://www.youtube.com/watch?v=cfRYp0nItZ8)... pretty dystopian?
| nopelynopington wrote:
| Oh this makes sense. chatGPT results have taken a nose dive in
| quality lately.
|
| It couldn't write a simple rename function for me yesterday,
| still buggy after seven attempts.
|
| I'm more and more convinced that they dumb down the core product
| when they plan to release a new version to make the difference
| seem bigger.
| wilg wrote:
| 99% chance that's confirmation bias
| nomel wrote:
| Sam tweeted that they're running out of computer. I think
| it's reasonable to think they may serve somewhat quantized
| models when out of capacity. It would be a rational business
| decision that would minimally disrupt lower tier ChatGPT
| users.
|
| Anecdotally, I've noticed what appears to be drops in
| quality, some days. When the quality drops, it responds in
| odd ways when asked what model it is.
| logicallee wrote:
| >It couldn't write a simple rename function for me yesterday,
| still buggy after seven attempts.
|
| I'm surprised and a bit nervous about that. We intend to
| bootstrap a large project with it!!
|
| Both ChatGPT 4o (fast) and ChatGPT o1 (a bit slower, deeper
| thinking) should easily be able to do this without fail.
|
| Where did it go wrong? Could you please link to your chat?
|
| About my project: I run the sovereign State of Utopia (will be
| at stateofutopia.com and stofut.com for short) which is a
| country based on the idea of state-owned, autonomous AI's that
| do all the work and give out free money, goods, and services to
| all citizens/beneficiaries. We've built a chess app (i.e. a
| free source of entertainment) as a proof of concept though the
| founder had to be in the loop to fix some bugs:
|
| https://taonexus.com/chess.html
|
| and a version that shows obvious blunders, by showing which
| squares are under attack:
|
| https://taonexus.com/blunderfreechess.html
|
| One of the largest and most complicated applications anyone can
| run is a web browser. We don't have a web browser built, but we
| do have a buggy minimal version of it that can load and
| minimally display some web pages, and post successsfully:
|
| https://taonexus.com/publicfiles/feb2025/84toy-toy-browser-w...
|
| It's about 1700 lines of code and at this point runs into the
| limitations of all the major engines. But it does run, can load
| some web pages and can post successfully.
|
| I'm shocked and surprised ChatGPT failed to get a rename
| function to work, in 7 attempts.
| skepticATX wrote:
| How many still believe that scaling up base models will lead to
| AGI?
| Mizza wrote:
| Sounds like it's a distill of O1? After R1, I don't care that
| much about non-reasoning models anymore. They don't even seem
| excited about it on the livestream.
|
| I want tiny, fast and cheap non-reasoning models I can use in
| APIs and I want ultra smart reasoning models that I can query a
| few times a day as an end user (I don't mind if it takes a few
| minutes while I refill a coffee).
|
| Oh, and I want that advanced voice mode that's good enough at
| transcription to serve as a babelfish!
|
| After that, I guess it's pretty much all solved until the robots
| start appearing in public.
| sebastiennight wrote:
| Probably not a distill of o1, since o1 is a reasoning model and
| GPT4.5 is not. Also, OpenAI has been claiming that this is a
| very large model (and it's 2.5x more expensive than even OG
| GPT-4) so we can assume it's the biggest model they've trained
| so far.
|
| They'll probably distill this one into GPT-4.5-mini or such,
| and have something faster and cheaper available soon.
| Mizza wrote:
| There are plenty of distills of reasoning models now, and
| they said in they livestream they used training data from
| "smaller models" - which is probably every model ever
| considering how expensive this one is.
| sebastiennight wrote:
| Knowledge distillation is literally _by definition_
| teaching a smaller model from a big one, not the opposite.
|
| Generating outputs from existing (therefore smaller) models
| to train the largest model of all time would simply be
| called "using synthetic data". These are not the same thing
| at all.
|
| Also, if you were to distill a reasoning model, the goal
| would be to get a (smaller) reasoning model because you're
| teaching your new model to mimic outputs that show a
| reasoning/thinking trace. E.G. that's what all of those
| "local" Deepseek models are: small LLama models distilled
| from the big R1 ; a process which "taught" Llama-7B to show
| reasoning steps before coming up with a final answer.
| eightysixfour wrote:
| It isn't even vaguely a distill of o1. The reasoning models
| are, from what we can tell, relatively small. This model is
| _massive_ and they probably scaled the parameter count to
| improve factual knowledge retention.
|
| They also mentioned developing some new techniques for training
| small models and then incorporating those into the larger model
| (probably to help scale across datacenters), so I wonder if
| they are doing a bit of what people _think_ MoE is, but isn 't.
| Pre-train a smaller model, focus it on specific domains, then
| use that to provide synthetic data for training the larger
| model on that domain.
| Mizza wrote:
| You can 'distill' with data from a smaller, better model into
| a larger, shittier one. It doesn't matter. This is what they
| said they did on the livestream.
| eightysixfour wrote:
| I have distilled models before, I know how it works. They
| may have used o1 or o3 to create some of the synthetic data
| for this one, but they clearly did not try and create any
| self-reflective reasoning in this model whatsoever.
| valine wrote:
| My impression is that it's a massive increase in the parameter
| count. This is likely the spiritual successor to GPT4 and would
| have been called GPT5 if not for the lackluster performance.
| The speculation is that there simply isn't enough data on the
| internet to support yet another 10x jump in parameters.
|
| O1-mini is a distill of O1. This definitely isn't the same
| thing.
| rvz wrote:
| You should really be paying attention to what DeepSeek AI open
| sources next.
|
| This announcement by OpenAI was already expected: [0]
|
| [0] https://x.com/sama/status/1889755723078443244
| brokensegue wrote:
| API price is crazy high. This model must be huge. Not sure this
| is practical
| jdprgm wrote:
| Wow you aren't kidding, 30x input price and 15x output price vs
| 4o is insane. The pricing on all AI API stuff changes so
| rapidly and is often so extreme between models it is all hard
| to keep track of and try to make value decisions. I would
| consider a 2x or 3x price increase quite significant, 30x is
| wild. I wonder how that even translates... there is no way the
| model size is 30 times larger right?
| jdprgm wrote:
| Per Altman on X: "we will add tens of thousands of GPUs next week
| and roll it out to the plus tier then". Meanwhile a month after
| launch rtx 5000 series is completely unavailable and hardly any
| restocks and the "launch" consisted of microcenters getting
| literally tens of cards. Nvidia really has basically abandoned
| consumers.
| apsec112 wrote:
| AI GPUs are bottlenecked mostly by high-bandwidth memory (HBM)
| chips and CoWoS (packaging tech used to integrate HBM with the
| GPU die), which are in short supply and aren't found in
| consumer cards at all
| jiggawatts wrote:
| You would think that by now they would have done something to
| ramp production capacity...
| bhouston wrote:
| Altman's claim and NVIDIA's consumer launch supply problems may
| be related - OpenAI may be eating up the GPU supply...
| _zoltan_ wrote:
| OpenAI is not purchasing consumer 5090s... :)
| BizarroLand wrote:
| No, but the supply constraints are part of what is driving
| the insane prices. Every chip they use for consumer grade
| instead of commercial grade is a potential loss of
| potential income.
| bangaladore wrote:
| Although you are correct, Nvidia is limited on total
| output. They can't produce 50XXs fast enough, and it's
| naive to think that isn't at least partially due to the
| wild amount of AI GPUs they are producing.
| sunaookami wrote:
| This seems very rushed because of DeepSeek's R1 and Anthropic's
| Claude 3.7 Sonnet. Pretty underwhelming, they didn't even show
| programming? In the livestream, they struggled to come up with
| reasons why I should prefer GPT-4.5 over GPT-4o or o1.
| apsec112 wrote:
| At least according to WSJ, they had planned to release it
| earlier but struggled to get the model quality up, especially
| relative to cost
| bhouston wrote:
| they do have coding benchmarks, I summarized them here:
| https://news.ycombinator.com/item?id=43197955
| bitshiftfaced wrote:
| This strikes me as the opposite of rushed. I get the impression
| that they've been sitting on this for a while and couldn't make
| it look as good as previous improvements. At some point they
| had to say, "welp here it is, now we can check that box and
| move on."
| Topfi wrote:
| Considering both this blog post and the livestream demos, I am
| underwhelmed. Having just finished the stream, I had a real "was
| that all" moment, which on one hand shows how spoiled I've gotten
| by new models impressing me, but on another feels like OpenAI
| really struggles to stay ahead of their competitors.
|
| What has been shown feels like it could be achieved using a
| custom system prompt on older versions of OpenAIs models, and I
| struggle to see anything here that truly required ground-up
| training on such a massive scale. Hearing that they were forced
| to spread their training across multiple data centers
| simultaneously, coupled with their recent release of SWE-Lancer
| [0] which showed Anthropic (Claude 3.5 Sonnet (new) to be exact)
| handily beating them, I was really expecting something more than
| "slightly more casual/shorter output", which again, I fail to see
| how that wasn't possible by prompting GPT-4o.
|
| Looking at pricing [1], I am frankly astonished.
|
| > Input: $75.00 / 1M tokens > Cached input: $37.50 / 1M tokens >
| Output: $150.00 / 1M tokens
|
| How could they justify that asking price? And, if they have some
| amazing capabilities that make a 30-fold pricing increase
| justifiable, why not show it? Like, OpenAI are many things, but I
| always felt they understood price vs performance incredibly well,
| from the start with gpt-3.5-turbo up to now with o3-mini, so this
| really baffles me. If GPT-4.5 can justify such immense cost in
| certain tasks, why hide that and if not, why release this at all?
|
| [0] https://github.com/openai/SWELancer-Benchmark
|
| [1] https://openai.com/api/pricing/
| Bjorkbat wrote:
| My first thought seeing this and looking at benchmarks was that
| if it wasn't for reasoning, then either pundits would be saying
| we've hit a plateau, or at the very least OpenAI is clearly in
| 2nd place to Anthropic in model performance.
|
| Of course we don't live in such a world, but I thought of this
| nonetheless because for all the connotations that come with a
| 4.5 moniker this is kind of underwhelming.
| uh_uh wrote:
| Pundits were saying that deep learning has hit a plateau even
| before the LLM boom.
| mvdtnz wrote:
| > How could they justify that asking price?
|
| They're still selling $1 for <$1. Like personal food delivery
| before it, consumers will eventually need to wake up to this
| fact - these things will get expensive, fast.
| spiderfarmer wrote:
| Let a thousand providers bloom.
| Ekaros wrote:
| I generally question how wide spread willingness to pay for
| the most expensive product is. And will most users of those
| who actually want AI go with ad ridden lesser models...
| vel0city wrote:
| I can just imagine Kraft having a subsidized AI model for
| recipe suggestions that adds Velveeta to everything.
| tmaly wrote:
| rethinking your comment "was that all" I am listening to the
| stream now and had a thought. Most of the new models that have
| come out in the past few weeks have been great at coding and
| logical reasoning. But 4o has been better at creative writing.
| I am wondering if 4.5 is going to be even better at creative
| writing than 4o.
| lasermike026 wrote:
| I would rather pay for 4.5 by the query.
| zaptrem wrote:
| GPT 4.5 pricing is insane: Price Input: $75.00 / 1M tokens Cached
| input: $37.50 / 1M tokens Output: $150.00 / 1M tokens
|
| GPT 4o pricing for comparison: Price Input: $2.50 / 1M tokens
| Cached input: $1.25 / 1M tokens Output: $10.00 / 1M tokens
|
| It sounds like it's so expensive and the difference in usefulness
| is so lacking(?) they're not even gonna keep serving it in the
| API for long:
|
| > GPT-4.5 is a very large and compute-intensive model, making it
| more expensive than and not a replacement for GPT-4o. Because of
| this, we're evaluating whether to continue serving it in the API
| long-term as we balance supporting current capabilities with
| building future models. We look forward to learning more about
| its strengths, capabilities, and potential applications in real-
| world settings. If GPT-4.5 delivers unique value for your use
| case, your feedback (opens in a new window) will play an
| important role in guiding our decision.
|
| I'm still gonna give it a go, though.
| MattSayar wrote:
| Input price difference: 4.5 is 30x more
|
| Output price difference:4.5 is 15x more
|
| In their model evaluation scores in the appendix, 4.5 is, on
| average, 26% better. I don't understand the value here.
| alwa wrote:
| If you ran the same query set 30x or 15x on the cheaper model
| (and compensated for all the extra tokens the reasoning model
| uses), would you be able to realize the same 26% quality gain
| in a machine-adjudicatible kind of way?
| j_maffe wrote:
| with a reasoning model you'd get better than both.
| mirekrusin wrote:
| Einstein's IQ = 3.5x chimpanzees IQs, right?
| redox99 wrote:
| 3.5x on a normal distribution with mean 100 and SD 15 is
| pretty insane. But I agree with your point, being 26%
| better at a certain benchmark could be a tiny difference,
| or an incredible improvement (imagine the hardest questions
| being Riemann hypothesis, P != NP, etc).
| minimaxir wrote:
| Sam Altman's explanation for the restriction is a bit fluffier:
| https://x.com/sama/status/1895203654103351462
|
| > bad news: it is a giant, expensive model. we really wanted to
| launch it to plus and pro at the same time, but we've been
| growing a lot and are out of GPUs. we will add tens of
| thousands of GPUs next week and roll it out to the plus tier
| then. (hundreds of thousands coming soon, and i'm pretty sure
| y'all will use every one we can rack up.)
| g-mork wrote:
| release blog post author: this is definitely a research
| preview
|
| ceo: it's ready
|
| the pricing is probably a mixture of dealing with GPU
| scarcity and intentionally discouraging actual users. I can't
| imagine the pressure they must be under to show they are
| releasing and staying ahead, but Altman's tweet makes it
| clear they aren't really ready to sell this to the general
| public yet.
| pk-protect-ai wrote:
| Yeap, that the thing, they are not ahead anymore. Not since
| last summer at least. Yes they have probably largest
| customer base, but their models are not the best for a
| while already.
| danenania wrote:
| Eh, I think o1-pro is by far the most capable model
| available right now in terms of pure problem solving.
| rvnx wrote:
| You can try Claude 3.7-Thinking and Grok 3 Think. 10
| times cheaper, as good, or very similar to o1-pro.
| danenania wrote:
| I haven't tried Grok yet so can't speak to that, but I
| find o1-pro is much stronger than 3.7-thinking for e.g.
| distributed systems and concurrency problems.
| chefandy wrote:
| I'm not an expert or anything, but from my vantage point,
| each passing release makes Altman's confidence look more
| aspirational than visionary, which is a really bad place to
| be with that kind of money tied up. My financial manager is
| pretty bullish on tech so I hope he is paying close attention
| to the way this market space is evolving. He's good at his
| job, a nice guy, and surely wears much more expensive
| underwear than I do-- I'd hate to see him lose a pair
| powering on his Bloomberg terminal in the morning one of
| these days.
| igor47 wrote:
| You're the one buying him the underwear. Don't index funds
| outperform managed investing? I think especially after
| accounting for fees, but possibly even after accounting
| that 50% of money managers are below average.
| chefandy wrote:
| He earns his undies. My returns are almost always
| modestly above index fund returns after his fees, though
| like last quarter, he's very upfront when they're not. He
| has good advice for pulling back when things are
| uncertain. I'm happy to delegate that to him.
| marcus0x62 wrote:
| A friend got taken in by a Ponzi scheme operator several
| years ago. The guy running it was known for taking his
| clients out to lavish dinners and events all the time.[0]
|
| After the scam came to light my friend said "if I knew I
| was paying for those dinners, I would have been fine with
| Denny's[1]"
|
| I wanted to tell him "you would have been paying for
| those dinners even if he wasn't outright stealing your
| money," but that seemed insensitive so I kept my mouth
| shut.
|
| 0 - a local steakhouse had a portrait of this guy drawn
| on the wall
|
| 1 - for any non-Americans, Denny's is a low cost diner-
| style restaurant.
| ProfessorLayton wrote:
| Not all investing is throwing cash at an index, though.
| There's other types of investing like direct indexing (to
| harvest losses), muni bonds, etc.
|
| Paying someone to match your risk profile and financial
| goals may be worth the fee, which as you pointed out is
| very measurable. YMMV though.
| fragmede wrote:
| Depends who's pitch deck you're reading. Warren Buffett
| didn't get rich waiting on index funds.
| Terr_ wrote:
| > each passing release makes Altman's confidence look more
| aspirational than visionary
|
| As an LLM cynic, I feel that point passed _long_ go,
| perhaps even before Altman claimed countries would start
| wars to conquer territory for its datacenters, or promoting
| the dream of a 7 T-for-trillion dollar investment.
|
| Alas, the market can remain irrational longer than I can
| remain solvent.
| chefandy wrote:
| That $7 trillion dollar ask pushed me from skeptical to
| full-on eye-roll emoji land-- the dude is clearly a
| narcissist with delusions of grandeur-- but it's getting
| _worse._ Considering the $200 pro subscription was
| significantly unprofitable before this model came out,
| imagine how _astonishingly expensive_ this model must be
| to run at many times that price.
| rebolek wrote:
| Bad news: Sam Altman runs the show.
| rvnx wrote:
| He is like the Elon Musk of OpenAI.
| sebastiennight wrote:
| I think it's fairer to compare it to the original GPT-4 which
| might the equivalent in term of "size" (though we don't have
| actual numbers for either).
|
| GPT-4: Input $30.00 / 1M tokens ; Output $60.00 / 1M tokens
|
| So 4.5 is 2.5x more expensive.
|
| I think they announced this as their last non-reasoning model,
| so it was maybe with the goal of stretching pre-training as far
| as they could, just to see what new capabilities would show up.
| We'll find out as the community gives it a whirl.
|
| I'm a Tier 5 org and I have it available already in the API.
| minimaxir wrote:
| The marginal costs for running a GPT-4-class LLM are much
| lower nowadays due to significant software and hardware
| innovations since then, so costs/pricing are harder to
| compare.
| sebastiennight wrote:
| Agreed, however it might make sense that a much-larger-
| than-GPT-4 LLM would also, at launch, be more expensive to
| run than the OG GPT-4 was at launch.
|
| (And I think this is probably also scarecrow pricing to
| discourage casual users from clogging the API since they
| seem to be too compute-constrained to deliver this at
| scale)
| jstummbillig wrote:
| Why would that be fairer? We can assume they did incorporate
| all learnings and optimizations they made post gpt-4 launch,
| no?
| spoaceman7777 wrote:
| There are some numbers on one of their Blackwell or Hopper
| info pages that notes the ability of their hardware in
| hosting an unnamed GPT model that is 1.8T params. My
| assumption was that it referred to GPT-4
|
| Sounds to me like GPT 4.5 likely requires a full Blackwell
| HGX cabinet or something, thus OpenAI's reference to needing
| to scale out their compute more (Supermicro only opened up
| their Blackwell racks for General Availability last month,
| and they're the prime vendor for water-cooled Blackwell
| cabinets right now, and have the ability to throw up a GPU
| mega-cluster in a few weeks, like they did for xAI/Grok)
| harlanlewis wrote:
| The price really is eye watering. At a glance, my first
| impression is this is something like Llama 3.1 405B, where the
| primary value may be realized in generating high quality
| synthetic data for training rather than direct use.
|
| I keep a little google spreadsheet with some charts to help
| visualize the landscape at a glance in terms of
| capability/price/throughput, bringing in the various index
| scores as they become available. Hope folks find it useful,
| feel free to copy and claim as your own.
|
| https://docs.google.com/spreadsheets/d/1foc98Jtbi0-GUsNySddv...
| bennyg wrote:
| This is an amazing spreadsheet - thank you for sharing!
| isoprophlex wrote:
| Thats... incredibly thorough. Wow. Thanks for sharing this.
| Philpax wrote:
| Holy shit, that's incredible. You should publicise this more!
| That's a fantastic resource.
| beklein wrote:
| They tried a while ago:
| https://news.ycombinator.com/item?id=40373284
|
| Sadly little people noticed...
| throwup238 wrote:
| Sadly _few_ people noticed.
|
| I don't normally cosplay as a grammar Nazi but in this
| case I feel like someone should stand up for the little
| people :)
| rebolek wrote:
| So you think that little people didn't notice? ;)
| dumpsterdiver wrote:
| A comma in the original comment would have made it pop
| even more:
|
| "Sadly, little people noticed."
|
| (queue a group of little people holding pitch forks
| (normal forks upon closer inspection))
| adinb wrote:
| I cannot overstate how good your shared spreadsheet is.
| Thanks again!
| jnd0 wrote:
| Thank you so much for sharing this!
| bglusman wrote:
| very impressive... also interested in your trip planner, it
| looks like invite only at the moment, but... would it be rude
| to ask for an invite?
| gwyllimj wrote:
| That is an amazing resource. Thanks for sharing!
| dumpsterdiver wrote:
| Nice, thank you for that (upvoted in appreciation). Regarding
| the absence of o1-Pro from the analysis, is that just because
| there isn't enough public information available?
| krwiseman wrote:
| Awesome spreadsheet. Would a 3D graph of fast, cheap & smart
| be possible?
| rendist wrote:
| Amazing, thank you so much for sharing this.
| kridsdale3 wrote:
| Hope you don't mind, I dumped your sheet in to GPT 4.5 and
| asked for it's take:
|
| 1. Capability vs. Throughput Correlation
| Insight: High composite-capability models (e.g., Claude 3.7
| Sonnet Thinking, GPT-4o) typically show moderate or lower
| throughput (<70 tokens/sec), while mid-tier models (DeepSeek,
| Mistral variants, Nova series) often deliver much higher
| throughput (>100 tokens/sec). Interpretation:
| Frontier models trade off throughput for sophisticated
| reasoning, likely due to heavier computational demands.
|
| 2. Cost Efficiency vs. Capability Plateau
| Insight: Models at the top capability (90-100) see a steep
| rise in cost per capability point, with diminishing cost-
| effectiveness above roughly 75 capability points.
| Example: Compare GPT-4o (100 capability, ~$0.085 per
| capability point) with DeepSeek V3 (74 capability, ~$0.002
| per capability point). Interpretation: Mid-
| capability models appear optimal for cost-sensitive
| applications.
|
| 3. Capability Consistency vs. Composite Capability
| Insight: Higher consistency scores (>85%) are generally
| linked to mid-to-high composite capability (60-80 range), yet
| the very top models (85+) can have moderate or lower
| consistency (e.g., GPT-4o at 81% vs. Claude 3.7 Sonnet at
| 57%). Interpretation: Even cutting-edge models
| may show variability across benchmarks due to task-specific
| optimizations.
|
| 4. Latency and Throughput Relationship
| Insight: Initial latency (first chunk response) has only a
| weak correlation with overall throughput. Some models with
| high sustained throughput exhibit high initial latency, while
| others with lower throughput can respond quickly.
| Interpretation: Optimizing for a rapid first response is a
| distinct engineering challenge from maximizing sustained
| throughput.
|
| 5. Open vs. Proprietary Models: Capability and Cost
| Efficiency Insight: Open models (e.g.,
| DeepSeek, Mistral, Llama variants) tend to cluster in a
| region of higher cost efficiency relative to capability
| compared to proprietary models (e.g., those from Anthropic,
| OpenAI, Google). Interpretation: Open-source
| offerings are closing the capability gap and forcing a
| rebalancing of cost structures.
|
| 6. Benchmark Correlations (Math vs. General Reasoning)
| Insight: Strong performance on general reasoning benchmarks
| (like BIG-BENCH-HARD, ARC Challenge) often aligns with high
| scores in mathematical reasoning (e.g., GSM8K, MATH).
| Interpretation: Math proficiency appears to be a good
| predictor of overall reasoning strength.
|
| 7. Throughput vs. Year of Release Insight:
| There is a noticeable trend where more recent models (2025)
| tend to offer higher throughput (>100 tokens/sec), even among
| mid-range capabilities. Interpretation: The
| market seems to be shifting toward models that emphasize
| operational efficiency alongside capability.
|
| 8. Long-context Performance vs. General Capability
| Insight: Models designed for long-context handling (as
| measured by benchmarks like ZeroSCROLLS, InfiniteBench)
| generally also exhibit high overall capability.
| Interpretation: The ability to process extended contexts
| remains a premium feature closely tied to overall model
| sophistication.
| swatcoder wrote:
| > We look forward to learning more about its strengths,
| capabilities, and potential applications in real-world
| settings. If GPT-4.5 delivers unique value for your use case,
| your feedback (opens in a new window) will play an important
| role in guiding our decision.
|
| "We don't really know what this is good for, but spent a lot of
| money and time making it and are under intense pressure to
| announce new things right now. If you can figure something out,
| we need you to help us."
|
| Not a confident place for an org trying to sustain a $XXXB
| valuation.
| tempaccount420 wrote:
| > "We don't really know what this is good for, but spent a
| lot of money and time making it and are under intense
| pressure to announce new things right now. If you can figure
| something out, we need you to help us."
|
| Where is this quote from?
| hotpocket777 wrote:
| It's not a quote. It is an interpretation or reading of a
| quote.
| cogman10 wrote:
| Perhaps even fed through an LLM ;)
| scythe wrote:
| I believe it's a "translation" in the sense of
| Wittgenstein's goal of philosophy:
|
| >My aim is: to teach you to pass from a piece of disguised
| nonsense to something that is patent nonsense.
| Nition wrote:
| Another great example on Hacker News is this old
| translation of Google's "Amazing Bet":
| https://news.ycombinator.com/item?id=12793033
| dd3boh wrote:
| I think it's supposed to be a translation of what OpenAI's
| quote means in real world terms.
| thih9 wrote:
| The quotation marks in the grandparent comment are scare
| (sneer) quotes and not actual quotation.
|
| https://en.m.wikipedia.org/wiki/Scare_quotes
|
| > Whether quotation marks are considered scare quotes
| depends on context because scare quotes are not visually
| different from actual quotations.
| riwsky wrote:
| Said the quiet part out loud! Or as we say these days,
| "transparently exposed the chain of thought tokens".
| porridgeraisin wrote:
| Lol, nice one
| Terr_ wrote:
| "I knew the dame was trouble the moment she walked into my
| office."
|
| "Uh... excuse me, Detective Nick Danger? I'd like to retain
| your services."
|
| "I waited for her to get the the point."
|
| "Detective, who are you talking to?"
|
| "I didn't want to deal with a client that was hearing
| voices, but money was tight and the rent was due. I
| pondered my next move."
|
| "Mr. Danger, are you... narrating out loud?"
|
| "Damn! My internal chain of thought, the key to my success
| --or at least, past successes--was leaking again. I
| rummaged for the familiar bottle of scotch in the drawer,
| kept for just such an occasion."
|
| ---
|
| But seriously: These "AI" products basically run on movie-
| scripts already, where the LLM is used to append more
| "fitting" content, and glue-code is periodically performing
| any lines or actions that arise in connection to the
| Helpful Bot character. Real humans are tricked into
| thinking the finger-puppet is a discrete entity.
|
| These new "reasoning" models are just switching the style
| of the movie script to _film noir_ , where the Helpful Bot
| character is making a layer of unvoiced commentary. While
| it may make the story more cohesive, it isn't a qualitative
| change in the kind of illusory "thinking" going on.
| kridsdale3 wrote:
| I don't know if it was you or someone else who made
| pretty much the same point a few days ago. But I still
| like it. It makes the whole thing a lot more fun.
| Terr_ wrote:
| https://news.ycombinator.com/context?id=43118925
|
| I've been banging that particular drum for a while on HN,
| and the mental-model still feels so intuitively strong to
| me that I'm starting to have doubts: "It's _too_ good, I
| must be wrong in some subtle and devastating way. "
| EA-3167 wrote:
| Maybe if they build a few more data centers, they'll be able
| to construct their machine god. Just a few more dedicated
| power plants, a lake or two, a few hundred billion more and
| they'll crack this thing wide open.
|
| And maybe Tesla is going to deliver truly full self driving
| tech any day now.
|
| And Star Citizen will prove to have been worth it along
| along, and Bitcoin will rain from the heavens.
|
| It's very difficult to remain charitable when people seem to
| always be chasing the new iteration of the same old thing,
| and we're expected to come along for the ride.
| alyandon wrote:
| And Star Citizen will prove to have been worth it along
| along
|
| Sounds like someone isn't happy with the 4.0 eternally
| incrementing "alpha" version release. :-D
|
| I keep checking in on SC every 6 months or so and still see
| the same old bugs. What a waste of potential. Fortunately,
| Elite Dangerous is enough of a space game to scratch my
| space game itch.
| 0x457 wrote:
| To be fAir, SC is trying to do things that no one else
| done in a context of a single game. I applaud their
| dedication, but I won't be buying JPGs of a ship for 2k.
| alyandon wrote:
| Yeah, they never should have expected to take an FPS game
| engine like CryEngine and expected to be able to modify
| it to work as the basis for a large scale space MMO game.
|
| Their backend is probably an async nightmare of
| replicated state that gets corrupted over time. Would
| explain why a lot of things seem to work more or less bug
| free after an update and then things fall to pieces and
| the same old bugs start showing up after a few weeks.
|
| And to be clear, I've spent money on SC and I've played
| enough hours goofing off with friends to have got my
| money's worth out of it. I'm just really bummed out about
| the whole thing.
| 0x457 wrote:
| Gonna go meta here for a bit, but I believe we going to
| get a fully working stable SC before we get fusion. "we"
| as in humanity, you and I might not be around when it's
| finally done.
| JohnMakin wrote:
| leave star citizen out of this :)
| philistine wrote:
| > And Star Citizen will prove to have been worth it along
| along
|
| Once they've implemented saccades in the eyeballs of the
| characters wearing helmets in spaceship millions of
| kilometres apart, then it will all have been worth it.
| sho_hn wrote:
| You have it all wrong. The end game is a scalable, reliable
| AI work force capable of finishing Star Citizen.
|
| At least this is the benchmark for super-human general
| intelligence that I propose.
| bodegajed wrote:
| Could this path lead to solving world hunger too? :)
| mattgreenrocks wrote:
| It's an honor to be dragged along so many ubermensch's
| Incredible Journeys.
| bloomingkales wrote:
| Star Citizen is a working model of how to do UBI. That
| entire staff of a thousand people is the test case.
| jodrellblank wrote:
| > "Early testing shows that interacting with GPT-4.5 feels
| more natural. Its broader knowledge base, improved ability to
| follow user intent, and greater "EQ" make it useful for tasks
| like improving writing, programming, and solving practical
| problems. We also expect it to hallucinate less."
|
| "Early testing doesn't show that it hallucinates less, but we
| expect that putting that sentence nearby will lead you to
| draw a connection there yourself".
| istjohn wrote:
| According to a graph they provide, it does hallucinate
| significantly less on at least one benchmark.
| jug wrote:
| It hallucinates at 37% on SimpleQA yeah, which is a set
| of very difficult questions inviting hallucinations.
| Claude 3.5 Sonnet (the June 2024 editiom, before October
| update and before 3.7) hallucinated at 35%. I think this
| is more of an indication of how behind OpenAI has been in
| this area.
| tmpz22 wrote:
| Are the benchmarks known ahead of time? Could the answer
| to the benchmarks be in the training data?
| llm_trw wrote:
| In general yes, bench mark pollution is a big problem and
| why only dynamic benchmarks matter.
| LeifCarrotson wrote:
| That's some top-tier sales work right there.
|
| I suck at and hate writing the mildly deceptive corporate
| puffery that seems to be in vogue. I wonder if GPT-4.5 can
| write that for me or if it's still not as good at it as the
| expert they paid to put that little gem together.
| zaptrem wrote:
| This seems like it should be attributed to better post
| training, not a bigger model.
| justspacethings wrote:
| The usage of "greater" is also interesting. It's like they
| are trying to say better, but greater is a geographic term
| and doesn't mean "better" instead it's closer to "wider" or
| "covers more area."
| lechatonnoir wrote:
| I'm all for skepticism of capabilities and cynicism about
| corporate messaging, but I really don't think there's an
| interpretation of the word "greater" in this context"
| that doesn't mean "higher" and "better".
| skissane wrote:
| > but greater is a geographic term and doesn't mean
| "better" instead it's closer to "wider" or "covers more
| area."
|
| You are confusing a specific geographical sense of
| "greater" (e.g. "greater New York") with the generic
| sense of "greater" which just means "more great". In "7
| is greater than 6", "greater" isn't geographic
|
| The difference between "greater" and "better", is
| "greater" just means "more than", without implying any
| value judgement-"better" implies the "more than" is a
| good thing: "The Holocaust had a greater death toll than
| the Armenian genocide" is an obvious fact, but only a
| horrendously evil person would use "better" in that
| sentence (excluding of course someone who accidentally
| misspoke, or a non-native speaker mixing up words)
| esafak wrote:
| GPT-4.5 may be an awesome model, some say!
| crazygringo wrote:
| > _We don 't really know what this is good for_
|
| Oh come on. Think how long of a gap there was between the
| first microcomputer and VisiCalc. Or between the start of the
| internet and social networking.
|
| First of all, it's going to take us 10 years to figure out
| how to use LLM's to their full productive potential.
|
| And second of all, it's going to take us collectively a long
| time to also figure out how much accuracy is necessary to pay
| for in which different applications. Putting out a higher-
| accuracy, higher-cost model for the market to try is an
| important part of figuring that out.
|
| With new disruptive technologies, companies aren't supposed
| to be able to look into a crystal ball and see the future.
| They're _supposed_ to try new things and see what the market
| finds useful.
| nyc_data_geek1 wrote:
| The Internet had plenty of very productive use cases before
| social networking, even from its most nascent origins.
| Spending billions building something on the assumption that
| someone else will figure out what it's good for, is not
| good business.
| crazygringo wrote:
| And LLM's already have tons of productive uses. The
| biggest ones are probably still waiting, though.
|
| But this is about one particular price/performance ratio.
|
| You need to build things before you can see how the
| market responds. You say it's "not good business" but
| that's entirely wrong. It's excellent business. It's the
| only way to go about it, in fact.
|
| Finding product-market fit is a process. Companies aren't
| omniscient.
| bigstrat2003 wrote:
| > And LLM's already have tons of productive uses.
|
| I disagree strongly with that. Right now they are fun
| toys to play with, but not useful tools, because they are
| not reliable. If and when that gets fixed, maybe they
| will have productive uses. But for right now, not so
| much.
| nyc_data_geek1 wrote:
| You go into this process with a perspective, you do not
| build a solution and then start looking for a problem.
| rsynnott wrote:
| Arguably social networking is older than the internet
| proper; USENET predates TCP/IP (though not ARPANet).
| nyrikki wrote:
| The TRS-80, Apple ][, and PET all came out in 1977,
| VisiCalc was released in 1979.
|
| Usenet, Bitnet, IRC, BBSs all predated the commercial
| internet, which are all forms of _Online_ social networks.
| mandevil wrote:
| ChatGPT had its initial public release November 30th, 2022.
| That's 820 days to today. The Apple II was first sold June
| 10, 1977, and Visicalc was first sold October 17, 1979,
| which is 859 days. So we're right about the same distance
| in time- the exact equal duration will be April 7th of this
| year.
|
| Going back to the very first commercially available
| microcomputer, the Altair 8800, that's four years and nine
| months to Visicalc release. This isn't a decade long
| process of figuring things out, it actually tends to move
| real fast.
| aylmao wrote:
| I generally agree with the idea of building things,
| iterating, and experimenting before knowing their full
| potential, but I do see why there's negative sentiment
| around this:
|
| 1. The first microcomputer predates VisiCalc, yes, but it
| doesn't predate the realization of what it could be useful
| for. The Micral was released in 1973. Douglas Engelbart
| gave "The Mother of All Demos" in 1968 [2]. It included
| things that wouldn't be commonplace for decades, like a
| collaborative real-time editor or video-conferencing.
|
| I wasn't yet born back then, but reading about the timeline
| of things, it sounds like the industry had a much more
| concrete and concise idea of what this technology would
| bring to everyone.
|
| "We look forward to learning more about its strengths,
| capabilities, and potential applications in real-world
| settings." doesn't inspire that sentiment for something
| that's already being marketed as "the beginning of a new
| era" and valued so exorbitantly.
|
| 2. I think as AI becomes more generally available, and
| "good enough" people (understandably) will be more
| skeptical of closed-source improvements that stem from
| spending big. Commoditizing AI is more clearly "useful", in
| the same way commoditizing computing was more clearly
| useful than just pushing numbers up.
|
| Again, I wasn't yet born back then, but I can imagine the
| announcement of Apple Macintosh with its 6MHz CPU and 128KB
| RAM was more exciting and had a bigger impact than the
| announcement of the Cray-2 with its 1.9GHz and +1GB memory.
|
| [1] https://en.wikipedia.org/wiki/Micral
|
| [2] https://en.wikipedia.org/wiki/The_Mother_of_All_Demos
| fsndz wrote:
| it's so over, pretraining is ngmi. maybe sam Altman was wrong
| after all ? https://www.lycee.ai/blog/why-sam-altman-is-wrong
| FpUser wrote:
| >"I also agree with researchers like Yann LeCun or Francois
| Chollet that deep learning doesn't allow models to
| generalize properly to out-of-distribution data--and that
| is precisely what we need to build artificial general
| intelligence."
|
| I think "generalize properly to out-of-distribution data"
| is too weak of criteria for general intelligence (GI). GI
| model should be able to get interested about some
| particular area, research all the known facts, derive new
| knowledge / create theories based upon said fact. If there
| is not enough of those to be conclusive: propose and
| conduct experiments and use the results to prove / disprove
| / improve theories. And it should be doing this constantly
| in real time on bazillion of "ideas". Basically model our
| whole society. Fat chance of anything like this happening
| in foreseeable future.
| amarcheschi wrote:
| I have a professor who founded a few companies, one of these
| was funded by gates after he managed to spoke with him and
| convinced him to give him money. This guy is goat, and he
| always tells us that we need to find solutions to problems,
| not to find problems to our solutions. It seems at openai
| they didn't get the memo this time
| serjester wrote:
| I suppose this was their final hurrah after two failed attempts
| at training GPT-5 with the traditional pre-training paradigm.
| Just confirms reasoning models are the only way forward.
| newfocogi wrote:
| I think this is the correct take. There are other axes to
| scale on AND I expect we'll see smaller and smaller models
| approach this level of pre-trained performance. But I believe
| massive pre-training gains have hit clearly diminished
| returns (until I see evidence otherwise).
| granzymes wrote:
| > Compared to OpenAI o1 and OpenAI o3-mini, GPT-4.5 is a more
| general-purpose, innately smarter model. We believe reasoning
| will be a core capability of future models, and that the two
| approaches to scaling--pre-training and reasoning--will
| complement each other. As models like GPT-4.5 become smarter
| and more knowledgeable through pre-training, they will serve
| as an even stronger foundation for reasoning and tool-using
| agents.
| jstummbillig wrote:
| What it confirms, I think, is, that we are going to need a
| _lot_ more chips.
| prisenco wrote:
| Or, possibly, we're stuck waiting for another theoretical
| breakthrough before real progress is made.
| resource0x wrote:
| breakthrough in biology
| georgemcbay wrote:
| Further confirmation, IMO, that the idea that any of this
| leads to anything close to AGI is people getting high on
| their own supply (in some cases literally).
|
| LLMs are a great tool for what is effectively collected
| knowledge search and summary (so long as you are willing to
| accept that you have to verify all of the 'knowledge' they
| spit back because they always have the ability to go off
| the rails) but they have been hitting the limits on how
| much better that can get without somehow introducing more
| real knowledge for close to 2 years now and everything
| since then is super incremental and IME mostly just
| benchmark gains and hype as opposed to actually being
| purely better.
|
| I personally don't believe that more GPUs solves this,
| like, at all. But its great for Nvidia's stock price.
| DannyBee wrote:
| Eh, no. More chips won't save this right now, or probably
| in the near future (IE barring someone sitting on a
| breakthrough right now).
|
| It just means either
|
| A. Lots and lots of hard work that get you a few percent at
| a time, but add up to a _lot_ over time.
|
| or
|
| B. Completely different approaches that people actually
| think about for a while rather than trying to incrementally
| get something done in the next 1-2 months.
|
| Most fields go through this stage. Sometimes more than once
| as they mature and loop back around :)
|
| Right now, AI seems bad at doing either - at least, from
| the outside of most of these companies, and watching open
| source/etc.
|
| While lots of little improvements seem to be released in
| lots of parts, it's rare to see anywhere that is collecting
| and aggregating them en masse and putting them in practice.
| It feels like for every 100 research papers, maybe 1 makes
| it into something in a way that anyone ends up using it by
| default.
|
| This could be because they aren't really even a few percent
| (which would be yet a different problem, and in some ways
| worse), or it could be because nobody has cared to, or ...
|
| I'm sure very large companies are doing a fairly reasonable
| job on this, because they historically do, but everyone
| else - even frameworks - it's still in the "here's a
| million knobs and things that may or may not help".
|
| It's like if compilers had no "O0/O1/O2/O3' at all and were
| just like "here's 16,283 compiler passes - you can put them
| in any order and amount you want". Thanks! I hate it!
|
| It's worse even because it's like this at every layer of
| the stack, whereas in this compiler example, this is just
| one layer.
|
| Additionally, everyone seems to rush half-baked things to
| try to get the next incremental improvement released and
| out the door because they think it will help them stay
| "sticky" or whatever. History does not suggest this is a
| good plan and even if it was a good plan in theory, it's
| pretty hard to lock people in with what exists right now.
| There isn't enough anyone cares about and rushing out half-
| baked crap is not helping that. mindshare doesn't really
| matter if no one cares about using _your_ product.
|
| Does anyone using these things truly feel locked into
| anyone's ecosystem at this point? Do they feel like they
| will be soon?
|
| I haven't met anyone who feels that way, even in corps
| spending tons and tons of money with these providers.
|
| The public companies - i can at least understand given the
| fickleness of public markets. That was supposed to be one
| of the serious benefit of staying private. So watching
| private companies do the same thing - it's just sort of
| mind-boggling.
|
| Hopefully they'll grow up soon, or someone who takes their
| time and does it right during one of the lulls will come
| and eat all of their lunches.
| usaar333 wrote:
| For OpenAI perhaps? Sonnet 3.7 without extended thinking is
| quite strong. Swe-bench scores tie o3
| DebtDeflation wrote:
| GPT 5 is likely just going to be a router model that decides
| whether to send the prompt to 4o, 4o mini, 4.5, o3, or o3
| mini.
| techorange wrote:
| I wonder how much money they're losing on it too even at those
| prices.
| nialv7 wrote:
| Looks like more signal that the scaling "law" is indeed
| faltering.
| ur-whale wrote:
| AI as it stands in 2025 is an amazing technology, but it is not
| a product _at all_.
|
| As a result, OpenAI simply does not have a business model, even
| if they are trying to convince the world that they do.
|
| My bet is that they're currently burning through other people's
| capital at an amazing rate, but that they are light-years from
| profitability
|
| They are also being chased by fierce competition and OpenSource
| which is very close behind. There simply is no moat.
|
| It will not end well for investors who sunk money in these
| large AI startups (unless of course they manage to find a
| Softbank-style mark to sell the whole thing to), but everyone
| will benefit from the progress AI will have made during the
| bubble.
|
| So, in the end, OpenAI will have, albeit very unwillingly,
| fulfilled their original charter of improving humanity's lot.
| jsheard wrote:
| > My bet is that they're currently burning through other
| people's capital at an amazing rate, but that they are light-
| years from profitability
|
| The Information leaked their internal projections a few
| months ago, and apparently their own estimates have them
| losing $44B between then and 2029 when they expect to finally
| turn a profit, maybe.
| j_maffe wrote:
| That's surprisingly small
| whiplash451 wrote:
| Except that if OpenAI goes bust, very little of what they did
| will actually be released to human kind.
|
| So their contribution was really to fuel a race for
| opensource (which they contributed little to). Pretty complex
| of an argument.
| emptysongglass wrote:
| I've been a Plus user for a long time now. My opinion is
| there is very much a ChatGPT suite of products that come
| together to make for a mostly delightful experience.
|
| Three things I use all the time:
|
| - Canvas for proofing and editing my article drafts before
| publishing. This has replaced an actual human editor for me.
|
| - Voice for all sorts of things, mostly for thinking out loud
| about problems or a quick question about pop culture, what
| something means in another language, etc. The Sol voice is so
| approachable for me.
|
| - GPTs I can use for things like D&D adventure summaries I
| need in a certain style every time without any manual
| prompting.
| vineyardmike wrote:
| > As a result, OpenAI simply does not have a business model,
| even if they are trying to convince the world that they do.
|
| They have a super popular subscription service. If they keep
| iterating on the product enough, they can lag on the models.
| The business is the product not the models and not the API.
| Subscriptions are pretty sticky when you start getting your
| data entrenched in it. I keep my ChatGPT subscription because
| it's the best app on Mac and already started to "learn me"
| through the memory and tasks feature.
|
| Their app experience is easily the best out of their
| competitors (grok, Claude, etc). Which is a clear sign they
| know that it's the product to sell. Things like DeepResearch
| and related are the way they'll make it a sustainable
| business - add value-on-top experiences which drive the
| differentiation over commodities. Gemini is the only
| competitor that compares because it's everywhere in Google
| surfaces. OpenAI's pro tier will surely continue to get
| better, I think more LLM-enabled features will continue to be
| a differentiator. The biggest challenge will be continuing
| distribution and new features requiring interfacing with
| third parties to be more "agentic".
|
| Frankly, I think they have enough strength in product with
| their current models today that even if model training
| stalled it'd be a valuable business.
| jcgrillo wrote:
| https://podcasts.apple.com/us/podcast/better-
| offline/id17305...
| nyarlathotep_ wrote:
| > AI as it stands in 2025 is an amazing technology, but it is
| not a product at all.
|
| Here I'm assuming "AI" to mean what's broadly called
| Generative AI (LLMs, photo, video generation)
|
| I genuinely am struggling to see what the product is too.
|
| The code assistant use cases are really impressive across the
| board (and I'm someone who was vocally against them less than
| a year ago), and I pay for Github CoPilot (for now) but I
| can't think of any offering otherwise to dispute your claim.
|
| It seems like companies are desperate to find a market fit,
| and shoving the words "agentic" everywhere doesn't inspire
| confidence.
|
| Here's the thing: I remember people lining up around the
| block for iPhone releases, XBox launches, hell even Grand
| Theft Auto midnight releases.
|
| Is there a market of people clamoring to use/get anything
| GenAI related?
|
| If any/all LLM services went down tonight, what's the impact?
| Kids do their own homework?
|
| JavaScript programmers have to remember how to write React
| components?
|
| Compare that with Google Maps disappearing, or similar.
|
| LLMs are in a position where they're forced onto people and
| most frankly aren't that interested. Did anyone ASK for
| Microsoft throwing some Copilot things all over their
| operating system? Does anyone want Apple Intelligence,
| really?
| jdprgm wrote:
| If it really costs them 30x more surely they must plan on
| putting pretty significant usage limits on any rollout to the
| Plus tier and if that is the case i'm not sure what the point
| is considering it seems primarily a replacement/upgrade for 4o.
|
| The cognitive overhead of choosing between what will be 6
| different models now on chatGPT and trying to map whether a
| query is "worth" using a certain model and worrying about
| hitting usage limits is getting kind of out of control.
| chollida1 wrote:
| > GPT 4.5 pricing is insane:
|
| > I'm still gonna give it a go, though.
|
| Seems like the pricing is pretty rational then?
| phito wrote:
| Not if people just try a few prompts then stop using it.
| chollida1 wrote:
| Sure but its in their best interest to lower it then and
| only then.
|
| OpenAI wouldn't be the first company to price something
| expensive when it first comes out to capitalize on people
| who are less price sensitive at first and then lower prices
| to capture a bigger audience.
|
| That's all pricing 101 as the saying goes.
| j_maffe wrote:
| If OAI are concerning themselves with collecting a few
| hundereds from a small group of individuals then they
| really have nothing better to do
| nyarlathotep_ wrote:
| How much of OAI's reported users are doing exactly this?
| raytopia wrote:
| Now the real question about AI automation starts. Is it cheaper
| to pay a human to do the task or a AI company?
| redox99 wrote:
| It still not smart enough to replace for example customer
| service.
| fragmede wrote:
| Humans have all sorts of issues you have to deal with. Being
| hungover, not sleeping well, having a personality, being late
| to work, not being able to work 24/7, very limited ability to
| copy them. If there's a soulless generic office-droidGPT that
| companies could hire that would never talk back and would do
| all sorts of menial work without needing breaks or to use the
| bathroom, I don't know that we humans stand a chance!
|
| I have a bunch of work that needs doing. I can do it myself,
| or I can hire one person to do it. I gotta train them and
| manage them and even after I train them theres still only
| going to be one of them, and it's subject to their
| availability. On the other hand, if I need to train an AI to
| do it, but I can copy that AI, and then spin them up/down
| like on demand computer in the cloud, and not feel remotely
| bad about spinning them down?
|
| It's definitely not there yet, but it's not hard to see the
| business case for it.
| crooked-v wrote:
| Doubly so with how good Claude 3.7 Sonnet is at $3 / 1M tokens.
| Hansenq wrote:
| GPT-4.5 is 15-30x more expensive than GPT-4o. Likely that much
| larger in terms of parameter count too. It's massive!!
|
| With more parameters comes more latent space to build a world
| model. No wonder its internal world model is so much better
| than previous SOTA
| wavemode wrote:
| This has been my suspicion for a long time - OpenAI have indeed
| been working on "GPT5", but training and running it is proving
| so expensive (and its actual reasoning abilities only
| marginally stronger than GPT4) that there's just no market for
| it.
|
| It points to an overall plateau being reached in the
| performance of the transformer architecture.
| goatlover wrote:
| Certainly hope so. The tech billionaires are little to
| excited to achieve AGI and replace the workforce.
| shoubidouwah wrote:
| TBH, with the safety/alignment paradigm we have, workforce
| replacement was not my top concern when we hit AGI. A pause
| / lull in capabilities would be hugely helpful so that we
| can figure how not to die along with the lightcone...
| hnuser123456 wrote:
| Is it inevitable to you that someone will create some
| kind of techno-god behemoth AI that will figure out how
| to optimally dominate an entire future light cone
| starting from the point in spacetime of its self-
| actualization? Borg or Cylons?
| hintymad wrote:
| > It sounds like it's so expensive and the difference in
| usefulness is so lacking(?) they're not even gonna keep serving
| it in the API for long
|
| I guess the rationale behind this is paying for the marginal
| improvement. Maybe the next few percent of improvement is so
| important to a business that the business is willing to pay a
| hefty premium.
| shawabawa3 wrote:
| I wonder if the pricing is partly to discourage distillation,
| if they suspect r1 was distilled from gpt 4o
| tomrod wrote:
| I can chew through 1MM tokens with a single standard (and
| optimized) call. This pricing is insane.
| MangoCoffee wrote:
| one of the problem seem to be there's no alternative to Nvidia
| ecosystem. (the gpu + CUDA).
| coliveira wrote:
| In other words, they want people to pay for the privilege of
| becoming beta testers....
| wiremine wrote:
| > GPT 4.5 pricing is insane: Price Input: $75.00 / 1M tokens
| Cached input: $37.50 / 1M tokens Output: $150.00 / 1M tokens
|
| > GPT 4o pricing for comparison: Price Input: $2.50 / 1M tokens
| Cached input: $1.25 / 1M tokens Output: $10.00 / 1M tokens
|
| Their examples don't seem 30x better. :-)
| ren_engineer wrote:
| hyperscalers in shambles, no clue why they even released this
| other than the fact they didn't want to admit they wasted an
| absurd amount of money for no reason
| hn_throwaway_99 wrote:
| The price is obviously 15-30x that of 4o, but I'd just posit
| that there are some use cases where it may make sense. It
| probably doesn't make sense for the "open-ended consumer facing
| chatbot" use case, but for other use cases that are fewer and
| higher value in nature, it could if it's abilities are
| considerably better than 4o.
|
| For example, there are now a bunch of vendors that sell
| "respond to RFP" AI products. The number of RFPs that any sales
| organization responds to is probably no more than a couple a
| week, but it's a very time-consuming, laborious process. But
| the payoff is obviously very high if a response results in a
| closed sale. So here paying 30x for marginally better
| performance makes perfect sense.
|
| I can think of a number of similar "high value, relatively low
| occurrence" use cases like this where the pricing may not be a
| big hindrance.
| Manouchehri wrote:
| Yeah, agreed.
|
| We're one of those types of customers. We wrote an OpenAI API
| compatible gateway that automatically batches stuff for us,
| so we get 50% off for basically no extra dev work in our
| client applications.
|
| I don't care about speed, I care about getting the right
| answer. The cost is fine as long as the output generates us
| more profit.
| kristofferR wrote:
| "GPT-4.5 is not a frontier model, but it is OpenAI's largest
| LLM, improving on GPT-4's computational efficiency by more than
| 10x."[1]
|
| I don't get it, it is supposedly much cheaper to run?
|
| [1] https://cdn.openai.com/gpt-4-5-system-card.pdf (page 7,
| bottom)
| acchow wrote:
| > It sounds like it's so expensive and the difference in
| usefulness is so lacking(?)
|
| The claimed hallucination rate is dropping from 61% to 37%.
| That's a "correct" rate increasing from 29% to 63%.
|
| Double the correct rate costs 15x the price? That seems absurd,
| unless you think about how mistakes compound. Even just 2 steps
| in and you're comparing a 8.4% correct rate vs 40%. 3 automated
| steps and it's 2.4% vs 25%.
| Xiol32 wrote:
| The example GPT-4.5 answers from the livestream are just... too
| excitable? Can't put my finger on it, but it feels like they're
| aimed towards little kids.
| jug wrote:
| It made me wonder how much of that was due to the system prompt
| too.
| virgildotcodes wrote:
| That presentation was super underwhelming. We got to watch them
| compare... the vibes? ... of 4.5 vs o1.
|
| No wonder Sam wasn't part of the presentation.
| Etheryte wrote:
| And to top it off, it costs $75.00 per 1M vibes.
| thomas34298 wrote:
| Sam tweeted "taking care of my kid in the hospital":
|
| https://x.com/sama/status/1895210655944450446
|
| Let's not assume that he's lying. Neither the presentation nor
| my short usage via the API blew me away, but to really evaluate
| it, you'd have to use it longer on a daily basis. Maybe that
| becomes a possiblity with the announced performance
| optimizations that would lower the price...
| jug wrote:
| It should've just been a web launch without video.
| bakugo wrote:
| API is literally 5 times more expensive than Claude 3 Opus, and
| it doesn't even seem to do anything impressive. What's the
| business strategy here?
| hidelooktropic wrote:
| I'm not sure that doing a live stream on this was the right way
| to go. I would've just quietly sent out a press release. I'm sure
| they have better things on the way.
| sebastiennight wrote:
| It is interesting that they are focusing a large part of this
| release on the model having a higher "EQ" (Emotional Quotient).
|
| We're far from the days of "this is not a person, we do not want
| to make it addictive" and getting a firm foot on the territory of
| "here's your new AI friend".
|
| This is very visible in the example comparing 4o with 4.5 when
| the user is complaining about failing a test, where 4o's response
| is what one would expect from a "typical AI response" with
| problem-solving bullets, and 4.5 is sending what you'd expect
| from a pal over instant messaging.
|
| It seems Anthropic and Grok have both been moving in this
| direction as well. Are we going to see an escalation of
| foundation models impersonating "a friendly person" rather than
| "a helpful assistant"?
|
| Personally I find this worrying and (as someone who builds upon
| SOTA model APIs) I really hope this behavior is not going to seep
| into API responses, or will at least be steerable through the
| system/developer prompt.
| og_kalu wrote:
| The whole robotic, monotone, helpful assistant thing was
| something these companies had to actively hammer in during the
| post-training stage. It's not really how LLMs will sound by
| default after pre-training.
|
| I guess they're caring less and less about that effort
| especially since it hurts the model in some ways like creative
| writing.
| sebastiennight wrote:
| If it's just a different choice during RLHF, I'll be curious
| to see what are the trade-offs in performance.
|
| The "buddy in a chat group" style answers do not make me feel
| like asking it for a story will make the story
| long/detailed/poignant enough to warrant the difference.
|
| I'll give it a try and compare on creative tasks.
| turnsout wrote:
| Or maybe they're just getting better at it, or developing
| better taste. After switching to Claude, I can't go back to
| ChatGPT's overly verbose bullet-point laden book reports
| every time I ask a question. I don't think that's pretraining
| --it's in the way OpenAI approaches tuning and prompting vs
| Anthropic.
| tmaly wrote:
| I would like to see a humor test. So far, I have not seen any
| model response that has made me laugh.
| sebastiennight wrote:
| The "roast" tools that have popped up (using either DeepSeek
| or o3-mini) are pretty funny.
|
| Eg. https://news.ycombinator.com/item?id=43163654
| jcims wrote:
| OK now that is some funny shit.
| AgentME wrote:
| My benchmark for this has been asking the model to write some
| tweets in the style of dril, a popular user who writes short
| funny tweets. Sometimes I include a few example tweets in the
| prompt too. Here's an example of results I got from Claude 3
| Opus and GPT 4 for this last year:
| https://bsky.app/profile/macil.tech/post/3kpcvicmirs2v. My
| opinion is that Claude's results were mostly bangers while
| GPT's were all a bit groanworthy. I need to try this again
| with the latest models sometime.
| tkgally wrote:
| How does the following stand-up routine by Claude 3.7 Sonnet
| work for you?
|
| https://gally.net/temp/20250225claudestandup2.html
| lurker9001 wrote:
| incredible
| turnsout wrote:
| If you like absurdist humor, go into the OpenAI playground,
| select 3.5-Turbo, and dial up the temperature to the point
| where the output devolves into garbled text after 500 tokens
| or so. The first ~200 tokens are in the freaking sweet spot
| of humor.
| amarcheschi wrote:
| Could someone post an example?
| rl3 wrote:
| Maybe it's rose-colored glasses, but 3.5 was really the
| golden era for LLM comedy. More modern LLMs can't touch it.
|
| Just ask it to write you a film screenplay involving some
| hard-ass 80s/90s action star and someone totally unrelated
| and opposite of that. The ensuring unhinged magic is
| unparalleled.
| jcims wrote:
| I built a little AI assistant to read my calendar and
| send me a summary of my day every morning. I told it to
| roast me and be funny with it.
|
| 3.5 was *way* better than anything else at that.
| nialv7 wrote:
| Well yeah, if the llm can keep you engaged and talking, that'll
| make them a lot more money; compared to if you just use it as a
| information retrieval tool in which case you are likely to
| leave after getting what you are looking for.
| TheAceOfHearts wrote:
| Since they offer a subscription, keeping you engaged just
| requires them to waste more compute. The ideal case would be
| that the LLM gives you a one shot correct response using as
| little compute as possible.
| sebastiennight wrote:
| In a subscription business, you don't want the user to use
| as few resources as possible. It's the wrong optimization
| to make.
|
| You want users to keep coming back as often as possible (at
| the lowest cost-per-run possible though). If they are not
| coming back they are not renewing.
|
| So, yes, it makes sense to make answers shorter to cut on
| compute cost (which these SMS-length replies could
| accomplish) but the main point of making the AI flirtatious
| or "concerned" is possibly the addictive factor of having a
| shoulder to cry on 24/7, one that does not call you on your
| BS and is always supportive... for just $20 a month
|
| The "one-shot correct response" to "I failed my exams"
| might be "Tough luck, try better next time" but if you do
| that, you will indeed use very little compute _because
| people will cancel the subscription and never come back_.
| johnthewise wrote:
| AI subscriptions are already very sticky . I can't
| imagine at least not paying for one, so I doubt they care
| about retention like the rest of us plebs do.
| nialv7 wrote:
| Plus level subscription has limits too, and Pro level costs
| 10x more - as long as Pro users don't use ChatGPT 10x more
| than Plus users on average, OpenAI can benefit. There's
| also the user retention factor.
| bredren wrote:
| Yes, the "personality" (vibe) of the model is a key qualitative
| attribute of gpt-4.5.
|
| I suspect this has something to do with shining light on an
| increased value prop in a dimension many people will appreciate
| since gains on quantitative comparison with other models were
| not notable enough to pop eyeballs.
| callc wrote:
| > We're far from the days of "this is not a person, we do not
| want to make it addictive" and getting a firm foot on the
| territory of "here's your new AI friend".
|
| That's a hard nope from me, when companies pull that move. I'll
| stick to my flesh and blood humans who still hallucinate but
| only rarely.
| orbital-decay wrote:
| Anthropic pretty much abandoned this direction after Claude 3,
| and said it wasn't what they wanted [1]. Claude 3.5+ is
| extremely dry and neutral, it doesn't seem to have the same
| training.
|
| _> Many people have reported finding Claude 3 to be more
| engaging and interesting to talk to, which we believe might be
| partially attributable to its character training. This wasn't
| the core goal of character training, however. Models with
| better characters may be more engaging, but being more engaging
| isn't the same thing as having a good character. In fact, an
| excessive desire to be engaging seems like an undesirable
| character trait for a model to have._
|
| [1] https://www.anthropic.com/research/claude-character
| CitizenTen wrote:
| GPT pro already has already rummored to be 100k users. You think
| GPT 4.5 will add to that even with the insane costs for corporate
| users?
| cristiancavalli wrote:
| What rumors? I looked and can't find something to substantiate
| that #
| andsoitis wrote:
| "this isn't a reasoning model and won't crush benchmarks."
|
| -- https://x.com/sama/status/1895203654103351462
| MaxPock wrote:
| One thing that Altman does extremely well is to over-promise and
| under-deliver.
| erulabs wrote:
| Finally a scaling wall? This is apparently (based on pricing)
| using about an order of magnitude more compute, and is only maybe
| 10% more intelligent. Ideally DeepSeeks optimizations help bring
| the costs way down, but do any AI researchers want to comment on
| if this changes the overall shape of the scaling curve?
| fpgaminer wrote:
| Seems on par with the existing scaling curve. If I had to
| speculate, this model would have been an internal-only model,
| but they're releasing it for PR. An optimized version with 99%
| of the performance for 1/10th the cost will come out later.
| j_maffe wrote:
| This is the shittiest PR move I've seen since the AI trend
| started.
| ein0p wrote:
| Now imagine this model (or an optimized/slightly downsized
| variant thereof) as a base for a "thinking" one.
| bparsons wrote:
| Are they saying that 4.5 has a 35% hallucination rate? That chart
| is a bit confusing.
| riku_iki wrote:
| its on that benchmark, which likely is very challenging.
| Seattle3503 wrote:
| The GPT-1 response to the example prompt "What was the first
| language?" got a chuckle out of me
| aldanor wrote:
| The question being, will we be chuckling at current models
| responses in 5-10y from now?
| joshuamcginnis wrote:
| I'm one week in on heavy grok usage. I didn't think I'd say this,
| but for personal use, I'm considering cancelling my OpenAI plan.
|
| The one thing I wish grok had was more separation of the UI from
| X itself. The interface being so coupled to X puts me off and
| makes it feel like a second-hand citizen. I like ChatGPTs
| minimalist UI.
| aldanor wrote:
| Theres grok.com which is standalone and with its own UI
| it wrote:
| There's also a standalone Grok app at least on iOS.
| richard_todd wrote:
| I find grok to be the best overall experience for the types of
| tasks I try to give AI (mostly: analyze pdf, perform and
| proofread OCR, translate Medieval Latin and Hebrew, remind me
| how to do various things in python or SwiftUI).
| ChatGPT/gemini/copilot all fight me occasionally, but grok just
| tries to help. And the hallucinations aren't as frequent, at
| least anecdotally.
| fzzzy wrote:
| Don't they have a standalone Grok app now? I thought I saw
| that. [edit] ah some sibling comments mention this as well
| selalipop wrote:
| I've been working on post-training models for tasks that require
| EQ, so it's validating to see OpenAI working towards that too.
|
| That being said, this is very expensive.
|
| - Input: $75.00 / 1M tokens
|
| - Cached input: $37.50 / 1M tokens
|
| - Output: $150.00 / 1M tokens
|
| One of the most interesting applications of models with higher EQ
| is personalized content generation, but the size and cost here
| are at odds with that.
| kgeist wrote:
| >GPT-4.5 is more succinct and conversational
|
| I wonder why they highlight it as an achievement when they could
| have simply tuned 4o to be more conversational and less like a
| bullet-point-style answer machine. They did something to 4o
| compared to the previous models which made the responses feel
| more canned.
| mvdtnz wrote:
| OpenAI doubling down on the American-style therapy-speak instead
| of focusing on usefulness. No thanks.
| wewewedxfgdf wrote:
| I feel like OpenAI is pursuing AGI when Anthropic/Claude is
| pursuing making AI awesome for practical things like coding.
|
| I only ever using OpenAI's coding now as a double check against
| Claude.
|
| Does OpenAI have their eyes on the ball?
| ls_stats wrote:
| >I feel like OpenAI is pursuing AGI
|
| I don't think so, the "AGI guy" was Ilya Sutskever, he is gone,
| he wanted to make OpenAI "less comercial", AGI is just a
| buzzword for Altmann.
| rakejake wrote:
| Right. A good chunk of the "old guard" is now gone - Ilya to
| SSI, Mira and a bunch of others to a new venture called
| Thinking Machines, Alec Radford etc. Remains to be seen if
| OpenAI will be the leader or if other players catch up.
| rakejake wrote:
| My usage has come down to mostly Claude (until I run out of
| free tier quota) and then Gemini. Claude is the best for code
| and Gemini 2.0 Flash is good enough while also being free (well
| considering how much data G has hoovered up over the years,
| perhaps not) and more importantly highly available.
|
| For simple queries like generating shell scripts for some
| plumbing, or doing some data munging, I go straight to Gemini.
| HarHarVeryFunny wrote:
| > My usage has come down to mostly Claude (until I run out of
| free tier quota) and then Gemini
|
| Yep, exactly same here.
|
| Gemini 2.0 Flash is extremely good, and I've yet to hit any
| usage limits with them - for heavy usage I just go to Gemini
| directly. For "talk to an expert" usage, Claude is hard to
| beat though.
| wayeq wrote:
| Claude still can't make real time web searches though for
| RAG, unless I'm missing something.
| resource0x wrote:
| Pursuing AGI? What method do they use to pursue something that
| no one knows what it is? They will keep saying they are
| pursuing AGI as long as there's a buyer for their BS.
| lblume wrote:
| Am I missing something, or do the results not even look that much
| better? Referring to the output quality, this just seems like a
| different prompting style and RLHF, not really an improved model
| at all.
| infinet wrote:
| Can it be self-hosted? Many institutions and organizations are
| hesitant to use AI because concerns of data leaking over chatbot.
| Open models, on the other hand, can be self-hosted. There is a
| deepseek arm race in other part of the world. Universities are
| racing to host their own deepseek. Hospitals, large businesses,
| local governments, even courts are deploying or showing interest
| in self-hosting deepseek.
| YetAnotherNick wrote:
| Do you know of any university that host Deepseek?
| moralestapia wrote:
| OpenAI has never released a single model that could be self-
| hosted.
|
| GPT-2? Maybe not even that one.
| taytus wrote:
| Who wants a model that is not reasoning? The older models are
| just fine.
| eightysixfour wrote:
| Seeing OpenAI and Anthropic go different routes here is
| interesting. It is worth moving past the initial knee jerk
| reaction of this model being unimpressive and some of the
| comments about "they spent a massive amount of money and had to
| ship something for it..."
|
| * Anthropic appears to be making a bet that a single paradigm
| (reasoning) can create a model which is excellent for all use
| cases.
|
| * OpenAI seems to be betting that you'll need an ensemble of
| models with different capabilities, working as a single system,
| to jump beyond what the reasoning models today can do.
|
| Based on all of the comments from OpenAI, GPT 4.5 is absolutely
| massive, and with that size comes the ability to store far more
| factual data. The scores in ability oriented things - like coding
| - don't show the kind of gains you get from reasoning models but
| the fact based test, SimpleQA, shows a pretty large jump and a
| dramatic reduction in hallucinations. You can imagine a scenario
| where GPT4.5 is coordinating multiple, smaller, reasoning agents
| and using its factual accuracy to enhance their reasoning, kind
| of like ruminating on an idea "feels" like a different process
| than having a chat with someone.
|
| I'm really curious if they're actually combining two things right
| now that could be split as well, EQ/communications, and factual
| knowledge storage. This could all be a bust, but it is an
| interesting difference in approaches none-the-less, and worth
| considering that OpenAI _could_ be right.
| nomel wrote:
| > OpenAI seems to be betting that you'll need an ensemble of
| models with different capabilities, working as a single system,
| to jump beyond what the reasoning models today can do.
|
| The high level block diagrams for tech always end up converging
| to those found in biological systems.
| eightysixfour wrote:
| Yeah, I don't know enough real neuroscience to argue either
| side. What I can say is I feel like this path is more like
| the way that I observe that I think, it _feels_ like there
| are different modes of thinking and processes in the brain,
| and it _seems_ like transformers are able to emulate at least
| two different versions of that.
|
| Once we figure out the frontal cortex & corpus callosum part
| of this, where we aren't calling other models over APIs
| instead of them all working in the same shared space, I have
| a feeling we'll be on to something pretty exciting.
| sebastiennight wrote:
| > * OpenAI seems to be betting that you'll need an ensemble of
| models with different capabilities, working as a single system,
| to jump beyond what the reasoning models today can do.
|
| Seems inaccurate as their most recent claim I've seen is that
| they expect this to be their last non-reasoning model, and are
| aiming to provide all capacities together in the future model
| releases (unifying the GPT-x and o-x lines)
|
| See this claim on TFA:
|
| > We believe reasoning will be a core capability of future
| models, and that the two approaches to scaling--pre-training
| and reasoning--will complement each other.
| eightysixfour wrote:
| From Sam's twitter:
|
| > After that, a top goal for us is to unify o-series models
| and GPT-series models by creating systems that can use all
| our tools, know when to think for a long time or not, and
| generally be useful for a very wide range of tasks.
|
| > In both ChatGPT and our API, we will release GPT-5 as a
| system that integrates a lot of our technology, including o3.
| We will no longer ship o3 as a standalone model.
|
| You could read this as unifying the models _or_ building a
| unified systems which coordinate multiple models. The second
| sentence, to me, implies that o3 will still exist, it just
| won 't be standalone, which matches the idea I shared above.
| sebastiennight wrote:
| Ah, great point. Yes, the wording here would imply that
| they're basically planning on building scaffolding around
| multiple models instead of having one more capable Swiss
| Army Knife model.
|
| I would feel a bit bummed if GPT-5 turned out not to be a
| model, but rather a "product".
| eightysixfour wrote:
| For me it depends on how the models are glued together.
| Connected by function calling and APIs? Probably meh...
|
| Somehow working together in the same latent space? That
| could be neat.
| tmpz22 wrote:
| I worry eliminating consumer choice will drive up prices
| for only a nominal gain in utility for most users.
| wongarsu wrote:
| Or the other way around: smaller reasoning models that can call
| out to GPT-4.5 to get their facts right.
| eightysixfour wrote:
| Maybe, I'm inclined to think OpenAI believes the way I laid
| it out though, specifically because of their focus on
| communication and EQ in 4.5. It seems like they believe the
| large, non-reasoning model, will be "front of house."
|
| Or they'll use some kind of trained router which sends the
| request to the one it thinks it should go to first.
| ilrwbwrkhv wrote:
| Question: explain thread first macros in clojure in a short
| paragraph
|
| GPT 4.5:
|
| In Clojure, the *thread-first macro* (`->`) is used to improve
| code readability by threading a value through a series of
| function calls. It takes an initial value and passes it as the
| first argument to the next function, then continues threading the
| result through subsequent functions. This eliminates deeply
| nested expressions and makes the code more readable.
|
| For example,
|
| ```clojure (-> 5 inc (* 2) str) ;; Equivalent to (str (* 2 (inc
| 5))) - "12" ```
|
| Here, `5` is incremented to `6`, then multiplied by `2` to get
| `12`, and finally converted to a string.
|
| ---
|
| Deepseek v3:
|
| In Clojure, the `->>` (thread-last) and `->` (thread-first)
| macros are used to improve code readability by allowing a more
| linear, pipeline-like flow of operations. The `->` (thread-first)
| macro takes an initial value and threads it as the first argument
| through a series of forms. For example, `(-> x (f) (g))`
| translates to `(g (f x))`. This is particularly useful when
| working with nested function calls, as it flattens the structure
| and makes the sequence of transformations clearer. It's commonly
| used with functions that take the threaded value as their first
| argument, such as collection operations or Java interop methods.
| datadrivenangel wrote:
| They tried, but doesn't look better?
| DaveMcMartin wrote:
| This feels more like a release they pushed out to keep the "hype"
| alive rather than something they were eager to share. Honestly,
| the results don't seem all that impressive, and considering the
| price, it just doesn't feel worth it.
| bla3 wrote:
| I wonder if we're starting to see the effects of the mass exodus
| a while ago.
| 42lux wrote:
| The announcements early on were relatively sincere and technical
| with papers and nice pages explaining the new models in easy
| language and now we get this marketing garbage. Probably the
| fastest enshitification I've seen.
| sky2224 wrote:
| Honestly, the most astounding part of this announcement is their
| comparison to o3-mini with QA prompts.
|
| EIGHTY PERCENT hallucination rate? Are you kidding me?
|
| I get that the model is meant to be used for logic and reasoning,
| but nowhere does OpenAI make this explicitly clear. A majority of
| users are going to be thinking, "oh newer is better," and pick
| that.
| jug wrote:
| Yeah it was an abysmal result (any 50%+ hallucination result in
| that bench is pretty bad) and worse than o1-mini in the
| SimpleQA paper. On that topic, Sonnet 3.5 "Old" hallucinates
| less than GPT-4.5, just for a bit of added perspective here.
| freediver wrote:
| The results for GPT - 4.5 are in for Kagi LLM benchmark too.
|
| It does crush our benchmark - time to make new? ;) - with
| performance similar of that of reasoning models. It does come at
| a great price both in cost and speed.
|
| A monster is what they created. But looking at the tasks it
| fails, some of them my 9 year old would solve. Still in this
| weird limbo space of super knowledge and low intelligence.
|
| May be remembered as the last the last of the 'big ones', can't
| imagine this will be a path for the future.
|
| https://help.kagi.com/kagi/ai/llm-benchmark.html
| theodorthe5 wrote:
| If Gemini 2 is the top in your benchmark, make sure to re-check
| your benchmark.
| shawabawa3 wrote:
| Gemini 2 pro is actually very impressive (maybe not for
| coding, haven't used it for that)
|
| Flash is pretty garbage but cheap
| istjohn wrote:
| Gemini 2.0 Pro is quite good.
| simonw wrote:
| If you want to try it out via their API you can run it through my
| LLM tool using uvx like this: uvx --with 'https:/
| /github.com/simonw/llm/archive/801b08bf40788c09aed617525287631031
| 2fe667.zip' \ llm -m gpt-4.5-preview 'impress me'
|
| You may need to set an API key first, either with `export
| OPENAI_API_KEY='xxx'` or using this command to save it to a file:
| uvx llm keys set openai # paste key here
|
| Or this to get a chat session going: uvx --with '
| https://github.com/simonw/llm/archive/801b08bf40788c09aed61752528
| 76310312fe667.zip' \ llm chat -m gpt-4.5-preview
|
| I'll probably have a proper release out later today. Details
| here: https://github.com/simonw/llm/issues/795
| ashu1461 wrote:
| Just curious, does this stream the output or renders all at
| once ?
| synapsomorphy wrote:
| Claude 3.6 (new 3.5) and 3.7 non-reasoning are much better at
| pretty much everything, and much cheaper. What's Anthropic's
| secret sauce?
| taytus wrote:
| They ship more focused on their mission than OpenAI.
| moralestapia wrote:
| Huh?
|
| Post benchmark links.
| film42 wrote:
| I think it's a classic expectations problem. OpenAI is neither
| _open_ nor is it releasing an _AGI_ model in the near future.
| But when you see a new major model drop, you can't help but
| ask, "how close is this to the promise of AGI they say is just
| around the corner?" Not even close. Meanwhile Anthropic is
| keeping their heads down, not playing the hype game, and
| letting the model speak for itself.
| anothermathbozo wrote:
| Anthropic's CEO said their technology would end all disease
| and expand our lifespans to 200 years. What on earth do you
| mean they're not playing the hype game?
| jampa wrote:
| First impression of GPT-4.5:
|
| 1. It is very very slow, for some applications where you want
| real time interactions is just not viable, the text attached
| below took 7s to generate with 4o, but 46s with GPT4.5
|
| 2. The style it writes is way better: it keeps the tone you ask
| and makes better improvements on the flow. One of my biggest
| complaints with 4o is that you want for your content to be more
| casual and accessible but GPT / DeepSeek wants to write like
| Shakespeare did.
|
| Some comparisons on a book draft: GPT4o (left) and GPT4.5
| (green). I also adjusted the spacing around the paragraphs, to
| better diff match. I still am wary of using ChatGPT to help me
| write, even with GPT 4.5, but the improvement is very noticeable.
|
| https://i.imgur.com/ogalyE0.png
| remus wrote:
| > It is very very slow
|
| Could that be partially due to a big spike in demand at launch?
| jampa wrote:
| Possibly, repeating the prompt I got a much higher speed,
| taking 20s on average now, which is much more viable. But
| that remains to be seen when more people start using this
| version in production.
| jedberg wrote:
| Oh yeah, that right side version is WAY better, and sounds much
| more like a human.
| MichaelZuo wrote:
| How does it compare with o1 and o3 preview?
| jampa wrote:
| o3 is okay for text checking but has issues following the
| prompt correctly, same as o1 and DeepSeek R1, I feel that I
| need to prompt smaller snippets with them.
|
| Here is the o3 vs a new run of the same text in GPT 4.5
|
| https://www.diffchecker.com/ZEUQ92u7/
| MichaelZuo wrote:
| Thanks, though it says o1 on the page, is that a typo?
| FergusArgyll wrote:
| I opened your link in a new tab and looked at it a couple
| minutes later. By then I forgot which was o and which was .5
|
| I honestly couldn't decide which I prefer
| niek_pas wrote:
| I definitely prefer the 4.5, but that might just be because
| it sounds 'less like ChatGPT', ironically.
| rl3 wrote:
| > _1. It is very very slow, ... below took 7s to generate with
| 4o, but 46s with GPT4.5_
|
| This is positively luxurious by o1-pro standards which I'd say
| _average_ 5 minutes. That said I totally agree even ~45s isn 't
| viable for real-time interactions. I'm sure it'll be optimized.
|
| Of course, my comparing it to the highest-end CoT model in
| [publicly-known] existence isn't entirely fair since they're
| sort of apples and oranges.
| philomath_mn wrote:
| I paid for pro to try `o1-pro` and I can't seem to find any
| use case to justify the insane inference time. `o3-mini-high`
| seems to do just as well in seconds vs. minutes.
| thfuran wrote:
| >One of my biggest complaints with 4o is that you want for your
| content to be more casual and accessible but GPT / DeepSeek
| wants to write like Shakespeare did.
|
| Well, maybe like a Sophomore's bumbling attempt to write like
| Shakespeare.
| ChiefNotAClue wrote:
| Right side, by a large margin. Better word choice and more
| natural flow. It feels a lot more human.
| kristianp wrote:
| How do the two versions match so closely? They have the same
| content in each paragraph, just worded slightly differently. I
| wouldn't expect them to write paragraphs that match in size and
| position like that.
| mchusma wrote:
| wow, openai really missed here. Reading the blog I thought like a
| minor, incremental minor catch up release for 4o. I thought "wow
| maybe this is cheaper than 4o so it will offset the pricing
| difference between this and something like Claude Sonnet 3.7 or
| Gemini 2.0 Flash both of which performs better. But its like
| 20x-100x more expensive!
|
| In other words, these performance stats with Gemini 2.0 Flash
| pricing looks reasonable. At these prices, zero usecases for
| anyone I think. This is a dead on arrival model.
| jasonjmcghee wrote:
| This is such as confusing release / announcement.
|
| It seems clearly worse than Claude Sonnet 3.7, yet costs 30x as
| much?
|
| What's going on with their pricing? I misread it as $7.5/M input
| and that that was very overpriced... then realized it was 10x
| that much!
| ashu1461 wrote:
| Is it worse than clause sonnet with reasoning enabled or
| disabled ?
| moffkalast wrote:
| OpenAI will do literally anything but ship GPT-5.
| saghm wrote:
| This is probably a dumb question, but are we just gonna be stuck
| on always having X.5 versions of GPT forever? If there's never an
| X.0, it feels like it's basically meaningless.
| xd1936 wrote:
| The gas costs $3.599
| dragonwriter wrote:
| There was 3, 3.5, 4, 4o, and then 4.5.
|
| There are things to complain about in that, but "its always X.5
| and never X.0" doesn't really seem to be one of them.
| JohnMakin wrote:
| Their example for 4.5 on "having a tough time":
|
| > I'm going through a tough time after failing a test > > Aw, I'm
| really sorry to hear that. Failing a test can feel pretty tough
| and discouraging, but remember, it's just one moment--not a
| reflection of who you are or your capabilities. Want to talk
| about what happened, or do you just need a distraction? I'm here
| either way.
|
| vs 4o:
|
| >I'm going through a tough time after failing a test > >I'm
| really sorry to hear that you're going through a tough time. Here
| are a few >things you might consider: > >1. Reflect and
| Understand: Try to analyze what might have gone wrong. Was it a
| >lack of preparation, time management during the test, or perhaps
| >misunderstanding the material? Understanding the root cause can
| help you address >it in the future. > >2. Seek Support: Talk to
| someone you trust--friends, family, or a school >counselor. They
| can offer a different perspective and sometimes just talking
| >about your feelings can be alleviating. > >3. Speak with Your
| Teacher or Professor: They can provide insights into where >you
| might have gone wrong and offer advice on how to improve. They
| might also >offer opportunities for make-up work or extra credit.
| > >4. Develop a Study Plan: Based on your reflections, create a
| plan to improve >your study habits. This might include setting
| regular study times, breaking down >material into manageable
| portions, or finding new study methods that work better >for you.
| > >5. Stay Positive and Forgive Yourself: Mistakes and setbacks
| are a part of life. >It's important to be kind to yourself and
| recognize that failure is a stepping >stone to success. > >6.
| Focus on the Bigger Picture: Remember that one test is just one
| part of your >educational journey. There will be many more
| opportunities to do well. > >If you need further support or
| resources, consider reaching out to educational >support services
| at your institution, or mental health resources if you're
| >feeling particularly overwhelmed. You're not alone in this, and
| things can get >better with time and effort.
|
| Is it just me or is the 4o response insanely better? I'm not the
| type of person to reach for a LLM for help about this kind of
| thing, but if I were, the 4o respond seems _vastly_ better to the
| point I 'm surprised they used that as their main "EQ" example.
| bitshiftfaced wrote:
| When people are in an emotional state, it's usually better to
| start by reacting with empathy and perspective-taking.
|
| Source: interacting with my SO.
| IMTDb wrote:
| 4o has a very strong artificial vibe. It feels a bit "autistic"
| (probably a bad analogy but couldn't find a better word to
| describe what I mean): you feel bad ? must say sorry then give
| a TODO list on how to feel better.
|
| 4.5 still feels a bit artificial but somehow also more
| emotionally connected. It removed the weird "bullet point lists
| of things to do" and focused on the emotional part; which is
| also longer than 4o
|
| If I am talking to a human I would definitely expect him/her to
| react more like 4.5 than like 4o. If the first sentence that
| comes out of their mouth after I explain them that I feel bad
| is "here is a list of things you might consider", I will find
| it strange. We can reach that point but it's usually after a
| bit more talk; human kinda need that process, and it feels like
| 4.5 understands that better than 4o.
|
| Now of course which one is "better" really depends on the
| context; what you expect of the model and how you intend to use
| is. Until now every single OpenAI update on the main series has
| always been a strict improvement over the previous model. Cost
| aside, there wasn't really any reason to keep using 3.5 when 4
| got released. This is not the case here; even assuming
| unlimited money you still might wanna select 4o in the dropdown
| sometimes instead of 4.5.
| torginus wrote:
| My 2 cents (disclaimer: I am talking out of my ass) here is why
| GPTs actually suck at fluid knowledge retrievel (which is kinda
| their main usecase, with them being used as knowledge engines) -
| they've mentioned that if you train it on 'Tom Cruise was born
| July 3, 1962', it won't be able to answer the question "Who was
| born on July 3, 1962", if you don't feed it this piece of
| information. It can't really internally corellate the information
| it has learned, unless you train it to, probably via synthethic
| data, which is what OpenAI has probably done, and that's the
| information score SimpleQA tries to measure.
|
| Probably what happened, is that in doing so, they had to scale
| either the model size or the training cost to untenable levels.
|
| In my experience, LLMs really suck at fluid knowledge retrieval
| tasks, like book recommendation - I asked GPT4 to recommend me
| some SF novels with certain characteristics, and what it spat out
| was a mix of stuff that didn't really match, and stuff that was
| really reaching - when I asked the same question on Reddit, all
| the answers were relevant and on point - so I guess there's still
| something humans are good for.
|
| Which is a shame, because I'm pretty sure relevant product
| recommendation is a many billion dollar business - after all
| that's what Google has built it's empire on.
| woah wrote:
| Perhaps you could use LLMs in a list ranking context to
| generate your scifi recommendations
| https://github.com/noperator/raink?tab=readme-ov-file
| staticman2 wrote:
| You make a good point: I think these LLM's have a strong bias
| towards recommending the most popular things in pop culture
| since they really only find the most likely tokens and report
| on that.
|
| So while they may have a chance of answering "What is this non
| mainstream novel about" they may be unable to recommend the
| novel since it's not a likely series of tokens in response to a
| request for a book recommendation.
| vel0city wrote:
| An LLM on its own isn't necessarily great for fluid knowledge
| retrieval, as in directly from its training data. But they're
| pretty good when you add RAG to it.
|
| For instance, asking Copilot "Who was born on July 3, 1962"
| gave the response:
|
| > One notable person born on July 3, 1962, is Tom Cruise, the
| famous American actor known for his roles in movies like Risky
| Business, Jerry Maguire, and Rain Man.
|
| > Are you a fan of his work?
|
| It cited this page:
|
| https://www.onthisday.com/date/1962/july/3
| jefffoster wrote:
| Does anyone have any intuition about the how reasoning improves
| based on the strength of the underlying model?
|
| I'm wondering whether this seemingly underwhelming bump on 4o
| magnifies when/if reasoning is added.
| porridgeraisin wrote:
| It is possible to understand the mechanism once you drop the
| anthropomorphisms.
|
| Each token output by an LLM involves one pass through the next-
| word predictor neural network. Each pass is a fixed amount of
| computation. Complexity theory hints to us that the problems
| which are "hard" for an LLM will need more compute than the
| ones which are "easy". Thus, the only mechanism through which
| an LLM can compute more and solve its "hard" problems is by
| outputting more tokens.
|
| You incentivise it to this end by human-grading its outputs
| ("RLHF") to prefer those where it spends time calculating
| before "locking in" to the answer. For example, you would
| prefer the output Ok let's begin... statement1
| => statement2 ... Thus, the answer is 5
|
| over The answer is 5. This is because....
|
| since in the first one, it has spent more compute before giving
| the answer. You don't in any way attempt to steer the extra
| computation in any particular direction. Instead, you simply
| reinforce preferred answers and hope that somewhere in that
| extra computation lies some useful computation.
|
| It turned out that such hope was well-placed. The DeepSeek
| R1-Zero training experiment showed us that if you apply this
| really generic form of learning (reinforcement learning)
| without _any_ examples, the model automatically starts
| outputting more and more tokens i.e "computing more".
| DeepseekMath was also a model trained directly with RL.
| Notably, the only signal given was whether the answer was right
| or not. No attention was paid to anything else. We even ignore
| the position of the answer in the sequence that we cared about
| before. This meant that it was possible to automatically grade
| the LLM without a human in the loop (since you're just checking
| answer == expected_answer). This is also why math problems were
| used.
|
| All this is to say, we get the most insight on what benefit
| "reasoning" adds by examining what happened when we applied it
| without training the model on any examples. Deepseek R1
| actually uses a few examples and then does the RL process on
| top of that, so we won't look at that.
|
| Reading the DeepseekMath paper[1], we see that the authors
| posit the following: As shown in Figure 7, RL
| enhances Maj@K's performance but not Pass@K. These
| findings indicate that RL enhances the model's overall
| performance by rendering the output distribution more
| robust, in other words, it seems that the improvement is
| attributed to boosting the correct response from TopK rather
| than the enhancement of fundamental capabilities.
|
| For context, Maj@K means that you mark the output of the LLM as
| correct only if the majority of the many outputs you sample are
| correct. Pass@K means that you mark it as correct even if just
| one of them is correct.
|
| So to answer your question, if you add an RL-based reasoning
| process to the model, it will improve simply because it will do
| more computation, of which a so-far-only-empirically-measured
| portion helps get more accurate answers on math problems. But
| outside that, it's purely subjective. If you ask me, I prefer
| claude sonnet for all coding/swe tasks over any reasoning LLM.
|
| [1] https://arxiv.org/pdf/2402.03300
| zone411 wrote:
| It significantly improves upon GPT-4o on my Extended NYT
| Connections Benchmark. 22.4 -> 33.7
| (https://github.com/lechmazur/nyt-connections).
| anotherpaulg wrote:
| GPT-4.5 Preview scored 45% on aider's polyglot coding benchmark
| [0]. OpenAI describes it as "good at creative tasks" [1], so
| perhaps it is not primarily intended for coding.
| 65% Sonnet 3.7, 32k think tokens (SOTA) 60% Sonnet 3.7, no
| thinking 48% DeepSeek V3 45% GPT 4.5 Preview <===
| 27% ChatGPT-4o 23% GPT-4o
|
| [0] https://aider.chat/docs/leaderboards/
|
| [1] https://platform.openai.com/docs/models#gpt-4-5
| doctoboggan wrote:
| I was waiting for your comment and wow... that's bad.
|
| I guess they are ceding the LLMs for coding market to
| Anthropic? I remember seeing an industry report somewhere and
| it claimed software development is the largest user of LLMs, so
| it seems weird to give up in this area.
| I_am_tiberius wrote:
| I assume they go all in "the new google" direction. Embedded
| ads coming soon I guess in the free version (chat.com).
| icemelt8 wrote:
| they are trying to copy Grok 3
| smcleod wrote:
| GPT 4.5 is insanely over price, it makes Anthropic look
| affordable!
| shshahshsusus wrote:
| brief and detailed summaries by chatgpt (4o):
|
| _Brief Summary (40-50 words)_
|
| OpenAI's GPT-4.5 is a research preview of their most advanced
| language model yet, emphasizing improved pattern recognition,
| creativity, and reduced hallucinations. It enhances unsupervised
| learning, has better emotional intelligence, and excels in
| writing, programming, and problem-solving. Available for ChatGPT
| Pro users, it also integrates into APIs for developers.
|
| _Detailed Summary (200 words)_
|
| OpenAI has introduced *GPT-4.5*, a research preview of its most
| advanced language model, focusing on *scaling unsupervised
| learning* to enhance pattern recognition, knowledge depth, and
| reliability. It surpasses previous models in *natural
| conversation, emotional intelligence (EQ), and nuanced
| understanding of user intent*, making it particularly useful for
| writing, programming, and creative tasks.
|
| GPT-4.5 benefits from *scalable training techniques* that improve
| its steerability and ability to comprehend complex prompts.
| Compared to GPT-4o, it has a *higher factual accuracy and lower
| hallucination rates*, making it more dependable across various
| domains. While it does not employ reasoning-based pre-processing
| like OpenAI o1, it complements such models by excelling in
| general intelligence.
|
| Safety improvements include *new supervision techniques*
| alongside traditional reinforcement learning from human feedback
| (RLHF). OpenAI has tested GPT-4.5 under its *Preparedness
| Framework* to ensure alignment and risk mitigation.
|
| *Availability*: GPT-4.5 is accessible to *ChatGPT Pro users*,
| rolling out to other tiers soon. Developers can also use it in
| *Chat Completions API, Assistants API, and Batch API*, with
| *function calling and vision capabilities*. However, it remains
| computationally expensive, and OpenAI is evaluating its long-term
| API availability.
|
| GPT-4.5 represents a *major step in AI model scaling*, offering
| *greater creativity, contextual awareness, and collaboration
| potential*.
| ripped_britches wrote:
| Obviously it's expensive and still I would prefer a reasoning
| model for coding.
|
| However for user facing applications like mine, this is an
| awesome step in the right direction for EQ / tone / voice.
| Obviously it will get distilled into cheaper open models very
| soon, so I'm not too worried about the price or even tokens per
| second.
| mkaic wrote:
| In a hilarious act of accidental satire, it seems that the AI-
| generated audio version of the post has a weird
| glitch/mispronunciation within the _first three words_ -- it
| struggles to say "GPT-4.5".
| advael wrote:
| It's sad that all I can think about this is that it's just
| another creep forward of the surveillance oligarchy
|
| I really used to get excited about ML in the wild and while there
| are much bigger problems right now it still makes me sad to have
| become so jaded about it
| i_love_retros wrote:
| Anyone really finding ai useful for coding?
|
| I'm finding it to make things up, get things wrong, ignore things
| I ask.
|
| Def not worried about losing my job to it.
| i_love_retros wrote:
| It gets confused if I give it 3 files - how is it going to scan
| a whole codebase and disparate systems and make correct
| changes.
|
| Pah! Don't believe the hype.
| twistslider wrote:
| I played around with Claude Code today, first time I've ever
| really been impressed by AI for coding.
|
| Tasked it with two different things, refactoring a huge
| function of around ~400 lines and creating some unit tests
| split into different files. The refactor was done flawlessly.
| The unit tests almost, only missed some imports.
|
| All I did was open it in the root of my project and prompt it
| with the function names. It's a large monolithic solution with
| a lot of subprojects. It found the functions I was talking
| about without me having to clarify anything. Cost was about $2.
| SkyPuncher wrote:
| Yes, massively.
|
| There's a learning curve to it, but it's worth literally every
| penny I spend on API calls.
|
| At worst, I'm no faster. At best, it's easily a 10x
| improvement.
|
| For me, one of the biggest benefits is talking about coding in
| natural language. It lowers my mental low and keeps me in a
| mental space where I'm more easily able to communicate with
| stakeholders holders.
| antirez wrote:
| In many ways I'm not an OpenAI fan (but I need to recognize their
| many merits). At the same time, I believe people are missing what
| they tried to do with GPT 4.5: it was needed and important to
| explore the pre-training scaling law in that direction. A gift to
| science, however selfist it could be.
| wewewedxfgdf wrote:
| GPT-2 was laugh out loud funny, rolling on the ground funny.
|
| I miss that - newer LLMs seem to have lost their sense of humor.
|
| On the other hand GPT-2's funny stories often veered into
| murdering everyone in the story and committing heinous crimes but
| that was part of the weird experience.
| kossTKR wrote:
| Totally agree, i think the gargantuan hidden pre prompts,
| censorship through reinforcement learning and whatever has
| killed most creativity.
|
| The newer models are incredible, but the tone is just soul
| sucking even when it tries to be "looser" in the later
| iterations.
| krackers wrote:
| Sydney is a glimpse at what an "unlobotomized" GPT-4 model
| would have been like.
| I_am_tiberius wrote:
| Not available in my Pro plan.
| highfrequency wrote:
| Overall take seems to be negative in the comments. But I see
| potential for a non-reasoning model that makes enough subtle
| tweaks in its tone that it is enjoyable to talk to instead of
| feeling like a summary of Wikipedia.
| Chance-Device wrote:
| And the AI stocks fell today.
|
| I'm sure it's unrelated.
| boznz wrote:
| So better than 4o but not good enough for a 5.0
| simonw wrote:
| I got gpt-4.5-preview to summarize this discussion thread so far
| (at 324 comments): hn-summary.sh 43197872 -m
| gpt-4.5-preview
|
| Using this script: https://til.simonwillison.net/llms/claude-
| hacker-news-themes...
|
| Here's the result:
| https://gist.github.com/simonw/5e9f5e94ac8840f698c280293d399...
|
| It took 25797 input tokens and 1225 input tokens, for a total
| cost (calculated using https://tools.simonwillison.net/llm-prices
| ) of $2.11! It took 154 seconds to generate.
| djhworld wrote:
| interesting summary but it's hard to gauge whether this is
| better/worse than just piping the contents into a much cheaper
| model.
| orbital-decay wrote:
| This looks like a first generation model to bootstrap future
| models from, not a competitive product at all. The knowledge
| cutoff is pretty old as well. (2023, seriously?)
|
| If they wanted to train it to have some character like Anthropic
| did with Claude 3... honestly I'm not seeing it, at least not in
| this iteration. Claude 3 was/is much much more engaging.
| GaggiX wrote:
| I imagine it will be used as a base for GPT-5 when it will be
| trained into a reasoning model, right now it probably doesn't
| make too much sense to use.
| dgfitz wrote:
| @sama, LLMs aren't going to create AGI. I realize you need to
| generate cash flow, this isn't the play.
|
| Sincerely, Me
___________________________________________________________________
(page generated 2025-02-27 23:00 UTC)