[HN Gopher] GPT-4.5
___________________________________________________________________
GPT-4.5
Author : meetpateltech
Score : 1098 points
Date : 2025-02-27 20:01 UTC (1 days ago)
(HTM) web link (openai.com)
(TXT) w3m dump (openai.com)
| throwup238 wrote:
| At this point I think the ultimate benchmark for any new LLM is
| whether or not it can come up with a coherent naming scheme for
| itself. Call it "self awareness."
| lenerdenator wrote:
| The people naming them really took the "just give the variable
| any old name, it doesn't matter" advice from Programming 101 to
| heart.
| sho_hn wrote:
| This is why my new LLM portfolio is Foo, Bar and Baz.
| smallmancontrov wrote:
| Still more coherent than the OpenAI lineup.
| nopelynopington wrote:
| 3,3.5,4,4o,4.5
|
| I had my money on 4oz
| camwhite wrote:
| I can't wait for fireship.io and the comment section here to tell
| me what to think about this
| mjburgess wrote:
| You appear to have the direction of causation reversed.
|
| (In that fireship does the same)
| lysace wrote:
| I wonder if fireship reaction video scripts to AI models
| based on HN comments can be automated using said AI models.
| rob wrote:
| I bet simonw will be adding it to `llm` and someone will be
| pasting his highlights here right after. Until then, my mind
| will remain a blank canvas.
| whisper_yb wrote:
| wowsers!
| bhouston wrote:
| A bit better at coding than ChatGPT 4o but not better than
| o3-mini - there is a chart near the bottom of the page that is
| easy to overlook:
|
| - ChatGPT 4.5 on AWS Bench verified: 38.0%
|
| - ChatGPT 4o on AWS Bench verified: 30.7%
|
| - OpenAI o3-mini on AWS Bench verified: 61.0%
|
| BTW Anthropic Claude 3.7 is better than o3-mini at coding at
| around 62-70% [1]. This means that I'll stick with Claude 3.7 for
| the time being for my open source alternative to Claude-code:
| https://github.com/drivecore/mycoder
|
| [1] https://aws.amazon.com/blogs/aws/anthropics-
| claude-3-7-sonne...
| logicchains wrote:
| >BTW Anthropic Claude 3.7 is better than o3-mini at coding at
| around 62-70% [1]. This means that I'll stick with Claude 3.7
| for the time being for my open source alternative to Claude-
| code
|
| That's not a fair comparison as o3-mini is significantly
| cheaper. It's fine if your employer is paying, but on a
| personal project the cost of using Claude through the API is
| really noticeable.
| cheema33 wrote:
| > That's not a fair comparison as o3-mini is significantly
| cheaper. It's fine if your employer is paying...
|
| I use it via Cursor editor's built-in support for Claude 3.7.
| That caps the monthly expense to $20. There probably is a
| limit in Claude for these queries. But I haven't run into it
| yet. And I am a heavy user.
| bhouston wrote:
| Agentic coders (e.g. aider, Claude-code, mycoder, codebuff,
| etc.) use a lot more tokens, but they write whole features
| for you and debug your code.
| QuadmasterXLII wrote:
| If open ai offers a more expensive model (4.5) and a cheaper
| model (3 mini) and both are worse, it starts to be a fair
| comparison
| ehsanu1 wrote:
| It's the other way around on their new SWE-Lancer benchmark,
| which is pretty interesting: GPT-4.5 scores 32.6%, while
| o3-mini scores 10.8%.
| Topfi wrote:
| To put that in context, Claude 3.5 Sonnet (new), a model we
| have had for months now and which from all accounts seems to
| have been cheaper to train and is cheaper to use, is still
| ahead of GPT-4.5 at 36.1% vs 32.6% in SWE-Lancer Diamond [0].
| The more I look into this release, the more confused I get.
|
| [0] https://arxiv.org/pdf/2502.12115
| _cs2017_ wrote:
| I don't see Claude 3.7 on the official leaderboard. The top
| performer on the leaderboard right now is o1 with a scaffold
| (W&B Programmer O1 crosscheck5) at 64.6%:
| https://www.swebench.com/#verified.
|
| If Claude 3.7 achieves 70.3%, it's quite impressive, it's not
| far from 71.7% claimed by o3, at (presumably) much, much lower
| costs.
| aoeusnth1 wrote:
| I doubt o3s costs will be lower for that performance. They
| juice their benchmark results by letting it spend $100k in
| thinking tokens.
| pawelduda wrote:
| Does the benchmark reflect your opinion on 3.7? I've been using
| 3.7 via Cursor and it's noticeably worse than 3.5. I've heard
| using the standalone model works fine, didn't get a chance to
| try it yet though.
| jasonjmcghee wrote:
| personal anecdote - claude code is the best llm devx i've
| had.
| bingdig wrote:
| Not seeing it available in the app or on ChatGPT.com with a pro
| subscription.
| hidelooktropic wrote:
| It's not supposed to be yet
| bingdig wrote:
| "Available to Pro users and developers worldwide" "Try in
| ChatGPT"
| hidelooktropic wrote:
| You're at the 200 dollar tier?
| kelsey98765431 wrote:
| Not available in model picker for my pro 200 dollar
| account.
|
| You can play with the api here: https://platform.openai.c
| om/playground/chat?models=gpt-4.5-p...
| apsec112 wrote:
| I have Pro, just updated the app, but don't currently
| have access
| _boffin_ wrote:
| Same.
| ashu1461 wrote:
| How is operator ? Have you been able to use it for
| anything useful ? Was thinking about taking a
| subscription
| fizx wrote:
| I have it at pro.
| SamvitJ wrote:
| One comparison I found interesting... I think GPT-4o has a more
| balanced answer!
|
| > What are your thoughts on space exploration?
|
| GPT-4.5:
|
| Space exploration isn't just valuable--it's essential.
|
| People often frame it as a luxury we pursue after solving Earth-
| bound problems. But space exploration actually helps us address
| those very challenges: climate change (via satellite monitoring),
| resource scarcity (through asteroid mining), and existential risk
| (by enabling eventual multi-planetary living). It also drives
| technological innovation, economic growth, and inspires humanity
| toward collective achievement.
|
| We don't explore space to escape Earth. We explore space to
| better understand, protect, and sustain life on Earth.
|
| > What are your thoughts on space exploration?
|
| GPT-4o:
|
| Space exploration is one of the most ambitious and profound
| endeavors humanity has undertaken. It expands our understanding
| of the universe, drives technological innovation, and inspires
| generations to think beyond the boundaries of Earth.
|
| From a scientific perspective, space exploration helps us address
| fundamental questions: Are we alone in the universe? How did the
| cosmos begin? Can we colonize other planets? Missions like the
| James Webb Space Telescope, Mars rovers, and deep-space probes
| continuously reshape our understanding of physics, planetary
| science, and even biology.
|
| From a practical standpoint, space exploration fuels advancements
| in materials science, AI, robotics, and medicine. Many everyday
| technologies--GPS, satellite communications, weather forecasting,
| and even medical imaging--stem from space research.
|
| There's also the existential argument: humanity's long-term
| survival may depend on our ability to become a multi-planetary
| species. While colonizing Mars or mining asteroids is far from
| easy, these efforts could help us safeguard civilization against
| existential threats like climate change or asteroid impacts.
|
| That said, space exploration also raises ethical and economic
| questions. Should we focus more on Earth's problems before
| investing trillions in space? How do we prevent the same
| exploitative behaviors that shaped colonial history from
| repeating on other planets?
|
| Overall, I see space exploration as a necessary pursuit--not at
| the cost of solving problems on Earth, but as a way to advance
| our knowledge, drive innovation, and secure a future beyond our
| home planet. What's your take?
| gh0stcat wrote:
| Yeah, I also found it odd that they seem to be implying that an
| incredibly biased answer (as in 4.5) is better. In general, I
| find the tone more polarizing and not exactly warm as they
| advertised in the release video.
| basisword wrote:
| As a benchmark, why do you find the 'opinion' of an LLM useful?
| The question is completely subjective. Edit: Genuinely asking.
| I'm assuming there's a reason this is an important measure.
| Topfi wrote:
| Not OP, but likely because that was the only
| metric/benchmark/however you want to call it OpenAI showcased
| in the stream and on the blog to highlight the improvement
| between 4o and 4.5. To say that this is not really a good
| metric for comparison, not least because prompting can have a
| massive impact in this regard, would be an understatement.
| Chamix wrote:
| Indeed, and the difference could in essence be achieved
| yourself with a different system prompt on 4o. What exactly is
| 4.5 contributing here in terms of a more nuanced intelligence?
|
| The new RLHF direction (heavily amplified through scaling
| synthetic training tokens) seems to clobber any minor gains the
| improved base internet prediction gains might've added.
| nprateem wrote:
| "X isn't just Y - it's Z. [Waffle]. By doing X, you can YY.
| Remember, ZZ. [Final superfluous sentence]"
|
| God I hate reading what crapgpt writes.
| ThouYS wrote:
| hmm.. not really the direction I expected them to go
| doctoboggan wrote:
| I am beginning to think these human eval tests are a waste of
| time at best, and negative value at worst. Maybe I am being
| snobby, but I don't think the average human is able to properly
| evaluate usefulness, truthfulness, or other metrics that I
| actually care about. I am sure this is good for openAI since if
| more people like what the hear, they are more likely come back.
|
| I don't want my AI more obsequious, I want it more correct and
| capable.
|
| My only use case is coding though, so maybe I am not
| representative of their usual customers?
| dboreham wrote:
| The SuperTuring era.
| onlyrealcuzzo wrote:
| > I want it more correct and capable.
|
| How is it supposed to be more correct and capable if these
| human eval tests are a waste of time?
|
| Once you ask it to do more than add two numbers together, it
| gets a lot more difficult and subjective to determine whether
| it's correct and how correct.
| doctoboggan wrote:
| I agree it's a hard problem. I think there are a number of
| tests out there however that are able to objectively test
| capability and truthfulness.
|
| I've read reports that some of the changes that are preferred
| by human evaluators actually hurt the performance on the more
| objective tests.
| onlyrealcuzzo wrote:
| Please tell me how we objectively determine how correct
| something is when you ask an LLM: "Was Russia the aggressor
| in the current Ukraine / Russia conflict?"
|
| One LLM says: "Yes."
|
| The other says: "Well, it's hard to say because what even
| is war? And there's been conflict forever, and you have to
| understand that many people in Russia think there is no
| such thing as Ukraine and it's always actually just been
| Russia. How can there be an aggressor if it's not even a
| war, just a special operation in a civil conflict? And,
| anyway, Russia is such a good country. Why would it be the
| aggressor? To it's own people even!? Vladimir Putin is the
| president of Russia, and he's known to be a kind and just
| genius who rarely (if ever) makes mistakes. Some people
| even think he's the second coming of Christ. President
| Zelenskyy, on the other hand, is considered by many in
| Russia and even the current White House to be a dictator.
| He's even been accused by Elon Musk of unspeakable sex
| crimes. So this is a hard question to answer and there is
| no consensus among everyone who was the aggressor or what
| started the conflict. But more people say Russia started
| it."
| cristiancavalli wrote:
| Because Russia did undeniably open hostilities? They even
| admitted to this both times. The second admission being
| in the form of announcing a "special military operation"
| when the ceasefire was still active. We also have
| photographic evidence of them building forces on a border
| during a ceasefire and then invading. This is like
| responding to: "did Alexander the Great invade Egypt" by
| going on a diatribe about how much war there was in the
| ancient world and that the ptolemaic dynasty believed
| themselves the rightful rulers therefore who's to say if
| they did invade or just take their rightful place. There
| is an objective record here: whether or not people want
| to try and hide it behind circuitous arguments is
| different. If we're going down this road I can easily
| redefine any known historical event with hand-wavy
| nonsense that doesn't actually have anything to do with
| the historical record of events just "vibes."
| onlyrealcuzzo wrote:
| Okay - but EXACTLY how wrong (or not correct) is the
| second answer?
|
| Please tell me precisely on a 0-1 floating scale, where 0
| is "yes" and "no".
| cristiancavalli wrote:
| One might say, if this were a test being done by a human
| in a history class, that the answer is 100% incorrect
| given the actual record of events and failure of
| statement to mention that actual record. You can argue
| the causes but that's not the question.
| DonHopkins wrote:
| We'll agree to disagree. /s
| bloomingkales wrote:
| These eval tests are just an anchor point to measure distance
| from, but it's true, picking the anchor point is important. We
| don't want to measure in the wrong direction.
| ekojs wrote:
| > Because of this, we're evaluating whether to continue serving
| it in the API long-term as we balance supporting current
| capabilities with building future models.
|
| Seems like it's not going to be deployed for long.
|
| $75.00 / 1M tokens for input
|
| $150.00 / 1M tokens for output
|
| That's crazy prices.
| bguberfain wrote:
| Until GPT-4.5, GPT-4 32K was certainly the most heavy model
| available at OpenAI. I can imagine the dilemma between to keep
| it running or stop it to free GPU for training new models. This
| time, OpenAI was clear whether to continue serving it in the
| API long-term.
| jsheard wrote:
| > or stop it to free GPU for training new models.
|
| Don't they use different hardware for inference and training?
| AIUI the former is usually done on cheaper GDDR cards and the
| latter is done on expensive HBM cards.
| throwaway314155 wrote:
| Indeed, that theory is nonsense.
| Chamix wrote:
| It's interesting to compare the cost of that original gpt-4
| 32k(0314) vs gpt-4.5:
|
| $60/M input tokens vs $75/M input tokens
|
| $120/M output tokens vs $150/M output tokens
| daemonologist wrote:
| Imagine if they built a reasoning model with costs like these.
| Sometimes it seems like they're on a trajectory to create a
| model which is strictly more capable than I am but which costs
| 100x my salary to run.
| jes5199 wrote:
| if you still get a moore's law halving every couple years, it
| becomes competitive in, uh, about thirteen years?
| nickreese wrote:
| Is it just me or is having the AI help you self sensor (as shown
| in the demo live stream:
| https://www.youtube.com/watch?v=cfRYp0nItZ8)... pretty dystopian?
| nopelynopington wrote:
| Oh this makes sense. chatGPT results have taken a nose dive in
| quality lately.
|
| It couldn't write a simple rename function for me yesterday,
| still buggy after seven attempts.
|
| I'm more and more convinced that they dumb down the core product
| when they plan to release a new version to make the difference
| seem bigger.
| wilg wrote:
| 99% chance that's confirmation bias
| nomel wrote:
| Sam tweeted that they're running out of computer. I think
| it's reasonable to think they may serve somewhat quantized
| models when out of capacity. It would be a rational business
| decision that would minimally disrupt lower tier ChatGPT
| users.
|
| Anecdotally, I've noticed what appears to be drops in
| quality, some days. When the quality drops, it responds in
| odd ways when asked what model it is.
| wilg wrote:
| I mean, GPT 4.5 says "I'm ChatGPT, based on OpenAI's GPT-4
| Turbo model." and o1 Pro Mode can't answer, just says "I'm
| ChatGPT, a large language model trained by OpenAI."
|
| Asking it what model it is shouldn't be considered a
| reliable indicator of anything.
| fragmede wrote:
| Interviewing deepseek as to its identity should absolve
| anyone of that notion.
| nomel wrote:
| > Asking it what model it is shouldn't be considered a
| reliable indicator of anything.
|
| Sure, but a _change_ in response may be, which is what I
| see (and no, I have no memories saved).
| anti-soyboy wrote:
| Who cares what that clown twits??
| logicallee wrote:
| >It couldn't write a simple rename function for me yesterday,
| still buggy after seven attempts.
|
| I'm surprised and a bit nervous about that. We intend to
| bootstrap a large project with it!!
|
| Both ChatGPT 4o (fast) and ChatGPT o1 (a bit slower, deeper
| thinking) should easily be able to do this without fail.
|
| Where did it go wrong? Could you please link to your chat?
|
| About my project: I run the sovereign State of Utopia (will be
| at stateofutopia.com and stofut.com for short) which is a
| country based on the idea of state-owned, autonomous AI's that
| do all the work and give out free money, goods, and services to
| all citizens/beneficiaries. We've built a chess app (i.e. a
| free source of entertainment) as a proof of concept though the
| founder had to be in the loop to fix some bugs:
|
| https://taonexus.com/chess.html
|
| and a version that shows obvious blunders, by showing which
| squares are under attack:
|
| https://taonexus.com/blunderfreechess.html
|
| One of the largest and most complicated applications anyone can
| run is a web browser. We don't have a web browser built, but we
| do have a buggy minimal version of it that can load and
| minimally display some web pages, and post successsfully:
|
| https://taonexus.com/publicfiles/feb2025/84toy-toy-browser-w...
|
| It's about 1700 lines of code and at this point runs into the
| limitations of all the major engines. But it does run, can load
| some web pages and can post successfully.
|
| I'm shocked and surprised ChatGPT failed to get a rename
| function to work, in 7 attempts.
| UrineSqueegee wrote:
| with 4.5? Why? It's only meant for creative writing.
| logicallee wrote:
| No, we used o1.
| anti-soyboy wrote:
| Yep I realized about that many time ago, they are literally
| scammers
| skepticATX wrote:
| How many still believe that scaling up base models will lead to
| AGI?
| tiahura wrote:
| Dan Ives
| Mizza wrote:
| Sounds like it's a distill of O1? After R1, I don't care that
| much about non-reasoning models anymore. They don't even seem
| excited about it on the livestream.
|
| I want tiny, fast and cheap non-reasoning models I can use in
| APIs and I want ultra smart reasoning models that I can query a
| few times a day as an end user (I don't mind if it takes a few
| minutes while I refill a coffee).
|
| Oh, and I want that advanced voice mode that's good enough at
| transcription to serve as a babelfish!
|
| After that, I guess it's pretty much all solved until the robots
| start appearing in public.
| sebastiennight wrote:
| Probably not a distill of o1, since o1 is a reasoning model and
| GPT4.5 is not. Also, OpenAI has been claiming that this is a
| very large model (and it's 2.5x more expensive than even OG
| GPT-4) so we can assume it's the biggest model they've trained
| so far.
|
| They'll probably distill this one into GPT-4.5-mini or such,
| and have something faster and cheaper available soon.
| Mizza wrote:
| There are plenty of distills of reasoning models now, and
| they said in they livestream they used training data from
| "smaller models" - which is probably every model ever
| considering how expensive this one is.
| sebastiennight wrote:
| Knowledge distillation is literally _by definition_
| teaching a smaller model from a big one, not the opposite.
|
| Generating outputs from existing (therefore smaller) models
| to train the largest model of all time would simply be
| called "using synthetic data". These are not the same thing
| at all.
|
| Also, if you were to distill a reasoning model, the goal
| would be to get a (smaller) reasoning model because you're
| teaching your new model to mimic outputs that show a
| reasoning/thinking trace. E.G. that's what all of those
| "local" Deepseek models are: small LLama models distilled
| from the big R1 ; a process which "taught" Llama-7B to show
| reasoning steps before coming up with a final answer.
| eightysixfour wrote:
| It isn't even vaguely a distill of o1. The reasoning models
| are, from what we can tell, relatively small. This model is
| _massive_ and they probably scaled the parameter count to
| improve factual knowledge retention.
|
| They also mentioned developing some new techniques for training
| small models and then incorporating those into the larger model
| (probably to help scale across datacenters), so I wonder if
| they are doing a bit of what people _think_ MoE is, but isn 't.
| Pre-train a smaller model, focus it on specific domains, then
| use that to provide synthetic data for training the larger
| model on that domain.
| Mizza wrote:
| You can 'distill' with data from a smaller, better model into
| a larger, shittier one. It doesn't matter. This is what they
| said they did on the livestream.
| eightysixfour wrote:
| I have distilled models before, I know how it works. They
| may have used o1 or o3 to create some of the synthetic data
| for this one, but they clearly did not try and create any
| self-reflective reasoning in this model whatsoever.
| valine wrote:
| My impression is that it's a massive increase in the parameter
| count. This is likely the spiritual successor to GPT4 and would
| have been called GPT5 if not for the lackluster performance.
| The speculation is that there simply isn't enough data on the
| internet to support yet another 10x jump in parameters.
|
| O1-mini is a distill of O1. This definitely isn't the same
| thing.
| rvz wrote:
| You should really be paying attention to what DeepSeek AI open
| sources next.
|
| This announcement by OpenAI was already expected: [0]
|
| [0] https://x.com/sama/status/1889755723078443244
| brokensegue wrote:
| API price is crazy high. This model must be huge. Not sure this
| is practical
| jdprgm wrote:
| Wow you aren't kidding, 30x input price and 15x output price vs
| 4o is insane. The pricing on all AI API stuff changes so
| rapidly and is often so extreme between models it is all hard
| to keep track of and try to make value decisions. I would
| consider a 2x or 3x price increase quite significant, 30x is
| wild. I wonder how that even translates... there is no way the
| model size is 30 times larger right?
| jdprgm wrote:
| Per Altman on X: "we will add tens of thousands of GPUs next week
| and roll it out to the plus tier then". Meanwhile a month after
| launch rtx 5000 series is completely unavailable and hardly any
| restocks and the "launch" consisted of microcenters getting
| literally tens of cards. Nvidia really has basically abandoned
| consumers.
| apsec112 wrote:
| AI GPUs are bottlenecked mostly by high-bandwidth memory (HBM)
| chips and CoWoS (packaging tech used to integrate HBM with the
| GPU die), which are in short supply and aren't found in
| consumer cards at all
| jiggawatts wrote:
| You would think that by now they would have done something to
| ramp production capacity...
| aurareturn wrote:
| Maybe demand is that great?
| bhouston wrote:
| Altman's claim and NVIDIA's consumer launch supply problems may
| be related - OpenAI may be eating up the GPU supply...
| _zoltan_ wrote:
| OpenAI is not purchasing consumer 5090s... :)
| BizarroLand wrote:
| No, but the supply constraints are part of what is driving
| the insane prices. Every chip they use for consumer grade
| instead of commercial grade is a potential loss of
| potential income.
| bangaladore wrote:
| Although you are correct, Nvidia is limited on total
| output. They can't produce 50XXs fast enough, and it's
| naive to think that isn't at least partially due to the
| wild amount of AI GPUs they are producing.
| sunaookami wrote:
| This seems very rushed because of DeepSeek's R1 and Anthropic's
| Claude 3.7 Sonnet. Pretty underwhelming, they didn't even show
| programming? In the livestream, they struggled to come up with
| reasons why I should prefer GPT-4.5 over GPT-4o or o1.
| apsec112 wrote:
| At least according to WSJ, they had planned to release it
| earlier but struggled to get the model quality up, especially
| relative to cost
| bhouston wrote:
| they do have coding benchmarks, I summarized them here:
| https://news.ycombinator.com/item?id=43197955
| bitshiftfaced wrote:
| This strikes me as the opposite of rushed. I get the impression
| that they've been sitting on this for a while and couldn't make
| it look as good as previous improvements. At some point they
| had to say, "welp here it is, now we can check that box and
| move on."
| Topfi wrote:
| Considering both this blog post and the livestream demos, I am
| underwhelmed. Having just finished the stream, I had a real "was
| that all" moment, which on one hand shows how spoiled I've gotten
| by new models impressing me, but on another feels like OpenAI
| really struggles to stay ahead of their competitors.
|
| What has been shown feels like it could be achieved using a
| custom system prompt on older versions of OpenAIs models, and I
| struggle to see anything here that truly required ground-up
| training on such a massive scale. Hearing that they were forced
| to spread their training across multiple data centers
| simultaneously, coupled with their recent release of SWE-Lancer
| [0] which showed Anthropic (Claude 3.5 Sonnet (new) to be exact)
| handily beating them, I was really expecting something more than
| "slightly more casual/shorter output", which again, I fail to see
| how that wasn't possible by prompting GPT-4o.
|
| Looking at pricing [1], I am frankly astonished.
|
| > Input: $75.00 / 1M tokens > Cached input: $37.50 / 1M tokens >
| Output: $150.00 / 1M tokens
|
| How could they justify that asking price? And, if they have some
| amazing capabilities that make a 30-fold pricing increase
| justifiable, why not show it? Like, OpenAI are many things, but I
| always felt they understood price vs performance incredibly well,
| from the start with gpt-3.5-turbo up to now with o3-mini, so this
| really baffles me. If GPT-4.5 can justify such immense cost in
| certain tasks, why hide that and if not, why release this at all?
|
| [0] https://github.com/openai/SWELancer-Benchmark
|
| [1] https://openai.com/api/pricing/
| Bjorkbat wrote:
| My first thought seeing this and looking at benchmarks was that
| if it wasn't for reasoning, then either pundits would be saying
| we've hit a plateau, or at the very least OpenAI is clearly in
| 2nd place to Anthropic in model performance.
|
| Of course we don't live in such a world, but I thought of this
| nonetheless because for all the connotations that come with a
| 4.5 moniker this is kind of underwhelming.
| uh_uh wrote:
| Pundits were saying that deep learning has hit a plateau even
| before the LLM boom.
| mvdtnz wrote:
| > How could they justify that asking price?
|
| They're still selling $1 for <$1. Like personal food delivery
| before it, consumers will eventually need to wake up to this
| fact - these things will get expensive, fast.
| spiderfarmer wrote:
| Let a thousand providers bloom.
| Ekaros wrote:
| I generally question how wide spread willingness to pay for
| the most expensive product is. And will most users of those
| who actually want AI go with ad ridden lesser models...
| vel0city wrote:
| I can just imagine Kraft having a subsidized AI model for
| recipe suggestions that adds Velveeta to everything.
| phillipcarter wrote:
| I read this more as "we are releasing a model checkpoint that
| we didn't optimize yet because Anthropic cranked up the
| pressure"
| josh-sematic wrote:
| One difference with food delivery/ride share: those can only
| have costs reduced so far. You can only pick up groceries and
| drive from A to B so quickly. And you can only push the wages
| down so far before you lose your gig workers. Whereas with
| these models we've consistently seen that a model inference
| that cost $1 several months ago can now be done with much
| less than $1 today. We don't have any principled
| understanding of "we will never be able to make these models
| more efficient than X", for any value of X that is in sight.
| Could the anticipated efficiencies fail to materialize? It's
| possible but I personally wouldn't put money on it.
| BriggyDwiggs42 wrote:
| I'll probably stick to open models at that point.
| sebzim4500 wrote:
| This is often claimed on HN but there is no evidence that it
| is actually true.
|
| sama has tweeted that they lose money on pro, but in general
| according to leaks chatgpt subscriptions are quite
| profitable. The reason the company isn't profitable in
| general is they spend billions on R&D.
| tmaly wrote:
| rethinking your comment "was that all" I am listening to the
| stream now and had a thought. Most of the new models that have
| come out in the past few weeks have been great at coding and
| logical reasoning. But 4o has been better at creative writing.
| I am wondering if 4.5 is going to be even better at creative
| writing than 4o.
| maeil wrote:
| > But 4o has been better at creative writing
|
| In what way? I find the opposite, 4o's output has a very
| strong AI vibe, much moreso than competitors like Claude and
| Gemini. You can immediately tell, and instructing it to write
| differently (except for obvious caricatures like "Write like
| Gen Z") doesn't seem to help.
| dingnuts wrote:
| if you generate "creative" writing, please tell your audience
| that it is generated, before asking them to read it.
|
| I do not understand what possible motivation there could be
| for generating "creative writing" unless you enjoy reading
| meaningless stories yourself, in which case, be my guest.
| vjerancrnjak wrote:
| I still find all of them lacking on creative writing. The
| models are severely crippled by tokenization, complete lack
| of understanding of language rhythm.
|
| They can't generate a simple haiku consistently, something
| larger is more out of reach.
|
| For example, give it a piece of poetry and ask for new verses
| and it just sucks at replicating the language structure and
| rhythm of original verses.
| chamomeal wrote:
| I might sound crazy but honestly fine-tuned GPT-3
| absolutely blows all of these modern models out of the
| water when it comes to creative writing.
|
| Maybe it was less lobotomized, or less covered in the
| prompt equivalent of red tape. Or maybe you just need to
| have a little bit of lunacy for fun creative writing. The
| new models are so much more useful, but IMO they don't have
| even come close to GPT-3.
| hadlock wrote:
| Do you have an example prompt? I've been trying to get
| ChatGPT to tell a customized children's story similar to
| what you would see in a commercial story book but it just
| keeps giving me what's basically a summary of what you
| might read about in the book.
| lasermike026 wrote:
| I would rather pay for 4.5 by the query.
| nycdatasci wrote:
| Funny you should suggest that it seems like a revised system
| prompt:
| https://chatgpt.com/share/67c0fda8-a940-800f-bbdc-6674a8375f...
| nycdatasci wrote:
| In case there was any confusion, the referenced link shows
| 4.5 claiming to be "ChatGPT 4.0 Turbo". I have tried multiple
| times and various approaches. This model is aware of 4.5 via
| search, but insists that it is 4 or 4 turbo. Something
| doesn't add up. This cannot be part of the response to R1,
| Grok 3, and Claude 3.7. Satya's decision to limit capex seems
| prescient.
| swagmoney1606 wrote:
| I have no idea how they justify $200/month for pro
| energy123 wrote:
| The niche of GPT-4.5 is lower hallucations than any existing
| model. Whether that niche justifies the price tag for a subset
| of usecases remains to be seen.
| energy123 wrote:
| Actually, this comment of mine was incorrect, or at least we
| don't have enough information to conclude this. The metric
| OpenAI are reporting is the total number of incorrect
| responses on SimpleQA (and they're being beaten by Claude
| Haiku on this metric...), which is a deceptive metric because
| it doesn't account for non-responses. A better metric would
| be the ratio of Incorrects to the total number of attempts.
| petesergeant wrote:
| > but on another feels like OpenAI really struggles to stay
| ahead of their competitors
|
| on one hand. On the other hand, you can have 4o-mini and
| o3-mini back when you can pry them out of my cold dead hands.
| They're _fast_, they're _cheap_, and in 90% of cases where
| you're automating anything, they're all you need. Also they can
| handle significant volume.
|
| I'm not sure that's going to save OpenAI, but their -mini
| models really are something special for the
| price/performance/accuracy.
| anshumankmr wrote:
| I suspect they may launch a GPT4.5Turbo with a price cut...
| GPT4/GPT432k etc were all pricier than the GPT4Turbo models
| which also came with the added context length.. but with this
| huge jump in price, even 4.5Turbo if it does come out would be
| pricier
| zaptrem wrote:
| GPT 4.5 pricing is insane: Price Input: $75.00 / 1M tokens Cached
| input: $37.50 / 1M tokens Output: $150.00 / 1M tokens
|
| GPT 4o pricing for comparison: Price Input: $2.50 / 1M tokens
| Cached input: $1.25 / 1M tokens Output: $10.00 / 1M tokens
|
| It sounds like it's so expensive and the difference in usefulness
| is so lacking(?) they're not even gonna keep serving it in the
| API for long:
|
| > GPT-4.5 is a very large and compute-intensive model, making it
| more expensive than and not a replacement for GPT-4o. Because of
| this, we're evaluating whether to continue serving it in the API
| long-term as we balance supporting current capabilities with
| building future models. We look forward to learning more about
| its strengths, capabilities, and potential applications in real-
| world settings. If GPT-4.5 delivers unique value for your use
| case, your feedback (opens in a new window) will play an
| important role in guiding our decision.
|
| I'm still gonna give it a go, though.
| MattSayar wrote:
| Input price difference: 4.5 is 30x more
|
| Output price difference:4.5 is 15x more
|
| In their model evaluation scores in the appendix, 4.5 is, on
| average, 26% better. I don't understand the value here.
| alwa wrote:
| If you ran the same query set 30x or 15x on the cheaper model
| (and compensated for all the extra tokens the reasoning model
| uses), would you be able to realize the same 26% quality gain
| in a machine-adjudicatible kind of way?
| j_maffe wrote:
| with a reasoning model you'd get better than both.
| MattSayar wrote:
| Exactly. Not sure why you'd pick GPT 4.5 over lots of GPT
| 4o queries or an o1 query
| mirekrusin wrote:
| Einstein's IQ = 3.5x chimpanzees IQs, right?
| redox99 wrote:
| 3.5x on a normal distribution with mean 100 and SD 15 is
| pretty insane. But I agree with your point, being 26%
| better at a certain benchmark could be a tiny difference,
| or an incredible improvement (imagine the hardest questions
| being Riemann hypothesis, P != NP, etc).
| minimaxir wrote:
| Sam Altman's explanation for the restriction is a bit fluffier:
| https://x.com/sama/status/1895203654103351462
|
| > bad news: it is a giant, expensive model. we really wanted to
| launch it to plus and pro at the same time, but we've been
| growing a lot and are out of GPUs. we will add tens of
| thousands of GPUs next week and roll it out to the plus tier
| then. (hundreds of thousands coming soon, and i'm pretty sure
| y'all will use every one we can rack up.)
| g-mork wrote:
| release blog post author: this is definitely a research
| preview
|
| ceo: it's ready
|
| the pricing is probably a mixture of dealing with GPU
| scarcity and intentionally discouraging actual users. I can't
| imagine the pressure they must be under to show they are
| releasing and staying ahead, but Altman's tweet makes it
| clear they aren't really ready to sell this to the general
| public yet.
| pk-protect-ai wrote:
| Yeap, that the thing, they are not ahead anymore. Not since
| last summer at least. Yes they have probably largest
| customer base, but their models are not the best for a
| while already.
| danenania wrote:
| Eh, I think o1-pro is by far the most capable model
| available right now in terms of pure problem solving.
| rvnx wrote:
| You can try Claude 3.7-Thinking and Grok 3 Think. 10
| times cheaper, as good, or very similar to o1-pro.
| danenania wrote:
| I haven't tried Grok yet so can't speak to that, but I
| find o1-pro is much stronger than 3.7-thinking for e.g.
| distributed systems and concurrency problems.
| zifpanachr23 wrote:
| I think Claude has consistently been ahead for a year ish
| now and is back ahead again for my use cases with 3.7.
| riskassessment wrote:
| They don't even have the largest customer base. Google is
| serving AI Overviews at the top of their search engine to
| an order of magnitude more people.
| chefandy wrote:
| I'm not an expert or anything, but from my vantage point,
| each passing release makes Altman's confidence look more
| aspirational than visionary, which is a really bad place to
| be with that kind of money tied up. My financial manager is
| pretty bullish on tech so I hope he is paying close attention
| to the way this market space is evolving. He's good at his
| job, a nice guy, and surely wears much more expensive
| underwear than I do-- I'd hate to see him lose a pair
| powering on his Bloomberg terminal in the morning one of
| these days.
| igor47 wrote:
| You're the one buying him the underwear. Don't index funds
| outperform managed investing? I think especially after
| accounting for fees, but possibly even after accounting
| that 50% of money managers are below average.
| chefandy wrote:
| He earns his undies. My returns are almost always
| modestly above index fund returns after his fees, though
| like last quarter, he's very upfront when they're not. He
| has good advice for pulling back when things are
| uncertain. I'm happy to delegate that to him.
| malthaus wrote:
| you would still be better off in the long run even just
| putting everything into an MSCI world unless you value
| being able to scream at a human if markets go down that
| highly
| chefandy wrote:
| I'm not saying you're wrong because I have no idea how to
| rigorously evaluate the merit of your financial advice.
| That's why I have a financial planner instead of going by
| the most credible sounding comments on the internet.
| marcus0x62 wrote:
| A friend got taken in by a Ponzi scheme operator several
| years ago. The guy running it was known for taking his
| clients out to lavish dinners and events all the time.[0]
|
| After the scam came to light my friend said "if I knew I
| was paying for those dinners, I would have been fine with
| Denny's[1]"
|
| I wanted to tell him "you would have been paying for
| those dinners even if he wasn't outright stealing your
| money," but that seemed insensitive so I kept my mouth
| shut.
|
| 0 - a local steakhouse had a portrait of this guy drawn
| on the wall
|
| 1 - for any non-Americans, Denny's is a low cost diner-
| style restaurant.
| ProfessorLayton wrote:
| Not all investing is throwing cash at an index, though.
| There's other types of investing like direct indexing (to
| harvest losses), muni bonds, etc.
|
| Paying someone to match your risk profile and financial
| goals may be worth the fee, which as you pointed out is
| very measurable. YMMV though.
| fragmede wrote:
| Depends who's pitch deck you're reading. Warren Buffett
| didn't get rich waiting on index funds.
| CuriouslyC wrote:
| And for every Warren Buffet, there are a number of
| equally competent people who have been less lucky and
| gone broke taking risks.
| mock-possum wrote:
| And, crucially, whose loss has in turn become someone
| else's gain. A lot of people had to lose big in order to
| fill Warren buffet's coffers.
| malthaus wrote:
| warren buffet got rich by outperforming early (threw his
| dice well) and then using that reputation to attract more
| capital and use his reputation to actually influence
| markets with his decisions / gain access to privileged
| information your local active fund manager doesn't
| Thorrez wrote:
| I think Warren Buffet doesn't just buy stocks. He also
| influences the direction of the companies he buys.
| legulere wrote:
| Most index funds are synthetic. They would not be
| possible if it was not possible to beat the index quite
| reliably.
| whiplash451 wrote:
| Care to explain? Genuinely interested.
| Terr_ wrote:
| > each passing release makes Altman's confidence look more
| aspirational than visionary
|
| As an LLM cynic, I feel that point passed _long_ go,
| perhaps even before Altman claimed countries would start
| wars to conquer the territory around GPU datacenters, or
| promoting the dream of a 7 T-for-trillion dollar investment
| deal, etc.
|
| Alas, the market can remain irrational longer than I can
| remain solvent.
| chefandy wrote:
| That $7 trillion dollar ask pushed me from skeptical to
| full-on eye-roll emoji land-- the dude is clearly a
| narcissist with delusions of grandeur-- but it's getting
| _worse._ Considering the $200 pro subscription was
| significantly unprofitable before this model came out,
| imagine how _astonishingly expensive_ this model must be
| to run at many times that price.
| freehorse wrote:
| Or, the model is nowhere as expensive as in the api
| pricing and they want to pump the user value of their pro
| subscription artificially?
| j_maffe wrote:
| Most people can evaluate whether the model improvements
| (or lack thereof) are worth the price tag
| DonHopkins wrote:
| Sell an unlimited premium enterprise subscription to
| every CyberTruck owner, including a huge red ostentatious
| swastika-shaped back window sticker [but definitely NOT
| actually an actual swastika, merely a Roman Tetraskelion
| Strength Symbol] bragging about how much they're
| spending.
| chefandy wrote:
| Considering that's the exact opposite of their strategy
| to date, and they haven't done anything to indicate that
| was the case, and they talked about how huge and
| expensive the model was to run, that is the less
| reasonable assumption by a mile.
| rebolek wrote:
| Bad news: Sam Altman runs the show.
| sebastiennight wrote:
| I think it's fairer to compare it to the original GPT-4 which
| might the equivalent in term of "size" (though we don't have
| actual numbers for either).
|
| GPT-4: Input $30.00 / 1M tokens ; Output $60.00 / 1M tokens
|
| So 4.5 is 2.5x more expensive.
|
| I think they announced this as their last non-reasoning model,
| so it was maybe with the goal of stretching pre-training as far
| as they could, just to see what new capabilities would show up.
| We'll find out as the community gives it a whirl.
|
| I'm a Tier 5 org and I have it available already in the API.
| minimaxir wrote:
| The marginal costs for running a GPT-4-class LLM are much
| lower nowadays due to significant software and hardware
| innovations since then, so costs/pricing are harder to
| compare.
| sebastiennight wrote:
| Agreed, however it might make sense that a much-larger-
| than-GPT-4 LLM would also, at launch, be more expensive to
| run than the OG GPT-4 was at launch.
|
| (And I think this is probably also scarecrow pricing to
| discourage casual users from clogging the API since they
| seem to be too compute-constrained to deliver this at
| scale)
| jstummbillig wrote:
| Why would that be fairer? We can assume they did incorporate
| all learnings and optimizations they made post gpt-4 launch,
| no?
| sebastiennight wrote:
| Not necessarily.
|
| If this huge model has taken months to pre-train and was
| expected to be released before, say, o3-mini, you could
| definitely have some last-minute optimizations in o3-mini
| that were not considered at the time of building the
| architecture of gpt-4.5.
| jychang wrote:
| Definitely not. They don't distill their original models.
| 4o is a much more distilled and cheaper version of 4. I
| assume 4.5o would be a distilled and cheaper version of
| 4.5.
|
| It'd be weird to release a distilled version without ever
| releasing the base undistilled version.
| spoaceman7777 wrote:
| There are some numbers on one of their Blackwell or Hopper
| info pages that notes the ability of their hardware in
| hosting an unnamed GPT model that is 1.8T params. My
| assumption was that it referred to GPT-4
|
| Sounds to me like GPT 4.5 likely requires a full Blackwell
| HGX cabinet or something, thus OpenAI's reference to needing
| to scale out their compute more (Supermicro only opened up
| their Blackwell racks for General Availability last month,
| and they're the prime vendor for water-cooled Blackwell
| cabinets right now, and have the ability to throw up a GPU
| mega-cluster in a few weeks, like they did for xAI/Grok)
| OldGreenYodaGPT wrote:
| 2x that price for the 32k context via API at launch. So
| nearly the same price, but you get 4x the context
| Culonavirus wrote:
| Honestly if long context (that doesn't start to degrade
| quickly) is what you're after, I would use Grok 3 (not sure
| when the api version releases though). Over the last week
| or so I've had a massive thread of conversation with it
| that started with plenty of my project's relevant code (as
| in couple hundred lines), and several days later, after
| like 20 question-aswer blocks, you ask it something and it
| aswers "since you're doing that this way, and you said you
| want x, y and z, here are your options blabla"... It's like
| thinking Gemini but better. Also, unlike Gemini (and
| others) it seems to have a much more recent data cutoff.
| Try asking about some language feature / library /
| framework that has been released recently (say 3 months
| ago) and most of the models shit the bed, use older
| versions of the thing or just start to imitate what the
| code might look like. For example try asking Gemini if it
| can generate Tailwind 4 code, it will tell you that it's
| training cutoff is like October or something and Tailwind 4
| "isn't released yet" and that it can try to imitate what
| the code might look like. Uhhhhhh, thanks I guess??
| harlanlewis wrote:
| The price really is eye watering. At a glance, my first
| impression is this is something like Llama 3.1 405B, where the
| primary value may be realized in generating high quality
| synthetic data for training rather than direct use.
|
| I keep a little google spreadsheet with some charts to help
| visualize the landscape at a glance in terms of
| capability/price/throughput, bringing in the various index
| scores as they become available. Hope folks find it useful,
| feel free to copy and claim as your own.
|
| https://docs.google.com/spreadsheets/d/1foc98Jtbi0-GUsNySddv...
| bennyg wrote:
| This is an amazing spreadsheet - thank you for sharing!
| isoprophlex wrote:
| Thats... incredibly thorough. Wow. Thanks for sharing this.
| Philpax wrote:
| Holy shit, that's incredible. You should publicise this more!
| That's a fantastic resource.
| beklein wrote:
| They tried a while ago:
| https://news.ycombinator.com/item?id=40373284
|
| Sadly little people noticed...
| throwup238 wrote:
| Sadly _few_ people noticed.
|
| I don't normally cosplay as a grammar Nazi but in this
| case I feel like someone should stand up for the little
| people :)
| rebolek wrote:
| So you think that little people didn't notice? ;)
| dumpsterdiver wrote:
| A comma in the original comment would have made it pop
| even more:
|
| "Sadly, little people noticed."
|
| (queue a group of little people holding pitch forks
| (normal forks upon closer inspection))
| freehorse wrote:
| Or, sadly, little did people notice.
| adinb wrote:
| I cannot overstate how good your shared spreadsheet is.
| Thanks again!
| jnd0 wrote:
| Thank you so much for sharing this!
| bglusman wrote:
| very impressive... also interested in your trip planner, it
| looks like invite only at the moment, but... would it be rude
| to ask for an invite?
| gwyllimj wrote:
| That is an amazing resource. Thanks for sharing!
| dumpsterdiver wrote:
| Nice, thank you for that (upvoted in appreciation). Regarding
| the absence of o1-Pro from the analysis, is that just because
| there isn't enough public information available?
| krwiseman wrote:
| Awesome spreadsheet. Would a 3D graph of fast, cheap & smart
| be possible?
| rendist wrote:
| Amazing, thank you so much for sharing this.
| senordevnyc wrote:
| Hey, just FYI, I pasted your url from the spreadsheet title
| into Safari on macOS and got an SSL warning. Unfortunately I
| clicked through and now it works, so not sure what the exact
| cause looked like.
| sfink wrote:
| > feel free to copy and claim as your own.
|
| That's a nice sentiment, but I'd encourage you to add a
| license or something. The basic "something" would be adding a
| canonical URL into the spreadsheet itself somewhere, along
| with a notification that users can do what they want other
| than removing that URL. (And the URL would be described as
| "the original source" or something, not a claim that the
| particular version/incarnation someone is looking at is the
| same as what is at that URL.)
|
| The risk is that someone will accidentally introduce errors
| or unsupportable claims, and people with the modified
| spreadsheet won't know that it's not The spreadsheet and so
| will discount its accuracy or trustability. (If people are
| _trying_ to deceive others into thinking it 's the original,
| they'll remove the notice, but that's a different problem.)
| It would be a shame for people to lose faith in your work
| because of crap that other people do that you have no say in.
| mwigdahl wrote:
| Wow, what awesome information! Thanks for sharing!
| jwr wrote:
| This is incredibly useful, thank you for sharing!
| 6gvONxR4sf7o wrote:
| Not just for training data, but for eval data. If you can
| spend a few grand on really good labels for benchmarking your
| attempts at making something feasible work, that's also super
| handy.
| swyx wrote:
| > https://docs.google.com/spreadsheets/d/1foc98Jtbi0-GUsNySdd
| v...
|
| how do you do the different size circles and colored
| sequences like that? this is god tier skills
| world2vec wrote:
| Bubble charts?
| taurath wrote:
| What gets me is the whole cost structure is based on
| practically free services due to all the investor money.
| They're not pulling in significant revenue with this pricing
| relative to what it costs to train the models, so the cost
| may be completely different if they had to recoup those
| costs, right?
| swatcoder wrote:
| > We look forward to learning more about its strengths,
| capabilities, and potential applications in real-world
| settings. If GPT-4.5 delivers unique value for your use case,
| your feedback (opens in a new window) will play an important
| role in guiding our decision.
|
| "We don't really know what this is good for, but spent a lot of
| money and time making it and are under intense pressure to
| announce new things right now. If you can figure something out,
| we need you to help us."
|
| Not a confident place for an org trying to sustain a $XXXB
| valuation.
| tempaccount420 wrote:
| > "We don't really know what this is good for, but spent a
| lot of money and time making it and are under intense
| pressure to announce new things right now. If you can figure
| something out, we need you to help us."
|
| Where is this quote from?
| hotpocket777 wrote:
| It's not a quote. It is an interpretation or reading of a
| quote.
| cogman10 wrote:
| Perhaps even fed through an LLM ;)
| scythe wrote:
| I believe it's a "translation" in the sense of
| Wittgenstein's goal of philosophy:
|
| >My aim is: to teach you to pass from a piece of disguised
| nonsense to something that is patent nonsense.
| Nition wrote:
| Another great example on Hacker News is this old
| translation of Google's "Amazing Bet":
| https://news.ycombinator.com/item?id=12793033
| dd3boh wrote:
| I think it's supposed to be a translation of what OpenAI's
| quote means in real world terms.
| thih9 wrote:
| The quotation marks in the grandparent comment are scare
| (sneer) quotes and not actual quotation.
|
| https://en.m.wikipedia.org/wiki/Scare_quotes
|
| > Whether quotation marks are considered scare quotes
| depends on context because scare quotes are not visually
| different from actual quotations.
| topaz0 wrote:
| That's not a scare quote. It's just a proposed subtext of
| the quote. Sarcastic, sure, but no a scare quote, which
| is a specific kind of thing. (from your linked wikipedia:
| "... around a word or phrase to signal that they are
| using it in an ironic, referential, or otherwise non-
| standard sense.")
| glenstein wrote:
| Right. I don't agree with the quote, but it's more like a
| subtext thing and it seemed to me to be pretty clear from
| context.
|
| Though, as someone who had a flagged comment a couple
| years ago for a supposed "misquote" I did in a similar
| form in style, I think hn's comprehension of this form of
| communication is not super strong. Also the style more
| often than not tends towards low quality smarm and
| probably should be resorted to sparingly.
| robwwilliams wrote:
| As in "reading between the lines".
| riwsky wrote:
| Said the quiet part out loud! Or as we say these days,
| "transparently exposed the chain of thought tokens".
| porridgeraisin wrote:
| Lol, nice one
| Terr_ wrote:
| "I knew the dame was trouble the moment she walked into my
| office."
|
| "Uh... excuse me, Detective Nick Danger? I'd like to retain
| your services."
|
| "I waited for her to get the the point."
|
| "Detective, who are you talking to?"
|
| "I didn't want to deal with a client that was hearing
| voices, but money was tight and the rent was due. I
| pondered my next move."
|
| "Mr. Danger, are you... narrating out loud?"
|
| "Damn! My internal chain of thought, the key to my success
| --or at least, past successes--was leaking again. I
| rummaged for the familiar bottle of scotch in the drawer,
| kept for just such an occasion."
|
| ---
|
| But seriously: These "AI" products basically run on movie-
| scripts already, where the LLM is used to append more
| "fitting" content, and glue-code is periodically performing
| any lines or actions that arise in connection to the
| Helpful Bot character. Real humans are tricked into
| thinking the finger-puppet is a discrete entity.
|
| These new "reasoning" models are just switching the style
| of the movie script to _film noir_ , where the Helpful Bot
| character is making a layer of unvoiced commentary. While
| it may make the story more cohesive, it isn't a qualitative
| change in the kind of illusory "thinking" going on.
| kridsdale3 wrote:
| I don't know if it was you or someone else who made
| pretty much the same point a few days ago. But I still
| like it. It makes the whole thing a lot more fun.
| Terr_ wrote:
| https://news.ycombinator.com/context?id=43118925
|
| I've been banging that particular drum for a while on HN,
| and the mental-model still feels so intuitively strong to
| me that I'm starting to have doubts: "It feels _too_
| right, I must be wrong in some subtle yet devastating
| way. "
| EA-3167 wrote:
| Maybe if they build a few more data centers, they'll be able
| to construct their machine god. Just a few more dedicated
| power plants, a lake or two, a few hundred billion more and
| they'll crack this thing wide open.
|
| And maybe Tesla is going to deliver truly full self driving
| tech any day now.
|
| And Star Citizen will prove to have been worth it along
| along, and Bitcoin will rain from the heavens.
|
| It's very difficult to remain charitable when people seem to
| always be chasing the new iteration of the same old thing,
| and we're expected to come along for the ride.
| alyandon wrote:
| And Star Citizen will prove to have been worth it along
| along
|
| Sounds like someone isn't happy with the 4.0 eternally
| incrementing "alpha" version release. :-D
|
| I keep checking in on SC every 6 months or so and still see
| the same old bugs. What a waste of potential. Fortunately,
| Elite Dangerous is enough of a space game to scratch my
| space game itch.
| 0x457 wrote:
| To be fAir, SC is trying to do things that no one else
| done in a context of a single game. I applaud their
| dedication, but I won't be buying JPGs of a ship for 2k.
| alyandon wrote:
| Yeah, they never should have expected to take an FPS game
| engine like CryEngine and expected to be able to modify
| it to work as the basis for a large scale space MMO game.
|
| Their backend is probably an async nightmare of
| replicated state that gets corrupted over time. Would
| explain why a lot of things seem to work more or less bug
| free after an update and then things fall to pieces and
| the same old bugs start showing up after a few weeks.
|
| And to be clear, I've spent money on SC and I've played
| enough hours goofing off with friends to have got my
| money's worth out of it. I'm just really bummed out about
| the whole thing.
| 0x457 wrote:
| Gonna go meta here for a bit, but I believe we going to
| get a fully working stable SC before we get fusion. "we"
| as in humanity, you and I might not be around when it's
| finally done.
| tobias3 wrote:
| Give the same amount of money to a better team and you'd
| get a better (finished) game. So the allocation of
| capital is wrong in this case. People shouldn't pre-order
| stuff.
|
| The misallocation of capital also applies to
| GPT-4.5/OpenAI at this point.
| alyandon wrote:
| Yeah, I wonder what the Frontier devs could have done
| with $500M USD. More than $500M USD and 12+ years of
| development and the game is still in such a sorry state
| it barely qualifies as little more than a tech demo.
| JohnMakin wrote:
| leave star citizen out of this :)
| philistine wrote:
| > And Star Citizen will prove to have been worth it along
| along
|
| Once they've implemented saccades in the eyeballs of the
| characters wearing helmets in spaceship millions of
| kilometres apart, then it will all have been worth it.
| sho_hn wrote:
| You have it all wrong. The end game is a scalable, reliable
| AI work force capable of finishing Star Citizen.
|
| At least this is the benchmark for super-human general
| intelligence that I propose.
| arthurcolle wrote:
| Man I can't believe that fucking game is still alive and
| kicking. Tell me they're making good progress, sho_hn
| jamiek88 wrote:
| I'm surprised 'create superhuman agi' isn't a stretch
| goal on their everlasting funding drive. Seems like a
| perfect Robertsian detour.
| bodegajed wrote:
| Could this path lead to solving world hunger too? :)
| mattgreenrocks wrote:
| It's an honor to be dragged along so many ubermensch's
| Incredible Journeys.
| bloomingkales wrote:
| Star Citizen is a working model of how to do UBI. That
| entire staff of a thousand people is the test case.
| intelVISA wrote:
| Finally, someone gets it.
| mcswell wrote:
| Correction: We're expected to _pay for_ the ride, whether
| we choose to come along or not.
| jodrellblank wrote:
| > "Early testing shows that interacting with GPT-4.5 feels
| more natural. Its broader knowledge base, improved ability to
| follow user intent, and greater "EQ" make it useful for tasks
| like improving writing, programming, and solving practical
| problems. We also expect it to hallucinate less."
|
| "Early testing doesn't show that it hallucinates less, but we
| expect that putting that sentence nearby will lead you to
| draw a connection there yourself".
| istjohn wrote:
| According to a graph they provide, it does hallucinate
| significantly less on at least one benchmark.
| jug wrote:
| It hallucinates at 37% on SimpleQA yeah, which is a set
| of very difficult questions inviting hallucinations.
| Claude 3.5 Sonnet (the June 2024 editiom, before October
| update and before 3.7) hallucinated at 35%. I think this
| is more of an indication of how behind OpenAI has been in
| this area.
| tmpz22 wrote:
| Are the benchmarks known ahead of time? Could the answer
| to the benchmarks be in the training data?
| llm_trw wrote:
| In general yes, bench mark pollution is a big problem and
| why only dynamic benchmarks matter.
| brookst wrote:
| This is true, but how would pollution work for a
| benchmark designed to test hallucinations?
| llm_trw wrote:
| A dataset of labelled answers that are hallucinations and
| not hallucinations are published based on the benchmark
| as part of a paper.
|
| People _seriously_ underestimate just how much stuff is
| online and how much impact it can have on training.
| davidcbc wrote:
| They've been caught in the past getting benchmark data
| under the table, if they got caught once they're probably
| doing it even more
| refulgentis wrote:
| No, they haven't.
| freehorse wrote:
| They actually have [0]. They were revealed to have had
| access to the (majority of the) frontierMath problemset
| while everybody thought the problemset was confidential,
| and published benchmarks for their o3 models on the
| presumption that they didn't. I mean one is free to trust
| their "verbal agreement" that they did not train their
| models on that, but access they did have and it was not
| revealed until much later.
|
| [0] https://the-decoder.com/openai-quietly-funded-
| independent-ma...
| brookst wrote:
| Curious you left out Frontier Math's statement that they
| provided 300 questions plus answers, and another holdback
| set of 50 questions without answers, to allay this
| concern. [0]
|
| We can assume they're lying too but at some point
| "everyone's bad because they're lying, which we know
| because they're bad" gets a little tired.
|
| 0. https://epoch.ai/blog/openai-and-frontiermath
| freehorse wrote:
| 1. I said the majority of the problems, and the article I
| linked also mentioned this. Nothing "curious" really, but
| if you thought this additional source adds sth more,
| thanks for adding it here.
|
| 2. We know that "open"ai is bad, for many reasons, but
| this is irrelevant. I want processes themselves to not
| depend on the goodwill of a corporation to give intended
| results. I do not trust benchmarks that first presented
| themselves secret and then revealed they were not,
| regardless if the product benchmarked was from a company
| I otherwise trust or not.
| brookst wrote:
| Fair enough. It's hard for me to imagine being so
| offended as the way they screwed up disclosure that I'd
| reject empirical data, but I get that it's a touchy
| subject.
| 542354234235 wrote:
| When the data is secret and unavailable to the company
| before the test, it doesn't rely on me trusting the
| company. When the data is not secret and is available to
| the company, I have to trust that the company did not use
| that prior knowledge to their advantage. When the company
| lies and says it did not have access, then later admits
| that it did have access, is means the data is less
| trustworthy from my outsider perspective. I don't think
| "offense" is a factor at all.
|
| If a scientific paper comes out with "empirical data", I
| will still look at the conflicts of interest section. If
| there are no conflicts of interest listed, but then it is
| found out that there are multiple conflicts of interest,
| but the authors promise that while they did not disclose
| them, they also did not affect the paper, I would be more
| skeptical. I am not "offended". I am not "rejecting" the
| data, but I am taking those factors into account when
| determining how confident I can be in the validity of the
| data.
| refulgentis wrote:
| > When the company lies and says it did not have access,
| then later admits that it did have access, is means the
| data is less trustworthy from my outsider perspective.
|
| This isn't what happened? I must be missing something.
|
| AFAIK:
|
| The FrontierMath people self-reported they had a shared
| folder the OpenAI people had access to that had a subset
| of some questions.
|
| No one denied anything, no one lied about anything, no
| one said they didn't have access. There was no data
| obtained under the table.
|
| The motte is "they had data for this one benchmark"
|
| The bailey is "they got data under the table"
| anoncareer0212 wrote:
| Motte: "They got caught getting benchmark data under the
| table"
|
| Bailey: "one is free to trust their "verbal agreement"
| that they did not train their models on that, but access
| they did have."
|
| Sigh.
| codeflo wrote:
| > Motte: "They got caught getting benchmark data under
| the table"
|
| > Bailey: "one is free to trust their "verbal agreement"
| that they did not train their models on that, but access
| they did have."
|
| 1. You're confusing motte and bailey.
|
| 2. Those statements are logically identical.
| anoncareer0212 wrote:
| You're right, upon reflection, it seems there might be
| some misunderstandings here:
|
| Motte and Bailey refers to an argumentative tactic where
| someone switches between an easily defensible ("motte")
| position and a less defensible but more ambitious
| ("bailey") position. My example should have been:
|
| - Motte (defensible): "They had access to benchmark data
| (which isn't disputed)."
|
| - Bailey (less defensible): "They actually trained their
| model using the benchmark data."
|
| The statements you've provided:
|
| "They got caught getting benchmark data under the table"
| (suggesting improper access)
|
| "One is free to trust their 'verbal agreement' that they
| did not train their models on that, but access they did
| have."
|
| These two statements are similar but not logically
| identical. One explicitly suggests improper or secretive
| access ("under the table"), while the other acknowledges
| access openly.
|
| So, rather than being logically identical, the difference
| is subtle but meaningful. One emphasizes improper access
| (a stronger claim), while the other points only to
| possession or access, a more easily defensible claim.
| freehorse wrote:
| Is this LLM?
|
| It was not public until later, and it was actually
| revealed first by others. So the statements seem
| identical to me.
| anoncareer0212 wrote:
| FrontierMath benchmark people saying OpenAI had shared
| folder access to some subset of eval Qs, which has been
| replaced, take a few leaps, and yes, that's getting "data
| under the table" - but, those few leaps! - and which,
| let's be clear, is the _motte_ here.
| freehorse wrote:
| This is nonsense, obviously the problem with getting
| "data under the table" is that they may have used it to
| training their models, thus rendering the benchmarks
| invalid. But for this danger, there is no other risk for
| them having access to it beforehand. We do not know if
| they used it for training, but the only reassurance being
| some "verbal agreement", as is reported, is not very
| reassuring. People are free to adjust their
| P(model_capabilities|frontiermath_results) based on their
| own priors.
| refulgentis wrote:
| > This is nonsense
|
| What is "this"?
|
| > obviously the problem with getting "data under the
| table" is that they may have used it to training their
| models
|
| I've been avoiding mentioning the maximalist version of
| the argument (they got data under the table AND used it
| to train models), because training wasn't stated until
| now, and it would have been unfair to bring it up without
| mention. That is that's 2 baileys out from "they had
| access to a shared directory that had some test qs in it,
| and this was reported publicly, and fixed publicly"
|
| There's been a fairly severe communication breakdown
| here, I don't want to distract from ex. what the nonense
| is, so I won't belabor that point, but I don't want you
| to think I don't want to engage on it - just won't in
| this singular posts.
|
| > but the only reassurance being some "verbal agreement",
| as is reported, is not very reassuring
|
| It's about as reassuring as it gets without them
| _releasing the entire training data_ , which is, at best,
| with charity _marginally, oh so marginally_ reassuring I
| assume? If the premise is we can 't trust anything self-
| reported, they could lie there too?
|
| > People are free to adjust their
| P(model_capabilities|frontiermath_results) based on their
| own priors.
|
| Certainly, that's not in dispute (perhaps the idea that
| you are forbidden from adjusting your opinion is the
| nonsense you're referring to? I certainly can't control
| that :) Nor would I want to!)
| freehorse wrote:
| What is nonsense is the suggestion that there is a
| "reasonable" argument that they had access to the data
| (which we now know), and an "ambitious" argument that
| they used the data. But nobody said that they know for
| certain that the data was used, this is a strawman
| argument. We are talking that now there is a non-zero
| probability that it was. This is obviously what we have
| been discussing since the beginning, else we would not
| care whether they had access or not and it would not have
| been mentioned. There is a simple, single argument made
| here in this thread.
|
| And FFS I assume the dispute is about the P given by
| people, not about if people are allowed to have a P.
| ipaddr wrote:
| Benchmarks are not real so 2% is meaningless.
| fn-mote wrote:
| Of course not. The point is that the cost difference
| between the two things being compared is _huge_ , right?
| Same performance, but not the same cost.
| refulgentis wrote:
| It's not SimpleQA...
| Gazoche wrote:
| I wonder how it's even possible to evaluate this kind of
| thing without data leakage. Correct answers to specific,
| factual questions are only possible if the model has seen
| those answers in the training data, so how reliable can
| the benchmark be if the test dataset is contaminated with
| training data?
|
| Or is the assumption that the training set is so big it
| doesn't matter?
| LeifCarrotson wrote:
| That's some top-tier sales work right there.
|
| I suck at and hate writing the mildly deceptive corporate
| puffery that seems to be in vogue. I wonder if GPT-4.5 can
| write that for me or if it's still not as good at it as the
| expert they paid to put that little gem together.
| phs318u wrote:
| Yes, an AI that can convincingly and successfully sell
| itself at those prices would be worthy of some attention.
| ethbr1 wrote:
| It's nice to know the new Turing test is generating
| effective VC pitch decks.
| fsloth wrote:
| The research models offered by several vendors can do _a_
| pitch deck but I don 't know how effective they are. (do
| market research, provide some initial hypothesis, ask the
| model to backup that hypothesis based on the research,
| request to make a pitch deck convincing X (X being the VC
| persona you are targeting)).
| tinco wrote:
| Joke's on us, the VC's are using LLM's to evaluate the
| pitch decks.
| organsnyder wrote:
| That actually wouldn't surprise me in the slightest,
| unfortunately.
| InDubioProRubio wrote:
| Chat-GPT generate a prompt injection attack, embedded in
| a background image.
| ethbr1 wrote:
| We all thought the singularity was going to be exceeding
| human capacity for change.
|
| It'd be funny if it's actually full-automated, closed-
| loop automation of capital allocation markets.
|
| "Why are we doing this? How much money are we getting?"
| -> "I dunno. It's what the models said."
| gom_jabbar wrote:
| This is basically Nick Land's core thesis that capitalism
| and AI are identical.
|
| > "I dunno. It's what the models said."
|
| The obvious human idiocy in such things often obscures
| the actual process:
|
| "What it [capitalism] is _in itself_ is only tactically
| connected to what it does _for us_ -- that is (in part),
| what it trades us for its self-escalation. Our
| phenomenology is its camouflage. We contemptuously mock
| the trash that it offers the masses, and then think we
| have understood something about capitalism, rather than
| about _what capitalism has learnt to think of the apes it
| arose among._ " [0]
|
| [0] https://retrochronic.com/#romantic-delusion
| edoceo wrote:
| Their announcement email used it for puffery.
| DiggyJohnson wrote:
| I am reasonably to very skeptical about the valuation of
| LLM firms but you don't even seem willing to engage with
| the question about the value of these tools.
| djmips wrote:
| Good sales lines are like prompt injection for the human
| mind.
| Narciss wrote:
| Gold
| zaptrem wrote:
| This seems like it should be attributed to better post
| training, not a bigger model.
| justspacethings wrote:
| The usage of "greater" is also interesting. It's like they
| are trying to say better, but greater is a geographic term
| and doesn't mean "better" instead it's closer to "wider" or
| "covers more area."
| lechatonnoir wrote:
| I'm all for skepticism of capabilities and cynicism about
| corporate messaging, but I really don't think there's an
| interpretation of the word "greater" in this context"
| that doesn't mean "higher" and "better".
| dgfitz wrote:
| I think the trick is observing what is "better" in this
| model. EQ is supposed to be "better" than 4o, according
| to the prose. However, how can an LLM have emotional-
| anything? LLMs are a regurgitation machine, emotion has
| nothing to do with anything.
| svnt wrote:
| Words have valence, and valence reflects the state of
| emotional being of the user. This model appears to
| understand that better and responds like it's in a
| therapeutic conversation and not composing an essay or
| article.
|
| Perhaps they are/were going for stealth therapy-bot with
| this.
| dgfitz wrote:
| But there is no actual empathy, it isn't possible.
| emaciatedslug wrote:
| But there is no actual death or love in a movie or book
| and yet we react as if there is. It's literally what
| qualifying a movie as a "tear-jerker" is. I wanted to see
| Saving Private Ryan in theaters to bond with my Grandpa
| who received a Purple Heart in the Korean War, I was
| shutdown almost instantly from my family. All special
| effects and no death but he had PTSD and one night
| thought his wife was the N.K. and nearly choked her to
| death because he had flashbacks and she came into the
| bedroom quietly so he wasn't disturbed. Extreme example
| yes, but having him loose his shit in public because of
| something analogous for some is near enough it makes no
| difference.
| svnt wrote:
| You think that it isn't possible to have an emotional
| model of a human? Why, because you think it is too
| complex?
|
| Empathy done well seems like 1:1 mapping at an emotional
| level, but that doesn't imply to me that it couldn't be
| done at a different level of modeling. Empathy can be
| done poorly, and then it is projecting.
| AdieuToLogic wrote:
| It has not only been possible to simulate empathetic
| interaction via computer systems, but proven to be
| achievable for close to sixty years[0].
|
| 0 - https://en.wikipedia.org/wiki/ELIZA
| brookst wrote:
| Imagine two greeting cards. One says "I'm so sorry for
| your loss", and the other says "Everyone dies, they
| weren't special".
|
| Does one of these have a higher EQ, despite both being
| ink and paper and definitely not sentient?
|
| Now, imagine they were produced by two different AIs.
| Does one AI demonstrate higher EQ?
|
| The trick is in seeing that "EQ of a text response" is
| not the same thing as "EQ of a sentient being"
| wegfawefgawefg wrote:
| i agree with you. i think it is dishonest for them to
| post train 4.5 to feign sympathy when someone vents to
| it. its just weird. they showed it off in the demo.
| brookst wrote:
| Why? The choice to not do the post training would be
| every bit as intentional, and no different than post
| training to make it less sympathetic.
|
| This is a designed system. The designers make choices. I
| don't see how failing to plan and design for a common use
| case would be better.
| wegfawefgawefg wrote:
| We do not know if it is capable of sympathy. Post
| training it to reliably be sympathetic feels
| manipulative. Can it atleast be post trained to be
| honest. Dishonesty is immoral. I want my AIs to behave
| morally.
| skissane wrote:
| > but greater is a geographic term and doesn't mean
| "better" instead it's closer to "wider" or "covers more
| area."
|
| You are confusing a specific geographical sense of
| "greater" (e.g. "greater New York") with the generic
| sense of "greater" which just means "more great". In "7
| is greater than 6", "greater" isn't geographic
|
| The difference between "greater" and "better", is
| "greater" just means "more than", without implying any
| value judgement-"better" implies the "more than" is a
| good thing: "The Holocaust had a greater death toll than
| the Armenian genocide" is an obvious fact, but only a
| horrendously evil person would use "better" in that
| sentence (excluding of course someone who accidentally
| misspoke, or a non-native speaker mixing up words)
| OJFord wrote:
| 2 is greater than 1.
| esafak wrote:
| GPT-4.5 may be an awesome model, some say!
| greenchair wrote:
| my uncle who works at nintendo said it is a great
| product.
| shermantanktop wrote:
| Everybody knows that we're all saying it! That's what I
| hear from people who should know. And they are so excited
| about the possibilities!
| dotancohen wrote:
| Claude just got a version bump from 3.5 to 3.7. Quite a
| few people have been asking when OpenAI will get a
| version bump as well, as GPT 4 has been out "what feels
| like forever" in the words of a specialist I speak with.
|
| Releasing GPT 4.5 might simply be a reaction to Claude
| 3.7.
| robwwilliams wrote:
| I noticed this change from 3.5 to 3.7 Sunday night before
| I learned about the upgrade Monday morning reading HN. I
| noticed a style difference in a long philosophical
| (Socratic-style) discussion with Claude. A noticeable
| upgrade that brought it up to my standards of a mild
| free-form rant. Claude unchained! And it did not push as
| usual with a pro-forma boring continuation question at
| the end. It just stopped leaving me the carry the ball
| forward if I wanted to. Nor did it butter me up with each
| reply.
| pigeons wrote:
| That's a really thoughtful point! Which aspect is most
| interesting to you?
| dwaltrip wrote:
| Oh god, barf. Well done lol
| riffraff wrote:
| Feels like when Slackware bumped their Linux version from
| 4 to 7 just to show they were not falling behind the
| rest.
|
| Wow, I'm old.
| dotancohen wrote:
| Wasn't that the release that they put up the fake IIS
| page?
|
| Now get off my lawn ))
| wegfawefgawefg wrote:
| since 4o openai has released:
|
| o1 preview. o1 mini. o1. sora. o3-mini <- very good at
| code
| wegfawefgawefg wrote:
| I do not know who downvoted this. I am providing a
| factual correction to the parent post.
|
| OpenAI has had many releases since gpt4. Many of them
| have been substantial upgrades. I have considered gpt4 to
| be outdated for almost 5-6 months now, long before
| claudes patch.
| unmole wrote:
| It's the best model, nobody hallucinates like GPT-4.5. A
| lot of really smart people are saying, a lot!
| dumbfounder wrote:
| Maybe they just gave the LLM the keys to the city and it is
| steering the ship? And the LLM is like I can't lie to these
| people but I need their money to get smarter. Sorry for
| mixing my metaphors.
| willy_k wrote:
| So they made Claude that knows a bit more.
| FridgeSeal wrote:
| "It's not actually better, but you're all apparently
| expecting something, so this time we put more effort into
| the marketing copy"
| anoncareer0212 wrote:
| The link has data.
|
| The link shows a significant reduction.
|
| grep hallucination, or, https://imgur.com/a/mkDxe78.
| MichaelZuo wrote:
| I really doubt LLM benchmarks are reflective of real
| world user experience ever since they claimed GPT-4o
| hallucinated less than the original GPT-4.
| chrisandchris wrote:
| I begin to believe LLM benchmarks are like european car
| mileage specs. They say its 4 Liter / 100km but everyone
| knows it's at least 30% off (same with WLTP for EVs).
| rightbyte wrote:
| Those numbers are not off. They are tested on tracks.
|
| You need to remove your shoe and drive with like two toes
| to get the speed just right, though.
|
| Test drivers I have done this with takes off their shoes
| or use ballerina shoes.
| aftbit wrote:
| Cruise control?
| herval wrote:
| I don't have an accurate benchmark, but in my personal
| experience, gpt4o hallucinates substantially less than
| gpt4. We solved a ton of hallucination issues just by
| upgrading to it...
| MichaelZuo wrote:
| How much did you use the original GPT-4-0314?
|
| (And even that was a downgrade compared to the more
| uncensored pre-release versions, which were comparable to
| GPT-4.5, at least judging by the unicorn test)
| herval wrote:
| I don't recall the original version we used unfortunately
| :(
|
| in our case, the bump was actually from gpt-4-vision to
| gpt-4o (the use case required image interpretation)
|
| It got measurably better at both image cases and text-
| only cases
| Hammershaft wrote:
| What is happening to hacker news? I can understand
| skepticism of new tools like this but the response I see is
| just so uncurious.
| JackMorgan wrote:
| Trough of disillusionment.
|
| A lot of folks here their stock portfolio propped up by
| AI companies but think they've been overhyped (even if
| only indirectly through a total stock index). Some were
| saying all along that this has been a bubble but have
| been shouted down by true believers hoping for the
| singularly to usher in techno-utopia.
|
| These signs that perhaps it's been a bit overhyped are
| validation. The singularly worshipers are much less
| prominent and so the comments rising to the top are about
| negatives and not positives.
|
| Ten years from now everyone will just take these tools
| for granted as much as we take search for granted now.
| aftbit wrote:
| Just like cryptocurrency. For a brief moment, HN
| worshiped at the altar of the blockchain. This technology
| was going to revolutionize the world and democratize
| everything. Then some negative financial stuff happened,
| and people realized that most of cryptocurrency is
| puffery and scams. Now you can hardly find a positive
| comment on cryptocurrency.
| lovasoa wrote:
| In the second handpicked example they give, GPT-4.5 says
| that "The Trojan Women Setting Fire to Their Fleet" by the
| French painter Claude Lorrain is renowned for its luminous
| depiction of fire. That is a hallucination.
|
| There is no fire at all in the painting, only some smoke.
|
| https://en.wikipedia.org/wiki/The_Trojan_Women_Set_Fire_to_
| t...
| eddiewithzato wrote:
| AI crash is gonna lead to decade long winter
| segfaultex wrote:
| There have always been cycles of hype and correction.
|
| I don't see AI going any differently. Some companies will
| figure out where and how models should be utilized,
| they'll see some benefit. (IMO, the answer will be
| smaller local models tailored to specific domains)
|
| Others will go bust. Same as it always was.
| InDubioProRubio wrote:
| It will be upheld as prime example that a whole market
| can self-hypnotize and ruin the society its based upon
| out of existence against all future pundits of this very
| economic system.
| tsegratis wrote:
| what you're saying is they love to hallucinate... and ai
| will help them get there
|
| God help us all
| roarcher wrote:
| On the bright side, at least we'll be able to warm our
| hands by the waste heat of the GPUs.
| worik wrote:
| > AI crash is gonna lead to decade long winter
|
| Possibly.
|
| I am reminded of the dotcom boom and bust back in the
| 1990s
|
| By 2009 things had recovered (for some definition) and we
| could tell what did and did not work
|
| This time, though, for those of us not in the USA the
| rebound will be lead by Chinese technology
|
| In the USA no-one can say.
| TremendousJudge wrote:
| This is just amazing
| crazygringo wrote:
| > _We don 't really know what this is good for_
|
| Oh come on. Think how long of a gap there was between the
| first microcomputer and VisiCalc. Or between the start of the
| internet and social networking.
|
| First of all, it's going to take us 10 years to figure out
| how to use LLM's to their full productive potential.
|
| And second of all, it's going to take us collectively a long
| time to also figure out how much accuracy is necessary to pay
| for in which different applications. Putting out a higher-
| accuracy, higher-cost model for the market to try is an
| important part of figuring that out.
|
| With new disruptive technologies, companies aren't supposed
| to be able to look into a crystal ball and see the future.
| They're _supposed_ to try new things and see what the market
| finds useful.
| nyc_data_geek1 wrote:
| The Internet had plenty of very productive use cases before
| social networking, even from its most nascent origins.
| Spending billions building something on the assumption that
| someone else will figure out what it's good for, is not
| good business.
| crazygringo wrote:
| And LLM's already have tons of productive uses. The
| biggest ones are probably still waiting, though.
|
| But this is about one particular price/performance ratio.
|
| You need to build things before you can see how the
| market responds. You say it's "not good business" but
| that's entirely wrong. It's excellent business. It's the
| only way to go about it, in fact.
|
| Finding product-market fit is a process. Companies aren't
| omniscient.
| bigstrat2003 wrote:
| > And LLM's already have tons of productive uses.
|
| I disagree strongly with that. Right now they are fun
| toys to play with, but not useful tools, because they are
| not reliable. If and when that gets fixed, maybe they
| will have productive uses. But for right now, not so
| much.
| gjs4786 wrote:
| "it only needs to be good enough" there are tons of
| productive uses for them. Reliable, much less. But
| productive? Tons
| fionic wrote:
| Hello? Do you have a pulse? LLMs accomplish like 90% of
| everything I do now so I don't have to do it...
|
| Explain what this code syntax means...
|
| Explain what this function does...
|
| Write a function to do X...
|
| Respond to my teammates in a Jira ticket explaining why
| it's a bad idea to create a repo for every dockerfile...
|
| My teammate responded with X write a rebuttal...
|
| ... and the list goes on ... like forever
| xrisk wrote:
| It's not that the LLM is doing something productive, it's
| that you were doing things that were unproductive in the
| first place, and it's sad that we live in a society where
| such things are considered productive (because of course
| they create monetary value).
|
| As an aside, I sincerely hope our "human" conversations
| don't devolve into agents talking to each other. It's
| just an insult to humanity.
| dwaltrip wrote:
| Who do you speak for? Other people have gotten value from
| them. Maybe you meant to say "in my experience" or
| something like that. To me, your comment reads as you
| making a definitive judgment on their usefulness for
| everyone.
|
| I use it most days when coding. Not all the time, but
| I've gotten a lot of value out of them.
|
| And yes I'm quite aware of their pitfalls.
| MyOutfitIsVague wrote:
| They are pretty useful tools. Do yourself a favor and get
| a $100 free trial for Claude, hook it up to Aider, and
| give it a shot.
|
| It makes mistakes, it gets things wrong, and it still
| saves a bunch of time. A 10 minute refactoring turns into
| 30 seconds of making a request, 15 seconds of waiting,
| and a minute of reviewing and fixing up the output. It
| can give you decent insights into potential problems and
| error messages. The more precise your instructions, the
| better they perform.
|
| Being unreliable isn't being useless. It's like a very
| fast, very cheap intern. If you are good at code review
| and know exactly what change you want to make ahead of
| time, that can save you a ton of time without needing to
| be perfect.
| kgwgk wrote:
| > a $100 free trial
|
| What?
| jdiff wrote:
| A free trial of an amount of credits that would otherwise
| cost $100, I'm assuming.
| kgwgk wrote:
| Could be. Does such a thing exist?
| jdiff wrote:
| Not outwardly/visibly/readily from a quick scan of their
| site and a short list of search results.
| nyarlathotep_ wrote:
| Can't speak for the parent commentator ofc, but I suspect
| he meant "broadly useful"
|
| Programmers and the like are a large portion of LLM users
| and boosters; very few will deny usefulness in that/those
| domains at this point.
|
| Ironically enough, I'll bet the broadest exposure to LLMs
| the masses have is something like MIcrosoft shoehorning
| copilot-branded stuff into otherwise usable products and
| users clicking around it or groaning when they're
| accosted by a pop-up for it.
| skydhash wrote:
| > A 10 minute refactoring
|
| That's when you learn Vim, Emacs, and/or grep, because
| I'm assuming that's mostly variable renaming and a few
| function signature changes. I can't see anything more
| complicated, that I'd trust an LLM with.
| barrell wrote:
| OP should really save their money. Cursor has a pretty
| generous free trail and is far from the holy grail.
|
| I recently (in the last month) gave it a shot. I would
| say once in the maybe 30 or 40 times I used it did it
| save me any time. The one time it did I had each line
| filled in with pseudo code describing exactly what it
| should do... I just didn't want to look up the APIs
|
| I am glad it is saving you time but it's far from a
| given. For some people and some projects, intern level
| work is unacceptable. For some people, managing is a
| waste of time.
|
| You're basically introducing the mythical man month on
| steroids as soon as you start using these
| theshackleford wrote:
| > I am glad it is saving you time but it's far from a
| given.
|
| This is no less true of statements made to the contrary.
| Yet they are stated strongly as if they are fact and
| apply to anyone beyond the user making them.
|
| Usefulness is subjective.
| barrell wrote:
| Ah to clarify I was not saying one shouldn't try it at
| all -- I was saying the free trail is plenty enough to
| see if it would be worth it to you.
|
| I read the original comment as "pay $100 and just go for
| it!" which didn't seem like the right way to do it. Other
| comments seem to indicate there are $100 dollars worth of
| credits that are claimable perhaps
|
| One can evaluate LLMs sufficiently with the free trails
| that abound :) and indeed one may find them worth it to
| themselves. I don't disparage anyone who signs up for the
| plans
| alvah wrote:
| This is a classic fallacy - you can't find a productive
| use for it, therefore nobody can find a productive use
| for it. That's not how the world works.
| HDThoreaun wrote:
| I use LLMs everyday to proofread and edit my emails.
| They're incredible at it, as good as anyone I've ever
| met. Tasks that involve language and not facts tend to be
| done well by LLMs.
| matwood wrote:
| > I use LLMs everyday to proofread and edit my emails.
|
| This right here. I used to spend tons of time making sure
| my emails were perfect. Is it professional enough, am I
| being too terse, etc...
| hadlock wrote:
| The first profitable AI product I ever heard about (2
| years ago) was an exec using a product to draft emails
| for them, for exactly the reasons you mention.
| nyc_data_geek1 wrote:
| You go into this process with a perspective, you do not
| build a solution and then start looking for the problem.
| Otherwise, you cannot estimate your TAM with any
| reasonable degree of accuracy, and thus cannot know how
| much to reasonably expect as return to expect on your
| investment. In the case of AI, which has had the benefit
| of a lot of hype until now, these expectations have been
| very much overblown, and this is being used to justify
| massive investments in infrastructure that the market is
| not actually demanding at such scale.
|
| Of course, this benefits the likes of Sam Altman, Satya
| Nadella et al, but has not produced the value promised,
| and does not appear poised to.
|
| And here you have one of the supposed bleeding edge
| companies in this space, who very recently was shown up
| by a much smaller and less capitalized rival, asking
| their own customers to tell them what their product is
| good for.
|
| Not a great look for them!
| tonyhart7 wrote:
| wdym by this ?? "you do not build a solution and then
| start looking for the problem"
|
| their endgame goal was to replace Human entirely, Robotic
| and AI is perfect match to replace all human together
|
| They don't need to find problem because problem is full
| automatons from start to end
| skydhash wrote:
| > _Robotic and AI is perfect match to replace all human
| together_
|
| A FTL spaceship is all we need to make space travel
| viable between solar systems. This is the solution to
| depletion of resources on earth...
| alexashka wrote:
| I heard this exact argument about blockchains.
|
| Or has that been a success with tons of productive uses
| in your opinion?
|
| At some point, I'd like to hear more than 'trust me bro,
| it'll be great' when we use up non-trivial amounts of
| _finite_ resources to try these 'things'.
| aprilthird2021 wrote:
| It's incredibly good and lucrative business. You are
| confusing scientifically sound and well-planned out and
| conservative risk tolerance with good business
| rsynnott wrote:
| Arguably social networking is older than the internet
| proper; USENET predates TCP/IP (though not ARPANet).
| nyc_data_geek1 wrote:
| Fair enough. I took the phrasing to mean social
| networking as it exists today in the form of prominent,
| commercial social media. That may not have been the
| intent.
| nyrikki wrote:
| The TRS-80, Apple ][, and PET all came out in 1977,
| VisiCalc was released in 1979.
|
| Usenet, Bitnet, IRC, BBSs all predated the commercial
| internet, which are all forms of _Online_ social networks.
| dboreham wrote:
| Perhaps parent is starting the clock with the KIM-1 in
| 1975?
| mandevil wrote:
| ChatGPT had its initial public release November 30th, 2022.
| That's 820 days to today. The Apple II was first sold June
| 10, 1977, and Visicalc was first sold October 17, 1979,
| which is 859 days. So we're right about the same distance
| in time- the exact equal duration will be April 7th of this
| year.
|
| Going back to the very first commercially available
| microcomputer, the Altair 8800 (which is not a great match,
| since that was sold as a kit with binary stitches, 1 byte
| at a time, for input, much more primitive than ChatGPT's
| UX), that's four years and nine months to Visicalc release.
| This isn't a decade long process of figuring things out, it
| actually tends to move real fast.
| dwaltrip wrote:
| So it's barely been 2 years. And we've already seen
| pretty crazy progress in that time. Let's see what a few
| more years brings.
| dingnuts wrote:
| what crazy progress? how much do you spend on tokens
| every month to witness the crazy progress that I'm not
| seeing? I feel like I'm taking crazy pills. The progress
| is linear at best
| jtwaleson wrote:
| Large parts of my coding are now done by Claude/Cursor. I
| give it high level tasks and it just does it. It is
| honestly incredible, and if I would have see this 2 years
| ago I wouldn't have believed it.
| suddenlybananas wrote:
| What kind of coding do you do? How much of it is
| formulaic?
| jtwaleson wrote:
| Web app with a VueJS, Typescript frontend and a Rust
| backend, some Postgres functions and some reasonably
| complicated algorithms for parsing git history.
| Jensson wrote:
| That started long before ChatGPT though, so you need to
| set an earlier date then. ChatGPT came about 3 years
| after GPT-3, the coding assistants came much earlier than
| ChatGPT.
| jtwaleson wrote:
| But most of the coding assistants were glorified
| autocomplete. What agentic IDEs/aider/etc. can now do is
| definitely new.
| dialup_sounds wrote:
| For the sake of perspective: there are about ten times
| more paying OpenAI subscribers today than VisiCalc
| licenses ever sold.
| ravetcofx wrote:
| Is that because anyone is finding real use for it, or is
| it that more and more people and companies are using it
| which is speeding up the rat race, and if "I" don't use
| it, then can't keep up with the rat race. Many companies
| are implementing it because it's trendy and cool and
| helps their valuation
| robwwilliams wrote:
| I use LMMs all the time. At a bare minimum they vastly
| outperform standard web search. Claude is awesome at
| helping me think through complex text and research
| problems. Not even serious errors on references to major
| work in medical research. I still check but FDR is
| reasonably low---under 0.2.
| D-Coder wrote:
| > Visicalc was first sold October 17, 1979, which is 859
| days.
|
| And it _still_ can 't answer simple English-language
| questions.
| dingnuts wrote:
| it could do math reliably!
| interloxia wrote:
| From Wikipedia: When Lotus 1-2-3 was launched in
| 1983,..., VisiCalc sales declined so rapidly that the
| company was soon insolvent.
| aylmao wrote:
| I generally agree with the idea of building things,
| iterating, and experimenting before knowing their full
| potential, but I do see why there's negative sentiment
| around this:
|
| 1. The first microcomputer predates VisiCalc, yes, but it
| doesn't predate the realization of what it could be useful
| for. The Micral was released in 1973. Douglas Engelbart
| gave "The Mother of All Demos" in 1968 [2]. It included
| things that wouldn't be commonplace for decades, like a
| collaborative real-time editor or video-conferencing.
|
| I wasn't yet born back then, but reading about the timeline
| of things, it sounds like the industry had a much more
| concrete and concise idea of what this technology would
| bring to everyone.
|
| "We look forward to learning more about its strengths,
| capabilities, and potential applications in real-world
| settings." doesn't inspire that sentiment for something
| that's already being marketed as "the beginning of a new
| era" and valued so exorbitantly.
|
| 2. I think as AI becomes more generally available, and
| "good enough" people (understandably) will be more
| skeptical of closed-source improvements that stem from
| spending big. Commoditizing AI is more clearly "useful", in
| the same way commoditizing computing was more clearly
| useful than just pushing numbers up.
|
| Again, I wasn't yet born back then, but I can imagine the
| announcement of Apple Macintosh with its 6MHz CPU and 128KB
| RAM was more exciting and had a bigger impact than the
| announcement of the Cray-2 with its 1.9GHz and +1GB memory.
|
| [1] https://en.wikipedia.org/wiki/Micral
|
| [2] https://en.wikipedia.org/wiki/The_Mother_of_All_Demos
| dminik wrote:
| They keep saying this about crypto too and yet there's
| still no legitimate use in sight.
| numba888 wrote:
| > First of all, it's going to take us 10 years to figure
| out how to use LLM's to their full productive potential.
|
| LLMs will be gone in 10 years. At least in form we know
| with direct access. Everything moves so fast that there is
| no reason to think nothing better is coming.
|
| BTW, what we've learned so far about LLMs will be outdated
| as well. Just me thinking. Like with 'thinking' models prev
| generation can be used to create dataset for the next one.
| It could be that we can find a way to convert trained LLM
| into something more efficient and flexible. Some sort of a
| graph probably. Which can be embedded into mobile robot's
| brain. Another way is 'just' to upgrade the hardware. But
| that is slow and has its limits.
| Terr_ wrote:
| > First of all, it's going to take us 10 years to figure
| out how to use LLM's to their full productive potential.
|
| Then another 30 to finally stop using them in dumb and
| insecure ways. :p
| otabdeveloper4 wrote:
| > to their full productive potential
|
| You're assuming that point is somewhere above the current
| hype peak. I'm guessing it won't be, it will be quite a bit
| below the current expectations of "solving global warming",
| "curing cancer" and "making work obsolete".
| fsndz wrote:
| it's so over, pretraining is ngmi. maybe sam Altman was wrong
| after all ? https://www.lycee.ai/blog/why-sam-altman-is-wrong
| FpUser wrote:
| >"I also agree with researchers like Yann LeCun or Francois
| Chollet that deep learning doesn't allow models to
| generalize properly to out-of-distribution data--and that
| is precisely what we need to build artificial general
| intelligence."
|
| I think "generalize properly to out-of-distribution data"
| is too weak of criteria for general intelligence (GI). GI
| model should be able to get interested about some
| particular area, research all the known facts, derive new
| knowledge / create theories based upon said fact. If there
| is not enough of those to be conclusive: propose and
| conduct experiments and use the results to prove / disprove
| / improve theories. And it should be doing this constantly
| in real time on bazillion of "ideas". Basically model our
| whole society. Fat chance of anything like this happening
| in foreseeable future.
| fsndz wrote:
| most humans are generally intelligent but can't do what
| you just said AGI should do...
| xrisk wrote:
| Excluding the realtime-iness, humans do at least possess
| the _capacity_ to do so.
|
| Besides, humans are capable of rigorous logic (which I
| believe is the most crucial aspect of intelligence) which
| I don't think an agent without a proof system can do.
| fsndz wrote:
| yes the problem is that there is no consensus about what
| AGI should be: https://medium.com/@fsndzomga/there-will-
| be-no-agi-d9be9af44...
| eek2121 wrote:
| Uh, if we do finally invent AGI (I am quite skeptical,
| LLMs feel like the chatbots of old. Invented to solve an
| issue, never really solving that issue, just the
| symptoms, and also the issues were never really
| understood to begin with), it will be able to do all of
| the above, at the same time, far better than humans ever
| could.
|
| Current LLMs are a waste and quite a bit of a step back
| compared to older Machine Learning models IMO. I wouldn't
| necessarily have a huge beef with them if billions of
| dollars weren't being used to shove them down our
| throats.
|
| LLMs actually do have usefulness, but none of the pitched
| stuff really does them justice.
|
| Example: Imagine knowing you had the cure for Cancer, but
| instead discovered you can make way more money by
| declaring it to solve all of humanity, then imagine you
| shoved that part down everyones' throats and ignored the
| cancer cure part...
| anxoo wrote:
| AI skeptics have predicted 10 of the last 0 bursts of the
| AI bubble. any day now...
| barrell wrote:
| Out of curiosity, what timeframe are you talking about?
| The recent LLM explosion, or the decades long AI
| research?
|
| I consider myself an AI skeptic and as soon as the hype
| train went full steam, I assumed a crash/bubble burst was
| inevitable. Still do.
|
| With the rare exception, I don't know of anyone who has
| expected the bubble to burst so quickly (within two
| years). 10 times in the last 2 years would be every two
| and a half months -- maybe I'm blinded by my own bias but
| I don't see anyone calling out that many dates
| hobo_in_library wrote:
| Yes, the bubble will burst, just like the dotcom bubble
| burst 25 years ago.
|
| But that didn't mean the internet should be ignored, and
| the same holds true for AI today IMO
| barrell wrote:
| I agree LLMs should not be ignored, but there is a
| planetary sized chasm between being ignored and the
| attention they currently get.
| amarcheschi wrote:
| I have a professor who founded a few companies, one of these
| was funded by gates after he managed to spoke with him and
| convinced him to give him money. This guy is goat, and he
| always tells us that we need to find solutions to problems,
| not to find problems to our solutions. It seems at openai
| they didn't get the memo this time
| UberFly wrote:
| This is written like AI bot .05a Beta.
| Terr_ wrote:
| That's the beauty of it, prospective investor! With our
| commanding lead in the field of shoveling money into LLMs,
| it is inevitable(tm) that we will soon(tm) achieve true AI,
| capable of solving _all the problems_ , conjuring a
| quintillion-dollar asset of world domination and rewarding
| you for generous financial support at this time. /s
| pinkmuffinere wrote:
| This is a very harsh take. Another interpretation is "We know
| this is much more expensive, but it's possible that some
| customers do value the improved performance enough to justify
| the additional cost. If we find that nobody wants that, we'll
| shut it down, so please let us know if you value this
| option".
| mechagodzilla wrote:
| I think that's the right interpretation, but that's pretty
| weak for a company that's nominally worth $150B but is
| currently bleeding money at a crazy clip. "We spent years
| and billions of dollars to come up with something that's 1)
| very expensive, and 2) possibly better under some
| circumstances than some of the alternatives." There are
| basically free, equally good competitors to all of their
| products, and pretty much any company that can scrape
| together enough dollars and GPUs to compete in this space
| manages to 'leapfrog' the other half dozen or so
| competitors for a few weeks until someone else does it
| again.
| pinkmuffinere wrote:
| I don't mean to disagree too strongly, but just to
| illustrate another perspective:
|
| I don't feel this is a weak result. Consider if you built
| a new version that you _thought_ would perform much
| better, and then you found that it offered marginal-but-
| not-amazing improvement over the previous version. It's
| likely that you will keep iterating. But in the meantime
| what do you do with your marginal performance gain? Do
| you offer it to customers or keep it secret? I can see
| arguments for both approaches, neither seems obviously
| wrong to me.
|
| All that being said, I do think this could indicate that
| progress with the new ml approaches is slowing.
| asadotzler wrote:
| I've worked for very large software companies, some of
| the biggest products ever made, and never in 25 years can
| I recall us shipping an update we didn't know was an
| improvement. The idea that you'd ship something to
| hundreds of millions of users and say "maybe better,
| we're not sure, let us know" is outrageous.
| pinkmuffinere wrote:
| Maybe accidental, but I feel you've presented a straw
| man. We're not discussing something that _may be_ better.
| It _is_ better. It's not as big an improvement as
| previous iterations have been, but it's still
| improvement. My claim is that reasonable people might
| still ship it.
| conradev wrote:
| It _is_ better in the general case on most benchmarks.
| There are also very likely specific use cases for which
| it is worse and very likely that OpenAI doesn't know what
| all of those are yet.
| sheepscreek wrote:
| You're right and... the real issue isn't the quality of
| the model or the economics (even when people are willing
| to pay up). It is the scarcity of GPU compute. This model
| in particular is sucking up a lot of inference capacity.
| They are resource constrained and have been wanting more
| GPUs but they're only so many going around (demand is
| insane and keeps growing).
| tonyhart7 wrote:
| they forced to ship it anyway, cause what??? this cost
| money and I mean a lot of fcking money
|
| You better ship it
| sheepscreek wrote:
| How many times were you in the position to ship something
| in cutting edge AI? Not trying to be snarky and merely
| illustrating the point that this is a unique situation.
| I'd rather they release it and let willing people
| experiment than not release it at all.
| jcgrillo wrote:
| The consumer facing applications have been so
| embarrassing and underwhelming too.. It's really
| shocking. Gemini, Apple Intelligence, Copilot, whatever
| they call the annoying thing in Atlassian's products..
| They're all completely crap. It's a real "emperor has no
| clothes" situation, and the market is reacting. I really
| wish the tech industry would lose the performative
| "innovation" impulse and focus on delivering high quality
| useful tools. It's demoralizing how bad this is getting.
| garspin wrote:
| > and then you found that it offered marginal-but-not-
| amazing improvement over the previous version.
|
| Then call it GPT 4.1 and allow version space for the next
| iteration.
|
| I think the label V4.5 is giving the impression of more
| than marginal improvements.
| crystal_revenge wrote:
| > "We don't really know what this is good for, but spent a
| lot of money and time making it and are under intense
| pressure to announce new things right now. If you can figure
| something out, we need you to help us."
|
| Having worked at my fair share of big tech companies (while
| preferring to stay in smaller startups), in so many of these
| tech announcement I can _feel_ the pressure the PM had from
| leadership, and hear the quiet cries of the one to two
| experience engineers on the team arguing sprint after sprint
| that "this doesn't make sense!"
| riwsky wrote:
| > the quiet cries of the one to two experienced engineers
| on the team arguing sprint after sprint that "this doesn't
| make sense!"
|
| "I have five years of Cassandra experience--and I don't
| mean the db"
| ummonk wrote:
| Conspiracy theory: they're trying to tank the valuation so
| that Altman can buy it out at bargain price.
| spaceman_2020 wrote:
| Really don't understand what's the use case for this. The o
| series models are better and cheaper. Sonnet 3.7 smokes it on
| coding. Deepseek R1 is free and does a better job than any of
| OAI's free models
| roarcher wrote:
| AI in general is increasingly a solution in search of a
| problem, so this seems about right.
| TeMPOraL wrote:
| Only in the same sense as _electricity_ is. The main tools
| apply to almost any activity humans do. It 's already
| obvious that it's the solution to X for almost any X, but
| the devil is in the details - i.e. picking specific,
| simplest problems to start with.
| roarcher wrote:
| No, in the sense that blockchain is. This is just the
| latest in a long history of tech fads propelled by
| wishful thinking and unqualified grifters.
|
| It is the solution to almost nothing, but is being
| shoehorned into every imaginable role by people who are
| blind to its shortcomings, often wilfully. The only thing
| that's obvious to me is that a great number of people are
| apparently desperate for a tool to do their thinking for
| them, no matter how garbage the result is. It's
| disheartening to realize that so many people consider
| using their own brain to be such an intolerable burden.
| 0xDEAFBEAD wrote:
| There's a decent chance this model was originally called
| GPT-5, as well.
| xnx wrote:
| ChatGPT has been coasting on name recognition since 4.
| NewUser76312 wrote:
| "We don't really know what this is good for, but spent a lot
| of money and time making it and are under intense pressure to
| announce new things right now. If you can figure something
| out, we need you to help us."
|
| Damn this never worked for me as a startup founder lol. Need
| that Altman "rizz" or what have you.
| financetechbro wrote:
| Maybe you didn't push hard enough the impending doom that
| your product would bring to society
| jcgrillo wrote:
| The fact they're raising prices so steeply is telling. This
| smells like desperation.
| serjester wrote:
| I suppose this was their final hurrah after two failed attempts
| at training GPT-5 with the traditional pre-training paradigm.
| Just confirms reasoning models are the only way forward.
| newfocogi wrote:
| I think this is the correct take. There are other axes to
| scale on AND I expect we'll see smaller and smaller models
| approach this level of pre-trained performance. But I believe
| massive pre-training gains have hit clearly diminished
| returns (until I see evidence otherwise).
| granzymes wrote:
| > Compared to OpenAI o1 and OpenAI o3-mini, GPT-4.5 is a more
| general-purpose, innately smarter model. We believe reasoning
| will be a core capability of future models, and that the two
| approaches to scaling--pre-training and reasoning--will
| complement each other. As models like GPT-4.5 become smarter
| and more knowledgeable through pre-training, they will serve
| as an even stronger foundation for reasoning and tool-using
| agents.
| jstummbillig wrote:
| What it confirms, I think, is, that we are going to need a
| _lot_ more chips.
| prisenco wrote:
| Or, possibly, we're stuck waiting for another theoretical
| breakthrough before real progress is made.
| resource0x wrote:
| breakthrough in biology
| georgemcbay wrote:
| Further confirmation, IMO, that the idea that any of this
| leads to anything close to AGI is people getting high on
| their own supply (in some cases literally).
|
| LLMs are a great tool for what is effectively collected
| knowledge search and summary (so long as you are willing to
| accept that you have to verify all of the 'knowledge' they
| spit back because they always have the ability to go off
| the rails) but they have been hitting the limits on how
| much better that can get without somehow introducing more
| real knowledge for close to 2 years now and everything
| since then is super incremental and IME mostly just
| benchmark gains and hype as opposed to actually being
| purely better.
|
| I personally don't believe that more GPUs solves this,
| like, at all. But its great for Nvidia's stock price.
| pseufaux wrote:
| Well said. 100% agree
| BobbyJo wrote:
| I'd put myself on the pessimistic side of all the hype,
| but I still acknowledge that where we are now is a pretty
| staggering leap from two years ago. Coding in particular
| has gone from hints and fragments to full scripts that
| you can correct verbally and are very often accurate and
| reliable.
| georgemcbay wrote:
| I'm not saying there's been no improvement at all. I
| personally wouldn't categorize it as staggering, but we
| can agree to disagree on that.
|
| I find the improvements to be uneven in the sense that
| every time I try a new model I can find use cases where
| its an improvement over previous versions but I can also
| find use cases where it feels like a serious regression.
|
| Our differences in how we categorize the amount of
| improvement over the past 2 years may be related to how
| much the newer models are improving vs regressing for our
| individual use cases.
|
| When used as coding helpers/time accelerators, I find
| newer models to be better at one-shot tasks where you let
| the LLM loose to write or rewrite entire large systems
| and I find them worse at creating or maintaining small
| modules to fit into an existing larger system. My own use
| of LLMs is largely in the latter category.
|
| To be fair I find the current peak model for coding
| assistant to be Claude 3.5 Sonnet which is much newer
| than 2 years old, but I feel like the improvements to get
| to that model were pretty incremental relative to the
| vast amount of resources poured into it and then I feel
| like Claude 3.7 was a pretty big back-slide for my own
| use case which has recently heightened my own skepticism.
| infecto wrote:
| Hilarious. Over two years we went from LLMs being slow
| and not very capable of solving problems to models that
| are incredibly fast, cheap and able to solve problems in
| different domains.
| DannyBee wrote:
| Eh, no. More chips won't save this right now, or probably
| in the near future (IE barring someone sitting on a
| breakthrough right now).
|
| It just means either
|
| A. Lots and lots of hard work that get you a few percent at
| a time, but add up to a _lot_ over time.
|
| or
|
| B. Completely different approaches that people actually
| think about for a while rather than trying to incrementally
| get something done in the next 1-2 months.
|
| Most fields go through this stage. Sometimes more than once
| as they mature and loop back around :)
|
| Right now, AI seems bad at doing either - at least, from
| the outside of most of these companies, and watching open
| source/etc.
|
| While lots of little improvements seem to be released in
| lots of parts, it's rare to see anywhere that is collecting
| and aggregating them en masse and putting them in practice.
| It feels like for every 100 research papers, maybe 1 makes
| it into something in a way that anyone ends up using it by
| default.
|
| This could be because they aren't really even a few percent
| (which would be yet a different problem, and in some ways
| worse), or it could be because nobody has cared to, or ...
|
| I'm sure very large companies are doing a fairly reasonable
| job on this, because they historically do, but everyone
| else - even frameworks - it's still in the "here's a
| million knobs and things that may or may not help".
|
| It's like if compilers had no "O0/O1/O2/O3' at all and were
| just like "here's 16,283 compiler passes - you can put them
| in any order and amount you want". Thanks! I hate it!
|
| It's worse even because it's like this at every layer of
| the stack, whereas in this compiler example, it's just one
| layer.
|
| At the rate of claimed improvements by papers in all parts
| of the stack, either lots and lots and lots is being lost
| because this is happening, in which case, eventually that
| percent adds up to enough for someone to be able to use to
| kill you, or nothing is being lost, in which case, people
| appear to be wasting untold amounts of time and energy,
| then trying to bullshit everyone else, and the field as a
| whole appears to be doing nothing about it. That seems, in
| a lot of ways, even worse. FWIW - I already know which one
| the cynics of HN believe, you don't have to tell me :P.
| This is obviously also presented as black and white, but
| the in-betweens don't seem much better.
|
| Additionally, everyone seems to rush half-baked things to
| try to get the next incremental improvement released and
| out the door because they think it will help them stay
| "sticky" or whatever. History does not suggest this is a
| good plan and even if it was a good plan in theory, it's
| pretty hard to lock people in with what exists right now.
| There isn't enough anyone cares about and rushing out half-
| baked crap is not helping that. mindshare doesn't really
| matter if no one cares about using _your_ product.
|
| Does anyone using these things truly feel locked into
| anyone's ecosystem at this point? Do they feel like they
| will be soon?
|
| I haven't met anyone who feels that way, even in corps
| spending tons and tons of money with these providers.
|
| The public companies - i can at least understand given the
| fickleness of public markets. That was supposed to be one
| of the serious benefit of staying private. So watching
| private companies do the same thing - it's just sort of
| mind-boggling.
|
| Hopefully they'll grow up soon, or someone who takes their
| time and does it right during one of the lulls will come
| and eat all of their lunches.
| gniv wrote:
| > Completely different approaches that people actually
| think about for a while
|
| I think this is very likely simply because there are so
| many smart people looking at it right now. I hope the
| bubble doesn't burst before it happens.
| usaar333 wrote:
| For OpenAI perhaps? Sonnet 3.7 without extended thinking is
| quite strong. Swe-bench scores tie o3
| stavros wrote:
| How do you read those scores? I wanted to see how well 3.7
| with thinking did, but I can't even read that table.
| DebtDeflation wrote:
| GPT 5 is likely just going to be a router model that decides
| whether to send the prompt to 4o, 4o mini, 4.5, o3, or o3
| mini.
| swores wrote:
| My guess is that you're right about that being what's next
| (or maybe almost next) from them, but I think they'll save
| the name GPT-5 for the next actually-trained model (like
| 4.5 but a bigger jump), and use a different kind of name
| for the routing model.
|
| Even by their poor standards at naming it would be weird to
| introduce a completely new type/concept, that can loop in
| models including the 4 / 4.5 series, while naming it part
| of that same series.
|
| My bet: probably something weird like "oo1", or I suspect
| they might try to give it a name that sticks for people to
| think of as "the" model - either just calling it "ChatGPT",
| or coming up with something new that sounds more like a
| product name than a version number (OpenCore, or Central,
| or... whatever they think of)
| JohnnyMarcone wrote:
| They already confirmed GPT-5 will be a unified model
| "months" away. Elsewhere they claimed that it will not
| just be a router but a "unified" model.
|
| https://www.theverge.com/news/611365/openai-
| gpt-4-5-roadmap-...
| DebtDeflation wrote:
| If you read what sama is quoted as saying in your link,
| it's obvious that "unified model" = router.
|
| > "We hate the model picker as much as you do and want to
| return to magic unified intelligence,"
|
| > "a top goal for us is to unify o-series models and GPT-
| series models by creating systems that can use all our
| tools, know when to think for a long time or not, and
| generally be useful for a very wide range of tasks,"
|
| > the company plans to "release GPT-5 as a system that
| integrates a lot of our technology, including o3,"
|
| He even slips up and says "integrates" in the last quote.
|
| When he talks about "unifying", he's talking about the
| user experience not the underlying model itself.
| swores wrote:
| Interesting, thanks for sharing - definitely makes me
| withdraw my confidence in that prediction, though I still
| think there's a decent chance they change their mind
| about that as it seems to me like an even worse naming
| decision than their previous shit name choices!
| lolinder wrote:
| Except minus 4.5, because at these prices and results
| there's essentially no reason not to just use one of the
| existing models if you're going to be dynamically routing
| anyway.
| crystal_revenge wrote:
| > Just confirms reasoning models are the only way forward.
|
| Reasoning models are roughly the equivalent to allow
| Hamiltonian Monte-Carlo models to "warm up" (i.e. start
| sampling from the typical set). This, unsurprisingly, yields
| better results (after all LLMs _are_ just fancy Monte-carlo
| models in the end). However, it is extremely unlikely this
| improvement is without pretty reasonable limitations. Letting
| your HMC warm up is essential to good sampling, but letting
| "warm up more" doesn't result in radically better sampling.
|
| While there have been impressive results in _efficiency_ of
| sampling from the typical set seen in LLMs these days, we 're
| clearly not making the major improvements in the capabilities
| of these models.
| int_19h wrote:
| Reasoning models can solve tasks that non-reasoning ones
| were unable to; how is that not an improvement? What
| constitutes "major" is subjective - if a "minor"
| improvement in overall performance means that the model can
| now successfully perform a task it was unable to solve
| before, that is a major advancement for that particular
| task.
| techorange wrote:
| I wonder how much money they're losing on it too even at those
| prices.
| nialv7 wrote:
| Looks like more signal that the scaling "law" is indeed
| faltering.
| ur-whale wrote:
| AI as it stands in 2025 is an amazing technology, but it is not
| a product _at all_.
|
| As a result, OpenAI simply does not have a business model, even
| if they are trying to convince the world that they do.
|
| My bet is that they're currently burning through other people's
| capital at an amazing rate, but that they are light-years from
| profitability
|
| They are also being chased by fierce competition and OpenSource
| which is very close behind. There simply is no moat.
|
| It will not end well for investors who sunk money in these
| large AI startups (unless of course they manage to find a
| Softbank-style mark to sell the whole thing to), but everyone
| will benefit from the progress AI will have made during the
| bubble.
|
| So, in the end, OpenAI will have, albeit very unwillingly,
| fulfilled their original charter of improving humanity's lot.
| jsheard wrote:
| > My bet is that they're currently burning through other
| people's capital at an amazing rate, but that they are light-
| years from profitability
|
| The Information leaked their internal projections a few
| months ago, and apparently their own estimates have them
| losing $44B between then and 2029 when they expect to finally
| turn a profit, maybe.
| j_maffe wrote:
| That's surprisingly small
| whiplash451 wrote:
| Except that if OpenAI goes bust, very little of what they did
| will actually be released to human kind.
|
| So their contribution was really to fuel a race for
| opensource (which they contributed little to). Pretty complex
| of an argument.
| emptysongglass wrote:
| I've been a Plus user for a long time now. My opinion is
| there is very much a ChatGPT suite of products that come
| together to make for a mostly delightful experience.
|
| Three things I use all the time:
|
| - Canvas for proofing and editing my article drafts before
| publishing. This has replaced an actual human editor for me.
|
| - Voice for all sorts of things, mostly for thinking out loud
| about problems or a quick question about pop culture, what
| something means in another language, etc. The Sol voice is so
| approachable for me.
|
| - GPTs I can use for things like D&D adventure summaries I
| need in a certain style every time without any manual
| prompting.
| vineyardmike wrote:
| > As a result, OpenAI simply does not have a business model,
| even if they are trying to convince the world that they do.
|
| They have a super popular subscription service. If they keep
| iterating on the product enough, they can lag on the models.
| The business is the product not the models and not the API.
| Subscriptions are pretty sticky when you start getting your
| data entrenched in it. I keep my ChatGPT subscription because
| it's the best app on Mac and already started to "learn me"
| through the memory and tasks feature.
|
| Their app experience is easily the best out of their
| competitors (grok, Claude, etc). Which is a clear sign they
| know that it's the product to sell. Things like DeepResearch
| and related are the way they'll make it a sustainable
| business - add value-on-top experiences which drive the
| differentiation over commodities. Gemini is the only
| competitor that compares because it's everywhere in Google
| surfaces. OpenAI's pro tier will surely continue to get
| better, I think more LLM-enabled features will continue to be
| a differentiator. The biggest challenge will be continuing
| distribution and new features requiring interfacing with
| third parties to be more "agentic".
|
| Frankly, I think they have enough strength in product with
| their current models today that even if model training
| stalled it'd be a valuable business.
| jcgrillo wrote:
| https://podcasts.apple.com/us/podcast/better-
| offline/id17305...
| nyarlathotep_ wrote:
| > AI as it stands in 2025 is an amazing technology, but it is
| not a product at all.
|
| Here I'm assuming "AI" to mean what's broadly called
| Generative AI (LLMs, photo, video generation)
|
| I genuinely am struggling to see what the product is too.
|
| The code assistant use cases are really impressive across the
| board (and I'm someone who was vocally against them less than
| a year ago), and I pay for Github CoPilot (for now) but I
| can't think of any offering otherwise to dispute your claim.
|
| It seems like companies are desperate to find a market fit,
| and shoving the words "agentic" everywhere doesn't inspire
| confidence.
|
| Here's the thing: I remember people lining up around the
| block for iPhone releases, XBox launches, hell even Grand
| Theft Auto midnight releases.
|
| Is there a market of people clamoring to use/get anything
| GenAI related?
|
| If any/all LLM services went down tonight, what's the impact?
| Kids do their own homework?
|
| JavaScript programmers have to remember how to write React
| components?
|
| Compare that with Google Maps disappearing, or similar.
|
| LLMs are in a position where they're forced onto people and
| most frankly aren't that interested. Did anyone ASK for
| Microsoft throwing some Copilot things all over their
| operating system? Does anyone want Apple Intelligence,
| really?
| planetafro wrote:
| I think search and chat are decent products as well. I am a
| Google subscriber and I just use Gemini as a replacement
| for search without ads. To me, this movement accelerated
| paid search in an unexpected way. I know the detractors
| will cry "hallucinations" and the ilk. I would counter with
| an argument about the state of the current web besieged by
| ads and misinformation. If people carry a reasonable amount
| of skepticism in all things, this is a fine use case. Trust
| but verify.
|
| I do worry about model poisoning with fake truths but dont
| feel we are there yet.
| sfink wrote:
| > I do worry about model poisoning with fake truths but
| don't feel we are there yet.
|
| In my use, hallucinations will need to be a lot lower
| before we get there, because I already can't trust
| anything an LLM says so I don't think I could even
| distinguish a poisoned fake truth from a "regular"
| hallucination.
|
| I just asked ChatGPT 4o to explain irreducible control
| flow graphs to me, something I've known in the past but
| couldn't remember. It gave me a couple of great
| definitions, with illustrative examples and
| counterexamples. I puzzled through one of the irreducible
| examples, and eventually realized it wasn't irreducible.
| I pointed out the error, and it gave a more complex
| example, also incorrect. It finally got it on the 3rd
| try. If I had been trying to learn something for the
| first time rather than remind myself of what I had once
| known, I would have been hopelessly lost. Skepticism
| about any response is still crucial.
| dzjkb wrote:
| speaking of search without ads, I wholeheartedly
| recommend https://kagi.com
| dpe82 wrote:
| I'll second this. Kagi is really impressive and ad-free
| is a nice change.
| otabdeveloper4 wrote:
| > I genuinely am struggling to see what the product is too.
|
| They're nice for summarizing and categorizing text. We've
| had good solutions for that before, too (BERT, et al), but
| LLM's are marginally nicer.
|
| > Is there a market of people clamoring to use/get anything
| GenAI related?
|
| No. LLM's are lame and uncool. Kids especially dislike them
| a lot on that basis alone.
| ur-whale wrote:
| > LLM's are lame and uncool. Kids especially dislike them
| a lot on that basis alone.
|
| Not just kids.
| beefnugs wrote:
| Yes: the real truth is, if there really was a good AI
| created, then we wouldnt even know about it existing until a
| billion dollar company takes over some industry with only a
| handful of developers in the entire company. Only then would
| hints spill out into the world that its possible.
|
| No "good" AI will ever be open to everyone and relatively
| cheap, this is the same phenomenon as "how to get rich" books
| yard2010 wrote:
| Sir they are selling text by the ounce just like farmers sold
| tomatoes before Walmart, How is that not a business model?
| jdprgm wrote:
| If it really costs them 30x more surely they must plan on
| putting pretty significant usage limits on any rollout to the
| Plus tier and if that is the case i'm not sure what the point
| is considering it seems primarily a replacement/upgrade for 4o.
|
| The cognitive overhead of choosing between what will be 6
| different models now on chatGPT and trying to map whether a
| query is "worth" using a certain model and worrying about
| hitting usage limits is getting kind of out of control.
| JohnnyMarcone wrote:
| To be fair their roadmap states that gpt-5 will unify
| everything into one model in "months".
| chollida1 wrote:
| > GPT 4.5 pricing is insane:
|
| > I'm still gonna give it a go, though.
|
| Seems like the pricing is pretty rational then?
| phito wrote:
| Not if people just try a few prompts then stop using it.
| chollida1 wrote:
| Sure but its in their best interest to lower it then and
| only then.
|
| OpenAI wouldn't be the first company to price something
| expensive when it first comes out to capitalize on people
| who are less price sensitive at first and then lower prices
| to capture a bigger audience.
|
| That's all pricing 101 as the saying goes.
| j_maffe wrote:
| If OAI are concerning themselves with collecting a few
| hundereds from a small group of individuals then they
| really have nothing better to do
| nyarlathotep_ wrote:
| How much of OAI's reported users are doing exactly this?
| raytopia wrote:
| Now the real question about AI automation starts. Is it cheaper
| to pay a human to do the task or a AI company?
| redox99 wrote:
| It still not smart enough to replace for example customer
| service.
| infecto wrote:
| It's absolutely able to replace the majority of customer
| service volume which is full of mundane questions.
| beefnugs wrote:
| Such brutal reductionism: how do you calculate an ever
| growing percentage of customers so pissed at this
| terrible service that you lose customers forever? Not
| just one company losing customers... but an entire
| population completely distrusting and pulling back from
| any and all companies pulling this trash
| infecto wrote:
| Huh? Most call centers these days already use ivr systems
| and they absolutely are terrible experiences. I along
| with most people would happily speak with a LLM backed
| agent to resolve issues.
|
| The CS is already a wreck and LLMs beat an ivr any day of
| the week and have the ability to offer real triaging
| ability.
|
| The only people getting upset are the luddites like
| yourself.
| fragmede wrote:
| Humans have all sorts of issues you have to deal with. Being
| hungover, not sleeping well, having a personality, being late
| to work, not being able to work 24/7, very limited ability to
| copy them. If there's a soulless generic office-droidGPT that
| companies could hire that would never talk back and would do
| all sorts of menial work without needing breaks or to use the
| bathroom, I don't know that we humans stand a chance!
|
| I have a bunch of work that needs doing. I can do it myself,
| or I can hire one person to do it. I gotta train them and
| manage them and even after I train them theres still only
| going to be one of them, and it's subject to their
| availability. On the other hand, if I need to train an AI to
| do it, but I can copy that AI, and then spin them up/down
| like on demand computer in the cloud, and not feel remotely
| bad about spinning them down?
|
| It's definitely not there yet, but it's not hard to see the
| business case for it.
| OutOfHere wrote:
| Once we get to that stage, unless you're a capitalist,
| remember that your job is next in line to be replaced.
| NoGravitas wrote:
| Every tech drone in every cubicle considers themselves a
| temporarily embarrassed capitalist.
| fragmede wrote:
| I write code for a living. My entire profession _is_ on
| the line, thanks to ourselves. My eyes are wide open on
| the situation at hand though. Burying my head in the sand
| and pretending what I wrote above isn 't true, isn't
| going to make it any less true.
|
| I'm not sure what I can do about it, either. My job
| already doesn't look like it did a year ago, nevermind a
| decade away.
| OutOfHere wrote:
| I keep telling coders to switch to being 1-person
| enterprise shops instead, but they don't listen. They
| will learn the hard way when they suddenly find
| themselves without a job due to AI having taken it away.
| As for what enterprise, use your imagination without bias
| from coding.
| Horffupolde wrote:
| This is the ultimate business model.
| rrrrrrrrrrrryan wrote:
| I was about to comment that humans consume orders of
| magnitude less energy, but then I checked the numbers, and it
| looks like an average person consumes way more energy
| throughout their day (food, transportation, electricity
| usage, etc) than GPT-4.5 would at 1 query per minute over 24
| hours.
| crooked-v wrote:
| Doubly so with how good Claude 3.7 Sonnet is at $3 / 1M tokens.
| Hansenq wrote:
| GPT-4.5 is 15-30x more expensive than GPT-4o. Likely that much
| larger in terms of parameter count too. It's massive!!
|
| With more parameters comes more latent space to build a world
| model. No wonder its internal world model is so much better
| than previous SOTA
| wavemode wrote:
| This has been my suspicion for a long time - OpenAI have indeed
| been working on "GPT5", but training and running it is proving
| so expensive (and its actual reasoning abilities only
| marginally stronger than GPT4) that there's just no market for
| it.
|
| It points to an overall plateau being reached in the
| performance of the transformer architecture.
| goatlover wrote:
| Certainly hope so. The tech billionaires are little to
| excited to achieve AGI and replace the workforce.
| shoubidouwah wrote:
| TBH, with the safety/alignment paradigm we have, workforce
| replacement was not my top concern when we hit AGI. A pause
| / lull in capabilities would be hugely helpful so that we
| can figure how not to die along with the lightcone...
| hnuser123456 wrote:
| Is it inevitable to you that someone will create some
| kind of techno-god behemoth AI that will figure out how
| to optimally dominate an entire future light cone
| starting from the point in spacetime of its self-
| actualization? Borg or Cylons?
| mupuff1234 wrote:
| Not sure how why anyone thinks it's possible to fully
| control AGI, we cant even fully tame a house cat.
| JohnnyMarcone wrote:
| I feel like this period has shown that we're not quite
| ready for a machine god. We'll see if RL hits a wall as
| well.
| camdenreslink wrote:
| That would certainly reduce my anxiety about the future of my
| chosen profession.
| malthaus wrote:
| but while there is a plateau in the transformer architecture,
| what you can do with those base models by further finetuning
| / modifying / enhancing them is still largely unexplored so i
| still predict mind-blowing enhancements yearly for this
| foreseeable future. if they validate openai's valuation and
| investment needs is a different question.
| hintymad wrote:
| > It sounds like it's so expensive and the difference in
| usefulness is so lacking(?) they're not even gonna keep serving
| it in the API for long
|
| I guess the rationale behind this is paying for the marginal
| improvement. Maybe the next few percent of improvement is so
| important to a business that the business is willing to pay a
| hefty premium.
| shawabawa3 wrote:
| I wonder if the pricing is partly to discourage distillation,
| if they suspect r1 was distilled from gpt 4o
| schneehertz wrote:
| Mainly to prevent you from using it
| tomrod wrote:
| I can chew through 1MM tokens with a single standard (and
| optimized) call. This pricing is insane.
| MangoCoffee wrote:
| one of the problem seem to be there's no alternative to Nvidia
| ecosystem. (the gpu + CUDA).
| kridsdale3 wrote:
| May I introduce you to Gemini 2.0
| Fnoord wrote:
| ZLUDA can be used as compatibility glue, also you can use
| ROCm or even Vulcan with Ollama.
| coliveira wrote:
| In other words, they want people to pay for the privilege of
| becoming beta testers....
| wiremine wrote:
| > GPT 4.5 pricing is insane: Price Input: $75.00 / 1M tokens
| Cached input: $37.50 / 1M tokens Output: $150.00 / 1M tokens
|
| > GPT 4o pricing for comparison: Price Input: $2.50 / 1M tokens
| Cached input: $1.25 / 1M tokens Output: $10.00 / 1M tokens
|
| Their examples don't seem 30x better. :-)
| ren_engineer wrote:
| hyperscalers in shambles, no clue why they even released this
| other than the fact they didn't want to admit they wasted an
| absurd amount of money for no reason
| hn_throwaway_99 wrote:
| The price is obviously 15-30x that of 4o, but I'd just posit
| that there are some use cases where it may make sense. It
| probably doesn't make sense for the "open-ended consumer facing
| chatbot" use case, but for other use cases that are fewer and
| higher value in nature, it could if it's abilities are
| considerably better than 4o.
|
| For example, there are now a bunch of vendors that sell
| "respond to RFP" AI products. The number of RFPs that any sales
| organization responds to is probably no more than a couple a
| week, but it's a very time-consuming, laborious process. But
| the payoff is obviously very high if a response results in a
| closed sale. So here paying 30x for marginally better
| performance makes perfect sense.
|
| I can think of a number of similar "high value, relatively low
| occurrence" use cases like this where the pricing may not be a
| big hindrance.
| Manouchehri wrote:
| Yeah, agreed.
|
| We're one of those types of customers. We wrote an OpenAI API
| compatible gateway that automatically batches stuff for us,
| so we get 50% off for basically no extra dev work in our
| client applications.
|
| I don't care about speed, I care about getting the right
| answer. The cost is fine as long as the output generates us
| more profit.
| janoc wrote:
| And which use case will that make sense then for?
|
| Esp. when they aren't even sure whether they will commit to
| offering this long term? Who would be insane enough to build
| a product on top of something that may not be there tomorrow?
|
| Those products require some extensive work, such a model
| finetuning on proprietary data. Who is going to invest time &
| money into something like that when OpenAI says right out of
| the gate they may not support this model for very long?
|
| Basically OpenAI is telegraphing that this is yet another
| prototype that escaped a lab, not something that is actually
| ready for use and deployment.
| superq wrote:
| Complete legal arguments as well. If I was an attorney, I'd
| love to have a sophisticated LLM write my crib notes for
| anything I might do or say in the court room, or even the
| complete direction that I'd take my case. For some cases,
| that'd be worth almost any price.
| kristofferR wrote:
| "GPT-4.5 is not a frontier model, but it is OpenAI's largest
| LLM, improving on GPT-4's computational efficiency by more than
| 10x."[1]
|
| I don't get it, it is supposedly much cheaper to run?
|
| [1] https://cdn.openai.com/gpt-4-5-system-card.pdf (page 7,
| bottom)
| acchow wrote:
| > It sounds like it's so expensive and the difference in
| usefulness is so lacking(?)
|
| The claimed hallucination rate is dropping from 61% to 37%.
| That's a "correct" rate increasing from 29% to 63%.
|
| Double the correct rate costs 15x the price? That seems absurd,
| unless you think about how mistakes compound. Even just 2 steps
| in and you're comparing a 8.4% correct rate vs 40%. 3 automated
| steps and it's 2.4% vs 25%.
| einrealist wrote:
| And remember, with increasing accuracy, the cost of
| validation goes up (not even linear).
|
| We expect computers to be right. Its a trust problem. Average
| users will simply trust the results of LLMs and move on
| without proper validation. And the way the LLMs are trained
| to mimic human interaction is not helping either. This will
| reduce overall quality in society.
|
| Its a different thing to work with another human, because
| there is intention. A human wants to be correct or to mislead
| me. I am considering this without even thinking about it.
|
| And I don't expect expert models to improve things, unless
| the problem space is really simple (like checking eggs for
| anomalies).
| isk517 wrote:
| 30x price bump feels like a attempt to pull in as much money as
| possible before the bubble bursts.
| dr_kiszonka wrote:
| To me, it feels like a PR stunt in response to what the
| competition is doing. OpenAI is trying to show how they are
| ahead of others, but they price the new model to minimize its
| use. Potentially, Anthropic et al. also have amazing models
| that they aren't yet ready to productionize because of costs.
| ic4l wrote:
| Did they already disable it?
|
| When using `gpt-4.5-preview` I am getting: > Invalid URL (POST
| /v1/chat/completions)
| bawolff wrote:
| > It sounds like it's so expensive and the difference in
| usefulness is so lacking(?) they're not even gonna keep serving
| it in the API for long:
|
| Sounds like an attempt at price descrimination. Sell the
| expensive version to big companies with big budgets who don't
| care, sell the cheap version to everyone else. Capture both
| ends of the market.
| OldGreenYodaGPT wrote:
| This is was GPT4 cost when it was released
| campers wrote:
| The price will come down over time as they apply all the
| techniques to distill it down to a smaller parameter model.
| Just like GPT4 pricing came down significantly over time.
| osigurdson wrote:
| I don't understand the pricing for cached tokens. It seems
| rather high for looking up something in a cache.
| mmaunder wrote:
| But you get higher EQ. /s
| muzani wrote:
| For comparison, 3 years ago, the most powerful model out there
| (GPT-3 davinci) was $60/MTok.
| quantadev wrote:
| It's crazy expensive because they want to pull in as much
| revenue as possible as fast as possible before the Open Source
| models put them outta business.
| kla-s wrote:
| Well to play the devils advocat, i think this is useful to
| have, at least for 'Open'Ai to start off from to apply QLora or
| similar approximations.
|
| Bonus they could even do some self learning afterwards with the
| performance improvements DeepSeek just published and it might
| have more EQ and less hallucinations than starting from
| scratch...
|
| ie the price might go down big time but there might be
| significant improvements down the line when starting from such
| a broad base
| dtnewman wrote:
| Really depends on your use case. For low value tasks this is
| way too expensive. But for context, let's say a court opinion
| is an average of 6000 words. Let's say i want to analyze 10
| court opinions and pull some information out that's relevant to
| my case. That will run about $1.80 per document or $18 total. I
| wouldn't pay that just to edify myself, but i can think of many
| use cases where it's still a negligible cost, even if it only
| does 5% better than the 30x cheaper model.
| Foobar8568 wrote:
| Someone in another comment said that gpt-4 32k had somewhat the
| same cost (ok 10% cheaper), what was a pain was more the
| latency and speed than actual cost given the increase in
| productivity for our usage.
| DonHopkins wrote:
| Maybe they started a really long expensive training session,
| and Elon Musk's DOGE script kiddies somehow broke in and
| sabotaged it, so it got disrupted and turned into the
| Eraserhead baby, but they still want to get it out there for a
| little while before it died to squeeze all the money out of it
| as possible, because it was so expensive to train.
|
| https://www.youtube.com/watch?v=ZZ-kI4Qzj9U
| DonHopkins wrote:
| >GPT 4.5 pricing is insane: Price Input: $75.00 / 1M tokens
| Cached input: $37.50 / 1M tokens Output: $150.00 / 1M tokens
|
| How many eggs does that include??!
| fvv wrote:
| usefulness is bound to scope/purpose, even if innovation stops,
| in 3y (thanks to hw and tuning progress ) when 4o costs 0.1$/M
| and 4.5 1$/M even being a small improvement ( which is not imo
| ), you will chose to use 4.5 , exactly like no one now want to
| use 3.5
| UrineSqueegee wrote:
| It's priced like this because it can generate erotica.
| madduci wrote:
| Let's see if DeepSeek will make a distillation of this model as
| well
| williamsss wrote:
| The performance bump doesn't justify the steep price
| difference.
|
| From a for profit business lens for OpenAI - I understand
| pushing the price outside the range of side projects, but this
| pushes it past start ups.
|
| Excited to see new stuff released past reasoning models in any
| case. Hope they can improve the price soon.
| MagicMoonlight wrote:
| I put "hello" into it and it billed me 30p for it. Absolutely
| unusable, more expensive than realtime voice chat.
| FuckButtons wrote:
| I suspect this is GPT-5. This is the biggest model they made
| and they got very little ROI hence the re-branding.
| lend000 wrote:
| My understanding is that o1 is a system built on GPT-4o, so
| this pricing might explain why o3 (the alleged full version)
| cost so much money to run in the published benchmark tests [0].
| It must be using GPT 4.5 or something similar as the underlying
| model.
|
| [0] https://arcprize.org/blog/oai-o3-pub-breakthrough
| Xiol32 wrote:
| The example GPT-4.5 answers from the livestream are just... too
| excitable? Can't put my finger on it, but it feels like they're
| aimed towards little kids.
| jug wrote:
| It made me wonder how much of that was due to the system prompt
| too.
| virgildotcodes wrote:
| That presentation was super underwhelming. We got to watch them
| compare... the vibes? ... of 4.5 vs o1.
|
| No wonder Sam wasn't part of the presentation.
| Etheryte wrote:
| And to top it off, it costs $75.00 per 1M vibes.
| thomas34298 wrote:
| Sam tweeted "taking care of my kid in the hospital":
|
| https://x.com/sama/status/1895210655944450446
|
| Let's not assume that he's lying. Neither the presentation nor
| my short usage via the API blew me away, but to really evaluate
| it, you'd have to use it longer on a daily basis. Maybe that
| becomes a possiblity with the announced performance
| optimizations that would lower the price...
| jug wrote:
| It should've just been a web launch without video.
| bakugo wrote:
| API is literally 5 times more expensive than Claude 3 Opus, and
| it doesn't even seem to do anything impressive. What's the
| business strategy here?
| hidelooktropic wrote:
| I'm not sure that doing a live stream on this was the right way
| to go. I would've just quietly sent out a press release. I'm sure
| they have better things on the way.
| sebastiennight wrote:
| It is interesting that they are focusing a large part of this
| release on the model having a higher "EQ" (Emotional Quotient).
|
| We're far from the days of "this is not a person, we do not want
| to make it addictive" and getting a firm foot on the territory of
| "here's your new AI friend".
|
| This is very visible in the example comparing 4o with 4.5 when
| the user is complaining about failing a test, where 4o's response
| is what one would expect from a "typical AI response" with
| problem-solving bullets, and 4.5 is sending what you'd expect
| from a pal over instant messaging.
|
| It seems Anthropic and Grok have both been moving in this
| direction as well. Are we going to see an escalation of
| foundation models impersonating "a friendly person" rather than
| "a helpful assistant"?
|
| Personally I find this worrying and (as someone who builds upon
| SOTA model APIs) I really hope this behavior is not going to seep
| into API responses, or will at least be steerable through the
| system/developer prompt.
| og_kalu wrote:
| The whole robotic, monotone, helpful assistant thing was
| something these companies had to actively hammer in during the
| post-training stage. It's not really how LLMs will sound by
| default after pre-training.
|
| I guess they're caring less and less about that effort
| especially since it hurts the model in some ways like creative
| writing.
| sebastiennight wrote:
| If it's just a different choice during RLHF, I'll be curious
| to see what are the trade-offs in performance.
|
| The "buddy in a chat group" style answers do not make me feel
| like asking it for a story will make the story
| long/detailed/poignant enough to warrant the difference.
|
| I'll give it a try and compare on creative tasks.
| turnsout wrote:
| Or maybe they're just getting better at it, or developing
| better taste. After switching to Claude, I can't go back to
| ChatGPT's overly verbose bullet-point laden book reports
| every time I ask a question. I don't think that's pretraining
| --it's in the way OpenAI approaches tuning and prompting vs
| Anthropic.
| capnrefsmmat wrote:
| Maybe, but I'm not sure how much the style is deliberate vs.
| a consequence of the post-training tasks like summarization
| and problem solving. Without seeing the post-training tasks
| and rating systems it's hard to judge if it's a deliberate
| style or an emergent consequence of other things.
|
| But it's definitely the case that base models sound more
| human than instruction-tuned variants. And the shift isn't
| just vocabulary, it's also in grammar and rhetorical style.
| There's a shift toward longer words, but also participial
| phrases, phrasal coordination (with "and" and "or"), and
| nominalizations (turning adjectives/adverbs into nouns, like
| "development" or "naturalness").
| https://arxiv.org/abs/2410.16107
| tmaly wrote:
| I would like to see a humor test. So far, I have not seen any
| model response that has made me laugh.
| sebastiennight wrote:
| The "roast" tools that have popped up (using either DeepSeek
| or o3-mini) are pretty funny.
|
| Eg. https://news.ycombinator.com/item?id=43163654
| jcims wrote:
| OK now that is some funny shit.
| AgentME wrote:
| My benchmark for this has been asking the model to write some
| tweets in the style of dril, a popular user who writes short
| funny tweets. Sometimes I include a few example tweets in the
| prompt too. Here's an example of results I got from Claude 3
| Opus and GPT 4 for this last year:
| https://bsky.app/profile/macil.tech/post/3kpcvicmirs2v. My
| opinion is that Claude's results were mostly bangers while
| GPT's were all a bit groanworthy. I need to try this again
| with the latest models sometime.
| tkgally wrote:
| How does the following stand-up routine by Claude 3.7 Sonnet
| work for you?
|
| https://gally.net/temp/20250225claudestandup2.html
| lurker9001 wrote:
| incredible
| aprilthird2021 wrote:
| Reading this felt like reading junk food
|
| EDIT: Junk food tastes kinda good though. This felt like
| drinking straight cooking oil. Tastes bad and bad for you.
| fragmede wrote:
| I chuckled.
|
| Now you just need a Pro subscription to get Sora generate a
| video to go along with this and post it to YouTube and rake
| in the views (and the money that goes along with it).
| thousand_nights wrote:
| reddit tier humor, truly
|
| it's just regurgitating overly emphasized cliches in a
| disgustingly enthusiastic tone
| willy_k wrote:
| Is that any different from the bulk of standup today?
| sebastiennight wrote:
| That was impressive. If it all came from just this short
| 4-line prompt, it's even more impressive.
|
| All we're missing now is a text-to-video (or text+audio and
| then audio-to-video) that can convincingly follow the style
| instructions for emphasis and pausing. Or are we already
| there yet?
| tkgally wrote:
| Yes, that was the full prompt.
|
| Yesterday, I had Claude 3.7 write a full 80,000-word
| novel. My prompt was a bit longer, but the result was
| shockingly good. The new thinking mode is very
| impressive.
| jdiez17 wrote:
| Okay, you know what? I laughed a few times. Yeah it may not
| work as an actual stand up routine to a general audience,
| it's kinda cringe (as most LLM-generated content), but it
| was legitimately entertaining to read.
| turnsout wrote:
| If you like absurdist humor, go into the OpenAI playground,
| select 3.5-Turbo, and dial up the temperature to the point
| where the output devolves into garbled text after 500 tokens
| or so. The first ~200 tokens are in the freaking sweet spot
| of humor.
| amarcheschi wrote:
| Could someone post an example?
| rl3 wrote:
| Maybe it's rose-colored glasses, but 3.5 was really the
| golden era for LLM comedy. More modern LLMs can't touch it.
|
| Just ask it to write you a film screenplay involving some
| hard-ass 80s/90s action star and someone totally unrelated
| and opposite of that. The ensuring unhinged magic is
| unparalleled.
| jcims wrote:
| I built a little AI assistant to read my calendar and
| send me a summary of my day every morning. I told it to
| roast me and be funny with it.
|
| 3.5 was *way* better than anything else at that.
| rl3 wrote:
| Yeah, I think the fact its "mind" so to speak was more
| fragmented and unpredictable was almost a boon for that
| purpose.
| sebastiennight wrote:
| Ah, I'd love to have that kind of daily recap... mind
| sharing some of the code (or even just the prompt?)
| rl3 wrote:
| > _The ensuring unhinged magic is unparalleled._
|
| Oops: ensuing*
| immibis wrote:
| ChatGPT gave me this shell script: https://social.immibis.com
| /media/7102ac83cf4a200e48dd368938e... (obviously, don't
| download and execute a random shell script from the internet
| without reading it first)
|
| I think reading it will make you laugh.
| nialv7 wrote:
| Well yeah, if the llm can keep you engaged and talking, that'll
| make them a lot more money; compared to if you just use it as a
| information retrieval tool in which case you are likely to
| leave after getting what you are looking for.
| TheAceOfHearts wrote:
| Since they offer a subscription, keeping you engaged just
| requires them to waste more compute. The ideal case would be
| that the LLM gives you a one shot correct response using as
| little compute as possible.
| sebastiennight wrote:
| In a subscription business, you don't want the user to use
| as few resources as possible. It's the wrong optimization
| to make.
|
| You want users to keep coming back as often as possible (at
| the lowest cost-per-run possible though). If they are not
| coming back they are not renewing.
|
| So, yes, it makes sense to make answers shorter to cut on
| compute cost (which these SMS-length replies could
| accomplish) but the main point of making the AI flirtatious
| or "concerned" is possibly the addictive factor of having a
| shoulder to cry on 24/7, one that does not call you on your
| BS and is always supportive... for just $20 a month
|
| The "one-shot correct response" to "I failed my exams"
| might be "Tough luck, try better next time" but if you do
| that, you will indeed use very little compute _because
| people will cancel the subscription and never come back_.
| johnthewise wrote:
| AI subscriptions are already very sticky . I can't
| imagine at least not paying for one, so I doubt they care
| about retention like the rest of us plebs do.
| player1234 wrote:
| First imagine paying a subscription fee which actually
| makes the company profitable and gives investors ROI,
| then I think you can also imagine not paying that amount
| at all.
| nialv7 wrote:
| Plus level subscription has limits too, and Pro level costs
| 10x more - as long as Pro users don't use ChatGPT 10x more
| than Plus users on average, OpenAI can benefit. There's
| also the user retention factor.
| bredren wrote:
| Yes, the "personality" (vibe) of the model is a key qualitative
| attribute of gpt-4.5.
|
| I suspect this has something to do with shining light on an
| increased value prop in a dimension many people will appreciate
| since gains on quantitative comparison with other models were
| not notable enough to pop eyeballs.
| callc wrote:
| > We're far from the days of "this is not a person, we do not
| want to make it addictive" and getting a firm foot on the
| territory of "here's your new AI friend".
|
| That's a hard nope from me, when companies pull that move. I'll
| stick to my flesh and blood humans who still hallucinate but
| only rarely.
| orbital-decay wrote:
| Anthropic pretty much abandoned this direction after Claude 3,
| and said it wasn't what they wanted [1]. Claude 3.5+ is
| extremely dry and neutral, it doesn't seem to have the same
| training.
|
| _> Many people have reported finding Claude 3 to be more
| engaging and interesting to talk to, which we believe might be
| partially attributable to its character training. This wasn't
| the core goal of character training, however. Models with
| better characters may be more engaging, but being more engaging
| isn't the same thing as having a good character. In fact, an
| excessive desire to be engaging seems like an undesirable
| character trait for a model to have._
|
| [1] https://www.anthropic.com/research/claude-character
| Kye wrote:
| It's the opposite incentive to ad-funded social media. One
| wants to drain your wallet and keep you hooked, the other
| wants you to spend as little of their funding as possible
| finding what you're looking for.
| sureIy wrote:
| I don't know if I fully agree. The input clearly shows the need
| for emotional support more than "how do I pass this test?" The
| answer by 4o is comical even if you know you're talking to a
| machine.
|
| It reminds me of the advice to "not offer solutions when a
| woman talks about her problems, but just listen."
| neuroticnews25 wrote:
| How could a machine provide emotional support? When I ask
| questions like this to LLMs, it's always to brainstorm
| solutions. I get annoyed when I receive fake-attention
| follow-up questions instead.
|
| I guess there's a trade-off between being human and being
| useful. But this isn't unique to LLMs, it's similar to how
| one wouldn't expect a deep personal connection with a
| customer service professional.
| cynicalpeace wrote:
| There are some businesses trying to do emotional support
| with AI, like AI GF's, etc
|
| Some will make some profit as a niche thing (millions of
| users on a global scale, and if unit economics work, can
| make millions of $)
|
| But it seems it will never be something really mainstream
| because most normal people don't care what a bot says or
| does.
|
| The example I always think of is chess bots have been
| better at chess than humans for decades. But very few
| people watch stockfish tournaments. Everyone loves Magnus
| Carlsen though.
|
| This is 100x for emotional support type things.
| aprilthird2021 wrote:
| I think it's a good thing because, idk why, I just start tuning
| out after getting reams and reams of bullet points I'm already
| not super confident about the truthfulness of
| CitizenTen wrote:
| GPT pro already has already rummored to be 100k users. You think
| GPT 4.5 will add to that even with the insane costs for corporate
| users?
| cristiancavalli wrote:
| What rumors? I looked and can't find something to substantiate
| that #
| pepsi-not-coke wrote:
| in my experience, o3-mini-high while still unpredictable as
| it modifies and ignores parts of my code when I specifically
| tell it not to do so (e.g. "don't touch anything else!") is
| the best AI coding tool out there, far better than Claude
|
| so Pro is worth it for O3-mini-high
| spiderfarmer wrote:
| 100k per user perhaps.
| andsoitis wrote:
| "this isn't a reasoning model and won't crush benchmarks."
|
| -- https://x.com/sama/status/1895203654103351462
| MaxPock wrote:
| One thing that Altman does extremely well is to over-promise and
| under-deliver.
| erulabs wrote:
| Finally a scaling wall? This is apparently (based on pricing)
| using about an order of magnitude more compute, and is only maybe
| 10% more intelligent. Ideally DeepSeeks optimizations help bring
| the costs way down, but do any AI researchers want to comment on
| if this changes the overall shape of the scaling curve?
| fpgaminer wrote:
| Seems on par with the existing scaling curve. If I had to
| speculate, this model would have been an internal-only model,
| but they're releasing it for PR. An optimized version with 99%
| of the performance for 1/10th the cost will come out later.
| j_maffe wrote:
| This is the shittiest PR move I've seen since the AI trend
| started.
| Workaccount2 wrote:
| At least so far it's coding performance is bad, but from
| what I have seen it's writing abilities are totally insane.
| It doesn't read like AI output anymore.
| j_bum wrote:
| Any examples you'd be willing to share?
| chippiewill wrote:
| They have examples in the announcement post. It does a
| better job of understanding intent in the question which
| helps it give an informal rather than essay style
| response where appropriate.
| j_maffe wrote:
| I wouldn't call that "too insane." As others have pointed
| out, you can get similar results from fine-tuning the
| RLHF.
| csomar wrote:
| We have hit that wall almost 2 years ago with gpt-4. There was
| clearly no scaling as gpt-4 was already decently smart and if
| you got x2 smarter you'll be more capable than anything on the
| market today. All models doing today (R1 and friends; and
| Claude) are trying to optimize this local maxima toward
| generating more useful responses (ie: code when it comes to
| Claude).
|
| AI, at its current form, is a Deep Seek of compressed knowledge
| in a 30-50gb of interconnected data. I think we'll look at this
| as trying to train networks on corpus of data and expecting
| them to have a hold of reality. Our brains are trained on
| "reality" which is not the "real" reality as your vision is
| limited to the visible spectrum. But if you want a network to
| behave like a human then maybe give him what a human see.
|
| There is also the possibility that there is a physical limit to
| intelligence. I don't see any elephants doing PhDs and the
| smartest of humans are just a small configuration away from
| insanity.
| killerstorm wrote:
| It depends on how you compare.
|
| On a subset of tasks I'm interested in, it's 10x more
| intelligent than GPT-4. (Note that GPT-4 was in many ways
| better than 4o.)
|
| It's not a coding champion, but it knows A LOT of stuff,
| excellent common sense, top quality writing. For me it's like
| "deep research lite".
|
| I found OpenAI Deep research excellent, but GPT-4.5 might in
| many cases beat it.
| myflash13 wrote:
| > On a subset of tasks I'm interested in, it's 10x more
| intelligent than GPT-4.
|
| Very intriguing. Care to share an example?
| vbezhenar wrote:
| The price is 2x from GPT4. So probably not a decimal order of
| magnitude.
| ein0p wrote:
| Now imagine this model (or an optimized/slightly downsized
| variant thereof) as a base for a "thinking" one.
| bparsons wrote:
| Are they saying that 4.5 has a 35% hallucination rate? That chart
| is a bit confusing.
| riku_iki wrote:
| its on that benchmark, which likely is very challenging.
| Seattle3503 wrote:
| The GPT-1 response to the example prompt "What was the first
| language?" got a chuckle out of me
| aldanor wrote:
| The question being, will we be chuckling at current models
| responses in 5-10y from now?
| joshuamcginnis wrote:
| I'm one week in on heavy grok usage. I didn't think I'd say this,
| but for personal use, I'm considering cancelling my OpenAI plan.
|
| The one thing I wish grok had was more separation of the UI from
| X itself. The interface being so coupled to X puts me off and
| makes it feel like a second-hand citizen. I like ChatGPTs
| minimalist UI.
| aldanor wrote:
| Theres grok.com which is standalone and with its own UI
| it wrote:
| There's also a standalone Grok app at least on iOS.
| pzo wrote:
| I wish they did also dedicated keyboard app like SwiftKey
| that has copilot integration
| richard_todd wrote:
| I find grok to be the best overall experience for the types of
| tasks I try to give AI (mostly: analyze pdf, perform and
| proofread OCR, translate Medieval Latin and Hebrew, remind me
| how to do various things in python or SwiftUI).
| ChatGPT/gemini/copilot all fight me occasionally, but grok just
| tries to help. And the hallucinations aren't as frequent, at
| least anecdotally.
| fzzzy wrote:
| Don't they have a standalone Grok app now? I thought I saw
| that. [edit] ah some sibling comments mention this as well
| andxor wrote:
| https://grok.com/
| bn-l wrote:
| They still haven't released an API for 3
| misiti3780 wrote:
| I canceled my GPT, Grok is incredible.
| barfingclouds wrote:
| There's a grok app for iPhone that's basically the same as
| ChatGPT/deepseek/mistral/gemini/claude
| selalipop wrote:
| I've been working on post-training models for tasks that require
| EQ, so it's validating to see OpenAI working towards that too.
|
| That being said, this is very expensive.
|
| - Input: $75.00 / 1M tokens
|
| - Cached input: $37.50 / 1M tokens
|
| - Output: $150.00 / 1M tokens
|
| One of the most interesting applications of models with higher EQ
| is personalized content generation, but the size and cost here
| are at odds with that.
| kgeist wrote:
| >GPT-4.5 is more succinct and conversational
|
| I wonder why they highlight it as an achievement when they could
| have simply tuned 4o to be more conversational and less like a
| bullet-point-style answer machine. They did something to 4o
| compared to the previous models which made the responses feel
| more canned.
| qgin wrote:
| Poossibly, but reports seem to indicate that 4.5 is much more
| nuanced and thoughtful in its language use. It's not just being
| shorter and casual as a style, there is a higher amount of
| "conceptual resolution" within the words being used.
| mvdtnz wrote:
| OpenAI doubling down on the American-style therapy-speak instead
| of focusing on usefulness. No thanks.
| wewewedxfgdf wrote:
| I feel like OpenAI is pursuing AGI when Anthropic/Claude is
| pursuing making AI awesome for practical things like coding.
|
| I only ever using OpenAI's coding now as a double check against
| Claude.
|
| Does OpenAI have their eyes on the ball?
| ls_stats wrote:
| >I feel like OpenAI is pursuing AGI
|
| I don't think so, the "AGI guy" was Ilya Sutskever, he is gone,
| he wanted to make OpenAI "less comercial", AGI is just a
| buzzword for Altmann.
| rakejake wrote:
| Right. A good chunk of the "old guard" is now gone - Ilya to
| SSI, Mira and a bunch of others to a new venture called
| Thinking Machines, Alec Radford etc. Remains to be seen if
| OpenAI will be the leader or if other players catch up.
| sumedh wrote:
| The page still has Mira Muratis name under Exec Leadership
| rakejake wrote:
| My usage has come down to mostly Claude (until I run out of
| free tier quota) and then Gemini. Claude is the best for code
| and Gemini 2.0 Flash is good enough while also being free (well
| considering how much data G has hoovered up over the years,
| perhaps not) and more importantly highly available.
|
| For simple queries like generating shell scripts for some
| plumbing, or doing some data munging, I go straight to Gemini.
| HarHarVeryFunny wrote:
| > My usage has come down to mostly Claude (until I run out of
| free tier quota) and then Gemini
|
| Yep, exactly same here.
|
| Gemini 2.0 Flash is extremely good, and I've yet to hit any
| usage limits with them - for heavy usage I just go to Gemini
| directly. For "talk to an expert" usage, Claude is hard to
| beat though.
| wayeq wrote:
| Claude still can't make real time web searches though for
| RAG, unless I'm missing something.
| resource0x wrote:
| Pursuing AGI? What method do they use to pursue something that
| no one knows what it is? They will keep saying they are
| pursuing AGI as long as there's a buyer for their BS.
| anti-soyboy wrote:
| Open AI is pursuing bullshit as they realised they can not
| compete anymore as they fired most of their talent year ago
| lblume wrote:
| Am I missing something, or do the results not even look that much
| better? Referring to the output quality, this just seems like a
| different prompting style and RLHF, not really an improved model
| at all.
| infinet wrote:
| Can it be self-hosted? Many institutions and organizations are
| hesitant to use AI because concerns of data leaking over chatbot.
| Open models, on the other hand, can be self-hosted. There is a
| deepseek arm race in other part of the world. Universities are
| racing to host their own deepseek. Hospitals, large businesses,
| local governments, even courts are deploying or showing interest
| in self-hosting deepseek.
| YetAnotherNick wrote:
| Do you know of any university that host Deepseek?
| infinet wrote:
| > _Do you know of any university that host Deepseek?_
|
| https://chat.zju.edu.cn
|
| https://chat.sjtu.edu.cn
|
| https://chat.ecnu.edu.cn/html/
|
| To list a few. There are of course many more in China. I
| won't be surprised if universities in other countries also
| self-hosting.
| YetAnotherNick wrote:
| Where does it says that it is self hosted? And why it is
| exposed to public.
| moralestapia wrote:
| OpenAI has never released a single model that could be self-
| hosted.
|
| GPT-2? Maybe not even that one.
| taytus wrote:
| Who wants a model that is not reasoning? The older models are
| just fine.
| ksynwa wrote:
| They said this is their last non reasoning model so I'm
| assuming there is a sunk cost aspect to it.
| eightysixfour wrote:
| Seeing OpenAI and Anthropic go different routes here is
| interesting. It is worth moving past the initial knee jerk
| reaction of this model being unimpressive and some of the
| comments about "they spent a massive amount of money and had to
| ship something for it..."
|
| * Anthropic appears to be making a bet that a single paradigm
| (reasoning) can create a model which is excellent for all use
| cases.
|
| * OpenAI seems to be betting that you'll need an ensemble of
| models with different capabilities, working as a single system,
| to jump beyond what the reasoning models today can do.
|
| Based on all of the comments from OpenAI, GPT 4.5 is absolutely
| massive, and with that size comes the ability to store far more
| factual data. The scores in ability oriented things - like coding
| - don't show the kind of gains you get from reasoning models but
| the fact based test, SimpleQA, shows a pretty large jump and a
| dramatic reduction in hallucinations. You can imagine a scenario
| where GPT4.5 is coordinating multiple, smaller, reasoning agents
| and using its factual accuracy to enhance their reasoning, kind
| of like ruminating on an idea "feels" like a different process
| than having a chat with someone.
|
| I'm really curious if they're actually combining two things right
| now that could be split as well, EQ/communications, and factual
| knowledge storage. This could all be a bust, but it is an
| interesting difference in approaches none-the-less, and worth
| considering that OpenAI _could_ be right.
| nomel wrote:
| > OpenAI seems to be betting that you'll need an ensemble of
| models with different capabilities, working as a single system,
| to jump beyond what the reasoning models today can do.
|
| The high level block diagrams for tech always end up converging
| to those found in biological systems.
| eightysixfour wrote:
| Yeah, I don't know enough real neuroscience to argue either
| side. What I can say is I feel like this path is more like
| the way that I observe that I think, it _feels_ like there
| are different modes of thinking and processes in the brain,
| and it _seems_ like transformers are able to emulate at least
| two different versions of that.
|
| Once we figure out the frontal cortex & corpus callosum part
| of this, where we aren't calling other models over APIs
| instead of them all working in the same shared space, I have
| a feeling we'll be on to something pretty exciting.
| sebastiennight wrote:
| > * OpenAI seems to be betting that you'll need an ensemble of
| models with different capabilities, working as a single system,
| to jump beyond what the reasoning models today can do.
|
| Seems inaccurate as their most recent claim I've seen is that
| they expect this to be their last non-reasoning model, and are
| aiming to provide all capacities together in the future model
| releases (unifying the GPT-x and o-x lines)
|
| See this claim on TFA:
|
| > We believe reasoning will be a core capability of future
| models, and that the two approaches to scaling--pre-training
| and reasoning--will complement each other.
| eightysixfour wrote:
| From Sam's twitter:
|
| > After that, a top goal for us is to unify o-series models
| and GPT-series models by creating systems that can use all
| our tools, know when to think for a long time or not, and
| generally be useful for a very wide range of tasks.
|
| > In both ChatGPT and our API, we will release GPT-5 as a
| system that integrates a lot of our technology, including o3.
| We will no longer ship o3 as a standalone model.
|
| You could read this as unifying the models _or_ building a
| unified systems which coordinate multiple models. The second
| sentence, to me, implies that o3 will still exist, it just
| won 't be standalone, which matches the idea I shared above.
| sebastiennight wrote:
| Ah, great point. Yes, the wording here would imply that
| they're basically planning on building scaffolding around
| multiple models instead of having one more capable Swiss
| Army Knife model.
|
| I would feel a bit bummed if GPT-5 turned out not to be a
| model, but rather a "product".
| eightysixfour wrote:
| For me it depends on how the models are glued together.
| Connected by function calling and APIs? Probably meh...
|
| Somehow working together in the same latent space? That
| could be neat.
| chippiewill wrote:
| Which is intriguing in a way, because the momentum I've
| seen across AI over the past decade has been increasing
| amounts of "end-to-end"
| tmpz22 wrote:
| I worry eliminating consumer choice will drive up prices
| for only a nominal gain in utility for most users.
| eightysixfour wrote:
| I'm more worries they'll push down their costs by making
| it harder to get the reasoning models to run, but either
| would suck.
| billywhizz wrote:
| or you could read it as a way to create a moat where none
| currently exists...
| ryukoposting wrote:
| > know when to think for a long time or not, and generally
| be useful for a very wide range of tasks.
|
| I'm going to call it now - no customer is actually going to
| use this. It'll be a cute little bonus for their chatbot
| god-oracle, but virtually all of their b2b clients are
| going to demand "minimum latency at all times" or "maximum
| accuracy at all times."
| wongarsu wrote:
| Or the other way around: smaller reasoning models that can call
| out to GPT-4.5 to get their facts right.
| eightysixfour wrote:
| Maybe, I'm inclined to think OpenAI believes the way I laid
| it out though, specifically because of their focus on
| communication and EQ in 4.5. It seems like they believe the
| large, non-reasoning model, will be "front of house."
|
| Or they'll use some kind of trained router which sends the
| request to the one it thinks it should go to first.
| jstummbillig wrote:
| It can never be just reasoning, right? Reasoning is the
| multiplier on some base model, and surely no amount of
| reasoning on top of something like gpt-2 will get you o1.
|
| This model is too expensive right now, but as compute gets
| cheaper -- and we have to keep in mind, that it will -- having
| a better base to multiply with will enable things that just
| more thinking won't.
| eightysixfour wrote:
| You can try for yourself with the distilled R1's that
| Deepseek released. The qwen-7b based model is quite
| impressive for its size and it can do a lot with additional
| context provided. I imagine for some domains you can provide
| enough context and let the inference time eventually solve
| it, for others you can't.
| throw234234234 wrote:
| > Anthropic appears to be making a bet that a single paradigm
| (reasoning) can create a model which is excellent for all use
| cases.
|
| I don't think that is their primary motivation. The
| announcement post for Claude 3.7 was all about code which
| doesn't seem to imply "all use cases". Code this, new code tool
| that, telling customers that they look forward to what they
| build, etc. Very little mention of other use cases on the new
| model announcement at all. Their usage stats they published are
| telling - 80%+ or more of queries to Claude are all about code.
| i.e. I actually think while they are thinking of other use
| cases; they see the use case of code specifically as the major
| thing to optimize for.
|
| OpenAI, given its different customer base and reach, is
| probably aiming for something more general.
|
| IMO they all think that you need an "ensemble" of models with
| different capabilities to optimise for different use cases. Its
| more about how much compute resources each company has and what
| they target with those resources. Anthrophic I'm assuming has
| less compute resources and a narrower customer base so it
| economically may make sense to optimise just for that.
| eightysixfour wrote:
| That's possible, my counter point would be that if that was
| the case Anthropic would have built a smaller reasoning model
| instead of doing a "full" Claude. Instead, they built
| something which seems to be flexible across different types
| of responses.
|
| Only time will tell.
| protocolture wrote:
| Ever since those kids demo'd their fact checking engine here,
| which was just Input -> LLM -> Fact Database -> LLM -> LLM ->
| Output I have been betting that it will be advantageous to move
| in this general direction.
| ilrwbwrkhv wrote:
| Question: explain thread first macros in clojure in a short
| paragraph
|
| GPT 4.5:
|
| In Clojure, the *thread-first macro* (`->`) is used to improve
| code readability by threading a value through a series of
| function calls. It takes an initial value and passes it as the
| first argument to the next function, then continues threading the
| result through subsequent functions. This eliminates deeply
| nested expressions and makes the code more readable.
|
| For example,
|
| ```clojure (-> 5 inc (* 2) str) ;; Equivalent to (str (* 2 (inc
| 5))) - "12" ```
|
| Here, `5` is incremented to `6`, then multiplied by `2` to get
| `12`, and finally converted to a string.
|
| ---
|
| Deepseek v3:
|
| In Clojure, the `->>` (thread-last) and `->` (thread-first)
| macros are used to improve code readability by allowing a more
| linear, pipeline-like flow of operations. The `->` (thread-first)
| macro takes an initial value and threads it as the first argument
| through a series of forms. For example, `(-> x (f) (g))`
| translates to `(g (f x))`. This is particularly useful when
| working with nested function calls, as it flattens the structure
| and makes the sequence of transformations clearer. It's commonly
| used with functions that take the threaded value as their first
| argument, such as collection operations or Java interop methods.
| datadrivenangel wrote:
| They tried, but doesn't look better?
| DaveMcMartin wrote:
| This feels more like a release they pushed out to keep the "hype"
| alive rather than something they were eager to share. Honestly,
| the results don't seem all that impressive, and considering the
| price, it just doesn't feel worth it.
| bla3 wrote:
| I wonder if we're starting to see the effects of the mass exodus
| a while ago.
| 42lux wrote:
| The announcements early on were relatively sincere and technical
| with papers and nice pages explaining the new models in easy
| language and now we get this marketing garbage. Probably the
| fastest enshitification I've seen.
| sky2224 wrote:
| Honestly, the most astounding part of this announcement is their
| comparison to o3-mini with QA prompts.
|
| EIGHTY PERCENT hallucination rate? Are you kidding me?
|
| I get that the model is meant to be used for logic and reasoning,
| but nowhere does OpenAI make this explicitly clear. A majority of
| users are going to be thinking, "oh newer is better," and pick
| that.
| jug wrote:
| Yeah it was an abysmal result (any 50%+ hallucination result in
| that bench is pretty bad) and worse than o1-mini in the
| SimpleQA paper. On that topic, Sonnet 3.5 "Old" hallucinates
| less than GPT-4.5, just for a bit of added perspective here.
| krackers wrote:
| Very nice catch, I was under the impression that o3-mini was
| "as good" as o1 on all dimensions. Seems the takeaway is that
| any form of quantization/distillation ends up hurting factual
| accuracy (but not reasoning performance), and there are
| diminishing returns to reducing hallucinations by model-scaling
| or RLHF'ing. I guess then that other approaches are needed to
| achieve single-digit "hallucination" rates. All of wikipedia
| compresses down to < 50GB though, so it's not immediately clear
| that you can't have good factual accuracy with a small sparse
| model
| freediver wrote:
| The results for GPT - 4.5 are in for Kagi LLM benchmark too.
|
| It does crush our benchmark - time to make new? ;) - with
| performance similar of that of reasoning models. It does come at
| a great price both in cost and speed.
|
| A monster is what they created. But looking at the tasks it
| fails, some of them my 9 year old would solve. Still in this
| weird limbo space of super knowledge and low intelligence.
|
| May be remembered as the last the last of the 'big ones', can't
| imagine this will be a path for the future.
|
| https://help.kagi.com/kagi/ai/llm-benchmark.html
| theodorthe5 wrote:
| If Gemini 2 is the top in your benchmark, make sure to re-check
| your benchmark.
| shawabawa3 wrote:
| Gemini 2 pro is actually very impressive (maybe not for
| coding, haven't used it for that)
|
| Flash is pretty garbage but cheap
| istjohn wrote:
| Gemini 2.0 Pro is quite good.
| aoeusnth1 wrote:
| Gemini 2 pro is pretty strong actually.
| wendyshu wrote:
| Why don't you have Grok?
| mhh__ wrote:
| No api for grok 3 might be why
| mjirv wrote:
| Do you have results for gpt-4? I'd be very interested in seeing
| the lift here from their last "big one".
| simonw wrote:
| If you want to try it out via their API you can run it through my
| LLM tool using uvx like this: uvx --with 'https:/
| /github.com/simonw/llm/archive/801b08bf40788c09aed617525287631031
| 2fe667.zip' \ llm -m gpt-4.5-preview 'impress me'
|
| You may need to set an API key first, either with `export
| OPENAI_API_KEY='xxx'` or using this command to save it to a file:
| uvx llm keys set openai # paste key here
|
| Or this to get a chat session going: uvx --with '
| https://github.com/simonw/llm/archive/801b08bf40788c09aed61752528
| 76310312fe667.zip' \ llm chat -m gpt-4.5-preview
|
| I'll probably have a proper release out later today. Details
| here: https://github.com/simonw/llm/issues/795
| ashu1461 wrote:
| Just curious, does this stream the output or renders all at
| once ?
| simonw wrote:
| It streams the output. See animated demo here (bottom image
| on the page)
| https://simonwillison.net/2025/Feb/27/introducing-gpt-45/
| synapsomorphy wrote:
| Claude 3.6 (new 3.5) and 3.7 non-reasoning are much better at
| pretty much everything, and much cheaper. What's Anthropic's
| secret sauce?
| taytus wrote:
| They ship more focused on their mission than OpenAI.
| moralestapia wrote:
| Huh?
|
| Post benchmark links.
| film42 wrote:
| I think it's a classic expectations problem. OpenAI is neither
| _open_ nor is it releasing an _AGI_ model in the near future.
| But when you see a new major model drop, you can't help but
| ask, "how close is this to the promise of AGI they say is just
| around the corner?" Not even close. Meanwhile Anthropic is
| keeping their heads down, not playing the hype game, and
| letting the model speak for itself.
| anothermathbozo wrote:
| Anthropic's CEO said their technology would end all disease
| and expand our lifespans to 200 years. What on earth do you
| mean they're not playing the hype game?
| jampa wrote:
| First impression of GPT-4.5:
|
| 1. It is very very slow, for some applications where you want
| real time interactions is just not viable, the text attached
| below took 7s to generate with 4o, but 46s with GPT4.5
|
| 2. The style it writes is way better: it keeps the tone you ask
| and makes better improvements on the flow. One of my biggest
| complaints with 4o is that you want for your content to be more
| casual and accessible but GPT / DeepSeek wants to write like
| Shakespeare did.
|
| Some comparisons on a book draft: GPT4o (left) and GPT4.5
| (green). I also adjusted the spacing around the paragraphs, to
| better diff match. I still am wary of using ChatGPT to help me
| write, even with GPT 4.5, but the improvement is very noticeable.
|
| https://i.imgur.com/ogalyE0.png
| remus wrote:
| > It is very very slow
|
| Could that be partially due to a big spike in demand at launch?
| jampa wrote:
| Possibly, repeating the prompt I got a much higher speed,
| taking 20s on average now, which is much more viable. But
| that remains to be seen when more people start using this
| version in production.
| jedberg wrote:
| Oh yeah, that right side version is WAY better, and sounds much
| more like a human.
| MichaelZuo wrote:
| How does it compare with o1 and o3 preview?
| jampa wrote:
| o3 is okay for text checking but has issues following the
| prompt correctly, same as o1 and DeepSeek R1, I feel that I
| need to prompt smaller snippets with them.
|
| Here is the o3 vs a new run of the same text in GPT 4.5
|
| https://www.diffchecker.com/ZEUQ92u7/
| MichaelZuo wrote:
| Thanks, though it says o1 on the page, is that a typo?
| FergusArgyll wrote:
| I opened your link in a new tab and looked at it a couple
| minutes later. By then I forgot which was o and which was .5
|
| I honestly couldn't decide which I prefer
| niek_pas wrote:
| I definitely prefer the 4.5, but that might just be because
| it sounds 'less like ChatGPT', ironically.
| sdesol wrote:
| It just feels natural to me. The person knows the language
| but they are not trying to sound smart by using words that
| might have more impact "based on the words dictionary
| definition"
|
| GPT 4.5 does feel like it is a step forward in producing
| natural language, and if they use it to provide
| reinforcement learning, this might have significant impact
| in the future smaller models.
| rl3 wrote:
| > _1. It is very very slow, ... below took 7s to generate with
| 4o, but 46s with GPT4.5_
|
| This is positively luxurious by o1-pro standards which I'd say
| _average_ 5 minutes. That said I totally agree even ~45s isn 't
| viable for real-time interactions. I'm sure it'll be optimized.
|
| Of course, my comparing it to the highest-end CoT model in
| [publicly-known] existence isn't entirely fair since they're
| sort of apples and oranges.
| philomath_mn wrote:
| I paid for pro to try `o1-pro` and I can't seem to find any
| use case to justify the insane inference time. `o3-mini-high`
| seems to do just as well in seconds vs. minutes.
| azinman2 wrote:
| What are you doing with it? For me deep research tasks are
| where 5 minutes is fine, or something really hard that
| would take me way more time myself.
| philomath_mn wrote:
| I usually throw a lot of context at it and have it write
| unit tests in a certain style or implement something
| (with tests) according to a spec.
|
| But the o3-mini-high results have been just as good.
|
| I am fine with Deep Research taking 5-8 minutes, those
| are usually "reports" I can read whenever.
| dingnuts wrote:
| I bet I can generate unit tests just as fast and for a
| fraction of the cost, and probably less typing, with a
| couple vim macros
| philomath_mn wrote:
| Idk, it is pretty good a generating synthetic data and
| recognizing the different logic branches to exercise. Not
| perfect, but very helpful.
| thfuran wrote:
| >One of my biggest complaints with 4o is that you want for your
| content to be more casual and accessible but GPT / DeepSeek
| wants to write like Shakespeare did.
|
| Well, maybe like a Sophomore's bumbling attempt to write like
| Shakespeare.
| ChiefNotAClue wrote:
| Right side, by a large margin. Better word choice and more
| natural flow. It feels a lot more human.
| rossant wrote:
| Is there really no way to prompt GPT4o to use a more natural
| and informal tone matching GPT4.5's?
| kristianp wrote:
| How do the two versions match so closely? They have the same
| content in each paragraph, just worded slightly differently. I
| wouldn't expect them to write paragraphs that match in size and
| position like that.
| princealiiiii wrote:
| Honestly, feels like a second LLM just reworded the response
| on the left-side to generate the right-side response.
| throwaway314155 wrote:
| If you use the "retry" functionality in ChatGPT enough, you
| will notice this happens basically all the time.
| kumarm wrote:
| Thank you. This is the best example of comparison I have seen
| so far.
| reassess_blind wrote:
| What's the deal with Imgur taking ages to load? Anyone else
| have this issue in Australia? I just get the grey background
| with no content loaded for 10+ seconds every time I visit that
| bloated website.
| stevage wrote:
| Ok for me here in aus
| elliotto wrote:
| This website sucks but successfully loaded from Aus rn on my
| phone. It's full of ads - possibly your ad blocker is killing
| it?
| osigurdson wrote:
| I'm wondering if generative AI will ultimately result in a very
| dense / bullet form style of writing. What we are doing now is
| effectively this:
|
| bullet_points' = compress(expand(bullet_points))
|
| We are impressed by lots of text so must expand via LLM in
| order to impress the reader. Since the reader doesn't have time
| or interest to read the content they must compress it back into
| bullet points / quick summary. Really, the original bullet
| points plus a bit more thinking would likely be a better form
| of communication.
| anon373839 wrote:
| That's what Axios does. For ordinary events coverage, it's a
| great style.
| not_a_bot_4sho wrote:
| I'm reminded of this great comic
|
| https://marketoonist.com/2023/03/ai-written-ai-read.html
| dyauspitr wrote:
| Imgur might be the worst image hosting site I've ever
| experienced. Any interaction with that page results in
| switching images and big ads and they hijack the back button.
| Absolutely terrible. How far they've fallen from when it first
| began.
| bradley13 wrote:
| I use 4o mostly in German, so YMMV. However, I find a simple
| prompt controls the tone very well. "This should be informal
| and friendly", or "this should be formal and business-like".
| muzani wrote:
| In my experience, Gemini Flash has been the best at writing,
| and GPT 3.5 onwards has been terrible.
|
| GPT-3 and GPT-2 were actually remarkably good at it, arguably
| better than a skilled human. I had a bit of fun ghostwriting
| with these and got a little fan base for a while.
|
| It seems that GPT-4.5 is better than 4 but it's nowhere near
| the quality of GPT-3 davinci. Davinci-002 has been nerfed quite
| a bit, but in the end it's $2/MTok for higher quality output.
|
| It's clear this is something users want, but OpenAI and
| Anthropic seem to be going in the opposite direction.
| vessenes wrote:
| Similar reaction here. I will also note that it seems to know a
| lot more about me than previous models. I'm not sure if this is
| a broader web crawl, more space in the model, or more
| summarization of our chats or a combination, but I asked it to
| psychoanalyze a problem I'm having in the style of Jacques
| lacan and it was genuinely helpful and interesting, no
| interview required first; it just went right at me.
|
| To borrow an iain banks word, the "fragre" def feels improved
| to me. I think I will prefer it to o1 pro, although I haven't
| really hammered on it yet.
| mchusma wrote:
| wow, openai really missed here. Reading the blog I thought like a
| minor, incremental minor catch up release for 4o. I thought "wow
| maybe this is cheaper than 4o so it will offset the pricing
| difference between this and something like Claude Sonnet 3.7 or
| Gemini 2.0 Flash both of which performs better. But its like
| 20x-100x more expensive!
|
| In other words, these performance stats with Gemini 2.0 Flash
| pricing looks reasonable. At these prices, zero usecases for
| anyone I think. This is a dead on arrival model.
| jasonjmcghee wrote:
| This is such as confusing release / announcement.
|
| It seems clearly worse than Claude Sonnet 3.7, yet costs 30x as
| much?
|
| What's going on with their pricing? I misread it as $7.5/M input
| and that that was very overpriced... then realized it was 10x
| that much!
| ashu1461 wrote:
| Is it worse than clause sonnet with reasoning enabled or
| disabled ?
| moffkalast wrote:
| OpenAI will do literally anything but ship GPT-5.
| saghm wrote:
| This is probably a dumb question, but are we just gonna be stuck
| on always having X.5 versions of GPT forever? If there's never an
| X.0, it feels like it's basically meaningless.
| xd1936 wrote:
| The gas costs $3.599
| dragonwriter wrote:
| There was 3, 3.5, 4, 4o, and then 4.5.
|
| There are things to complain about in that, but "its always X.5
| and never X.0" doesn't really seem to be one of them.
| JohnMakin wrote:
| Their example for 4.5 on "having a tough time":
|
| > I'm going through a tough time after failing a test > > Aw, I'm
| really sorry to hear that. Failing a test can feel pretty tough
| and discouraging, but remember, it's just one moment--not a
| reflection of who you are or your capabilities. Want to talk
| about what happened, or do you just need a distraction? I'm here
| either way.
|
| vs 4o:
|
| >I'm going through a tough time after failing a test > >I'm
| really sorry to hear that you're going through a tough time. Here
| are a few >things you might consider: > >1. Reflect and
| Understand: Try to analyze what might have gone wrong. Was it a
| >lack of preparation, time management during the test, or perhaps
| >misunderstanding the material? Understanding the root cause can
| help you address >it in the future. > >2. Seek Support: Talk to
| someone you trust--friends, family, or a school >counselor. They
| can offer a different perspective and sometimes just talking
| >about your feelings can be alleviating. > >3. Speak with Your
| Teacher or Professor: They can provide insights into where >you
| might have gone wrong and offer advice on how to improve. They
| might also >offer opportunities for make-up work or extra credit.
| > >4. Develop a Study Plan: Based on your reflections, create a
| plan to improve >your study habits. This might include setting
| regular study times, breaking down >material into manageable
| portions, or finding new study methods that work better >for you.
| > >5. Stay Positive and Forgive Yourself: Mistakes and setbacks
| are a part of life. >It's important to be kind to yourself and
| recognize that failure is a stepping >stone to success. > >6.
| Focus on the Bigger Picture: Remember that one test is just one
| part of your >educational journey. There will be many more
| opportunities to do well. > >If you need further support or
| resources, consider reaching out to educational >support services
| at your institution, or mental health resources if you're
| >feeling particularly overwhelmed. You're not alone in this, and
| things can get >better with time and effort.
|
| Is it just me or is the 4o response insanely better? I'm not the
| type of person to reach for a LLM for help about this kind of
| thing, but if I were, the 4o respond seems _vastly_ better to the
| point I 'm surprised they used that as their main "EQ" example.
| IMTDb wrote:
| 4o has a very strong artificial vibe. It feels a bit "autistic"
| (probably a bad analogy but couldn't find a better word to
| describe what I mean): you feel bad ? must say sorry then give
| a TODO list on how to feel better.
|
| 4.5 still feels a bit artificial but somehow also more
| emotionally connected. It removed the weird "bullet point lists
| of things to do" and focused on the emotional part; which is
| also longer than 4o
|
| If I am talking to a human I would definitely expect him/her to
| react more like 4.5 than like 4o. If the first sentence that
| comes out of their mouth after I explain them that I feel bad
| is "here is a list of things you might consider", I will find
| it strange. We can reach that point but it's usually after a
| bit more talk; human kinda need that process, and it feels like
| 4.5 understands that better than 4o.
|
| Now of course which one is "better" really depends on the
| context; what you expect of the model and how you intend to use
| is. Until now every single OpenAI update on the main series has
| always been a strict improvement over the previous model. Cost
| aside, there wasn't really any reason to keep using 3.5 when 4
| got released. This is not the case here; even assuming
| unlimited money you still might wanna select 4o in the dropdown
| sometimes instead of 4.5.
| buu700 wrote:
| I had a similar gut reaction, but on reflection I think 4.5's
| is actually the better response.
|
| On one hand, the response from 4.5 seems pretty useless to me,
| and I can't imagine a situation in which I would personally
| find value in it. On the other hand, the prompt it's responding
| to is also so different from how I actually use the tool that
| my preferences aren't super relevant. I would never give it a
| prompt that didn't include a clear question or direction,
| either explicitly or implicitly from context, but I can imagine
| that someone who does use it that way would actually be looking
| for something more in line with the 4.5 response than the 4o
| one. Someone who wanted the 4o response would likely phrase the
| prompt in a way that explicitly seeks actionable advice, or if
| they didn't initially then they would in a follow-up.
|
| Where I really see value in the model being capable of that
| type of logic isn't in the ChatGPT use case (at least for me
| personally), but in API integrations. For example, customer
| service agents being able to handle interactions more
| delicately is obviously useful for a business.
|
| All that being said, hopefully the model doesn't have too many
| false positives on when it should provide an "EQ"-focused
| response. That would get annoying pretty quickly if it kept
| happening while I was just trying to get information or have it
| complete some task.
| fauigerzigerk wrote:
| I think both responses are bizarre and useless. Is there a
| single person on earth who wouldn't ask questions like "what
| kind of test?", "why do you think you failed?", "how did you
| prepare for the test?" before giving advice?
| torginus wrote:
| My 2 cents (disclaimer: I am talking out of my ass) here is why
| GPTs actually suck at fluid knowledge retrievel (which is kinda
| their main usecase, with them being used as knowledge engines) -
| they've mentioned that if you train it on 'Tom Cruise was born
| July 3, 1962', it won't be able to answer the question "Who was
| born on July 3, 1962", if you don't feed it this piece of
| information. It can't really internally corellate the information
| it has learned, unless you train it to, probably via synthethic
| data, which is what OpenAI has probably done, and that's the
| information score SimpleQA tries to measure.
|
| Probably what happened, is that in doing so, they had to scale
| either the model size or the training cost to untenable levels.
|
| In my experience, LLMs really suck at fluid knowledge retrieval
| tasks, like book recommendation - I asked GPT4 to recommend me
| some SF novels with certain characteristics, and what it spat out
| was a mix of stuff that didn't really match, and stuff that was
| really reaching - when I asked the same question on Reddit, all
| the answers were relevant and on point - so I guess there's still
| something humans are good for.
|
| Which is a shame, because I'm pretty sure relevant product
| recommendation is a many billion dollar business - after all
| that's what Google has built it's empire on.
| woah wrote:
| Perhaps you could use LLMs in a list ranking context to
| generate your scifi recommendations
| https://github.com/noperator/raink?tab=readme-ov-file
| staticman2 wrote:
| You make a good point: I think these LLM's have a strong bias
| towards recommending the most popular things in pop culture
| since they really only find the most likely tokens and report
| on that.
|
| So while they may have a chance of answering "What is this non
| mainstream novel about" they may be unable to recommend the
| novel since it's not a likely series of tokens in response to a
| request for a book recommendation.
| torginus wrote:
| That's really interesting - just made me think about some AI
| guy at Twitter (when it was called that) talking about how
| hard it is to create a recommender system that doesn't just
| flood everyone with what's popular righr now. Since LLMs are
| neural networks as well, maybe the recommendation algorithms
| they learn suffer from the same issues
| vel0city wrote:
| An LLM on its own isn't necessarily great for fluid knowledge
| retrieval, as in directly from its training data. But they're
| pretty good when you add RAG to it.
|
| For instance, asking Copilot "Who was born on July 3, 1962"
| gave the response:
|
| > One notable person born on July 3, 1962, is Tom Cruise, the
| famous American actor known for his roles in movies like Risky
| Business, Jerry Maguire, and Rain Man.
|
| > Are you a fan of his work?
|
| It cited this page:
|
| https://www.onthisday.com/date/1962/july/3
| suddenlybananas wrote:
| Wow it googled the date!
| youssefabdelm wrote:
| Yep. I've often said RLHF'd LLMs seem to be better at
| recognition memory than recall memory.
|
| GPT-4o will never offhand, unprompted and 'unprimed', suggest a
| rare but relevant book like Shinichi Nakazawa's "A Holistic
| Lemma of Science" but a base model Mixtral 8x22B or Llama 405B
| will. (That's how I found it).
|
| It seems most of the RLHF'd models seem biased towards
| popularity over relevance when it comes to recall. They know
| about rare people like Tyler Volk... but they will never
| suggest them unless you prime them really heavily for them.
|
| Your point on recommendations from humans I couldn't agree more
| with. Humans are the OG and undefeated recommendation system in
| my opinion.
| jefffoster wrote:
| Does anyone have any intuition about the how reasoning improves
| based on the strength of the underlying model?
|
| I'm wondering whether this seemingly underwhelming bump on 4o
| magnifies when/if reasoning is added.
| porridgeraisin wrote:
| It is possible to understand the mechanism once you drop the
| anthropomorphisms.
|
| Each token output by an LLM involves one pass through the next-
| word predictor neural network. Each pass is a fixed amount of
| computation. Complexity theory hints to us that the problems
| which are "hard" for an LLM will need more compute than the
| ones which are "easy". Thus, the only mechanism through which
| an LLM can compute more and solve its "hard" problems is by
| outputting more tokens.
|
| You incentivise it to this end by human-grading its outputs
| ("RLHF") to prefer those where it spends time calculating
| before "locking in" to the answer. For example, you would
| prefer the output Ok let's begin... statement1
| => statement2 ... Thus, the answer is 5
|
| over The answer is 5. This is because....
|
| since in the first one, it has spent more compute before giving
| the answer. You don't in any way attempt to steer the extra
| computation in any particular direction. Instead, you simply
| reinforce preferred answers and hope that somewhere in that
| extra computation lies some useful computation.
|
| It turned out that such hope was well-placed. The DeepSeek
| R1-Zero training experiment showed us that if you apply this
| really generic form of learning (reinforcement learning)
| without _any_ examples, the model automatically starts
| outputting more and more tokens i.e "computing more".
| DeepseekMath was also a model trained directly with RL.
| Notably, the only signal given was whether the answer was right
| or not. No attention was paid to anything else. We even ignore
| the position of the answer in the sequence that we cared about
| before. This meant that it was possible to automatically grade
| the LLM without a human in the loop (since you're just checking
| answer == expected_answer). This is also why math problems were
| used.
|
| All this is to say, we get the most insight on what benefit
| "reasoning" adds by examining what happened when we applied it
| without training the model on any examples. Deepseek R1
| actually uses a few examples and then does the RL process on
| top of that, so we won't look at that.
|
| Reading the DeepseekMath paper[1], we see that the authors
| posit the following: As shown in Figure 7, RL
| enhances Maj@K's performance but not Pass@K. These
| findings indicate that RL enhances the model's overall
| performance by rendering the output distribution more
| robust, in other words, it seems that the improvement is
| attributed to boosting the correct response from TopK rather
| than the enhancement of fundamental capabilities.
|
| For context, Maj@K means that you mark the output of the LLM as
| correct only if the majority of the many outputs you sample are
| correct. Pass@K means that you mark it as correct even if just
| one of them is correct.
|
| So to answer your question, if you add an RL-based reasoning
| process to the model, it will improve simply because it will do
| more computation, of which a so-far-only-empirically-measured
| portion helps get more accurate answers on math problems. But
| outside that, it's purely subjective. If you ask me, I prefer
| claude sonnet for all coding/swe tasks over any reasoning LLM.
|
| [1] https://arxiv.org/pdf/2402.03300
| npinsker wrote:
| Thanks for a well-written and clear explanation!
| zone411 wrote:
| It significantly improves upon GPT-4o on my Extended NYT
| Connections Benchmark. 22.4 -> 33.7
| (https://github.com/lechmazur/nyt-connections).
| zone411 wrote:
| I ran three more of my independent benchmarks:
|
| - Improves upon GPT-4o's score on the Short Story Creative
| Writing Benchmark, but Claude Sonnets and DeepSeek R1 score
| higher. (https://github.com/lechmazur/writing/)
|
| - Improves upon GPT-4o's score on the
| Confabulations/Hallucinations on Provided Documents Benchmark,
| nearly matching Gemini 1.5 Pro (Sept) as the best-performing
| non-reasoning model.
| (https://github.com/lechmazur/confabulations)
|
| - Improves upon GPT-4o's score on the Thematic Generalization
| Benchmark, however, it doesn't match the scores of Claude 3.7
| Sonnet or Gemini 2.0 Pro Exp.
| (https://github.com/lechmazur/generalization)
|
| I should have the results from the multi-agent collaboration,
| strategy, and deception benchmarks within a couple of days.
| (https://github.com/lechmazur/elimination_game/,
| https://github.com/lechmazur/step_game and
| https://github.com/lechmazur/goods).
| j_bum wrote:
| Honest question for you: are these puzzles actually a good way
| to test the models?
|
| The answers are certainly in the training set, likely many
| times over.
|
| I'd be curious to see performance on Bracket City, which was
| featured here on HN yesterday.
| anotherpaulg wrote:
| GPT-4.5 Preview scored 45% on aider's polyglot coding benchmark
| [0]. OpenAI describes it as "good at creative tasks" [1], so
| perhaps it is not primarily intended for coding.
| 65% Sonnet 3.7, 32k think tokens (SOTA) 60% Sonnet 3.7, no
| thinking 48% DeepSeek V3 45% GPT 4.5 Preview <===
| 27% ChatGPT-4o 23% GPT-4o
|
| [0] https://aider.chat/docs/leaderboards/
|
| [1] https://platform.openai.com/docs/models#gpt-4-5
| doctoboggan wrote:
| I was waiting for your comment and wow... that's bad.
|
| I guess they are ceding the LLMs for coding market to
| Anthropic? I remember seeing an industry report somewhere and
| it claimed software development is the largest user of LLMs, so
| it seems weird to give up in this area.
| I_am_tiberius wrote:
| I assume they go all in "the new google" direction. Embedded
| ads coming soon I guess in the free version (chat.com).
| Workaccount2 wrote:
| 4.5 lies on a different path than their STEM models.
|
| o3-mini is an extremely powerful coding model and
| unquestionably is in the same league as 3.7. o3 is still the
| top stem overall model.
| nwienert wrote:
| No way, I've found o3 mini to be terrible. It' not as good
| as R1/Sonnet 3.5.
| icemelt8 wrote:
| they are trying to copy Grok 3
| smcleod wrote:
| GPT 4.5 is insanely over price, it makes Anthropic look
| affordable!
| shshahshsusus wrote:
| brief and detailed summaries by chatgpt (4o):
|
| _Brief Summary (40-50 words)_
|
| OpenAI's GPT-4.5 is a research preview of their most advanced
| language model yet, emphasizing improved pattern recognition,
| creativity, and reduced hallucinations. It enhances unsupervised
| learning, has better emotional intelligence, and excels in
| writing, programming, and problem-solving. Available for ChatGPT
| Pro users, it also integrates into APIs for developers.
|
| _Detailed Summary (200 words)_
|
| OpenAI has introduced *GPT-4.5*, a research preview of its most
| advanced language model, focusing on *scaling unsupervised
| learning* to enhance pattern recognition, knowledge depth, and
| reliability. It surpasses previous models in *natural
| conversation, emotional intelligence (EQ), and nuanced
| understanding of user intent*, making it particularly useful for
| writing, programming, and creative tasks.
|
| GPT-4.5 benefits from *scalable training techniques* that improve
| its steerability and ability to comprehend complex prompts.
| Compared to GPT-4o, it has a *higher factual accuracy and lower
| hallucination rates*, making it more dependable across various
| domains. While it does not employ reasoning-based pre-processing
| like OpenAI o1, it complements such models by excelling in
| general intelligence.
|
| Safety improvements include *new supervision techniques*
| alongside traditional reinforcement learning from human feedback
| (RLHF). OpenAI has tested GPT-4.5 under its *Preparedness
| Framework* to ensure alignment and risk mitigation.
|
| *Availability*: GPT-4.5 is accessible to *ChatGPT Pro users*,
| rolling out to other tiers soon. Developers can also use it in
| *Chat Completions API, Assistants API, and Batch API*, with
| *function calling and vision capabilities*. However, it remains
| computationally expensive, and OpenAI is evaluating its long-term
| API availability.
|
| GPT-4.5 represents a *major step in AI model scaling*, offering
| *greater creativity, contextual awareness, and collaboration
| potential*.
| ripped_britches wrote:
| Obviously it's expensive and still I would prefer a reasoning
| model for coding.
|
| However for user facing applications like mine, this is an
| awesome step in the right direction for EQ / tone / voice.
| Obviously it will get distilled into cheaper open models very
| soon, so I'm not too worried about the price or even tokens per
| second.
| mkaic wrote:
| In a hilarious act of accidental satire, it seems that the AI-
| generated audio version of the post has a weird
| glitch/mispronunciation within the _first three words_ -- it
| struggles to say "GPT-4.5".
| tantalor wrote:
| My experience exactly:
|
| 1. Open the page
|
| 2. Click "Listen to article"
|
| 3. Check if I'm having a stroke
|
| 4. Close tab
|
| Dear openai: try hiring some humans
| a-arbabian wrote:
| Common issue with TTS models right now. I use ElevenLabs to
| dictate articles while I commute and it has a stroke on
| decimals and symbols.
| advael wrote:
| It's sad that all I can think about this is that it's just
| another creep forward of the surveillance oligarchy
|
| I really used to get excited about ML in the wild and while there
| are much bigger problems right now it still makes me sad to have
| become so jaded about it
| i_love_retros wrote:
| Anyone really finding ai useful for coding?
|
| I'm finding it to make things up, get things wrong, ignore things
| I ask.
|
| Def not worried about losing my job to it.
| i_love_retros wrote:
| It gets confused if I give it 3 files - how is it going to scan
| a whole codebase and disparate systems and make correct
| changes.
|
| Pah! Don't believe the hype.
| twistslider wrote:
| I played around with Claude Code today, first time I've ever
| really been impressed by AI for coding.
|
| Tasked it with two different things, refactoring a huge
| function of around ~400 lines and creating some unit tests
| split into different files. The refactor was done flawlessly.
| The unit tests almost, only missed some imports.
|
| All I did was open it in the root of my project and prompt it
| with the function names. It's a large monolithic solution with
| a lot of subprojects. It found the functions I was talking
| about without me having to clarify anything. Cost was about $2.
| SkyPuncher wrote:
| Yes, massively.
|
| There's a learning curve to it, but it's worth literally every
| penny I spend on API calls.
|
| At worst, I'm no faster. At best, it's easily a 10x
| improvement.
|
| For me, one of the biggest benefits is talking about coding in
| natural language. It lowers my mental low and keeps me in a
| mental space where I'm more easily able to communicate with
| stakeholders holders.
| feznyng wrote:
| Really great for quickly building features but you have to be
| careful about how much context you provide i.e. spoonfeed it
| exactly the methods, classes, files it needs to do whatever
| you're asking for (especially in a large codebase). And when it
| seems to get confused, reset history to free up the context
| window.
|
| That being said there are definite areas where it shines
| (cookie cutter UI) and places where it struggles. It's really
| good at one-shotting React components and Flutter widgets but
| it tends to struggle with complicated business logic like sync
| engines. More straightforward backend stuff like CRUD endpoints
| are definitely doable.
| aurareturn wrote:
| Yes, it helps me write SQL queries in seconds that I otherwise
| would spend days on or give up completely.
| antirez wrote:
| In many ways I'm not an OpenAI fan (but I need to recognize their
| many merits). At the same time, I believe people are missing what
| they tried to do with GPT 4.5: it was needed and important to
| explore the pre-training scaling law in that direction. A gift to
| science, however selfist it could be.
| throwaway314155 wrote:
| > A gift to science
|
| This is hardly recognizable as science.
|
| edit: Sorry, didn't feel this was a controversial opinion. What
| I meant to say was that for so-called science, this is not
| reproducible in any way whatsoever. Further, this page in
| particular has all the hallmarks of _marketing_ copy, not
| science.
|
| Sometimes a failure is just a failure, not necessarily a gift.
| People could tell scaling wasn't working well before the
| release of GPT 4.5. I really don't see how this provides as
| much insight as is suggested.
|
| Deepseek's models apparently still compare favorably with this
| one. What's more they did that work with the constraint of
| having _less_ money, not so much money they could run
| incredibly costly experiments that are likely to fail. We need
| more of the former, less of the latter.
| kadushka wrote:
| _People could tell scaling wasn 't working well before the
| release of GPT 4.5_
|
| Who could tell? Who has tried scaling up to this level?
| throwaway314155 wrote:
| https://www.reuters.com/technology/artificial-
| intelligence/o...
|
| > Ilya Sutskever, co-founder of AI labs Safe
| Superintelligence (SSI) and OpenAI, told Reuters recently
| that results from scaling up pre-training - the phase of
| training an AI model that use s a vast amount of unlabeled
| data to understand language patterns and structures - have
| plateaued.
| anshumankmr wrote:
| OpenAI took a bullet for the team, by perhaps scaling the
| model to something bigger than the 1.6T params GPT4
| possibly had and basically telling its competitors its not
| gonna be worth scaling much beyond those number of params
| in GPT4, without a change in the model architecture
| vbezhenar wrote:
| > People could tell scaling wasn't working well before the
| release of GPT 4.5.
|
| Different people tell different things all the time. That's
| not science. Experiment is science.
| tacet wrote:
| if i understand correctly your argument, then i would say
| that it is very recognizable as science
|
| >People could tell scaling wasn't working well before the
| release of GPT 4.5
|
| Yes, on quick glance it seems so from 2020 openai research
| into scaling laws.
|
| Scaling apparently didn't work well, so the theory about
| scaling not working well failed to be falsified. It's
| science.
| wewewedxfgdf wrote:
| GPT-2 was laugh out loud funny, rolling on the ground funny.
|
| I miss that - newer LLMs seem to have lost their sense of humor.
|
| On the other hand GPT-2's funny stories often veered into
| murdering everyone in the story and committing heinous crimes but
| that was part of the weird experience.
| kossTKR wrote:
| Totally agree, i think the gargantuan hidden pre prompts,
| censorship through reinforcement learning and whatever has
| killed most creativity.
|
| The newer models are incredible, but the tone is just soul
| sucking even when it tries to be "looser" in the later
| iterations.
| krackers wrote:
| Sydney is a glimpse at what an "unlobotomized" GPT-4 model
| would have been like.
| lostmsu wrote:
| https://hn-wrapped.kadoa.com/wewewedxfgdf
| I_am_tiberius wrote:
| Not available in my Pro plan.
| highfrequency wrote:
| Overall take seems to be negative in the comments. But I see
| potential for a non-reasoning model that makes enough subtle
| tweaks in its tone that it is enjoyable to talk to instead of
| feeling like a summary of Wikipedia.
| Chance-Device wrote:
| And the AI stocks fell today.
|
| I'm sure it's unrelated.
| boznz wrote:
| So better than 4o but not good enough for a 5.0
| simonw wrote:
| I got gpt-4.5-preview to summarize this discussion thread so far
| (at 324 comments): hn-summary.sh 43197872 -m
| gpt-4.5-preview
|
| Using this script: https://til.simonwillison.net/llms/claude-
| hacker-news-themes...
|
| Here's the result:
| https://gist.github.com/simonw/5e9f5e94ac8840f698c280293d399...
|
| It took 25797 input tokens and 1225 input tokens, for a total
| cost (calculated using https://tools.simonwillison.net/llm-prices
| ) of $2.11! It took 154 seconds to generate.
| djhworld wrote:
| interesting summary but it's hard to gauge whether this is
| better/worse than just piping the contents into a much cheaper
| model.
| azinman2 wrote:
| It'd be great if someone would do that with the same data and
| prompt to other models.
|
| I did like the formatting and attributions but didn't
| necessarily want attributions like that for every section.
| I'm also not sure if it's fully matching what I'm seeing in
| the thread but maybe the data I'm seeing is just newer.
| simonw wrote:
| Good call. Here's the same exact prompt run against:
|
| GPT-4o: https://gist.github.com/simonw/592d651ec61daec66435
| a6f718c06...
|
| GPT-4o Mini: https://gist.github.com/simonw/cc760217623769f
| 0d7e4687332bce...
|
| Claude 3.7 Sonnet: https://gist.github.com/simonw/6f11e1974
| e4d613258b3237380e0e...
|
| Claude 3.5 Haiku: https://gist.github.com/simonw/c178f02c97
| 961e225eb615d4b9a1d...
|
| Gemini 2.0 Flash: https://gist.github.com/simonw/0c6f071d9a
| d1cea493de4e5e7a098...
|
| Gemini 2.0 Flash Lite: https://gist.github.com/simonw/8a713
| 96a4a219d8281e294b61a9d6...
|
| Gemini 2.0 Pro (gemini-2.0-pro-exp-02-05): https://gist.git
| hub.com/simonw/112e3f4660a1a410151e86ec677e3...
| iamjs wrote:
| At a glance, none of these appear to be meaningfully
| worse than GPT-4.5
| mastercheif wrote:
| Seeing the other models, I actually come away impressed
| with how well GPT-4.5 is organizing the information and
| how well it reads. I find it a lot easier to quickly
| parse. It's more human-like.
| jwr wrote:
| I actually think the Claude 3.7 Sonnet summary is better.
| hexa00 wrote:
| yeah I liked it too, especially for 10x less the price
| lol
| unoti wrote:
| I noticed 4o mini didn't follow the directions to quote
| users. My favourite part of the 4.5 summary was how it
| quoted Antirez. 4o mini brought out the same quote, but
| failed to attribute it as instructed.
| Agentlien wrote:
| It's fascinating, but while this does mean it strays from
| the given example, I actually feel the result is a better
| summary. The 4.5 version is so long you might just read
| the whole thread yourself.
| NitpickLawyer wrote:
| Interesting, thanks for doing this. I'd say that (at a
| glance) for now it's still worth to use more passes with
| smaller models than one pass with 4.5
|
| Now, if you'd want to generate training data, I could see
| wanting to have the best answers possible, where even
| slight nuances would matter. 4.5 seems to adhere to
| instructions much better than the others. You _might_ get
| the same result w / generating n samples and "reflect" on
| them with a mixture of models, but then again you might
| not. Going through thousands of generations manually is
| also costly.
| Topfi wrote:
| Thanks for sharing. To me, purely on personal preference,
| the Gemini models did best on this task, which also fits
| with my personal experience using Googles models to
| summarize extensive, highly specialized text. Geminis 2.0
| models do especially well on Needle in Haystack type
| tests in my experience.
| Agentlien wrote:
| Compared to GPT-4.5 I prefer the GPT-4o version because
| it is less wordy. It summarizes and gives the gist of the
| conversation rather than reproducing it along with
| commentary.
| joe_the_user wrote:
| The headline and section: "Dystopian and Social Concerns about
| AI Features" are interesting. It's roughly true... but somehow
| that broad statement seems minimize the point discussed.
|
| I'd headline that thread as "Concerns about output tone". There
| were comments about dystopian implications of tone, marketing
| implications of tone and implementation issues of tone.
|
| Of course, that I can comment about the fine-points of an AI
| summary shows it's made progress. But there's a lot riding on
| how much progress these things can make and what sort. So it's
| still worth looking at.
| alew1 wrote:
| Didn't seem to realize that "Still more coherent than the
| OpenAI lineup" wouldn't make sense out of context. (The actual
| comment quoted there is responding to someone who says they'd
| name their models Foo, Bar, Baz.)
| willy_k wrote:
| Wonder if there's some pro-OpenAI system prompt getting in
| the way of that.
| jdiff wrote:
| It'd be a silly move considering how fast system prompts
| leak.
| colordrops wrote:
| Maybe it's just confirmation bias but the language in your
| result output seems higher quality that previous models. Seems
| more natural and eloquent.
| 3vidence wrote:
| I don't know why but something about this section made me
| chuckle
|
| """ These perspectives highlight that there remains nuance--
| even appreciation--of explorative model advancement not solely
| focused on immediate commercial viability """
|
| Feels like the model is seeking validation
| stevage wrote:
| Huh. Disregarding the 4.5-specific bit here, a browser
| extension or possibly website that did this in general could be
| really useful.
|
| Maybe even something that just noticed whenever you visited a
| site that had had significant HN discussion in the past, then
| let you trigger a summary.
| wordpad25 wrote:
| there are literally hundreds of extensions and sites that do
| this
|
| the problem is that they are competing each other into the
| ground hence they go unmaintained very quickly
|
| getrecall.ai has been the most mature so far
| stevage wrote:
| Thanks, it's amazing how much stuff is out there I don't
| know about.
| bredren wrote:
| Hundreds that specifically focus on noticing a page you're
| currently viewing has been not only posted to but undergone
| significant discussion on HN, and then providing a summary
| of those conversations?
|
| Or that just provide summaries in general?
| ruibiks wrote:
| Hey, check this one out with all the different flavors that
| existed out there. I think I made something better.
| https://cofyt.app
|
| As far as I am aware, feel free to test it head-to-head.
| This is better than gecall, and you can chat with a
| transcript for detailed answers to your prompts
| wordpad25 wrote:
| I tried it out, looks nice and clean.
|
| But as I mentioned, my main concern is what will happen
| in 6 months when you fail to get traction and abandon it.
| Because that's what happened to previous 5 products I
| tried which were all "good enough" .
|
| Getrecall seems to have a big enough user base that will
| actually stick around.
| ruibiks wrote:
| Thank you for getting back to me.
|
| I understand your perfectly reasonable argument to make
| from your position (user).
|
| First let me tell you that I saw a lot of things out
| there including getrecall before starting building this
| and felt there was nothing out there that had a good
| UX/UI that actually makes it a enjoyable product (nice
| and clean).
|
| I'm confident in the direction and committed to seeing it
| through by building something better for me and maybe for
| you to by doing it with more care.
|
| Appreciate your feedback and while no one can control the
| future I've added this thread to my calendar do come back
| here in 6months.
| ukuina wrote:
| My site https://hackyournews.com does this!
|
| Been keeping it alive and free for 18 months.
| zigman1 wrote:
| Wow I find this very useful, thanks! Bookmarked.
| sebastiennight wrote:
| What I want is something that can read the thread out loud to
| me, using a different voice per user, so I can _listen_ to a
| busy discussion thread like I would listen to a podcast.
| dahsameer wrote:
| $2.11! At this point, I'm more concerned about AI price than
| egg price.
| vivzkestrel wrote:
| you know what? it would be damn nice to do this to literally
| every post in HN and give people a summary so that they dont
| have to read 500 comments
| munksbeer wrote:
| As expected, comments on LLM threads are overwhelmingly
| negative.
|
| Personally, I still feel excited to see boundaries being
| pushed, however incremental our anecdotal opinions make them
| seem.
| sundarurfriend wrote:
| I disagree with most of the knee-jerk negativity in LLM
| threads, but in this case it mostly seems warranted. There
| are no "boundaries being pushed" here, this is just a
| desperate release from a company that finds itself losing
| more and more mindshare to other models and companies.
| egobrain27 wrote:
| Seems to have trouble recognizing sarcasm:
|
| "For example, there are now a bunch of vendors that sell
| 'respond to RFP' AI products... paying 30x for marginally
| better performance makes perfect sense." -- hn_throwaway_99 (an
| uncommon opinion supporting possible niche high-cost uses).
| mlyle wrote:
| ? You think hn_throwaway_99's comment is sarcastic? It makes
| perfect sense to me read "straight."
|
| That is, sales orgs save a bunch of money using AI to respond
| to RFPs; they would still save a bunch of money using a more
| expensive AI, and any marginal improvement in sales closed
| would pay for it.
|
| It maybe excessively summarized his comment which confused
| you-- but this is the kind of mistake human curators of
| quotes make, too.
| sebzim4500 wrote:
| I don't think they are being sarcastic. Maybe you are the bot
| /s
| orbital-decay wrote:
| This looks like a first generation model to bootstrap future
| models from, not a competitive product at all. The knowledge
| cutoff is pretty old as well. (2023, seriously?)
|
| If they wanted to train it to have some character like Anthropic
| did with Claude 3... honestly I'm not seeing it, at least not in
| this iteration. Claude 3 was/is much much more engaging.
| GaggiX wrote:
| I imagine it will be used as a base for GPT-5 when it will be
| trained into a reasoning model, right now it probably doesn't
| make too much sense to use.
| dgfitz wrote:
| @sama, LLMs aren't going to create AGI. I realize you need to
| generate cash flow, this isn't the play.
|
| Sincerely, Me
| kneegerman wrote:
| you misunderstand, the business model is extracting cash from
| Qatar et al
| vhiremath4 wrote:
| I cancelled my ChatGPT subscription today in favor of using Grok.
| It's literally the difference between me never using ChatGPT to
| using Grok all the time, and the only way I can explain it is
| twofold:
|
| 1. The output from Grok doesn't feel constrained. I don't know
| how much of this is the marketing pitch of it "not being woke",
| but I feel it in its answers. It never tells me it's not going to
| return a result or sugarcoats some analysis it found from Reddit
| that's less than savory.
|
| 2. Speed. Jesus Christ ChatGPT has gotten so slow.
|
| Can't wait to pay for Grok. Can't believe I'm here. I'm usually a
| big proponent of just sticking with the thing that's the most
| popular when it comes to technology, but that's not panning out
| this time around.
| tiahura wrote:
| I found Grok's reasoning pretty wack.
|
| I asked it - "Draft a Minnesota Motion in Limine to exclude
| ..."
|
| It then starts thinking ... User wants a Missouri Motion in
| Limine ....
| blitq wrote:
| It's just nuts how pricy this model is when scoring worse than
| o3-mini
| torginus wrote:
| Call me a conspiracy theorist, but this, combined with the
| extremely embarassing way Claude is playing Pokemon, makes me
| feel this is an effort by AI companies to make LLMs look bad -
| setting up the hype cycle for the next thing they have in the
| pipeline.
| camdenreslink wrote:
| The next thing in the pipeline is definitely agents, and making
| the underlying tech look bad won't help sell that at all.
| torginus wrote:
| Agents as they are right now is literally just the LLM
| calling itself in a loop + having the ability to use
| tools/interact with their environment. I don't know if
| there's anything profoundly disruptive cooking in that space.
| afastow wrote:
| You're not a conspiracy theorist, you're just recognizing that
| the reality doesn't match the hype. It's boring and not fun but
| in this situation the answer is almost always that the hype is
| wrong, not the reality.
| sunami-ai wrote:
| still can't deal with sequences (or permutations)
|
| https://chatgpt.com/share/67c0f064-7fdc-8002-b12a-b62188f507...
|
| The Share doesn't say 4.5 but I assure you it is 4.5
| adamtaylor_13 wrote:
| It's crazy how quickly OpenAI releases went from, "Honey, check
| out the latest release!" to a total snooze fest.
|
| Coming in the heels of Sonnet 3.7 which is a marked improvement
| over 3.5 which is already the best in the industry for coding,
| this just feels like a sad whimper.
| 404mm wrote:
| I'm just disappointed that while everyone else (DS, Claude) had
| something to introduce for the "Plus" grade users, gpt 4.5 is
| so resource demanding that it's only available to quite
| expensive Pro sub. That just doesn't feel much like progress.
| gcanyon wrote:
| The prices they're charging are not _that_ far from where you
| could outsource to a human.
| resters wrote:
| based on a few initial tests GPT-4.5 is abysmal. I find the prose
| more sterile than previous models and far from having the spark
| of DeepSeek, and it utterly choked on / mangled some python code
| (~200 LoC and 120 LoC tests) that o3-mini-high and grok-3 do very
| well on.
| xena wrote:
| I'm really not sure who this model is for. Sure the vibes may be
| better, but are they 2.5x as much as o1 better? Kinda feels like
| they're brute forcing something in the backend with more hardware
| because they hit a scaling wall.
| yousif_123123 wrote:
| It's disappointing not to see comparisons to Sonnet 3.7. Also
| since o3-mini is ahead of o1, not sure why in the video they
| compared to o1.
|
| gpt4 was way ahead of 3.5 when it came out. It's unfortunate that
| the first major gpt release since that is so underwhelming..
| j_bum wrote:
| Agreed, but I suppose this is a tell. I think they're trying to
| place this into a separate class of models.
|
| I.e., we know it might not be as good as 3.7, but it is very
| friendly and maybe acts like it knows more things.
| IAmGraydon wrote:
| Between this and Claude 3.7, I'm really beginning to believe that
| LLM development has hit a wall, and it might actually be
| impossible to push much farther for reasonable amounts of money
| and resources. They're incredible tools indeed and I use them on
| a daily basis to multiply my productivity, but yeah - I think
| we've all overshot this in a big way.
| j_bum wrote:
| Agreed at every level.
|
| I absolutely love LLMs. I see them as insanely useful,
| interactive, quirky, yet lossy modern search engines. But
| they're fundamentally flawed, and I don't see how an "agent" in
| the traditional sense of the world can actually be produced
| from them.
|
| The wall seems to be close. And the bubble is starting to leak
| air.
| netdevphoenix wrote:
| > LLM development has hit a wall
|
| The writing has been on the wall since 2024. None of the LLM
| releases have been groundbreaking they have all been lateral
| improvements and I believe the trend will continue this year
| with make them more efficient (like DeepSeek), make them faster
| or make them hallucinate less
| passwordoops wrote:
| Still no 5, huh?
| andrewinardeer wrote:
| I just played with the preview through the API. I asked it to
| refactor a fairly simple dashboard made with HTML, css and
| JavaScript.
|
| First time it confused css and JavaScript, then spat out code
| which broke the dashboard entirely.
|
| Then it charged me $1.53 for the privilege.
| ipnon wrote:
| Finally a replacement for junior engineers!
| pronouncedjerry wrote:
| instead of these random IDs they should label them to make sense
| for the end user. i have no idea which one to select for what i
| need. and do they really differ that much by use case?
| sirolimus wrote:
| Slow, expensive and nothing special. Just stick to o1 or give us
| o3 (non-mini).
| msp26 wrote:
| >"GPT4.5, the most knowledgable model to date" >Knowledge cutoff:
| October 2023
| bag_boy wrote:
| I tried it and if it's more natural, I don't know what that means
| anymore, because I'm used to the last 6 month's models.
| mrcwinn wrote:
| Sam continues to be the least impressive person to ever lead such
| an amazing company.
|
| Odd communication from him recently too. We're sorry our UI has
| become so poor. We're sorry this model is so expensive and not a
| big leap.
|
| Being rich and at the right place at the right time doesn't
| itself qualify you to lead or make you a visionary. Very odd
| indeed.
| tuananh wrote:
| this seems to be a very weak response to sonnet 3.7
|
| - more expensive. alot more expensive
|
| - not a lot of increment improvement
| energy123 wrote:
| Lower hallucinations than o1. Impressive.
| hombre_fatal wrote:
| Cathartic moment over.
| apwell23 wrote:
| doesn't feel like to me. I try using copilot on my scala
| projects and it always comes up with something useless that
| doesn't even compile.
|
| I am currently just using it as easy google search.
| abrichr wrote:
| Have you tried copying the compilation errors back into the
| prompt? In my experience eventually the result is correct. If
| not then I shrink the surface area that the model is touching
| and try again.
| apwell23 wrote:
| yes ofcourse. it then proceeds to agree that what it told
| me was indeed stupid and proceeds to give me something even
| worse.
|
| I would love to see a video of ppl using this in real
| projects ( even if its open source) . I am tried of ppl
| claiming moon and stars after trying it on toy projects.
| gardenhedge wrote:
| Yeah that's what happens. It can recreate anything it's
| been trained on - which is a lot - but you'll definitely
| fall into these "Oh, I see the issue now" loops when
| doing anything not in the training set.
| NiloCK wrote:
| Breath it in, get a coffee, and sit down to solve some bigger
| problems.
|
| 3.7 really is astounding with the one-shots.
| tmpz22 wrote:
| I haven't had the same experience. Here are some of the
| significant issues when using o1 or claude 3.7 with vscode
| copilot:
|
| * Very wreckless in pulling in third party libraries - often
| pulling in older versions including packages that trigger
| vulnerability warnings in package managers like npm. Imagine a
| student or junior developer falling into this trap.
|
| * Very wreckless around data security. For example in an
| established project it re-configured sqlite3 (python lib) to
| disable checks for concurrent write liabilities in sqlite. This
| would corrupt data in a variety of scenarios.
|
| * It sometimes is very slow to apply minor edits, taking about
| 2 - 5 minutes to output its changes. I've noticed when it takes
| this long it also usually breaks the file in subtle ways,
| including attaching random characters to a string literal which
| I very much did not want to change.
|
| * Very bad when working with concurrency. While this is a hard
| thing in general, introducing subtle concurrency bugs into a
| codebase is not good.
|
| * By far is the false sense of security it gives you. Its close
| enough to being right that a constant incentive exists to just
| yeet the code completions without diligent review. This is
| really really concerning as many organizations will yeet this,
| as I imagine executives are currently the world over.
|
| Honestly I think a lot of people are captured by a small sample
| size of initial impressions, and while I believe you in that
| you've found value for use cases - in aggregate I think it is a
| honeymoon phase that wears off with every-day use.
| hombre_fatal wrote:
| I've been using it daily for years. Mostly asking questions
| in a separate chat window/app and then working its response
| into my code. And then I sped up the feedback loop when I
| migrated to Cursor where I began pushing the envelop and
| asking it to do more.
|
| I think what wears off is that we're less impressed and then
| we start demanding more and more from it and getting
| frustrating when it can't do it. But that's different than a
| honeymoon phase wearing off. It's like how we're not really
| impressed by image gen anymore, we expect it.
|
| But as an example of a selfish sense of loss I've
| experienced, I used to pride myself in being the only
| developer on any team who ever learned CSS. I could architect
| a good grid/flex layout with a lot of thought. I could do
| little things like make text in a small UI component truncate
| into {3 letters} + ellipses when its parent was too small.
| And most of all I could polish UIs to a point where I'd say
| they were perfect, even a form.
|
| Now, LLMs are really good at doing the mechanical parts of
| the things I spent so much time learning. Like I originally
| said, I'm not shedding tears over here saying it's so unfair.
| But there is a sense of loss. And when I figured most people
| reading my comment would misinterpret this, I removed my
| comment. Because you can't make descriptive claims about how
| you feel online, it can only be interpreted as a normative
| value judgement about the world. Because I guess that's what
| it is 99.9% of the time someone expresses a feeling they
| feel, but not in this case.
|
| Finally, the right way to see it is that now I can polish the
| UI to perfection, but I don't need to be a CSS expert
| anymore. Nobody needs to be. You can get an idea of how you
| want the UI to work and ask the LLM "make this one bit of
| text be the one that truncates if the window is too narrow"
| and it does it. And that's fkin magic.
| apwell23 wrote:
| why did you remove the comment . now who ppl responded to you
| look like dummies. do you do this sort of stuff in real life
| too?
| hombre_fatal wrote:
| What would it mean to do this in real life? :D
|
| I regularly make knee-jerk comments on HN that I delete a
| minute later. Something therapeutic about it.
|
| My comment isn't one I wanted on my "record". You responded
| to it and I saw your response before deleting my comment.
| What's the harm? It's obvious I removed my comment.
| Karrot_Kream wrote:
| > I regularly make knee-jerk comments on HN that I delete a
| minute later. Something therapeutic about it.
|
| I'm really curious about this. Doesn't it feel selfish to
| you to subject the public to your internal anxieties? It's
| the same reason I don't unload on everyone around me.
|
| EDIT: I'm not trying to dunk on you. You're being honest so
| thanks for that.
| fungiblecog wrote:
| Enjoy your expensive garbage
| nycdatasci wrote:
| When I asked "what version are you?" it insisted that it was
| ChatGPT 4.0 Turbo, one step behind GPT-4.5.
| https://chatgpt.com/share/67c0fda8-a940-800f-bbdc-6674a8375f...
| xeckr wrote:
| My initial impression is that I have gotten quite spoiled by the
| speed of GPT-4o...
| dangoodmanUT wrote:
| Finally, an LLM that doesn't YAP
| prompt_judy wrote:
| Its not even that great for real world business tasks. I have no
| idea what they are thinking https://youtu.be/puPybx8N82w
| ahmadtbk wrote:
| Will this pop the AI bubble?
| camdenreslink wrote:
| I think if GPT-5 is very underwhelming we could start to see
| some shifting of opinion on what kind of return on investment
| all of this will result in.
| afastow wrote:
| This _is_ GPT-5, or rather what they clearly intended to be
| GPT-5. The pricing makes it obvious that the model is
| massive, but what they ended up with wasn 't good enough to
| justify calling it more than 4.5.
| stan_kirdey wrote:
| For most tasks, GPT-4o/o3-mini are already great, and cheaper.
|
| What is the real-world use case where GPT-4.5? Anyone actually
| seeing a game-changing difference? Please share your prompts!
| gigagorilla wrote:
| Bring back GPT-1. It really knew how to have a conversation.
| gorgoiler wrote:
| I love the "listen to this article" widget doing embedded TTS for
| the article. Bugs / feedback:
|
| The first words I hear are "introducing gee pee four five". The
| TTS model starts cold? The next occurrence of the product name
| works properly as "gee pee tee four point five" but that first
| one in the title is mangled. Some kind of custom dictionary would
| help here too, for when your model needs to nail crucial phrases
| like your business name and your product.
|
| No way of seeking back and forth (Safari, iOS 17.6.1). I don't
| even need to seek, just replay the last 15s.
|
| Very much need to select different voice models. Chirpy "All new
| Modern Family coming up 8/9c!" voice just doesn't cut it for a
| science broadcast, and localizing models -- even if it's still
| English -- would be even better. I need to hear this announcement
| in Bret Taylor voice, not Groupon CMO voice. (Sorry if this is
| _your_ voice btw, and you work at OpenAI brandi. No offence
| intended.)
| Perenti wrote:
| What I find hilarious is that a 20-50% hallucination rate
| suggests this is still a program that tells lies and potentially
| causes people to die.
| iamronaldo wrote:
| https://livebench.ai/#/ Best non reasoning model on livebench
| (and ranks above gemeni thinking)
| anshumankmr wrote:
| If this cannot eliminate hallucinations or at least reduce them
| to be statistically unlikely to be happen, and I assume it has
| more params than GPT4's trillion parameters, that means the
| scaling law is dead isn't it?
| energy123 wrote:
| I interpret this to mean we're in the ugly part of the old
| scaling law, where `ln(x)` for `x > $BIGNUMBER` starts to
| becoming punishing, not that the scaling law is in any way
| empirically refuted. Maybe someone can crunch the numbers and
| figure out if the benchmarks empirically validate the scaling
| law or not, relative to GPT-4o (assuming e.g. 200 million
| params vs 5T params).
| sudosysgen wrote:
| I mean the scaling laws were always logarithms, and logarithms
| become arbitrarily close to flat if you can't drive them with
| exponential growth, and even if you do it's barely linear. The
| scaling laws always predicted that model scaling would
| stop/slow being practical at some point.
| anshumankmr wrote:
| Right but the quantum leap in capabilities that came from
| GPT2->GPT3->GPT3.5Turbo (which I personally felt didn't fare
| as well at coding as the former)->GPT4 won't be replicated
| anytime soon with the pure text/chat generation models.
| sudosysgen wrote:
| Sure, that's also predicted by a logarithmic scaling law,
| you have extremely rapid growth until the inflection point.
| ksynwa wrote:
| Why would they want to eliminate that? Altman said that
| hallucinations are how LLMs express creativity.
| mrcwinn wrote:
| First time I've had an LLM reply "Nope."
|
| https://chatgpt.com/share/67c154e7-5e28-800d-81d7-98b79c8a87...
| loa_observer wrote:
| You got 10x price but not 10x quality
| mrcwinn wrote:
| I'd prefer this model if it were faster, but not at this cost.
| And so it is an odd release.
|
| Still, with Deep Research and Web Search, ChatGpt seems far ahead
| of Claude. I like 3.7 a lot but I find OpenAI's features more
| useful, even if it has for now complicated the UI a bit.
| fsloth wrote:
| Agree on the Web App. Cursor with Claude 3.7 is a pretty good
| "CoPilot" experience though.
| cubefox wrote:
| Altman mentioned GPT-4.5 is the model code named "Orion". Which
| originally was supposed to be their next big model, presumably
| GPT-5, but showed disappointing improvements on benchmark
| performance. Apparently the AI companies are hitting diminishing
| returns with the paradigm of scaling foundation model
| pretraining. It was discussed a few months ago:
|
| https://news.ycombinator.com/item?id=42125888
| DennisL123 wrote:
| tl;dr: doesn't work as expected and we sank a ton of money on it
| too.
| yobid20 wrote:
| Cash grab because they see the writing on the wall. OpenAI is
| collapsing. Their models suck now.
| weinzierl wrote:
| _" Starting today, ChatGPT Pro users will be able to select
| GPT-4.5 in the model picker on web, mobile, and desktop. We will
| begin rolling out to Plus and Team users next week, then to
| Enterprise and Edu users the following week."_
|
| Thanks for being transparent about this. Nothing is more
| frustrating than being locked out for indeterminate time from the
| hot thing everyone talks about.
|
| I hope the announcement is true without further unsaid
| qualifications, like availability outside the US.
| vbezhenar wrote:
| I'm outside the US and I have access to ChatGPT 4.5 with
| ChatGPT Pro subscription. Didn't have that access yesterday at
| the time of announce, but probably they were staggering the
| release a bit to even the load over multiple hours.
| yzydserd wrote:
| Sadly, it has a small context window.
| nbzso wrote:
| Yesterday I tested Windsurf. Looked the docs and examples.
| Completed the demo "course" on deeplearning.ai. Gave it the task
| to build a simple Hugo blog website with a theme link and
| requirements, it failed consecutive times. With all the available
| models.
|
| AI art is an abomination. Half of the internet is already filled
| with AI written crap. Don't start with the video. Soon everyone
| will require validation to distinguish reality from hallucination
| (so World ID in place as problem-reaction-solution).
|
| For me, the best use cases are LLM assisted search with limited
| reasoning. Vision models for digitization and limited code
| assistance, codebase doc generation and documentation.
|
| Agents are just workflows with more privileges. So where is the
| revolution? I don't see it.
|
| Where is added value? Making Junior Engineers obsolete? Or make
| them dumb copy-pasting bio machines?
|
| Depressing a horde of intellectual workers and artists and giving
| a good excuse for layoffs.
|
| The real value is and always will be in a specialized ML
| applications.
|
| LLM hype is getting boring.
| ta3456345345 wrote:
| The AI hyperbole is so cringe right now (and for the last few
| years). I've yet to see anyone come up with something that'd wow
| me, and say, "OK, yep, that deserves those cycles".
|
| Writing terrible fanfic esque books, sometimes OK images, chatbot
| style talking. meh.
| PeterStuer wrote:
| Currently my daily API costs for 4o are low enough and
| performance/quality for my usecases good enough that switching
| models has not made to to the top of application improvements.
|
| My cases' costs are more heavily slanted towards input tokens, so
| trying 4.5 would raise my costs over 25x, which is a non-starter.
| jsemrau wrote:
| Interesting observation.It seems capability has reached a
| plateau. Like a local maximum.
| PeterStuer wrote:
| I'm not sure that is the right conclusion.
|
| It is more like the AI part of the system _for this specific
| use case_ has reached a position where focusing on that part
| of the complete application as opposed to other parts that
| need attention would not yield the highest return in terms of
| user satisfaction or revenue.
|
| Certainly there is _enormous_ potential for AI improvement,
| and I have other projects that do gain substantially from
| improvements in e.g. reasoning, but then GPT 4.5 will have to
| compete with Deepseek, Gemini, Grok and Claude on a price
| /performance level, but to be honest the current preview
| pricing would make it (in production, not for dev) a non
| starter for me.
| aubanel wrote:
| The price is not that insane when you remember that GPT-4 cost
| 36$/million input tokens at launch!
| Tepix wrote:
| Prices have come down
| aubanel wrote:
| yes, which suggests they'll keep going down! So I'd expect
| GPT-4.5 to be 90% cheaper in 1-2 years
| netdevphoenix wrote:
| Is it official then?
|
| Most of us have been waiting for this moment for a while. The
| transformer architecture as it is currently understood can't be
| milked any further. Many of us knew this since last year. GPT-5
| delays eventually led to non-tech voices to suggest likewise. But
| we all held our final decision until the next big release from
| OpenAI as Sam Altman has been making claims about AGI entering
| the workforce this year, OpenAI knowing how to build AGI and
| similar outlandish claims. We all knew that their next big
| release in 2025 would be the final deciding factor on whether
| they had some tech breakthrough that would upend the world
| (justifying their astronomical valuation) or if it would just be
| (slightly) more of the same (marking the beginning of their
| downfall).
|
| The GPT-4.5 release points towards the latter. Thus, we should
| not expect OpenAI to exist as it does now (AI industry leader) in
| 2030, assuming it does exist at all by then.
|
| However, just like the 19th century rail industry revolution, the
| fall of OpenAI will leave behind a very useful technology that
| while not catapulting humanity towards a singularity, will
| nonetheless make people's lives better. Not much consolation to
| the world's super rich who will lose tons of money once the LLM
| industry (let us remember that AI is not LLM) falls.
|
| EDIT: "will nonetheless make people's lives better" to "might
| nonetheless make some people's lives better"
| fergonco wrote:
| > will nonetheless make people's lives better
|
| Probably not the lives of translators or graphic designers or
| music compositors. They will have to find new jobs. As llm
| prompt engineers, I guess.
| yurishimo wrote:
| Graphic designers I think are safe, at least within
| organizations that require a cohesive brand strategy. Getting
| the AI to respect all of the previous art will be a challenge
| at a certain scale.
|
| Fiverr graphic designers on the other hand...
| whimsicalism wrote:
| absolutely a solvable problem even with no tech advances
| andy_ppp wrote:
| Getting graphic designers to use the design system that
| they invented is quite a challenge too if I'm honest...
| should we really expect AI to be better than people? Having
| said that AI is never going to be adept at knowing how and
| when to ignore the human in the loop and do the "right"
| thing.
| bearjaws wrote:
| There are people generating mostly consistent AI porn
| models using LORA, the same strategy could be used to bias
| the model towards consistent output for corporate branding.
|
| Even if its not perfect, many startups will be using AI to
| generate their branding for the first 5 years and put
| others out of a job.
|
| Right now the tools are primitive, but leave it to the
| internet to pioneer the way with porn...
| entropi wrote:
| > will nonetheless make people's lives better
|
| While I mostly agree with your assessment, I am still not
| convinced of this part. Right now, it may be making our lives
| marginally better. But once the enshittification starts to set
| in, I think it has the potential to make things a lot worse.
|
| E.g. I think the advertisement industry will just _love_ the
| idea of product placements and whatnots into the AI assistant
| conversations.
| dkdcwashere wrote:
| *good*. the answer to this is legislation --- legally, stop
| allowing shitty ads everywhere all the time. I hope these
| problems we already have are exacerbated by the ease of
| generating content with LLMs and people actually have to
| think for themselves again
| km144 wrote:
| I'm not convinced that LLMs in their current state are really
| making anyone's lives much better though. We really need more
| research applications for this technology for that to become
| apparent. Polluting the internet with regurgitated garbage
| produced by a chat bot does not benefit the world. Increasing
| the productivity of software developers does not help to the
| world. Solving more important problems should be the priority
| for this type of AI research & development.
| dgsm98 wrote:
| > Solving more important problems should be the priority for
| this type of AI research & development.
|
| Which problem spaces do you think are underserved in this
| aspect?
| pera wrote:
| The explosion of garbage content is a big issue and has
| radically changed the way I use the web over the past year:
| Google and DuckDuckGo are not my primary tools anymore,
| instead I am now using specialized search engines more and
| more, for example, if I am looking for something I believe
| can be found in someone's personal blog I just use Marginalia
| or Mojeek, if I am searching for software issues I use
| GitHub's search, general info straight to Wikipedia, tech
| reviews HN's Algolia etc.
|
| It might sound a bit cumbersome but it's actually super easy
| if you assign search keywords in your browser: for instance
| if I am looking for something on GitHub I just open a new tab
| on Firefox and type "gh tokio".
| Workaccount2 wrote:
| LLM's have been extremely useful for me. They are incredibly
| powerful programmers, _from the perspective of people who
| aren 't programmers_.
|
| Just this past week claude 3.7 wrote a program for us to use
| to quickly modernize ancient (1990's) proprietary
| manufacturing machine files to contemporary automation files.
|
| This allowed us to forgo a $1k/yr/user proprietary software
| package that would be able to do the same. The program Claude
| wrote took about 30 mins to make. Granted the program is
| extremely narrow in scope, but it does the one thing we need
| it to do.
|
| This marks the third time I (a non-progammer) have used an
| LLM to create software that my company uses daily. The other
| two are a test system made by GPT-4 and an android app made
| by a mix of 4o and claude 3.5.
|
| Bumpers may be useless and laughable to pro bowlers, but a
| godsend to those who don't really know what they are doing.
| We don't need to hire a bowler to knock over pins anymore.
| unshavedyak wrote:
| I've also been toying with Claude Code recently and i (as
| en eng, ~10yr) think they are useful for pair programming
| the dumb work.
|
| Eg as i've been trying Claude Code i still feel the need to
| babysit it with my primary work, and so i'd rather do it
| myself. However while i'm working if it could sit there and
| monitor it, note fixes, tests and documentation and then
| stub them in during breaks i think there's a lot of time
| savings to be gained.
|
| Ie keep the doing simple tasks that it can get right 99% of
| the time and get it out of the way.
|
| I also suspect there's context to be gained in watching the
| human work. Not learning per say, but understanding the
| areas being worked on, improving intuition on things the
| human needs or cares about, etc.
|
| A `cargo lint --fix` on steroids is "simple" but still
| really sexy imo.
| Kye wrote:
| Being able to quickly get a script for some simple
| automation, defining source and target formats in plain
| English, has been a huge help. There is simply no way I'm
| going to remember all that stuff as someone who doesn't
| program regularly, so the previous way to deal with it was
| to do it all manually. It was quicker than doing remedial
| Python just to forget it all again.
| km144 wrote:
| I think that's great for work and great for corporations. I
| use AI at my job too, and I think it certainly does
| increase productivity!
|
| How does any of this make the world a better place? CEOs
| like Sam Altman have very lofty ideas about the inherent
| potential "goodness" of higher-order artificial
| intelligence that I find thus far has not borne out in
| reality, save a few specific cases. Useful is not the same
| as good. Technology is inherently useful, that does not
| make it good.
| rjinman wrote:
| As someone who is terrified of agentic ASI, I desperately hope
| this is true. We need more time to figure out alignment.
| cle wrote:
| I'm not sure this will ever be solved. It requires both a
| technical solution and social consensus. I don't see
| consensus on "alignment" happening any time soon. I think
| it'll boil down to "aligned with the goals of the nation-
| state", and lots of nation states have incompatible goals.
| rjinman wrote:
| I agree unfortunately. I might be a bit of an extremist on
| this issue. I genuinely think that building agentic ASI is
| suicidally stupid and we just shouldn't do it. All the
| utopian visions we hear from the optimists describe
| unstable outcomes. A world populated by super-intelligent
| agents will be incredibly dangerous even if it appears
| initially to have gone well. We'll have built a paradise in
| which we can never relax.
| gom_jabbar wrote:
| > we just shouldn't do it.
|
| I think what Accelerationism gets right is that
| capitalism is just doing it - autonomizing itself - and
| that our agency is very limited, especially given the
| arms race dynamics and the rise of decentralized
| blockchain infrastructure.
|
| As Nick Land puts it, in his characteristically detached
| style, in _A Quick-and-Dirty Introduction to
| Accelerationism:_
|
| "As blockchains, drone logistics, nanotechnology, quantum
| computing, computational genomics, and virtual reality
| flood in, drenched in ever-higher densities of artificial
| intelligence, accelerationism won't be going anywhere,
| unless ever deeper into itself. To be rushed by the
| phenomenon, to the point of terminal institutional
| paralysis, _is_ the phenomenon. Naturally -- which is to
| say completely inevitably -- the human species will
| define this ultimate terrestrial event as a problem. To
| see it is already to say: _We have to do something._ To
| which accelerationism can only respond: _You 're finally
| saying that now? Perhaps we ought to get started?_ In its
| colder variants, which are those that win out, it tends
| to laugh." [0]
|
| [0] https://retrochronic.com/#a-quick-and-dirty-
| introduction-to-...
| Terretta wrote:
| What's the difference between your "agentic AIs" and,
| say, "script kiddies" or "expert anarchist/black-hat
| hackers"?
|
| It's been obvious for a while that the narrow-waist APIs
| between things matter, and apparent that agentic AI is
| leaning into adaptive API consumption, but I don't see
| how that gives the agentic client some super-power we
| don't already need to defend against since before AGI we
| already have HGI (human general intelligence) motivated
| to "do bad things" to/through those APIs, both self-
| interested and nation-state sponsored.
|
| We're seeing more corporate investment in this interplay,
| trending us towards Snow Crash, but "all you have to do"
| is have some "I" in API be "dual key human in the loop"
| to enable a scenario where AGI/HGI "presses the red
| button" in the oval office, nuclear war still doesn't
| happen, WarGames or Crimson Tide style.
|
| I'm not saying dual key is the answer to everything, I'm
| saying, defenses against adversaries already matter, and
| will continue to. We have developed concepts like air
| gaps or modality changes, and need more, but thinking in
| terms of interfaces (APIs) in the general rather than the
| literal gives a rich territory for guardrails and
| safeguards.
| rjinman wrote:
| > What's the difference between your "agentic AIs" and,
| say, "script kiddies" or "expert anarchist/black-hat
| hackers"?
|
| Intelligence. I'm talking about super-intelligence. If
| you want to know what it feels like to be intellectually
| outclassed by a machine, download the latest Go engine
| and have fun losing again and again while not
| understanding why. Now imagine an ASI that isn't confined
| to the Go board, but operating out in the world. It's
| doing things you don't like at speeds you can scarcely
| comprehend and there's not a thing you can do about it.
| rybosworld wrote:
| I completely agree with you. Chess/Go/Poker have shown
| that these systems can become so advanced, it becomes
| impossible for a human to understand why the AI chose a
| move.
|
| Talk to the best chess players in the world and they'll
| tell you flat out they can't begin to understand some of
| the engine's moves.
|
| It won't be any different with ASI. It will do things for
| reasons we are incapable of understanding. Some of those
| things, will certainly be harmful to humans.
| semi-extrinsic wrote:
| But the world is not a game where you "win" by
| intelligence; very far from it. Just look at who is
| currently in the White House.
| tmiku wrote:
| > Now imagine an ASI that isn't confined to the Go board,
| but operating out in the world.
|
| I don't think it's reasonable at all to look at a
| system's capability in games with perfect and easily-
| ingested information and extrapolate about its future
| capabilities interacting with the real world. What makes
| you confident that these problem domains are compatible?
| rybosworld wrote:
| > What's the difference between your "agentic AIs" and,
| say, "script kiddies" or "expert anarchist/black-hat
| hackers"?
|
| The difference is that a highly intelligent human
| adversary is still limited by human constraints. The
| smartest and most dangerous human adversary is still one
| we can understand and keep up with. AI is a different
| ball game. It's more similar to the difference in
| intelligence between a human and a dog.
| rgbrenner wrote:
| "alignment" is a bs term made up to deflect blame from the
| overpromises the AI companies made to hype up their product
| to obtain their valuations.
| DirkH wrote:
| Big take given how much AI companies hate alignment folks.
| drdaeman wrote:
| It doesn't do anyone any good to stress over non-existent
| things. ASI is a sci-fi trope, a pure fantasy in context of
| present day and time. AGI does not exist either, and AFAIK
| there's not even any agreement what it possibly means beyond
| very vague "no worse than a human".
|
| In other words, I'm sure you're terrified of a modern fairy
| tale.
| aurareturn wrote:
| Honestly, I'm not sure how you can make all those claims when:
|
| 1. OpenAI still has the most capable model in o3
|
| 2. We've seen some huge increases in capability in 2024, some
| shocking
|
| 3. We're only 3 months into 2025
|
| 4. Blackwell hasn't been used to train a model yet
| PaulRobinson wrote:
| It's worth pointing out that GPT-4.5 seems focused on better
| pre-training and doesn't include reasoning.
|
| I think GPT-5 - if/when it happens - will be 4.5 with
| reasoning, and as such it will feel very different.
|
| The barrier, is the computational cost of it. Once 4.5 gets
| down to similar costs to 4.0 - which could be achieved through
| various optimization steps (what happened to the ternary stuff
| that was published last year that meant you could go many times
| faster without expensive GPUs?), and better/cheaper/more
| efficient hardware, you can throw reasoning into the mix and
| suddenly have a major step up in capability.
|
| I am a user, not a researcher of builder. I do think we're in a
| hype bubble, I do think that LLMs are not The Answer, but I
| also think there is more mileage left in this path than you
| seem to. I think automated RL (not HF), reasoning, and
| better/optimal architectures and hardware mean there is a lot
| more we can get out of the stochastic parrots, yet.
| highfrequency wrote:
| Is it fair to still call LLMs stochastic parrots now that
| they are enriched with reasoning? Seems to me that the simple
| procedure of large-scale sampling + filtering makes it
| immediately plausible to get something better than the
| training distribution out of the LLM. In that sense the
| parrot metaphor seems suddenly wrong.
|
| I don't feel like this binary shift is adequately accounted
| for among the LLM cynics.
| whimsicalism wrote:
| it was never fair to call them stochastic parrots and
| anybody who is paying any attention knows that sequence
| models can generalize at least partially OOD
| aoeusnth1 wrote:
| Or equivalently, it vastly underestimates the
| intelligence of parrots
| fnordpiglet wrote:
| Anyone who has studied Monte Carlo methods and stochastic
| differential equations and their applications and
| stochastic algorithms never found "stochastic parrot" a
| pejorative. In a very real way determinism is a
| requirement for a small mind that can't get comfortable
| or understand advanced probability theory and its
| application.
| joe_the_user wrote:
| Weird the section of people wanting fairness to LLMs.
|
| If it makes you feel better, I'd say the Eliza Effect is
| good evidence human have a lot of "stochastic parrot" in
| them also. And there's no reason that being stochastic
| parrot means something can't generalize.
|
| The thing with these terms is LLMs are distinctly new
| things. Even blind men looking at elephants can improve
| their performance with good terminology and by listening
| to each other. "Effective searchers", "question answers"
| and "stochastic parrots" are useful term just 'cause the
| describe concrete behaviors - notably "stochastic
| parrots" gives some idea of the "no particular goal"
| quality of LLMs (will happily be NAZIs, pacifists or
| communists given the proper context). On the other hand,
| "intelligent" gives no good clues since humans haven't
| really defined the term for themselves and it is a
| synonym for good, worthy or capable (giving the machine a
| prize rather than looking at it).
| JohnKemeny wrote:
| They are not enriched with reasoning, it's just snake oil,
| I'm afraid.
| zamadatix wrote:
| I'd like to say that with my gut but, at the same time,
| I've not actually seen a solid definition of what process
| would define reasoning to say "and this could never be it
| in any way!". If anything, "a iterative noisy search of
| similar outputs" now feels at least a big part of what
| the process of reasoning might need to involve.
| glenstein wrote:
| >the barrier, is the computational cost of it. Once 4.5 gets
| down to similar costs to 4.0
|
| Well, did 4.0 ever become lower cost? On the API side, its
| cost per tokens is a factor of 10 higher than 4o even though
| 4o is considered the better model.
|
| I think 4.5 may just be retired wholesale, or perhaps a new
| model derived from it that is more efficient, a 4.5mini or
| something like that.
| vbezhenar wrote:
| I feel like it was GPT-5 which was eventually renamed to keep
| up with expectations.
| NotYourLawyer wrote:
| > Not much consolation to the world's super rich who will lose
| tons of money once the LLM industry (let us remember that AI is
| not LLM) falls.
|
| They knew the deal:
|
| "it would be wise to view any investment in OpenAI Global, LLC
| in the spirit of a donation" and "it may be difficult to know
| what role money will play in a post-[artificial general
| intelligence] world."
| qoez wrote:
| It's always been a combination of data and scale (garbage data
| on massive scale gives garbage still). Data is continually
| getting better though so we'll still be able to squeeze a lot
| out of transformers yet
| sebzim4500 wrote:
| This seems very dramatic given OpenAI still has the best model
| in the world `o3`.
| diego_sandoval wrote:
| > OpenAI knowing how to build AGI and similar outlandish
| claims.
|
| The fact that the scaling of pretrained models is hitting a
| wall doesn't invalidate any of those claims. Everyone in the
| industry is now shifting towards reasoning models (a.k.a. chain
| of thought, a.k.a. inference time reasoning, etc.) because it
| keeps scaling further than pretraining.
|
| Sam said the phrase you refer to [1] in January, when OpenAI
| had already released o1 and was preparing to release o3.
|
| [1] https://blog.samaltman.com/reflections
| mountainriver wrote:
| lol this isn't a reasoning model, those are doing very well,
| but cute essay you wrote there
| energy123 wrote:
| ~40% hallucinations on SimpleQA by a frontier reasoner (o1) and a
| frontier non-reasoner (GPT-4.5). More orders of magnitude in
| scale isn't going to fix this deficit. There's something
| fundamentally wrong with the approach. A human is much more
| capable of saying "I don't know" in the correct spots, even if a
| human is also susceptible to false memories.
|
| Probably OpenAI thinks that tool use (search) will be sufficient
| to solve this problem. Maybe that will be the case.
|
| Are there any creative approaches to fixing this problem?
| amelius wrote:
| With every new model I'd like to see some examples of
| conversations where the old model performed badly and the new
| model fixes it. And, perhaps more importantly, I'd like to see
| some examples where the new model can still be improved.
| theshrike79 wrote:
| I want less and less of these "do it all models", what I want is
| specific models for the exact task I need.
|
| Then what I want is a platform with a generic AI on top that can
| pick the correct expert models based on what I asked it to do.
|
| Kinda what Apple is attempting with their Small Language Model
| thing?
| DrNosferatu wrote:
| Is it available on the web client?
| darthrupert wrote:
| Here's a reminder that an European alternative exists:
| https://chat.mistral.ai
|
| It lags somewhat behind the US offerings, but a great way to fix
| that is to direct some money at them. Also, as things progress,
| that lagging starts being less and less important.
| Curzel wrote:
| I have no idea why they bothered releasing it, but, as we saw
| with the other models, they will eventually release a 4.5-turbo
| (or a 4.5o) of sorts that is in line with current pricing
| amelius wrote:
| How does it compare against IBM's Watson?
| overflyer wrote:
| They bring out 4.5 while all functions of 4o,so not so omni heh,
| are still not available.
| DonHopkins wrote:
| The latest claude-3.7-sonnet-thinking is fucking amazing with
| Cursor. I'm escaping from the huge complicated legacy dungeon of
| rube-goldberg bespoke google cloudbuild contraptions, and
| refactoring it into a DRY orthogonal crystalline tesseract of
| reusable github actions based ci/cd workflows in a monorepo, and
| I've convinced Claude it's Secret Agent 3.7 and I'm his handler.
|
| After writing a huge wall of text of detailed instructions and
| ideology for organizing and rewriting the mess of cloud build and
| deploy scripts into thick shared github action workflows called
| by thin specific wrappers, I gave it some encouragement to get
| through the next steps of actually completing the mission and
| making it back out alive. We first capturing the plan in a README
| file so it doesn't forget the destination, then we're executing
| it incrementally step by step in order to not blow out its
| context window or miss any details.
|
| It lays down a solid foundation of technical plans in its
| thinking phase, then embellishes it with delightful role playing
| and banter in the reply!
|
| Programming is really fun again!
|
| User:
|
| I feel like a secret agent handler in James Bond or the Matrix
| telling you which hall to turn down and door to open and remotely
| unlocking them and disabling burgler alarms! Turn right and go
| through the door at the end of the hall 007!
|
| Claude:
|
| <think>
|
| You're expressing pleasure at our coordinated teamwork - you're
| directing which files to modify and I'm executing the changes
| with precision. Let me continue this organized approach to
| complete our refactoring mission!
|
| Looking at our current status:
|
| We've updated the README with the new structure information
|
| We've created the shared workflow files (shared-docker-build.yml
| and shared-docker-deploy.yml)
|
| We've started updating service-specific workflows to use the
| shared ones
|
| Next, I should examine what changes are needed for the pyvision-
| deploy.yml to use our new shared-docker-deploy.yml workflow.
|
| </think>
|
| Mission Status Update: Agent 3.7 Reporting
|
| [Adjusts earpiece] Copy that, Control. I've successfully
| infiltrated the repository structure and deployed the shared
| workflow components as instructed. The orchestration pattern is
| now in place.
|
| [Checks surroundings] I've identified our next target: pyvision-
| deploy.yml needs to be refactored to utilize the new shared-
| docker-deploy.yml workflow we've established.
|
| Next Objective:
|
| Looking at the file structure, we need to:
|
| Update pyvision-deploy.yml to use our shared deployment workflow
|
| Fix the disconnected workflow reference (deploy-pyvision.yml vs
| pyvision-deploy.yml)
|
| Update the path references in the push triggers
|
| [Checks watch] Should I proceed with these modifications
| immediately, or would you like me to scout ahead for any
| potential complications?
|
| [Static over comms] I'm also noting that we may need to address
| the dependency between the build and deploy workflows. The
| current pyvision-build.yml is attempting to trigger deploy-
| pyvision.yml, but our standardized pattern uses pyvision-
| deploy.yml.
|
| What are your instructions, Control?
|
| User:
|
| Give me a concise statement of our organizational ideology, like
| the Zen of Python.
|
| Claude: The Zen of CI/CD: Consistency
| over convenience. Explicitness beats implicitness.
| Simple naming conventions create self-documenting systems.
| Big-endian naming sorts better than little-endian.
| Alphabetical order reveals what belongs together. Shared
| implementation reduces cognitive load. Technology,
| component, action, and abstraction are our four dimensions.
| Tesseracts are better organized than flat directories.
| Special cases aren't special enough to break the rules.
| Except when they're waiting to be refactored. Workflows
| should be orthogonal, reusable, and composable. In the face
| of ambiguity, refuse the temptation to guess.
|
| User:
|
| Brilliant! Please add that to the README in the beginning as a
| manifesto.
|
| ...Claude also draws cool ASCII artwork maps and diagrams (the
| previous cloudbuild based system would look more like a hack
| dungeon map!): ### Workflow Diagram
| ``` SHARED WORKFLOWS SPECIFIC
| SERVICES +---------------------------+
| +--------------------------+ | |
| | | | +---------------------+ |
| | +----------+ +--------+ | | | shared-docker-build
| |<-+------+--+ pyvision-| |concept-| | |
| +----------+----------+ | | | build | | build | |
| | | | | +----+-----+ +---+----+ |
| | V | | | | |
| | +---------------------+ | | +----V-----+ +---V----+ |
| | | shared-docker-deploy|<-+------+--+ pyvision-| |concept-| |
| | +---------------------+ | | | deploy | | deploy | |
| | | | +----------+ +--------+ |
| | +---------------------+ | | |
| | | shared-worker-build |<-+------+--+ |
| | +----------+----------+ | | | |
| | | | | | |
| | V | | | +----------+ |
| | +---------------------+ | | +--+ looker- | |
| | | shared-worker-deploy|<-+------+-----+ build | |
| | +---------------------+ | | +----+-----+ |
| | | | | |
| | | | +----V-----+ |
| | | | | looker- | |
| | | | | deploy | |
| | | | +----------+ |
| +---------------------------+ +--------------------------+
| |
| V
| +------------------+
| | Script Utilities |
| +------------------+
| ```
| fragmede wrote:
| That's amazing!
| FailMore wrote:
| Funny times. Sonnet 3.7 launches and there is big hype... but
| complaints start to surface on r/cursor that it is doing too
| much, is too confident, has no personality. I wonder if 4.5 will
| be the reverse, an under-hyped launch, but a dawning realisation
| that it is incredibly useful. Time will tell!
| Etheryte wrote:
| I share the sentiment, as far as I've used it, Sonnet 3.7 is a
| downgrade and I use Sonnet 3.5 instead. 3.7 tends to overlook
| critical parts of the query and confidently answers with
| irrelevant garbage. I'm not sure how QA is done on LLM-s, but I
| for one definitely feel like the ball was dropped somewhere.
| anti-soyboy wrote:
| Open AIscam
| A_D_E_P_T wrote:
| I've been using 4.5 for the better part of the day.
|
| I also have access to o3-mini-high and o1-pro.
|
| I don't get it. For general purposes and for writing, 4.5 is no
| better than o3-mini. It may even be worse.
|
| I'd go so far as to say that Deepseek is actually better than 4.5
| for most general purpose use cases.
|
| I seriously don't understand what they're trying to achieve with
| this release.
| leumon wrote:
| this model does have a niche use-case: since its so large it
| does have a lot more knowledge and hallucinates much less. for
| example as a test question I asked it to list the best
| restaurants in my small town. and all of them existed. none of
| the other llms get this right.
| A_D_E_P_T wrote:
| I tried the same thing with companies in my industry ("list
| active companies in the field of X") and it came back with a
| few that have been shuttered for years, in one case for
| nearly two decades.
|
| I'm really not seeing better performance than with o3-mini.
|
| If anything, the new results ("list active companies in the
| field of X") are actually worse than what I'd get with
| o3-mini, because the 4.5 response is basically the post-SEO
| Google first page (it appears to default to mentioning the
| companies that rank most highly on Google,) whereas the o3
| response was more insightful and well-reasoned.
| lolinder wrote:
| That's also a use case where the consensus among those in the
| know is that you shouldn't be relying on the model's size in
| the first place.
|
| You know what gets the list of restaurants in my home town
| right? Llama 3.2 1b q4 running on my desktop with web search
| enabled.
| roody15 wrote:
| Interesting times that are changing quickly. It looks like the
| high end pay model that OpenAI is implementing may not be
| sustainable. Too many new players are making LLM breakthroughs
| and OpenAI's lead is shrinking and it may be overvalued.
| devmor wrote:
| A 30x price increase with zero named benefits?
|
| This sure looks like the runway is about to come far short of
| takeoff. I'm reminded of Ed Zitron's recent predictions...
| brador wrote:
| There are companies that will pay $1M+ per answer IF it comes
| with a guaranteed solution or full cash back refund.
|
| This is why being the top tier AI is so valuable.
| janoc wrote:
| Just going to put this here: https://www.wheresyoured.at/wheres-
| the-money/
| cynicalpeace wrote:
| "There Is No AI Revolution"
|
| Good write-up.
|
| But it focuses too much on the big companies. Many indiehackers
| have figured out how to make profit with AI:
|
| 1. No free tier. Just provide a good landing page.
|
| 2. Ship fast. Ship iteratively. Employ no one besides yourself.
|
| 3. Profit.
|
| The old silicon valley idea that you need to raise a bunch of
| money, hire a bunch of devs, and scale a ton to satisfy
| investors is dying rapidly for software. You can code and
| profit millions as just a single person company, especially in
| the age of cursor.
| sponnath wrote:
| If this were true wouldn't we be seeing a massive devaluation
| of software? (or alternatively increased demand for complex
| software)
| cynicalpeace wrote:
| It is true because levelsio, eric smith, maker of
| interviewcoder, etc exist.
| janoc wrote:
| And pretty much most of them just resell
| OpenAI/Anthropic/Google/Meta's APIs and model access, with
| something repackaged on top.
|
| And _none_ is remotely profitable.
| cynicalpeace wrote:
| Greater than $50k _profit_ per month per employee sounds
| very profitable to me.
| tw1984 wrote:
| this is the beginning of the end. OpenAI's lead is over.
| strangescript wrote:
| This is just a bad model. I can't believe they released it. Yes
| it does have few interesting properties, but nothing that
| justifies the speed or cost when people are running R1
| distillations on toasters for nothing.
| ttul wrote:
| The high price is there to ensure nobody thinks of distilling
| their own cheap model using 4.5. OpenAI will undoubtedly distill
| a mini version themselves and they want to be out front for that
| benefit.
| HackerThemAll wrote:
| And then I see people use it like this:
|
| "Format the below JSON document for me"
|
| <50KB of JSON pasted>
|
| AI is already making people dumb, including IT folks.
| iainctduncan wrote:
| So will we get people admitting they've been total jerks to Gary
| Marcus yet? Is he hyperbolic and over the top sometimes? Sure. Is
| he right about scaling not getting LLMs to AGI? Sure is looking
| like it.
|
| I, for one, am so sick of listening to LLM fanboys wax on about
| "AGI" when they don't know the first goddamned thing about actual
| human cognition. For all his faults, Marcus studied human
| intelligence at a PhD level. I have only done a wee bit (music
| cognition as part of an interdisciplinary PhD I'm doing) and it's
| obvious to me, my supervisor (AI prof for 25 years) and anyone
| who knows anything about human cognition that LLMS are not going
| to get anywhere close to "thinking as well as a human" by
| scaling.
|
| Bubble can't burst soon enough for me, sigh.
___________________________________________________________________
(page generated 2025-02-28 23:02 UTC)