[HN Gopher] GPT-5.2
       ___________________________________________________________________
        
       GPT-5.2
        
       https://platform.openai.com/docs/guides/latest-model  System card:
       https://cdn.openai.com/pdf/3a4153c8-c748-4b71-8e31-aecbde944...
        
       Author : atgctg
       Score  : 1154 points
       Date   : 2025-12-11 18:04 UTC (1 days ago)
        
 (HTM) web link (openai.com)
 (TXT) w3m dump (openai.com)
        
       | villgax wrote:
       | Marginal gains for exorbitantly pricey and closed model.....
        
       | sfmike wrote:
       | Everything is still based on 4 4o still right? is a new model
       | training just too expensive? They can consult deepseek team maybe
       | for cost constrained new models.
        
         | verdverm wrote:
         | Apparently they have not had a successful pre training run in
         | 1.5 years
        
           | fouronnes3 wrote:
           | I want to read a short scify story set in 2150 about how,
           | mysteriously, no one has been able to train a better LLM for
           | 125 years. The binary weights are studied with unbelievably
           | advanced quantum computers but no one can really train a new
           | AI from scratch. This starts cults, wars and legends and
           | ultimately (by the third book) leads to the main protagonist
           | learning to code by hand, something that no human left alive
           | still knows how to do. Could this be the secret to making a
           | new AI from scratch, more than a century later?
        
             | armenarmen wrote:
             | I'd read it!
        
             | barrenko wrote:
             | Monsieur, if I may offer a vaaaguely similar story on how
             | things may progress https://www.owlposting.com/p/a-body-
             | most-amenable-to-experim...
        
             | verdverm wrote:
             | You can ask 2025 Ai to write such a book, it's happy to
             | comply and may or may not actually write the book
             | 
             | https://www.pcgamer.com/software/ai/i-have-been-fooled-
             | reddi...
        
             | WhyOhWhyQ wrote:
             | There's a scifi short story about a janitor who knows how
             | to do basic arithmetic and becomes the most important
             | person in the world when some disaster happens. Of course
             | after things get set up again due to his expertise, he
             | becomes low status again.
        
               | bradfitz wrote:
               | I had to go look that up! I assume that's
               | https://en.wikipedia.org/wiki/The_Feeling_of_Power ? (Not
               | a janitor, but "a low grade Technician"?)
        
               | WhyOhWhyQ wrote:
               | Hmm it could be a false memory, since this was almost 15
               | years ago, but I really do remember it differently than
               | the text of 'Feeling of Power'.
        
             | ssl-3 wrote:
             | Sounds good.
             | 
             | Might sell better with the protagonist learning iron age
             | leatherworking, with hides tanned from cows that were grown
             | within earshot, as part of a process of finding the real
             | root of the reason for why any of us ever came to be in the
             | first place. This realization process culminates in the
             | formation of a global, unified steampunk BDSM movement and
             | a wealth of new diseases, and then: Zombies.
             | 
             | (That's the end. Zombies are always the end.)
        
               | wafflemaker wrote:
               | Sorry, but compared with the parent, my money is in you
               | ssl-3. Do you get better results from prompting by being
               | more poetic?
        
               | ssl-3 wrote:
               | > Do you get better results from prompting by being more
               | poetic?
               | 
               | Is that yet-another accusation of having used the bot?
               | 
               | I don't use the bot to write English prose. If something
               | I write seems particularly great or poetic or something,
               | then that's just me: I was in the right mood, at the
               | right time, with the right idea -- and with the right
               | audience.
               | 
               | When it's bad or fucked-up, then that's also just me. I
               | most-assuredly fuck up plenty.
               | 
               | They can't all be zingers. I'm fine with that.
               | 
               | ---
               | 
               | I do use the hell out of the bot for translating my ideas
               | (and the words that I use to express them) into languages
               | that I can't speak well, like Python, C, and C++. But
               | that's very different. (And at least so far I haven't
               | shared any of those bot outputs with the world at all,
               | either.)
               | 
               | So to take your question very literally: No, I don't get
               | better results from prompting being more poetic. The
               | responses to my prompts don't improve by those prompts
               | being articulate or poetic.
               | 
               | Instead, I've found that I get the best results from the
               | bot fastest by carrying a big stick, and using that stick
               | to hammer and welt it into compliance.
               | 
               | Things can get rather irreverent in my interactions with
               | the bot. Poeticism is pretty far removed from any of that
               | business.
        
               | wafflemaker wrote:
               | No. I just genuinely liked your style, and didn't notice
               | previous posts by you. I haven't yet learned to look at
               | names on hn, it's mostly anonymous posts for me. No snark
               | here. And was also genuinely curious if better writing
               | style yields better results.
               | 
               | I've observed that using proper grammar gives slightly
               | better answers. And using more "literacy"(?) kind of
               | language in prompts sometimes gives better answers and
               | sometimes just more interesting ones, when bots try to
               | follow my style.
               | 
               | Sorry for using the word poetic, I'm travelling and sleep
               | deprived and couldn't find the proper word, but didn't
               | want to just use "nice" instead either.
        
               | ssl-3 wrote:
               | It's all good. I'm largely "face-blind", myself, in that
               | I don't often recognize others in person or online --
               | which is certainly not to say that I think I'm
               | particularly memorable myself.
               | 
               | As to the bot: Man, I beat the bot to death. It's pretty
               | brutal.
               | 
               | I'm profane and demanding because that's the most _terse_
               | language I know how to construct in English.
               | 
               | When I set forth to have the bot do a thing for me, the
               | slowest part of the process that I can improve on my part
               | is the quantity of the words that I use.
               | 
               | I can type fast and think fast, but my one-letter-at-a-
               | time response to the bot is usually the only part that
               | that I can make a difference with. So I tend to be very
               | terse.
               | 
               | "a+b=c, you fuck!" is certainly terse, unambiguous, and
               | fast to type, so that's my usual style.
               | 
               | Including the emphatic "you fuck!" appendage seems to
               | stir up the context more than without. Its inclusion or
               | omission is a dial that can be turned.
               | 
               | Meanwhile: "I have some reservations about the proposed
               | implementation. Might it be possible for you to revise it
               | so as to be in a different form? As previously discussed,
               | it is my understanding that a+b=c. Would you like to try
               | again to implement a solution that incorporates this
               | understanding?" is very slow to write.
               | 
               | They both get similar results. One method is faster for
               | me than the other, just because I can only type so fast.
               | The operative function of the statement is ~the same
               | either way.
               | 
               | (I don't owe the bot anything. It isn't alive. It is just
               | a computer running a program. I could work harder to be
               | more polite, empathetic, or cordial, but: It's just code
               | running on a box somewhere in a datacenter that is
               | raising my electric rate and making the RAM for my next
               | system upgrade very expensive. I don't owe it anything,
               | much less politeness or poeticism.
               | 
               | Relatedly, my inputs at the _bash_ prompt on my home
               | computer are also very terse. For instance I don 't have
               | any desire or ability to be polite to _bash_ ; I just
               | issue commands like _ls_ and _awk_ and _grep_ without any
               | filler-words or pleasantries. The bot is no different to
               | me.
               | 
               | When I want something particularly poetic or verbose as
               | output from the bot, I simply command it to be that way.
               | 
               | It's just a program.)
        
               | astrange wrote:
               | This is somewhat similar to a Piers Anthony series that I
               | suspect noone has ever read except for me.
               | 
               | What was with that guy anyway.
        
             | georgefrowny wrote:
             | An software version of Asimov's Holmes-Ginsbook device?
             | https://sfwritersworkshop.org/node/1232
             | 
             | I feel like there was a similar one about software, but it
             | might have been mathematics (also Asimov: The Feeling of
             | Power)
        
           | ijl wrote:
           | What kind of issues could prevent a company with such
           | resources from that?
        
             | verdverm wrote:
             | Drama if I had to pick the symptom most visible from the
             | outside.
             | 
             | A lot of talent left OpenAI around that time, most notably
             | in this regard would be Ilya in May '24. Remember that time
             | Ilya and the board ousted Sam only to reverse it almost
             | immediately?
             | 
             | https://arstechnica.com/information-
             | technology/2024/05/chief...
        
         | Wowfunhappy wrote:
         | I thought whenever the knowledge cutoff increased that meant
         | they'd trained a new model, I guess that's completely wrong?
        
           | brokencode wrote:
           | Typically I think, but you could pre-train your previous
           | model on new data too.
           | 
           | I don't think it's publicly known for sure how different the
           | models really are. You can improve a lot just by improving
           | the post-training set.
        
           | rockinghigh wrote:
           | They add new data to the existing base model via continuous
           | pre-training. You save on pre-training, the next token
           | prediction task, but still have to re-run mid and post
           | training stages like context length extension, supervised
           | fine tuning, reinforcement learning, safety alignment ...
        
             | astrange wrote:
             | Continuous pretraining has issues because it starts
             | forgetting the older stuff. There is some research into
             | other approaches.
        
         | elgatolopez wrote:
         | Where did you get that from? Cutoff date says august 2025.
         | Looks like a newly pretrained model
        
           | SparkyMcUnicorn wrote:
           | If the pretraining rumors are true, they're probably using
           | continued pretraining on the older weights. Right?
        
           | FergusArgyll wrote:
           | > This stands in sharp contrast to rivals: OpenAI's leading
           | researchers have not completed a successful full-scale pre-
           | training run that was broadly deployed for a new frontier
           | model since GPT-4o in May 2024, highlighting the significant
           | technical hurdle that Google's TPU fleet has managed to
           | overcome.
           | 
           | - https://newsletter.semianalysis.com/p/tpuv7-google-takes-
           | a-s...
           | 
           | It's also plainly obvious from using it. The "Broadly
           | deployed" qualifier is presumably referring to 4.5
        
         | catigula wrote:
         | The irony is that Deepseek is still running with a distilled 4o
         | model.
        
           | blovescoffee wrote:
           | Source?
        
       | zamadatix wrote:
       | https://openai.com/index/introducing-gpt-5-2/
        
       | system2 wrote:
       | "Investors are putting pressure, change the version number
       | now!!!"
        
         | exe34 wrote:
         | I'm quite sad about the S-curve hitting us hard in the
         | transformers. For a short period, we had the excitement of "ooh
         | if GPT-3.5 is so good, GPT-4 is going to be amazing! ooh GPT-4
         | has sparks of AGI!" But now we're back to version inflation for
         | inconsequential gains.
        
           | verdverm wrote:
           | 2025 is the year most Big AI released their first real
           | thinking models
           | 
           | Now we can create new samples and evals for more complex
           | tasks to train up the next gen, more planning, decomp,
           | context, agentic oriented
           | 
           | OpenAI has largely fumbled their early lead, exciting stuff
           | is happening elsewhere
        
           | ToValueFunfetti wrote:
           | Take this all with a grain of salt as it's hearsay:
           | 
           | From what I understand, nobody has done any real scaling
           | since the GPT-4 era. 4.5 was a bit larger than 4, but not as
           | much as the orders of magnitude difference between 3 and 4,
           | and 5 is smaller than 4.5. Google and Anthropic haven't gone
           | substantially bigger than GPT-4 either. Improvements since 4
           | are almost entirely from reasoning and RL. In 2026 or 2027,
           | we should see a model that uses the current datacenter
           | buildout and actually scales up.
        
             | snovv_crash wrote:
             | Datacenter capacity is being snapped up for inference too
             | though.
        
             | Leynos wrote:
             | 4.5 is widely believed to be an order of magnitude larger
             | than GPT-4, as reflected in the API inference cost. The
             | problem is the quantity of parameters you can fit in the
             | memory of one GPU. Pretty much every large GPT model from 4
             | onwards has been mixture of experts, but for a 10 trillion
             | parameter scale model, you'd be talking a lot of experts
             | and a lot of inter-GPU communication.
             | 
             | With FP4 in the Blackwell GPUs, it should become much more
             | practical to run a model of that size at the deployment
             | roll-out of GPT-5.x. We're just going to have to wait for
             | the GBx00 systems to be physically deployed at scale.
        
           | JanSt wrote:
           | I don't feel the S-curve at all yet. Still an exponential for
           | me
        
             | exe34 wrote:
             | With a very long doubling time?
        
           | gessha wrote:
           | Because it will take thousands of underpaid researchers
           | random searching through solution space to get to the next
           | improvement, not 2-3 companies pressed to monetize and
           | enshittify their product before money runs out. That and
           | winning more hardware lotteries.
        
             | astrange wrote:
             | Underpaid? OpenAI!? It's pretty good I think.
             | 
             | https://www.levels.fyi/companies/openai/salaries/software-
             | en...
        
               | gessha wrote:
               | I'm talking about grad students, not OpenAI researchers.
        
       | tabletcorry wrote:
       | Slight increase in model cost, but looks like benefits across the
       | board to match.                 gpt-5.2 $1.75 $0.175 $14.00
       | gpt-5.1 $1.25 $0.125 $10.00
        
         | llmslave wrote:
         | They probably just beefed up compute run time on the what is
         | the same underlying model
        
         | jtbayly wrote:
         | 40% increase is not "slight."
        
           | credit_guy wrote:
           | Not the OP, but I think "slight" here is in relation to
           | Anthropic and Google. Claude Opus 4.5 comes at $25/MT
           | (million tokens), Sonnet 4.5 at $22.5/MT, and Gemini 3 at
           | $18/MT. GPT 5.2 at $14/MT is still the cheapest.
        
             | deaux wrote:
             | Your numbers are very off.                 $25 - Opus 4.5
             | $15 - Sonnet 4.5       $14 - GPT 5.2       $12 - Gemini 3
             | Pro
             | 
             | Even if you're including input, your numbers are still off.
        
               | credit_guy wrote:
               | I used the pricing for long context (>200k) in all cases.
               | I personally use AI as coding assistants, like lots of
               | other people, and as such, hitting and exceeding 200k is
               | quite the norm. The numbers you are showing are for <200k
               | context length.
        
         | commandar wrote:
         | In particular, the API pricing for GPT-5.2 Pro has me wondering
         | what on earth the possible market for that model is beyond
         | getting to claim a couple of percent higher benchmark
         | performance in press releases.
         | 
         | >Input:
         | 
         | >$21.00 / 1M tokens
         | 
         | >Output:
         | 
         | >$168.00 / 1M tokens
         | 
         | That's the most "don't use this" pricing I've seen on a model.
         | 
         | https://openai.com/api/pricing/
        
           | reactordev wrote:
           | Less an issue if your company is paying
        
             | rvnx wrote:
             | Even less an issue when OpenAI provides you free credits
        
           | arthurcolle wrote:
           | gpt-4-32k pricing was originally $60.00 / $120.00.
        
           | Leynos wrote:
           | Someone on Reddit reported that they were charged $17 for one
           | prompt on 5-pro. Which suggests around 125000 reasoning
           | tokens.
           | 
           | Makes me feel guilty for spamming pro with any random
           | question I have multiple times a day.
        
           | asgraham wrote:
           | Those prices seem geared toward people who are completely
           | price insensitive, who just want "the best" at any cost. If
           | the margins on that premium model are as high as they should
           | be, it's a smart business move to give them what they want.
        
           | wahnfrieden wrote:
           | Pro solves many problems for me on first try that the other
           | 5.1 models are unable to after many iterations. I don't pay
           | API pricing but if I could afford it I would in some cases
           | for the much higher context window it affords when a problem
           | calls for it. I'd rather spend some tens of dollars to solve
           | a problem than grind at it for hours.
        
           | aimanbenbaha wrote:
           | Last year o3 high did 88% on ARC-AGI 1 at more than
           | $4,000/task. This model at its X high configuration scores
           | 90.5% at just $11,64 per task.
           | 
           | General intelligence has ridiculously gotten less expensive.
           | I don't know if it's because of compute and energy
           | abundance,or attention mechanisms improving in efficiency or
           | both but we have to acknowledge the bigger picture and
           | relative prices.
        
             | commandar wrote:
             | Sure, but the reason I'm confused by the pricing is that
             | the pricing doesn't exist in a vacuum.
             | 
             | Pro _barely_ performs better than Thinking in OpenAI 's
             | published numbers, but comes at ~10x the price with an
             | explicit disclaimer that it's slow on the order of minutes.
             | 
             | If the published performance numbers are accurate, it seems
             | like it'd be incredibly difficult to justify the premium.
             | 
             | At least on the surface level, it looks like it exists
             | mostly to juice benchmark claims.
        
               | rvnx wrote:
               | It could be using the same early trick of Grok (at least
               | in the earlier versions) that they boot 10 agents who
               | work on the problem in parallel and then get a consensus
               | on the answer. This would explain the price and the
               | latency.
               | 
               | Essentially a newbie trick that works really well but not
               | efficient, but still looking like it's amazing
               | breakthrough.
               | 
               | (if someone knows the actual implementation I'm curious)
        
         | anvuong wrote:
         | In what world is that a slight increase?
        
       | meetpateltech wrote:
       | GPT-5.2 System Card PDF:
       | https://cdn.openai.com/pdf/3a4153c8-c748-4b71-8e31-aecbde944...
        
         | dang wrote:
         | Thanks, we'll put that in the toptext as well.
        
       | josalhor wrote:
       | From GPT 5.1 Thinking:
       | 
       | ARC AGI v2: 17.6% -> 52.9%
       | 
       | SWE Verified: 76.3% -> 80%
       | 
       | That's pretty good!
        
         | verdverm wrote:
         | We're also in benchmark saturation territory. I heard it
         | speculated that Anthropic emphasizes benchmarks less in their
         | publications because internally they don't care about them
         | nearly as much as making a model that works well on the day-to-
         | day
        
           | quantumHazer wrote:
           | Seems pretty false if you look at the model card and web site
           | of Opus 4.5 that is... (check notes) their latest model.
        
             | verdverm wrote:
             | Building a good model generally means it will do well on
             | benchmarks too. The point of the speculation is that
             | Anthropic is not focused on benchmaxxing which is why they
             | have models people like to use for their day-to-day.
             | 
             | I use Gemini, Anthropic stole $50 from me (expired and kept
             | my prepaid credits) and I have not forgiven them yet for
             | it, but people rave about claude for coding so I may try
             | the model again through Vertex Ai...
             | 
             | The person who made the speculation I believe was more
             | talking about blog posts and media statements than model
             | cards. Most ai announcements come with benchmark touting,
             | Anthropic supposedly does less / little of this in their
             | announcements. I haven't seen or gathered the data to know
             | what is truth
        
               | elcritch wrote:
               | You could try Codex cli. I prefer it over Claude code
               | now, but only slightly.
        
               | verdverm wrote:
               | No thanks, not touching anything Oligarchy Altman is
               | behind
        
           | Mistletoe wrote:
           | How do you measure whether it works better day to day without
           | benchmarks?
        
             | standardUser wrote:
             | Subscriptions.
        
               | mrguyorama wrote:
               | Ah yes, humans are famously empirical in their behavior
               | and we definitely do not have direct evidence of the
               | "best" sports players being much more likely than the
               | average to be superstitious or do things like wear "lucky
               | underwear" or buy right into scam bracelets that "give
               | you more balance" using a holographic sticker.
        
             | bulbar wrote:
             | Manually labeling answers maybe? There exist a lot of
             | infrastructure built around and as it's heavily used for 2
             | decades and it's relatively cheap.
             | 
             | That's still benchmarking of course, but not utilizing any
             | of the well known / public ones.
        
             | verdverm wrote:
             | Internal evals, Big AI certainly has good, proprietary
             | training and eval data, it's one reason why their models
             | are better
        
               | aydyn wrote:
               | Then publish the results of those internal evals. Public
               | benchmark saturation isn't an excuse to be un-
               | quantitative.
        
               | verdverm wrote:
               | How would published numbers be useful without knowing
               | what the underlying data being used to test and evaluate
               | them are? They are proprietary for a reason
               | 
               | To think that Anthropic is not being intentional and
               | quantitative in their model building, because they care
               | less for the saturated benchmaxxing, is to miss the
               | forest for the trees
        
               | aydyn wrote:
               | Do you know everything that exists in public benchmarks?
               | 
               | They can give a description of what their metrics are
               | without giving away anything proprietary.
        
               | verdverm wrote:
               | I'd recommend watching Nathan Lambert's video he dropped
               | yesterday on Olmo 3 Thinking. You'll learn there's a lot
               | of places where even descriptions of proprietary testing
               | regimes would give away some secret sauce
               | 
               | Nathan is at Ai2 which is all about open sourcing the
               | process, experience, and learnings along the way
        
               | aydyn wrote:
               | Thanks for the reference I'll check it out. But it doesnt
               | really take away from the point I am making. If a level
               | of description would give away proprietary information,
               | then go one level up to a more vague description. How to
               | describe things to a proper level is more of a social
               | problem than a technical one.
        
           | brokensegue wrote:
           | how do you quantitatively measure day-to-day quality? only
           | thing i can think is A/B tests which take a while to evaluate
        
             | verdverm wrote:
             | more or less this, but also synthetic
             | 
             | if you think about GANs, it's all the same concept
             | 
             | 1. train model (agent)
             | 
             | 2. train another model (agent) to do something interesting
             | with/to the main model
             | 
             | 3. gain new capabilities
             | 
             | 4. iterate
             | 
             | You can use a mix of both real and synthetic chat sessions
             | or whatever you want your model to be good at. Mid/late
             | training seems to be where you start crafting personality
             | and expertises.
             | 
             | Getting into the guts of agentic systems has me believing
             | we have quite a bit of runway for iteration here,
             | especially as we move beyond single model / LLM training. I
             | still need to get into what all is de jour in the RL / late
             | training, that's where a lot of opportunity lies from my
             | understanding so far
             | 
             | Nathan Lambert
             | (https://bsky.app/profile/natolambert.bsky.social) from Ai2
             | (https://allenai.org/) & RLHF Book (https://rlhfbook.com/)
             | has a really great video out yesterday about the experience
             | training Olmo 3 Think
             | 
             | https://www.youtube.com/watch?v=uaZ3yRdYg8A
        
           | HDThoreaun wrote:
           | Arc-AGI is just an iq test. I don't see the problem with
           | training it to be good at iq tests because that's a skill
           | that translates well.
        
             | CamperBob2 wrote:
             | Exactly. In principle, at least, the only way to overfit to
             | Arc-AGI is to actually _be_ that smart.
             | 
             | Edit: if you disagree, try actually TAKING the Arc-AGI 2
             | test, then post.
        
               | npinsker wrote:
               | Completely false. This is like saying being good at chess
               | is equivalent to being smart.
               | 
               | Look no farther than the hodgepodge of independent teams
               | running cheaper models (and no doubt thousands of their
               | own puzzles, many of which surely overlap with the
               | private set) that somehow keep up with SotA, to see how
               | impactful proper practice can be.
               | 
               | The benchmark isn't particularly strong against gaming,
               | especially with private data.
        
               | CamperBob2 wrote:
               | _Completely false. This is like saying being good at
               | chess is equivalent to being smart._
               | 
               | No, it isn't. Go take the test yourself and you'll
               | understand how wrong that is. Arc-AGI is intentionally
               | unlike any other benchmark.
        
               | fwip wrote:
               | Took a couple just now. It seems like a straight-forward
               | generalization of the IQ tests I've taken before,
               | reformatted into an explicit grid to be a little bit
               | friendlier to machines.
               | 
               | Not to humble-brag, but I also outperform on IQ tests
               | well beyond my actual intelligence, because "find the
               | pattern" is fun for me and I'm relatively good at visual-
               | spatial logic. I don't find their ability to measure
               | 'intelligence' very compelling.
        
               | CamperBob2 wrote:
               | Given your intellectual resources -- which you've
               | successfully used to pass a test that is _designed_ to be
               | easy for humans to pass while tripping up AI models --
               | why not use them to suggest a better test? The people who
               | came up with Arc-AGI were not actually morons, but I 'm
               | sure there's room for improvement.
               | 
               | What would be an example of a test for machine
               | intelligence that you would accept? I've already
               | suggested one (namely, making up more of these sorts of
               | tests) but it'd be good to get some additional opinions.
        
               | fwip wrote:
               | Dunno :) I'm not an expert at LLMs or test design, I just
               | see a lot of similarity between IQ tests and these
               | questions.
        
               | mrandish wrote:
               | ARC-AGI was designed specifically for evaluating deeper
               | reasoning in LLMs, including being resistant to LLMs
               | 'training to the test'. If you read Francois' papers,
               | he's well aware of the challenge and has done valuable
               | work toward this goal.
        
               | npinsker wrote:
               | I agree with you. I agree it's valuable work. I totally
               | disagree with their claim.
               | 
               | A better analogy is: someone who's never taken the AIME
               | might think "there are an infinite number of math
               | problems", but in actuality there are a relatively small,
               | enumerable number of techniques that are used repeatedly
               | on virtually all problems. That's not to take away from
               | the AIME, which is quite difficult -- but not infinite.
               | 
               | Similarly, ARC-AGI is much more bounded than they seem to
               | think. It correlates with intelligence, but doesn't imply
               | it.
        
               | keeda wrote:
               | Maybe I'm misinterpreting your point, but this makes it
               | seem that your standard for "intelligence" is "inventing
               | entirely new techniques"? If so, it's a bit extreme,
               | because to a first approximation, _all_ problem solving
               | is combining and applying existing techniques in novel
               | ways to new situations.
               | 
               | At the point that you are inventing entirely new
               | techniques, you are usually doing groundbreaking work.
               | Even groundbreaking work in one field is often inspired
               | by techniques from other fields. In the limit,
               | discovering truly new techniques often requires
               | discovering new principles of reality to exploit, i.e.
               | research.
               | 
               | As you can imagine, this is very difficult and hence
               | rather uncommon, typically only accomplished by a handful
               | of people in any given discipline, i.e way above the
               | standards of the general population.
               | 
               | I feel like if we are holding AI to those standards, we
               | are talking about not just AGI, but artificial super-
               | intelligence.
        
               | yovaer wrote:
               | > but in actuality there are a relatively small,
               | enumerable number of techniques that are used repeatedly
               | on virtually all problems
               | 
               | IMO/AIME problems perhaps, but surely that's too narrow a
               | view for all of mathematics. If solving conjectures were
               | simply a matter of trying a standard range of techniques
               | enough times, then there would be a lot fewer open
               | problems around than what's the case.
        
               | esafak wrote:
               | I would not be so sure. You can always prep to the test.
        
               | HDThoreaun wrote:
               | How do you prep for arc agi? If the answer is just "get
               | really good at pattern recognition" I do not see that as
               | a negative at all.
        
               | ben_w wrote:
               | It can be not-negative without being sufficient.
               | 
               | Imagine that pattern recognition is 10% of the problem,
               | and we just don't know what the other 90% is yet.
               | 
               | Streetlight effect for "what is intelligence" leads to
               | all the things that LLMs are now demonstrably good at...
               | and yet, the LLMs are somehow missing a lot of stuff and
               | we have to keep inventing new street lights to search
               | underneath:
               | https://en.wikipedia.org/wiki/Streetlight_effect
        
               | HDThoreaun wrote:
               | I dont think many people are saying 100% arc-agi 2 is
               | equivalent to AGI(names are dumb as usual). Its just the
               | best metric I have found, not the final answer. Spatial
               | reasoning is an important part of intelligence even if it
               | doesnt encompass all of it.
        
               | jimbokun wrote:
               | Is it different every time? Otherwise the training could
               | just memorize the answers.
        
               | CamperBob2 wrote:
               | The models never have access to the answers for the
               | private set -- again, at least in principle. Whether
               | that's actually true, I have no idea.
               | 
               | The idea behind Arc-AGI is that you can train all you
               | want on the answers, because knowing the solution to one
               | problem isn't helpful on the others.
               | 
               | In fact, the way the test works is that the model is
               | given several examples of worked solutions for each
               | problem class, and is then required to infer the
               | underlying rule(s) needed to solve a different instance
               | of the same type of problem.
               | 
               | That's why comparing Arc-AGI to chess or other
               | benchmaxxing exercises is completely off base.
               | 
               | (IMO, an even better test for AGI would be "Make up some
               | original Arc-AGI problems.")
        
               | FergusArgyll wrote:
               | It's very much a vision test. The reason all the models
               | don't pass it easily is only because of the vision
               | component. It doesn't have much to do with reasoning at
               | all
        
               | ACCount37 wrote:
               | With this kind of thing, the tails ALWAYS come apart, in
               | the end. They come apart later for more robust tests, but
               | "later" isn't "never", far from it.
               | 
               | Having a high IQ helps a lot in chess. But there's a
               | considerable "non-IQ" component in chess too.
               | 
               | Let's assume "all metrics are perfect" for now. Then,
               | when you score people by "chess performance"? You
               | wouldn't see the people with the highest intelligence
               | ever at the top. You'd get people with pretty high
               | intelligence, but extremely, hilariously strong chess-
               | specific skills. The tails came apart.
               | 
               | Same goes for things like ARC-AGI and ARC-AGI-2. It's an
               | interesting metric (isomorphic to the progressive matrix
               | test? usable for measuring human IQ perhaps?), but no
               | metric is perfect - and ARC-AGI is biased heavily towards
               | spatial reasoning specifically.
        
             | fwip wrote:
             | It is very similar to an IQ test, with all the attendant
             | problems that entails. Looking at the Arc-AGI problems, it
             | seems like visual/spatial reasoning is just about the only
             | thing they are testing.
        
           | stego-tech wrote:
           | These models still consistently fail the only benchmark that
           | matters: if I give you a task, can you complete it
           | successfully without making shit up?
           | 
           | Thus far they all fail. Code outputs don't run, or variables
           | aren't captured correctly, or hallucinations are stated as
           | factual rather than suspect or "I don't know."
           | 
           | It's 2000's PC gaming all over again ("gotta game the
           | benchmark!").
        
             | verdverm wrote:
             | I'm not sure, here's my anecdotal counter example, was able
             | to get gemini-2.5-flash, in two turns, to understand and
             | implement something I had done separately first, and it
             | found another bug (also that I had fixed, but forgot was in
             | this path)
             | 
             | That I was able to have a flash model replicate the same
             | solution I had, to two problems in two turns, it's just the
             | opposite experience of your consistency argument. I'm using
             | tasks I've already solved as the evals while developing my
             | custom agentic setup (prompts/tools/envs). They are able to
             | do more of them today then they were even 6-12 months ago
             | (pre-thinking models).
             | 
             | https://bsky.app/profile/verdverm.com/post/3m7p7gtwo5c2v
        
               | stego-tech wrote:
               | And therein lies the rub for why I still approach this
               | technology with caution, rather than charge in full steam
               | ahead: _variable outputs based on immensely variable
               | inputs_.
               | 
               | I read stories like yours all the time, and it encourages
               | me to keep trying LLMs from almost all the major vendors
               | (Google being a noteworthy exception while I try and get
               | off their platform). I _want_ to see the magic others
               | see, but when my IT-brain starts digging in the guts of
               | these things, I'm always disappointed at how unstructured
               | and random they ultimately are.
               | 
               | Getting back to the benchmark angle though, we're firmly
               | in the era of benchmark gaming - hence my quip about
               | these things failing "the only benchmark that matters." I
               | _meant_ for that to be interpreted along the lines of,
               | "trust your own results rather than a spreadsheet matrix
               | of other published benchmarks", but I clearly missed the
               | mark in making that clear. That's on me.
        
               | verdverm wrote:
               | I mean more the guts of the agentic systems. Prompts,
               | tool design, state and session management, agent transfer
               | and escalation. I come from devops and backend dev, so
               | getting in at this level, where LLMs are tasked and
               | composed, is more interesting.
               | 
               | If you are only using provider LLM experiences, and not
               | something specific to coding like copilot or Claude code,
               | that would be the first step to getting the magic as you
               | say. It is also not instant. It takes time to learn any
               | new tech, this one has a above average learning curve,
               | despite the facade and hype of how it should just be
               | magic
               | 
               | Once you find the stupid shit in the vendor coding
               | agents, like all us it/devops folks do eventually, you
               | can go a level down and build on something like the ADK
               | to bring your expertise and experience to the building
               | blocks.
               | 
               | For example, I am now implementing environments for
               | agents based on container layers and Dagger, which
               | unlocks the ability to cheaply and reproducible clone
               | what one agent was doing and have a dozen variations
               | iterate on the next turn. Real useful for long term
               | training data and evals synth, but also for my own
               | experimentation as I learn how to get better at using
               | these things. Another thing I did was change how
               | filesystem operations look to the agent, in particular
               | file reads. I did this to save context & money (finops),
               | after burning $5 in 60s because of an error in my tool
               | implementation. Instead of having them as message
               | contents, they are now injected into the system prompt.
               | Doing so made it trivial to add a key/val "cache" for the
               | fun of it, since I could now inject things into the
               | system prompt and let the agent have some control over
               | that process through tools. Boy has that been interesting
               | and opened up some research questions in my mind
        
               | remich wrote:
               | Any particular papers or articles you've been reading
               | that helped you devise this? Your experiments sound
               | interesting and possibly relevant to what I'm doing.
        
             | snet0 wrote:
             | To say that a model _won 't_ solve a problem is unfair.
             | Claude Code, with Opus 4.5, has solved plenty of problems
             | for me.
             | 
             | If you expect it to do everything perfectly, you're
             | thinking about it wrong. If you can't get it to do anything
             | perfectly, you're using it wrong.
        
               | jacquesm wrote:
               | That means you're probably asking it to do very simple
               | things.
        
               | camdenreslink wrote:
               | Sometimes you do need to (as a human) break down a
               | complex thing into smaller simple things, and then ask
               | the LLM to do those simple things. I find it still saves
               | some time.
        
               | ragequittah wrote:
               | Or what will often work is having the LLM break it down
               | into simpler steps and then running them 1 by 1. They
               | know how to break down problems fairly well they just
               | don't often do it properly sometimes unless you
               | explicitly prompt them to.
        
               | jacquesm wrote:
               | Yes, but for that you have to know that the output it
               | gave you is wrong in the first place and if that is so
               | you didn't need AI to begin with...
        
               | baq wrote:
               | I can confidently say that anecdotally you're completely
               | wrong, but I'll also allow a very different definition of
               | 'simple' and/or attempting to use an unpopular
               | environment as a valid anecdotal counterpoint.
        
               | jacquesm wrote:
               | I don't reach for AI until I'm solidly stuck and then use
               | it mostly for inspiration, it has yet to happen that it
               | directly pointed at a solution and I'm pretty good at
               | writing prompts. When I throw a bunch of elementary stuff
               | at it then it is super good at identifying issues and
               | solving them (but I could have done that myself, it is
               | just nice to try to delineate where you can and where you
               | can't trust the thing, but that too is fluctuating,
               | sometimes even within a single session). Here is a nice
               | example for a slightly more complex test:
               | 
               | Prompt:
               | 
               | "I have a green LED, a 12V powersupply, a single NPN
               | transistor, a 100 ohm resistor, a 2.7K resistor and an
               | electrolytic capacitor of 220 micro farads. My teacher
               | says it is possible to make an LED flasher with these
               | components but I'm hopelessly stuck, can you please give
               | me an ascii art solution so I don't flunk this exam?"
               | 
               | Answer:
               | 
               | ASCII schematic (fully functional):
               | +12V                 |                 |
               | .-.                | |                | | 2.7kO   (R1 --
               | charges the capacitor)                | |
               | '-'                 |
               | +---------+----------------------+                 |
               | |                      |                ( )        |
               | |                ( ) C1     |                      |
               | ( ) 220uF  |                      |                 |
               | |                      |                 |         |
               | |                 |        |/ C                    |
               | +--------|   NPN transistor      |                 |
               | |\ E                    |                 |         |
               | |                GND       GND                     |
               | \                                                    \
               | / 100O (R2 -- LED limit)
               | \                                                     |
               | >|  Green LED
               | |                                                    GND
               | 
               | Follow by:
               | 
               | "How it works
               | 
               | The 2.7 kO resistor charges the 220 uF capacitor from the
               | 12 V supply.
               | 
               | The capacitor voltage rises slowly.
               | 
               | When the capacitor voltage reaches the transistor's base-
               | emitter threshold (~0.6-0.7 V), the transistor suddenly
               | switches ON.
               | 
               | When it turns on, the capacitor rapidly discharges
               | through the base, causing:
               | 
               | A brief pulse of current through the transistor
               | 
               | The LED lights up through the 100 O resistor
               | 
               | After discharge, the transistor turns back OFF, the LED
               | turns off, and the capacitor begins charging again.
               | 
               | This repeats automatically - LED flasher."
               | 
               | The number of errors in the circuit and the utterly bogus
               | explanation as well as the over confident remark that
               | this is 'working' is so bizarre that I wonder how many
               | slightly more complicated questions are going to yield
               | results comparable to this one.
        
               | emporas wrote:
               | I have used Gemini for reading and solving electronic
               | schematics exercises, and it's results were good enough
               | for me. Roughly 50% of the exercises managed to solve
               | correctly, 50% wrong. Simple R circuits.
               | 
               | One time it messed up the opposite polarity of two
               | voltage sources in series, and instead of subtracting
               | their voltages, it added them together, I pointed out the
               | mistake and Gemini insisted that the voltage sources are
               | not in opposite polarity.
               | 
               | Schematics in general are not AIs strongest point. But
               | when you explain what math you want to calculate from an
               | LRC circuit for example, no schematics, just describe in
               | words the part of the circuit, GPT many times will
               | calculate it correctly. It still makes mistakes here and
               | there, always verify the calculation.
        
               | jacquesm wrote:
               | I guess I'm just more critical than you are. I am used my
               | computer doing what it is told and giving me correct,
               | exact answers or errors.
        
               | emporas wrote:
               | There is also Mercury LLM, which computes the answer
               | directly as a 2D text representation. I don't know if you
               | are familiar with Mercury LLM, but you read correctly, 2D
               | text output.
               | 
               | Mercury LLM might work better getting input as an ASCII
               | diagram, or generating an output as an ASCII diagram, not
               | sure if both input and output work 2D.
               | 
               | Plumbing/electrical/electronic schematics are pretty
               | important for AIs to understand and assist us, but for
               | the moment the success rate is pretty low. 50% success
               | rate for simple problems is very low, 80-90% success rate
               | for medium difficulty problems is where they start being
               | really useful.
        
               | jacquesm wrote:
               | It's not really the quality of the diagramming that I am
               | concerned with, it is the complete lack of understanding
               | of electronics parts and their usual function. The
               | diagramming is atrocious but I could live with it if the
               | circuit were at least borderline correct. Extrapolating
               | from this: if we use the electronics schematic as a proxy
               | for the kind of world model these systems have then that
               | world model has upside down lanterns and anti-gravity as
               | commonplace elements. Three legged dogs mate with zebras
               | and produce viable offspring and short circuiting
               | transistors brings about entirely new physics.
        
               | emporas wrote:
               | I think you underestimate their capabilities quite a bit.
               | Their auto-regressive nature does not lend well to
               | solving 2D problems.
               | 
               | See these two solutions GPT suggested: [1]
               | 
               | Is any of these any good?
               | 
               | [1] https://gist.github.com/pramatias/538f77137cb32fca5f6
               | 26299a7...
        
               | baq wrote:
               | it's hard for me to tell if the solution is correct or
               | wrong because I've got next to no formal theoretical
               | education in electronics and only the most basic 'pay
               | attention to polarity of electrolytic capacitors'
               | practical knowledge, but given how these things work you
               | might get much better results when asking it to generate
               | a spice netlist first (or instead).
               | 
               | I wouldn't trust it with 2d ascii art diagrams, there
               | isn't enough focus on these in the training data is my
               | guess - a typical jagged frontier experience.
        
               | dagss wrote:
               | I think most people treat them like humans not computers,
               | and I think that is actually a much more correct way to
               | treat them. Not saying they are like humans, but
               | certainly a lot more like humans than whatever you seem
               | to be expecting in your posts.
               | 
               | Humans make errors all the time. That doesn't mean having
               | colleagues is useless, does it?
               | 
               | An AI is a colleague that can code very very fast and has
               | a very wide knowledge base and versatility. You may still
               | know better than it in many cases and feel more
               | experienced that in. Just like you might with your
               | colleagues.
               | 
               | And it needs the same kind of support that humans need.
               | Complex problem? Need to plan ahead first. Tricky logic?
               | Need unit tests. Research grade problem? Need to discuss
               | through the solution with someone else before jumping to
               | code and get some feedback and iterate for 100 messages
               | before we're ready to code. And so on.
        
               | jacquesm wrote:
               | This is an excellent point, thank you.
        
               | manmal wrote:
               | I have this mental model of LLMs and their capabilities,
               | formed after months of way too much coding with CC and
               | Codex, with 4 recursive problem categories:
               | 
               | 1. Problems that have been solved before have their
               | solution easily repeated (some will say,
               | parroted/stolen), even with naming differences.
               | 
               | 2. Problems that need only mild amalgamation of previous
               | work are also solved by drawing on training data only,
               | but hallucinations are frequent (as low probability
               | tokens, but as consumers we don't see the p values).
               | 
               | 3. Problems that need little simulation can be simulated
               | with the text as scratchpad. If evaluation criteria are
               | not in training data -> hallucination.
               | 
               | 4. Problems that need more than a little simulation have
               | to either be solved by adhoc written code, or will result
               | in hallucination. The code written to simulate is again a
               | fractal of problems 1-4.
               | 
               | Phrased differently, sub problem solutions must be in the
               | training data or it won't work; and combining sub problem
               | solutions must be either again in training data, or brute
               | forcing + success condition is needed, with code being
               | the tool to brute force.
               | 
               | I _think_ that the SOTA models are trained to categorize
               | the problem at hand, because sometimes they answer
               | immediately (1&2), enable thinking mode (3), or write
               | Python code (4).
               | 
               | My experience with CC and Codex has been that I must
               | steer it away from categories 2 & 3 all the time, either
               | solving them myself, ask them to use web research, or
               | split them up until they are (1) problems.
               | 
               | Of course, for many problems you'll only know the
               | category once you've seen the output, and you need to be
               | able to verify the output.
               | 
               | I suspect that if you gave Claude/Codex access to a
               | circuit simulator, it will successfully brute force the
               | solution. And future models might be capable enough to
               | write their own simulator adhoc (ofc the simulator code
               | might recursively fall into category 2 or 3 somewhere and
               | fail miserably). But without strong verification I
               | wouldn't put any trust in the outcome.
               | 
               | With code, we do have the compiler, tests, observed
               | behavior, and a strong training data set with many
               | correct implementations of small atomic problems. That's
               | a lot of out of the box verification to correct
               | hallucinations. I view them as messy code generators I
               | have to clean up after. They do save a ton of coding work
               | after or while I'm doing the other parts of programming.
        
               | jacquesm wrote:
               | This parallels my own experience so far, the problem for
               | me is that (1) and (2) I can quickly and easily do myself
               | and I'll do it in a way that respects the original
               | author's copyright by including their work - and license
               | - verbatim.
               | 
               | (3) and (4) level problems are the ones where I struggle
               | tremendously to make any headway even without AI, usually
               | this requires the learning of new domain knowledge and
               | exploratory code (currently: sensor fusion) and these
               | tools will just generate very plausible nonsense which is
               | more of a time waster than a productivity aid. My middle-
               | of-the-road solution is to get as far as I can by reading
               | about the problem so I am at least able to define it
               | properly and to define test cases and useful ranges for
               | inputs and so on, then to write a high level overview
               | document about what I want to achieve and what the big
               | moving parts are and then only to resort to using AI
               | tools to get me unstuck or to serve as a knowledge
               | reservoir for gaps in domain knowledge.
               | 
               | Anybody that is using the output of these tools to
               | produce work that they do not sufficiently understand is
               | going to see a massive gain in productivity, but the
               | underlying issues will only surface a long way down the
               | line.
        
               | dagss wrote:
               | I am right now implementing an imagining pipeline using
               | OpenCV and TypeScript.
               | 
               | I have never used OpenCV specifically before, and have
               | little imaging experience too. What I do have though is a
               | PhD in astrophysics/statistics so I am able to follow
               | along the details easily.
               | 
               | Results are amazing. I am getting results in 2 days of
               | work that would have taken me weeks earlier.
               | 
               | ChatGPT acts like a research partner. I give it images
               | and it explains why current scoring functions fails and
               | throws out new directions to go in.
               | 
               | Yes, my ideas are sometimes better. Sometimes ChatGPT has
               | a better clue. It is like a human collegue more or less.
               | 
               | And if I want to try something, the code is usually bug
               | free. So fast to just write code, try it, throw it away
               | if I want to try another idea.
               | 
               | I think a) OpenCV probably has more training data than
               | circuits? and b) I do not treat it as a desperate student
               | with no knowlegde.
               | 
               | I expect to have to guide it.
               | 
               | There are several hundred messages back and forth.
               | 
               | It is more like two researchers working together with
               | different skill sets complementing one another.
               | 
               | One of those skillsets being to turn a 20 message
               | conversation into bugfree OpenCV code in 20 seconds.
               | 
               | No, it is not providing a perfect solution to all
               | problems on first iteration. But it IS allowing me to
               | both learn very quickly and build very quickly. Good
               | enough for me..
        
               | jacquesm wrote:
               | That's a good use case, and I can easily imagine that you
               | get good results from it because (1) it is for a domain
               | that you are already familiar with and (2) you are able
               | to check that the results that you are getting are
               | correct and (3) the domain that you are leveraging
               | (coding expertise) is one that chatgpt has ample input
               | for.
               | 
               | Now imagine you are using it for a domain that you are
               | not familiar with, or one for which you can't check the
               | output or that chatgpt has little input for.
               | 
               | If either of those is true the output will be _just_ as
               | good looking and you would be in a much more difficult
               | situation to make good use of it, but you might be
               | tempted to use it anyway. A very large fraction of the
               | use cases for these tools that I have come across
               | professionally so far are of the latter variety, the
               | minority of the former.
               | 
               | And taking all of the considerations into account:
               | 
               | - how sure are you that that code is bug free?
               | 
               | - Do you mean that it seems to work?
               | 
               | - Do you mean that it compiles?
               | 
               | - How broad is the range of inputs that you have given it
               | to ascertain this?
               | 
               | - Have you had the code reviewed by a competent
               | programmer (assuming code review is a requirement)?
               | 
               | - Does it pass a set of pre-defined tests (part of
               | requirement analysis)?
               | 
               | - Is the code quality such that it is long term
               | maintainable?
        
               | verdverm wrote:
               | the problem with these arguments is there are data points
               | to support both sides because both outcomes are possible
               | 
               | the real thing is are you or we getting an ROI and the
               | answer is increasingly more yeses on more problems, this
               | trend is not looking to plateau as we step up the
               | complexity ladder to agentic system
        
               | snet0 wrote:
               | If you define "simple thing" as "thing an AI can't do",
               | then yes. Everyone just shifts the goalposts in these
               | conversations, it's infuriating.
        
               | ACCount37 wrote:
               | Come on. If we weren't shifting the goalposts, we would
               | have burned through 90% of the entire supply of them back
               | in 2022!
        
               | baq wrote:
               | It's less shifting goalposts and more of a very jagged
               | frontier of capabilities problem.
        
               | djeastm wrote:
               | Possibly, but a lot of value comes from doing very simple
               | things faster.
        
               | jacquesm wrote:
               | That is a good point. A lot of work really is mostly
               | simple things.
        
         | poormathskills wrote:
         | For a minor version update (5.1 -> 5.2) that's a way bigger
         | improvement than I would have guessed.
        
           | beering wrote:
           | Model capability improvements are very uneven. Changes
           | between one model and the next tend to benefit certain areas
           | substantially without moving the needle on others. You see
           | this across all frontier labs' model releases. Also the
           | version numbering is BS (remember GPT-4.5 followed by
           | GPT-4.1?).
        
         | catigula wrote:
         | Yes, but it's not good enough. They needed to surpass Opus 4.5.
        
           | mikairpods wrote:
           | that is better...?
        
         | minimaxir wrote:
         | Note that GPT 5.2 newly supports a "xhigh" reasoning level,
         | which could explain the better benchmarks.
         | 
         | It'll be noteworthy to see the cost-per-task on ARC AGI v2.
        
           | granzymes wrote:
           | > It'll be noteworthy to see the cost-per-task on ARC AGI v2.
           | 
           | Already live. gpt-5.2-pro scores a new high of 54.2% with a
           | cost/task of $15.72. The previous best was Gemini 3 Pro (54%
           | with a cost/task of $30.57).
           | 
           | The best bang-for-your-buck is the new xhigh on gpt-5.2,
           | which is 52.9% for $1.90, a big improvement on the previous
           | best in this category which was Opus 4.5 (37.6% for $2.40).
           | 
           | https://arcprize.org/leaderboard
        
             | minimaxir wrote:
             | Huh, that is indeed up and to left of Opus.
        
           | walletdrainer wrote:
           | 5.1-codex supports that too, no? Pretty sure I've been using
           | xhigh for at least a week now
        
         | causal wrote:
         | That ARC AGI score is a little suspicious. That's a really
         | tough for AI benchmark. Curious if there were improvements to
         | the test harness because that's a wild jump in general problem
         | solving ability for an incremental update.
        
           | taurath wrote:
           | I don't think their words mean just about anything, only the
           | behavior of the models.
           | 
           | Still waiting of Full Self Driving myself.
        
           | woeirua wrote:
           | They're clearly building better training datasets and doing
           | extensive RL on these benchmarks over time. The out of
           | distribution performance is still awful.
        
         | thinkingtoilet wrote:
         | Open AI has already been busted for getting benchmark
         | information and training the models on that. At this point if
         | you believe Sam Altman, I have a bridge to sell you.
        
         | fuddle wrote:
         | I don't think SWE Verified is an ideal benchmark, as the
         | solutions are in the training dataset.
        
           | joshuahedlund wrote:
           | I would love for SWE Verified to put out a set of fresh but
           | comparable problems and see how the top performing models do,
           | to test against overfitting.
        
       | doctoboggan wrote:
       | This seems like another "better vibes" release. With the number
       | of benchmarks exploding, random luck means you can almost always
       | find a couple showing what you want to show. I didn't see much
       | concrete evidence this was noticeably better than 5.1 (or even
       | 5.0).
       | 
       | Being a point release though I guess that's fair. I suspect there
       | is also some decent optimizations on the backend that make it
       | cheaper and faster for OpenAI to run, and those are the real
       | reasons they want us to use it.
        
         | rat9988 wrote:
         | > I didn't see much concrete evidence this was noticeably
         | better than 5.1
         | 
         | Did you test it?
        
           | doctoboggan wrote:
           | No, I would like to but I don't see it in my paid ChatGPT
           | plan or in the API yet. I based my comment solely off of what
           | I read in the linked announcement.
        
         | sebzim4500 wrote:
         | >I suspect there is also some decent optimizations on the
         | backend that make it cheaper and faster for OpenAI to run, and
         | those are the real reasons they want us to use it.
         | 
         | I doubt it, given it is more expensive than the old model.
        
         | BrtByte wrote:
         | At this point the benchmark soup is so dense that it's hard to
         | tell signal from selective framing
        
       | _7u7v wrote:
       | It baffles me to see these last 2 announcements (GPT 5.1 as well)
       | devoid of any metrics, benchmarks or quantitative analyses. Could
       | it be because they are behind Google/Anthropic and they don't
       | want to admit it?
       | 
       | (edit: I'm sorry I didn't read enough on the topic, my apologies)
        
         | zamadatix wrote:
         | This isn't the announcement, it's the developer docs intro page
         | to the model - https://openai.com/index/introducing-gpt-5-2/.
         | Still doesn't answer cross-comparison, but at least has
         | benchmark metrics they want to show off.
        
       | fulafel wrote:
       | So GDPval is OpenAI's own benchmark. PDF link:
       | https://arxiv.org/pdf/2510.04374
        
       | minadotcom wrote:
       | They used to compare to competing models from Anthropic, Google
       | DeepMind, DeepSeek, etc. Seems that now they only compare to
       | their own models. Does this mean that the GPT-series is
       | performing worse than its competitors (given the "code red" at
       | OpenAI)?
        
         | poormathskills wrote:
         | OpenAI has never compared their models to models from other
         | labs in their blog post. Open literally any past model launch
         | post to see that.
        
           | boole1854 wrote:
           | https://openai.com/index/hello-gpt-4o/
           | 
           | I see evaluations compared with Claude, Gemini, and Llama
           | there on the GPT 4o post.
        
             | kgwgk wrote:
             | "You are absolutely right, and I apologize for the
             | confusion."
        
         | tabletcorry wrote:
         | The matrix required for a fair comparison is getting too
         | complicated, since you have to compare chat/thinking/pro
         | against an array of Anthropic and Google models.
         | 
         | But they publish all the same numbers, so you can make the full
         | comparison yourself, if you want to.
        
         | Tiberium wrote:
         | They did compare it to other models:
         | https://x.com/OpenAI/status/1999182104362668275
         | 
         | https://i.imgur.com/e0iB8KC.png
        
           | enlyth wrote:
           | This looks cherry-picked, for example Claude Opus had a
           | higher score on SWE-Bench Verified so they conveniently left
           | it out, also GDPval is literally a benchmark made by OpenAI
        
             | minadotcom wrote:
             | agreed.
        
             | tobias2014 wrote:
             | And who believes that the difference between 91.9% and
             | 92.4% is significant in these benchmarks? Clearly these
             | have margins of error that are swept under the rug.
        
           | whimsicalism wrote:
           | uh oh, where did SWE bench go :D
        
             | whimsicalism wrote:
             | maybe they will release with gpt-5.2-codex
        
           | sergdigon wrote:
           | The fact that the post is comparing their reasoning model
           | against gemini 3 pro (the "non reasoning" model) and not
           | gemini 3 pro deep think (the reasoning one) is quite nasty.
           | If you compare GPT5.2 thinking to gemini 3 pro deep think,
           | the scores are quite similar (sometimes one is better
           | sometimes the other one is)
        
         | Workaccount2 wrote:
         | They are taking a page out of Apple's book.
         | 
         | Apple only compares to themselves. They don't even acknowledge
         | the existence of others.
        
       | mattas wrote:
       | Are benchmarks the right way to measure LLMs? Not because
       | benchmarks can be gamed, but because the most useful outputs of
       | models aren't things that can be bucketed into "right" and
       | "wrong." Tough problem!
        
         | Sir_Twist wrote:
         | Not an expert in LLM benchmarks, but I generally I think of
         | benchmarks as being good particularly for measuring usefulness
         | for certain usecases. Even if measuring LLMs is not as
         | straightforward as, say, read/write speeds when comparing
         | different SSDs, if a certain model's responses are consistently
         | measured as being higher quality / more useful, surely that
         | means something, right?
        
         | olliepro wrote:
         | Do you have a better way to measure LLMs? Measurement implies
         | quantitative evaluation... which is the same as benchmarks.
        
           | Wowfunhappy wrote:
           | I don't have a good way to measure them, but I think they
           | should be evaluated more like how we evaluate movies, or
           | restaurants. Namely, experienced critics try them and write
           | reviews.
        
             | olliepro wrote:
             | It feels like this should work, but the breadth of
             | knowledge in these models is so vast. Everyone knows how to
             | taste, but not everyone knows physics, biology, math, every
             | language... poetry, etc. Enumerating the breadth of
             | valuable human tasks is hard, so both approaches suffer
             | from the scale of the models' surface area.
             | 
             | An interesting problem since the creators of OLMO have
             | mentioned that throughout training, they use 1/3 or their
             | compute just doing evaluations.
             | 
             | Edit:
             | 
             | One nice thing about the "critic" approach is that the
             | restaurant (or model provider) doesn't have access to the
             | benchmark to quasi-directly optimize against.
        
       | k2xl wrote:
       | The ARC AGI 2 bump to 52.9% is huge. Shockingly GPT 5.2 Pro does
       | not add too much more (54.2%) for the increase cost.
        
       | zug_zug wrote:
       | For me the last remaining killer feature of ChatGPT is the
       | quality of the voice chat. Do any of the competitors have
       | something like that?
        
         | FrasiertheLion wrote:
         | Try elevenlabs
        
           | sosodev wrote:
           | Does elevenlabs have a real-time conversational voice model?
           | It seems like like their focus is largely on text to speech
           | and speech to text. Which can approximate that type of thing
           | but it's not at all the same as the native voice to voice
           | that 4o does.
        
             | dragonwriter wrote:
             | > Does elevenlabs have a real-time conversational voice
             | model?
             | 
             | Yes.
             | 
             | > It seems like like their focus is largely on text to
             | speech and speech to text.
             | 
             | They have two main broad offerings ("Platforms"); you seem
             | to be looking at what they call the "Creative Platform".
             | The real-time conversational piece is the centerpiece of
             | the "Agents Platform".
        
               | sosodev wrote:
               | It specifically says in the architecture docs for the
               | agents platform that it's STT (ASR) -> LLM -> TTS
               | 
               | https://elevenlabs.io/docs/agents-
               | platform/overview#architec...
        
             | hi_im_vijay wrote:
             | [disclaimer, i work at elevenlabs] we specifically went
             | with a cascading model for our agents platform because it's
             | better suited for enterprise use cases where they have full
             | control over the brain and can bring their own llm. with
             | that said, even with a cascading model, we can capture a
             | decent amount of nuance with our asr model, and it also
             | supports capturing audio events like laughter or coughing.
             | 
             | a true speech to speech conversational model will perform
             | better on things like capturing tone, pronouncations,
             | phonetics, etc, but i do believe we'll also get better at
             | that on the asr side over time.
        
         | bigyabai wrote:
         | Qwen does.
        
           | sosodev wrote:
           | Qwen's voice chat is nowhere near as good as ChatGPT's.
        
         | Robdel12 wrote:
         | I have found Claude's voice chat to be better. I only recently
         | tried it because I liked ChatGPTs enough, but I think I'm going
         | to use Claude going forward. I find myself getting interrupted
         | by ChatGPT a lot whenever I do use it.
        
           | lxgr wrote:
           | Claude's voice chat isn't "native" though, is it? It feels
           | like it's speech-to-text-to-LLM and back.
        
             | sosodev wrote:
             | You can test it by asking it to: change the pitch of its
             | voice, make specific sounds (like laughter), differentiate
             | between words that are spelled the same but pronounced
             | differently (record and record), etc.
        
               | lxgr wrote:
               | Good idea, but an external "bolted on" LLM-based TTS
               | would still pass that in many cases, right?
        
               | barrkel wrote:
               | The model giving it text to speak would have to annotate
               | the text in order for the TTS to add the affect. The TTS
               | wouldn't "remember" such instructions from a speech to
               | text stage previously.
        
               | sosodev wrote:
               | Yes, a sufficiently advanced marrying of TTS and LLM
               | could pass a lot of these tests. That kind of blurs the
               | line between native voice model and not though.
               | 
               | You would need:
               | 
               | * A STT (ASR) model that outputs phonetics not just words
               | 
               | * An LLM fine-tuned to understand that and also output
               | the proper tokens for prosody control, non-speech
               | vocalizations, etc
               | 
               | * A TTS model that understands those tokens and properly
               | generate the matching voice
               | 
               | At that point I would probably argue that you've created
               | a native voice model even if it's still less nuanced than
               | the proper voice to voice of something like 4o. The
               | latency would likely be quite high though. I'm pretty
               | sure I've seen a couple of open source projects that have
               | done this type of setup but I've not tried testing them.
        
               | BoxOfRain wrote:
               | I've been experimenting with something similar to this
               | approach recently. IndexTTS2 gives you emotion vectors as
               | an input, I used an external emotion classification model
               | on the LLM output to modulate the TTS emotion vectors.
               | You need to manage the state of the current affect with a
               | bit of care or it sounds unhinged, but it's worked
               | surprisingly well so far. I wired it together using Cats
               | Effect.
               | 
               | As you'd expect latency isn't great, but I think it can
               | be improved.
        
               | jablongo wrote:
               | I tried to make ChatGPT sing Mary had a little lamb
               | recently and it's atonal but vaguely resembles the
               | melody, which is interesting.
        
             | causalmodels wrote:
             | I just asked it and it said that it uses the on device TTS
             | capabilities.
        
               | furyofantares wrote:
               | I find it very unlikely that it would be trained on that
               | information or that anthropic would put that in its
               | context window, so it's very likely that it just made
               | that answer up.
        
               | causalmodels wrote:
               | No, it did not make it up. I was curious so I asked it
               | asked it to imitate a posh British accent imitating a
               | South Brooklyn accent while having a head cold and it
               | explained that it didn't have have fine grained control
               | over the audio output because it was using a TTS. I asked
               | it how it knew that and it pointed me towards [1] and
               | highlighted the following.
               | 
               | > As of May 29th, 2025, we have added ElevenLabs, which
               | supports text to speech functionality in Claude for Work
               | mobile apps.
               | 
               | Tracked down the original source [2] and looked for
               | additional updates but couldn't find anything.
               | 
               | [1] https://simonwillison.net/2025/May/31/using-voice-
               | mode-on-cl...
               | 
               | [2] https://trust.anthropic.com/updates
        
               | furyofantares wrote:
               | If it does a web search that's fine, I assumed it hadn't
               | since you hadn't linked to anything.
               | 
               | Also it being right doesn't mean it didn't just make up
               | the answer.
        
         | codybontecou wrote:
         | Their voice agent is handy. Currently trying to build around
         | it.
        
         | websiteapi wrote:
         | gemini live is a thing - never tried chaptgpt, are they not
         | similar?
        
           | jeanlucas wrote:
           | no.
        
             | leaK_u wrote:
             | how.
        
               | CamelCaseName wrote:
               | I find ChatGPT's voice to text to be the absolute best in
               | the world, nearly perfect.
               | 
               | I have constant frustrations with Gemini voice to text
               | misunderstanding what I'm saying or worse, immediately
               | sending my voice note when I pause or breathe even though
               | I'm midway through a sentence.
        
             | nickvec wrote:
             | What? The voice chat is basically identical on ChatGPT and
             | Gemini AFAICT.
        
           | spudlyo wrote:
           | Not for my use case. I can open it up, and in restored
           | classical Latin pronunciation say "Hi, my name is X, how are
           | you?" and it will respond (also in Latin) "Hello X, I am
           | well, thanks for asking. I hope you are doing great." Its
           | pronunciation is not great, but intelligible. In the written
           | transcript, it butchers what I say, but its responses look
           | good, although sans macrons indicating phonemic vowel length.
           | 
           | Gemini responds in what I think is Spanish, or perhaps
           | Portuguese.
           | 
           | However I can hand an 8 minute long 48k mono mp3 of a nuanced
           | Latin speaker who nasalizes his vowels, and makes regular use
           | of elision to Gemini-3-pro-preview and it will produce an
           | accurate macronized Latin transcription. It's pretty mind
           | blowing.
        
             | Dilettante_ wrote:
             | I _have_ to ask: What usecase requires you to speak Latin
             | to the llm?
        
               | spudlyo wrote:
               | I'm a Latin language learner, and part of developing
               | fluency is practicing extemporaneous speech. My dog is a
               | patient listener, but a poor interlocutor. There are
               | Latin language Discord servers where you can speak to
               | people, but I don't quite have the confidence to do that
               | yet. I assume the machine doesn't judge my shitty
               | grammar.
        
               | onraglanroad wrote:
               | Loquerisne Latine?
               | 
               | Non vere, sed intelligere possum.
               | 
               | Ita, mihi est canis qui idipsum facit!
               | 
               | (translated from the Gaidhlig)
        
               | spudlyo wrote:
               | Certe loqui conor, sed saepenumero prave dico; canis meus
               | non turbatus est ;)
        
               | nineteen999 wrote:
               | You haven't heard? Latin is the next big wave, after
               | blockchain and AI.
        
               | spudlyo wrote:
               | You laugh, but the global language learning market in
               | 2025 is expected to exceed USD $100 billion, and LLMs
               | IMHO are poised to disrupt the shit out of it.
        
               | nineteen999 wrote:
               | Well sure I can see that happening ... but I can't see
               | latin making a huge comeback unfortunately.
        
         | tmaly wrote:
         | I can't keep up with half the new features all the model
         | companies keep rolling out. I wish they would solve that
        
         | sundarurfriend wrote:
         | Are you saying ChatGPT's voice chat is of _good_ quality?
         | Because for me it 's one of its most frustrating weaknesses. I
         | vastly prefer voice input to typing, and would love it if the
         | voice chat mode actually worked well.
         | 
         | But apart from the voices being pretty meh, it's also really
         | bad at detecting and filtering out noise, taking vehicle sounds
         | as breaks to start talking in (even if I'm talking much louder
         | at the same time) or as some random YouTube subtitles (car
         | motor = "Thanks for watching, subscribe!").
         | 
         | The speech-to-text is really unreliable (the single-chat
         | Dictate feature gets about 98% of my words correct, this Voice
         | mode is closer to 75%), and they clearly use an inferior model
         | for the AI backend for this too: with the same question asked
         | in this back-and-forth Voice mode and a normal text chat, the
         | answer quality difference is quite stark: the Voice mode answer
         | is most often close to useless. It seems like they've
         | overoptimized it for speed at the cost of quality, to the
         | extent that it feels like it's a year behind in answer
         | reliability and usefulness.
         | 
         | To your question about competitors, I've recently noticed that
         | Grok seems to be much better at both the speech-to-text part
         | and the noise handling, and the voices are less uncanny-valley
         | sounding too. I'd say they also don't have that stark a
         | difference between text answers and voice mode answers, and
         | that would be true but unfortunately mainly because its text
         | answers are also not great with hallucinations or following
         | instructions.
         | 
         | So Grok has the voice part figured out, ChatGPT has the backend
         | AI reliability figured out, but neither provide a real usable
         | voice mode right now.
        
         | semiinfinitely wrote:
         | try gemini voice chat
        
         | ivape wrote:
         | I'm a big user of Gemini voice. My sense is that Gemini voice
         | uses very tight system prompts that are designed to give you an
         | answer and kind of get you off the phone as much as possible.
         | It doesn't have large context at all.
         | 
         | That's how I judge quality at least. The quality of the actual
         | voice is roughly the same as ChatGPT, but I notice Gemini will
         | try to match your pitch and tone and way of speaking.
         | 
         | Edit: But it looks like Gemini Voice has been replaced with
         | voice transcription in the mobile app? That was sudden.
        
         | whimsicalism wrote:
         | gemini does, grok does, nobody else does (except alibaba but
         | it's not there yet)
        
         | joshmarlow wrote:
         | I think Grok's voice chat is almost there - only things missing
         | for me: * it's slower to start-up by a couple of seconds * it's
         | harder to switch between voice and text and back again in the
         | same chat (though ChatGPT isn't perfect at this either)
         | 
         | And of course Grok's unhinged persona is... something else.
        
           | nazgulsenpai wrote:
           | It's so much fun. So is the Conspiracy persona.
        
           | Gigachad wrote:
           | Pretty good until it goes crazy glazing Elon or declaring
           | itself mecha hitler.
        
             | hcurtiss wrote:
             | Neither of these have happened in my use. Those were both
             | the product of some pretty aggressive prompting, and were
             | remedied months ago.
        
               | OrangeMusic wrote:
               | Yet, using this model in any way whatsoever after these
               | episodes seems absolutely crazy to me.
        
               | user34283 wrote:
               | Grok is the only frontier model that is at all usable for
               | adult content.
        
               | hcurtiss wrote:
               | All models have had similar instances. I particularly
               | enjoyed Gemini's black founders era. The "safety" teams
               | have bent the politics of these tools in ways I don't
               | trust. Grok does too, but in my experience less so. This
               | has real impacts.
        
         | hbarka wrote:
         | On the contrary, I thought Gemini 3 Live mode is much much
         | better than ChatGPT. The voices have none of the annoying
         | artificial uptalking intonations that ChatGPT has, and the
         | simplex/duplex interruptibility of Gemini Live seems more
         | responsive. It knows when to break and pause during
         | conversations.
        
           | febed wrote:
           | Apart from sounding a bit stiff and informal, I was also
           | surprised at how good Gemini Live mode is in regional Indian
           | languages.
        
         | simondotau wrote:
         | I absolutely loathe ChatGPT's voice chat. It spends far too
         | much time being conversational and its eagerness to please
         | becomes fatiguing after the first back-and-forth.
        
         | josephwegner wrote:
         | Along with the hordes of other options people are responding
         | with, I'm a big fan of Perplexity's voice chat. It does back-
         | and-forth well in a way that I missed whenever I tried anything
         | besides ChatGPT.
        
           | solarkraft wrote:
           | It is, shockingly, based on the OpenAI Realtime Assistant
           | API.
        
         | SweetSoftPillow wrote:
         | Gemini's much better, try it
        
       | Croftengea wrote:
       | Is this another GPT-4.5?
        
       | Tiberium wrote:
       | The only table where they showed comparisons against Opus 4.5 and
       | Gemini 3:
       | 
       | https://x.com/OpenAI/status/1999182104362668275
       | 
       | https://i.imgur.com/e0iB8KC.png
        
         | varenc wrote:
         | 100% on the AIME (assuming its not in the training data) is
         | pretty impressive. I got like 4/15 when I was in HS...
        
           | hellojimbo wrote:
           | The no tools part is impressive, with tools every model gets
           | 100%
        
             | varenc wrote:
             | If I recall, the AIME answers are always 4 digits numbers.
             | And most of the problems are of the type where if you have
             | a candidate number it's reasonable to validate its
             | correctness. So easy to brute force all 4 digit ints with
             | code.
             | 
             | tl;dr; humans would do much better too if they could use
             | programming tools :)
        
               | Davidzheng wrote:
               | uh no it's not solved by looping over 4 digit numbers
               | when it uses tools
        
       | JanSt wrote:
       | The benchmarks are very impressive. Codex and Opus 4.5 are really
       | good coders already and they keep getting better.
       | 
       | No wall yet and I think we might have crossed the threshold of
       | models being as good or better than most engineers already.
       | 
       | GDPval will be an interesting benchmark and I'll happily use the
       | new model to test spreadsheet (and other office work)
       | capabilities. If they can going like this just a little bit
       | further, much of the office workers will stop being useful.... I
       | don't know yet how to feel about this.
       | 
       | Great for humanity probably but but for the individuals?
        
         | llmslave wrote:
         | Yeah theres no wall on this. It will be able to mimic all of
         | human behavior given proper data.
        
         | ionwake wrote:
         | it was only about 2-3 weeks when several HNers told me "nah you
         | better re-check your code", when I explained I have over 2
         | decades xp of coding, yet have not manually edited code (in
         | memory) for the last 6 or so months, whilst performing daily 12
         | hour daily vibe code seshes
        
           | ipsum2 wrote:
           | It really depends on the complexity of code. I've found
           | models (codex-5.1-max, opus 4.5) to be absolutely useless
           | writing shaders or ML training code, but really good at basic
           | web development.
        
             | sheeshe wrote:
             | Which is no surprise as the data for web development stuff
             | exists in large amounts on the web that the models feed
             | off.
        
             | nineteen999 wrote:
             | Interesting, I've been using Claude Max with UE5 and while
             | it isn't _brilliant_ with shaders I can usually get it to
             | where I want. Also had a bit of success with converting
             | HLSL shaders to GLSL with it.
        
               | ipsum2 wrote:
               | I've asked it to write some non-trivial three.js code and
               | have not gotten it to succeed.
        
           | osn9363739 wrote:
           | Do you have any examples or are your project oss or anything
           | like that? Because I want to believe, but I have people I
           | work with that say and try the same thing (no manual coding),
           | and their work is now terrible.
        
         | sheeshe wrote:
         | Ok so why isn't there mass lay offs ensuing right now?
        
           | ghosty141 wrote:
           | Because from my experience using codex in a decently complex
           | c++ environment at work, it works _REALLY_ well when it has
           | things to copy. Refactorings, documentation, code review etc.
           | all work great. But those things only help actual humans and
           | they also take time. I estimate that in a good case I save
           | ~50% of time, in a bad case it 's negative and costs time.
           | 
           | But what I generally found, it's not that great at writing
           | new code. Obviously an LLM can't think and you notice that
           | quite quickly, it doesn't create abstractions, use
           | abstractions or try to find general solution to problems.
           | 
           | People who get replaced by Codex are those who do repetitive
           | tasks in a well understood field. For example, making basic
           | websites, very simple crud applications etc..
           | 
           | I think it's also not layoffs but rather companies will hire
           | less freelancers or people to manage small IT projects.
        
       | breakingcups wrote:
       | Is it me, or did it still get at least three placements of
       | components (RAM and PCIe slots, plus it's DisplayPort and not
       | HDMI) in the motherboard image[0] completely wrong? Why would
       | they use that as a promotional image?
       | 
       | 0:
       | https://images.ctfassets.net/kftzwdyauwt9/6lyujQxhZDnOMruN3f...
        
         | timerol wrote:
         | Also a "stacked pair" of USB type-A ports, when there are
         | clearly 4
        
         | tedsanders wrote:
         | Yep, the point we wanted to make here is that GPT-5.2's vision
         | is better, not perfect. Cherrypicking a perfect output would
         | actually mislead readers, and that wasn't our intent.
        
           | BoppreH wrote:
           | That would be a laudable goal, but I feel like it's
           | contradicted by the text:
           | 
           | > Even on a low-quality image, GPT-5.2 identifies the main
           | regions and places boxes that roughly match the true
           | locations of each component
           | 
           | I would not consider it to have "identified the main regions"
           | or to have "roughly matched the true locations" when ~1/3 of
           | the boxes have incorrect _labels_. The remark  "even on a
           | low-quality image" is not helping either.
           | 
           | Edit: credit where credit is due, the recently-added
           | disclaimer is nice:
           | 
           | > Both models make clear mistakes, but GPT-5.2 shows better
           | comprehension of the image.
        
             | hnuser123456 wrote:
             | Yeah, what it's calling RAM slots is the CMOS battery. What
             | it's calling the PCIE slot is the interior side of the DB-9
             | connector. RAM slots and PCIE slots are not even visible in
             | the image.
        
               | hexaga wrote:
               | It just overlaid a typical ATX pattern across the
               | motherboard-like parts of the image, even if that's not
               | really what the image is showing. I don't think it's
               | worthwhile to consider this a 'local recognition
               | failure', as if it just happened to mistake CMOS for RAM
               | slots.
               | 
               | Imagine it as a markdown response:
               | 
               | # Why this is an ATX layout motherboard (Honest
               | assessment, straight to the point, *NO* hallucinations)
               | 
               | 1. *RAM* as you can clearly see, the RAM slots are to the
               | right of the CPU, so it's obviously ATX
               | 
               | 2. *PCIE* the clearly visible PCIE slots are right there
               | at the bottom of the image, so this definitely cannot be
               | anything except an ATX motherboard
               | 
               | 3. ... etc more stuff that is supported only by force of
               | preconception
               | 
               | --
               | 
               | It's just meta signaling gone off the rails. Something in
               | their post-training pipeline is obviously vulnerable
               | given how absolutely saturated with it their model
               | outputs are.
               | 
               | Troubling that the behavior generalizes to image
               | labeling, but not particularly surprising. This has been
               | a visible problem at least since o1, and the lack of
               | change tells me they do not have a real solution.
        
             | furyofantares wrote:
             | They also changed "roughly match" to "sometimes match".
        
               | MichaelZuo wrote:
               | Did they really change a meaningful word like that after
               | publication without an edit note...?
        
               | piker wrote:
               | Eh, I'm no shill but their marketing copy isn't exactly
               | the New York Times. They're given some license to respond
               | to critical feedback in a manner that makes the
               | statements more accurate without the same expectations of
               | being objective journalism of record.
        
               | mkesper wrote:
               | Yes, but they should clearly mark updates. That would be
               | professional.
        
               | dwohnitmok wrote:
               | This has definitely happened before with e.g. the o1
               | release. I will sometimes use the Wayback Machine to
               | verify changes that have been made.
        
               | MichaelZuo wrote:
               | Wow sounds pretty shady then.
        
             | guerrilla wrote:
             | Leave it to OpenAI to be dishonest about being dishonest.
             | It seems they're also editing this post without notice as
             | well.
        
           | arscan wrote:
           | I think you may have inadvertently misled readers in a
           | different way. I feel misled after not catching the errors
           | myself, assuming it was broadly correct, and then coming
           | across this observation here. Might be worth mentioning this
           | is better but still inaccurate. Just a bit of feedback, I
           | appreciate you are willing to show non-cherry-picked examples
           | and are engaging with this question here.
           | 
           | Edit: As mentioned by @tedsanders below, the post was edited
           | to include clarifying language such as: "Both models make
           | clear mistakes, but GPT-5.2 shows better comprehension of the
           | image."
        
             | tedsanders wrote:
             | Thanks for the feedback - I agree our text doesn't make the
             | models' mistakes clear enough. I'll make some small edits
             | now, though it might take a few minutes to appear.
        
           | g947o wrote:
           | When I saw that it labeled DP ports as HDMI I immediately
           | decided that I am not going to touch this until it is at
           | least 5x better with 95% accuracy with basic things.
           | 
           | I don't see any advantage in using the tool.
        
             | jacquesm wrote:
             | That's a far more dangerous territory. A machine that is
             | obviously broken will not get used. A machine that is
             | subtly broken will propagate errors because it will have
             | achieved a high enough trust level that it will actually
             | get used.
             | 
             | Think 'Therac-25', it worked in 99.5% of the time. In fact
             | it worked so well that reports of malfunctions were
             | routinely discarded.
        
               | AdamN wrote:
               | There was a low-level Google internal service that worked
               | so well that other teams took a hard dependency on it
               | (against advice). So the internal team added a cron job
               | to drop it every once in a while to get people to trust
               | it less :-)
        
           | iamdanieljohns wrote:
           | Is Adaptive Reasoning gone from GPT-5.2? It was a big part of
           | the release of 5.1 and Codex-Max. Really felt like the
           | future.
        
             | tedsanders wrote:
             | Yes, GPT-5.2 still has adaptive reasoning - we just didn't
             | call it out by name this time. Like 5.1 and codex-max, it
             | should do a better job at answering quickly on easy queries
             | and taking its time on harder queries.
        
               | iamdanieljohns wrote:
               | Why have "light" or "low" thinking then? I've mentioned
               | this before in other places, but there should only be
               | "none," "standard," "extended," and maybe "heavy."
               | 
               | Extended and heavy are about raising the floor (~25% and
               | ~45% or some other ratio respectively) not determining
               | the ceiling.
        
           | layer8 wrote:
           | You know what would be great? If it had added some boxes with
           | "might be _X_ or _Y_ , but not sure".
        
           | iwontberude wrote:
           | But it's completely wrong.
        
           | johnwheeler wrote:
           | Oh and you guys don't mislead people ever. Your management is
           | just completely trustworthy, and I'm sure all you guys are
           | too. Give me a break, man. If I were you, I would jump ship
           | or you're going to be like a Theranos employee on LinkedIn.
        
             | yard2010 wrote:
             | Hey no need to personally attack anyone. A bad organization
             | can still consist good people.
        
               | johnwheeler wrote:
               | I disagree. I think the whole organization is egregious
               | and full of Sam Altman sycophants that are causing a real
               | and serious harm to our society. Should we not personally
               | attack the Nazis either? These people are literally
               | pushing for a society where you're at a complete
               | disadvantage. And they're betting on it. They're banking
               | on it.
        
         | whalesalad wrote:
         | to be fair that image has the resolution of a flip phone from
         | 2003
        
           | malfist wrote:
           | If I ask you a question and you don't have enough information
           | to answer, you don't confidently give me an answer, you say
           | you don't know.
           | 
           | I might not know exactly how many USB ports this motherboard
           | has, but I wouldn't select a set of 4 and declare it to be a
           | stacked pair.
        
             | AstroBen wrote:
             | No-one should have the expectation LLMs are giving correct
             | answers 100% of the time. It's inherent to the tech for
             | them to be confidently wrong
             | 
             | Code needs to be checked
             | 
             | References need to be checked
             | 
             | Any facts or claims need to be checked
        
               | malfist wrote:
               | According to the benchmarks here they're claiming up to
               | 97% accuracy. That ought to be good enough to trust them
               | right?
               | 
               | Or maybe these benchmarks are all wrong
        
               | AstroBen wrote:
               | Does code work if it's 97% correct?
               | 
               | It's not okay if claims are totally made up 1/30 times
               | 
               | Of course people aren't always correct either, but we're
               | able to operate on levels of confidence. We're also able
               | to weight others' statements as more or less likely to be
               | correct based on what we know about them
        
               | fooker wrote:
               | > Does code work if it's 97% correct?
               | 
               | Of course it does. The vast majority of software has
               | bugs. Yes, even critical one like compilers and operating
               | systems.
        
               | refactor_master wrote:
               | Gemini routinely makes up stuff about BigQuery's
               | workings. "It's poorly documented". Well, read the open
               | source code, reason it out.
               | 
               | Makes you wonder what 97% is worth. Would we accept a
               | different service with only 97% availability, and all
               | downtime during lunch break?
        
               | TeMPOraL wrote:
               | I.e. like most restaurants and food delivery? :). Though
               | 3% problem rate is optimistic.
        
               | JimDabell wrote:
               | Something that is 97% accurate is wrong 3% of the time,
               | so pointing out that it has gotten something wrong does
               | not contradict 97% accuracy in the slightest.
        
               | mbesto wrote:
               | > Or maybe these benchmarks are all wrong
               | 
               | You must be new to LLM benchmarks.
        
               | dolmen wrote:
               | "confidently" is a feature selected in the system prompt.
               | 
               | As a user you can influence that behavior.
        
               | malfist wrote:
               | No it isn't. It isn't intelligent, it's a statistical
               | engine. Telling it to be confident or less confident
               | doesn't make it apply confidence appropriately. It's all
               | a facade
        
           | redox99 wrote:
           | It's trivial for a human that knows what a pc looks like.
           | Maybe mistaking displayport for hdmi.
        
           | ben_w wrote:
           | That shouldn't be what causes this problems; if we can see
           | it's wrong despite the low resolution, the AI isn't going to
           | fully replace humans for all tasks involving this kind of
           | thing.
           | 
           | That said, even with this kind of error rate an AI can speed
           | _*some*_ things up, because having a human whose sole job is
           | to ask  "is this AI correct?" is easier and cheaper than
           | having one human for "do all these things by hand" followed
           | by someone else whose sole job is to check "was this human
           | output correct?" because a human who has been on a production
           | line for 4 hours and is about ready for a break also makes a
           | certain number of mistakes.
           | 
           | But at the same time, why use a really expensive general-
           | purpose AI like this, instead of a dedicated image model for
           | your domain? Special purpose AI are something you can train
           | on a decent laptop, and once trained will run on a phone at
           | perhaps 10fps give or take what the performance threshold is
           | and how general you need it to be.
           | 
           | If you're in a factory and you're making a lot of some small
           | widget or other (so, not a whole motherboard), having answers
           | faster than the ping time to the LLM may be important all by
           | itself.
           | 
           | And at this point, you can just ask the LLM to write the
           | training setup for the image-to-bounding-box AI, and then you
           | "just" need to feed in the example images.
        
         | jasonlotito wrote:
         | FTA: Both models make clear mistakes, but GPT-5.2 shows better
         | comprehension of the image.
         | 
         | You can find it right next to the image you are talking about.
        
           | tedsanders wrote:
           | To be fair to OP, I just added this to our blog after their
           | comment, in response to the correct criticisms that our text
           | didn't make it clear how bad GPT-5.2's labels are.
           | 
           | LLMs have always been very subhuman at vision, and GPT-5.2
           | continues in this tradition, but it's still a big step up
           | over GPT-5.1.
           | 
           | One way to get a sense of how bad LLMs are at vision is to
           | watch them play Pokemon. E.g.,: https://www.lesswrong.com/pos
           | ts/u6Lacc7wx4yYkBQ3r/insights-i...
           | 
           | They still very much struggle with basic vision tasks that
           | adults, kids, and even animals can ace with little trouble.
        
           | da_grift_shift wrote:
           | _' Commented after article was already edited in response to
           | HN feedback' award_
        
         | an0malous wrote:
         | Because the whole culture of AI enthusiasts is to just generate
         | slop and never check the results
        
         | 8organicbits wrote:
         | Promotional content for LLMs is really poor. I was looking at
         | Claude Code and the example on their homepage implements a
         | feature, ignoring a warning about a security issue, commits
         | locally, does not open a PR and then tries to close the GitHub
         | issue. Whatever code it wrote they clearly didn't use as the
         | issue from the prompt is still open. Bizarre examples.
        
         | fumeux_fume wrote:
         | General purpose LLMs aren't very good with generating bounding
         | boxes, so with that context, this is actually seen as decent
         | performance for certain use cases.
        
         | dolmen wrote:
         | Not that bad compared to product images seen on AliExpress.
        
         | tennisflyi wrote:
         | You seen the charts on their last release? They obviously don't
         | check - too rich
        
         | az226 wrote:
         | And here is Gemini 3:
         | https://media.licdn.com/dms/image/v2/D5610AQH7v9MtrZxxug/ima...
        
           | saejox wrote:
           | This is very impressive. Google really is ahead
        
             | pietz wrote:
             | They are definitely ahead in multi modality and I'd argue
             | they have been for a long time. Their image understanding
             | was already great, when their core LLM was still terrible.
        
           | FinnKuhn wrote:
           | This is genuinly impressive. The OpenAI equivalent is less
           | detailed AND less correct.
        
           | Lionga wrote:
           | When OpenAI Marketing Material is actually showing how far
           | Gemini3 is ahead...
        
       | dandiep wrote:
       | Still no GPT 5.x fine tuning?
       | 
       | I emailed support a while back to see if there was an early
       | access program (99.99% sure the answer is yes). This is when I
       | discovered that their support is 100% done by AI and there is no
       | way to escalate a case to a human.
        
         | jazzyjackson wrote:
         | Containment breach is going to occur from a disgruntled
         | customer convincing the customer service bot it needs to get a
         | hold of a supervisor
        
           | dandiep wrote:
           | I tried this!
           | 
           | Me: Is it possible get early access to the GPT 5 fine tuning?
           | It is much better at non-English languages, which is the
           | majority of what we do, so it would be greatly beneficial to
           | us.
           | 
           | Bot: I'm an AI support agent. I don't have information about
           | early access to GPT-5 fine-tuning or unannounced features.
           | Currently, there is no public information or program for
           | early access to GPT-5 fine-tuning. Please keep an eye on
           | official OpenAI channels for updates. Let me know if you have
           | any questions about current features or products!
           | 
           | Me: Can you ask someone from the team if I can be part of the
           | early access? I know there is no public information, but
           | there is an early access program and I would like to be part
           | of it.
           | 
           | Bot: I'm an AI support agent and do not have the ability to
           | forward requests or make inquiries about early access
           | programs with the OpenAI team. If early access becomes
           | available, it will be announced through official channels.
           | Let me know if you have any other questions about OpenAI
           | products or current fine-tuning options.
           | 
           | Me: How do I get in touch with the open ai team?
           | 
           | Bot: I'm an AI support agent and can only assist with OpenAI
           | products or services through this support channel. There is
           | no separate contact method for the OpenAI team beyond
           | official support and public announcements. If you have
           | questions about current products or features, I'm here to
           | help
           | 
           | Should have tried a jailbreak maybe...
        
             | jazzyjackson wrote:
             | Yea you might have to convince it someone is in mortal
             | danger or something
        
               | int_19h wrote:
               | You're talking about a model that will notoriously refuse
               | to say a racist word even if you tell it that it's the
               | only way to prevent a nuclear apocalypse.
        
       | orliesaurus wrote:
       | I told all my friends to upgrade or they're not my friends
       | anymore /s
        
       | ImprobableTruth wrote:
       | An almost 50% price increase. Benchmarks look nice, but 50% more
       | nice...?
        
         | arnaudsm wrote:
         | #1 models are usually priced at 2x more than the competition,
         | and they often decrease the price right when they lose the
         | crown.
        
           | wewtyflakes wrote:
           | There are too few examples to say this is a trend. There have
           | been counterexamples of top models actually lowering the
           | pricing bar (gpt-5, gpt-3.5-turbo, some gemini releases were
           | even totally free [at first]).
        
       | sigmar wrote:
       | Are there any specifics about how this was trained? Especially
       | when 5.1 is only a month old. I'm a little skeptical of
       | benchmarks these days and wish they put this up on llmarena
       | 
       | edit: noticed 5.2 is ranked in the webdev arena (#2 tied with
       | gemini-3.0-pro), but not yet in text arena (last update 22hrs
       | ago)
        
         | kouteiheika wrote:
         | Unfortunately there are never any real specifics about how any
         | of their models were trained. It's OpenAI we're talking about
         | after all.
        
         | emp17344 wrote:
         | I'm extremely skeptical because of all those articles claiming
         | OpenAI was freaking out about Gemini - now it turns out they
         | just casually had a better model ready to go? I don't buy it.
        
           | tempaccount420 wrote:
           | They had to rush it out, I'm sure the internal safety folks
           | are not happy about it.
        
           | Workaccount2 wrote:
           | I (and others) have a strong suspicion that they can modulate
           | models intelligence in almost real time by adjusting
           | quantization and thinking time.
           | 
           | It seems if anyone wants, they can really gas a model up in
           | the moment and back it off after the hype wave.
        
             | bamboozled wrote:
             | Yeah I've noticed with Claude, around the time of the Opus
             | 4.5 release, at least for a few days, Sonnet 4.5 was just
             | dumb, but it seems temporary. I feel that redirected
             | resources to Opus.
        
             | qeternity wrote:
             | Quantization is not some magical dial you can just turn. In
             | practice you basically have 3 choices: fp16, fp8 and fp4.
             | 
             | Also thinking time means more tokens which costs more
             | especially at the API level where you are paying per token
             | and would be trivially observable.
             | 
             | There is basically no evidence that either of these are
             | occurring in the way you suggest (boosting up and down).
        
               | Workaccount2 wrote:
               | API users probably wouldn't be affected since they are
               | paying in full. Most people complaining are free users,
               | followed by $20/mo users.
        
           | bamboozled wrote:
           | It's very inline with their PR strategy, or lack of.
        
           | robots0only wrote:
           | how do you know this is a better model? I wouldn't take any
           | of the numbers at face value especially when all they have
           | done is more/better post-training and thus the base pre-
           | trained model capabilities is still the same. The model may
           | just elicit some of the benchmark capabilities better. You
           | really need to spend time using the model to come to any
           | reliable conclusions.
        
       | DeathArrow wrote:
       | Pricing is the same?
        
         | tedsanders wrote:
         | ChatGPT pricing is the same. API pricing is +40% per token,
         | though greater token efficiency means that cost per task is not
         | always that much higher. On some agentic evals we actually saw
         | costs per task go down with GPT-5.2. It really depends on the
         | task though; your mileage may vary.
        
           | ComputerGuru wrote:
           | How long have you been previewing 5.2?
        
       | johnsutor wrote:
       | https://platform.openai.com/docs/models/gpt-5.2 More information
       | on the price, context window, etc.
        
       | gkbrk wrote:
       | Is this the "Garlic" model people have been hyping? Or are we not
       | there yet?
        
         | 0x457 wrote:
         | Garlic will be released 2026Q1.
        
       | coolfox wrote:
       | the halving of error rates for image inputs is pretty awesome,
       | this makes it far more practical for issues where it isn't easy
       | to input all the needed context. when I get lazy I'll just
       | shift+win+s the problem and ask one of the chatbots to solve it.
        
       | xd1936 wrote:
       | > While GPT-5.2 will work well out of the box in Codex, we expect
       | to release a version of GPT-5.2 optimized for Codex in the coming
       | weeks.
       | 
       | https://openai.com/index/introducing-gpt-5-2/
        
         | jstummbillig wrote:
         | > For coding tasks, GPT-5.1-Codex-Max is a faster, more
         | capable, and more token-efficient coding variant
         | 
         | Hm, yeah, strange. You would not be able to tell, looking at
         | every chart on the page. Obviously not a gotcha, they put it on
         | the page themselves after all, but how does that make sense
         | with those benchmarks?
        
           | tempaccount420 wrote:
           | Coding requires a mindset shift that the -codex fine-tunes
           | provide. Codex will do all kinds of weird stuff like poking
           | in your ~/.cargo ~/go etc. to find docs and trying out code
           | in isolation, these things definitely improve capability.
        
             | dmos62 wrote:
             | The biggest advantage of codex variants, for me, is
             | terseness and reduced sicophany. That, and presumably
             | better adherence to requested output formats.
        
           | deaux wrote:
           | Looks like they removed that line.
        
           | baq wrote:
           | Codex talks _much_ less than the standard variant, especially
           | between tool calls.
        
         | k_bx wrote:
         | gpt-5.2 is already present in codex at this moment
        
       | preetamjinka wrote:
       | It's actually more expensive than GPT-5.1. I've gotten used to
       | prices going down with each latest model, but this time it's gone
       | up.
       | 
       | https://platform.openai.com/docs/pricing
        
         | Handy-Man wrote:
         | It also seems much more "smarter" though
        
         | PhilippGille wrote:
         | Gemini 3 Pro Preview also got more expensive than 2.5 Pro.
         | 
         | 2.5 Pro: $1.25 input, $10 output (million tokens)
         | 
         | 3 Pro Preview: $2 input, $12 output (million tokens)
        
           | TechDebtDevin wrote:
           | Literally no difference in productivity from a free/ <0.50c
           | output OpenRouter model. All these > $1.00+ per mm output are
           | literal scams. No added value to the world.
        
             | wahnfrieden wrote:
             | 5.1 Pro is great
        
               | manmal wrote:
               | I struggle to see where Pro is better than 5.x with
               | Thinking. Actually prefer the latter.
        
               | wahnfrieden wrote:
               | Many problems where latter spins its wheel and Pro gets
               | it in one go, for me. You need to give Pro full files as
               | context and you need to fit within its ~60k (I forget
               | exactly) silent context window if using via ChatGPT.
               | Don't have it make edits directly, have it give the
               | execution plan back to Codex
        
         | moralestapia wrote:
         | Previous model's prices usually go down, but their flagship has
         | always been the most expensive one.
        
           | moralestapia wrote:
           | Wtf, why would this be downvoted?
           | 
           | I'm adding context and what I stated is provably true.
        
         | endorphine wrote:
         | Reading this comment, it just occurred to me that we're still
         | in the first phase of the enshittification process.
        
         | kingstnap wrote:
         | Flagship models have rarely being cheaper, and especially not
         | on release day. Only a few cases of this really.
         | 
         | Notable exceptions are Deepseek 3.2 and Opus 4.5 and GPT 3.5
         | Turbo.
         | 
         | The price drops usually are the form of flash and mini models
         | being really cheap and fast. Like when we got o4 mini or 2.0
         | flash which was a particularly significant one.
        
           | n2d4 wrote:
           | That's not true.                   > Notable exceptions are
           | Deepseek 3.2 and Opus 4.5 and GPT 3.5 Turbo.
           | 
           | And GPT-4o, GPT-4.1, and GPT-5. Almost every OpenAI release
           | got cheaper on a per-input-token basis.
        
         | deaux wrote:
         | Getting more expensive has been the trend for the closed
         | weights frontier models. See Gemini 3 Pro vs 2.5 Pro. Also see
         | Gemini 2.5 Flash vs 2.0 Flash. The only thing that got cheaper
         | recently was Opus 4.5 vs Opus 4.
        
       | ComputerGuru wrote:
       | Wish they would include or leak more info about what this is,
       | exactly. 5.1 was just released, yet they are claiming big
       | improvements (on benchmarks, obviously). Did they purposely not
       | release the best they had to keep some cards to play in case of
       | Gemini 3 success or is this a tweak to use more time/tokens to
       | get better output, or what?
        
       | Ninjinka wrote:
       | Man this was rushed, typo in the first section:
       | 
       | > Unlike the previous GPT-5.1 model, GPT-5.2 has new features for
       | managing what the model "knows" and "remembers to improve
       | accuracy.
        
         | petercooper wrote:
         | Also, did they mention these features? I was looking out for it
         | but got to the end and missed it.
         | 
         | (No, I just looked again and the new features listed are around
         | verbosity, thinking level and the tool stuff rather than memory
         | or knowledge.)
        
       | gigatexal wrote:
       | So how much better is it than opus or Gemini ?
        
       | HardCodedBias wrote:
       | Huge fan that Gemini-3 prompted OAI to ship this.
       | 
       | Competition works!
       | 
       | GDPval seems particularly strong.
       | 
       | I wonder why they held this back.
       | 
       | 1) Maybe this is uneconomical ?
       | 
       | 2) Did the safety somehow hold back the company ?
       | 
       | looking forward to the internet trying this and posting their
       | results over the next week or two.
       | 
       | COMPETITION!
        
         | mrandish wrote:
         | > I wonder why they held this back.
         | 
         | IMHO, I doubt they were holding much back. Obviously, they're
         | always working on 'next improvements' and rolled what was done
         | enough into this but I suspect the real difference here is
         | throwing significantly more compute (hence investor capital) at
         | improving the quality - right now. How much? While the cost is
         | currently staying the same for most users, the API costs seem
         | to be ~40% higher.
         | 
         | The impetus was the serious threat Gemini 3 poses. Perception
         | about ChatGPT was starting to shift, people were speculating
         | that maybe OAI is more vulnerable than assumed. This caused
         | Altman to call an all-hands "Code Red" two weeks ago,
         | triggering a significant redeployment of priorities, resources
         | and people. I think this launch is the first 'stop the
         | perceptual bleeding' result of the Code Red. Given the timing,
         | I think this is mostly akin to overclocking a CPU or running an
         | F1 race car engine too hot to quickly improve performance - at
         | the cost of being unsustainable and unprofitable. To placate
         | serious investor concerns, OAI has recently been trying to
         | gradually work toward making current customers profitable (or
         | at least less unprofitable). I think we just saw the effort to
         | reduce the insane burn rate go out the window.
        
       | Jackson__ wrote:
       | Funny that, their front page demo has a mistake. For the waves
       | simulation, the user asks:
       | 
       | >- The UI should be calming and realistic.
       | 
       | Yet what it did is make a sleek frosted glass UI with rounded
       | edges. What it should have done is call a wellness check on the
       | user on suspicion of a co2 leak leading to delirium.
        
       | jasonthorsness wrote:
       | Does anyone have it yet in ChatGPT? I'm still on 5.1 :(.
        
         | mudkipdev wrote:
         | No, but it's already in codex
        
         | FergusArgyll wrote:
         | > We deploy GPT-5.2 gradually to keep ChatGPT as smooth and
         | reliable as we can; if you don't see it at first, please try
         | again later.
        
         | jasonthorsness wrote:
         | I have it now
        
       | FergusArgyll wrote:
       | > Additionally, on our internal benchmark of junior investment
       | banking analyst spreadsheet modeling tasks--such as putting
       | together a three-statement model for a Fortune 500 company with
       | proper formatting and citations, or building a leveraged buyout
       | model for a take-private--GPT 5.2 Thinking's average score per
       | task is 9.3% higher than GPT-5.1's, rising from 59.1% to 68.4%.
       | 
       | Confirming prior reporting about them hiring junior analysts
        
       | zhyder wrote:
       | Big knowledge cutoff jump from Sep 2024 to Aug 2025. How'd they
       | pull that off for a small point release, which presumably hasn't
       | done a fresh pre-training over the web?
       | 
       | Did they figure out how to do more incremental knowledge updates
       | somehow? If yes that'd be a huge change to these releases going
       | forward. I'd appreciate the freshness that comes with that
       | (without having to rely on web search as a RAG tool, which isn't
       | as deeply intelligent, as is game-able by SEO).
       | 
       | With Gemini 3, my only disappointment was 0 change in knowledge
       | cutoff relative to 2.5's (Jan 2025).
        
         | throwaway314155 wrote:
         | > which presumably hasn't done a fresh pre-training over the
         | web
         | 
         | What makes you think that?
         | 
         | > Did they figure out how to do more incremental knowledge
         | updates somehow?
         | 
         | It's simple. You take the existing model and continue
         | pretraining with newly collected data.
        
           | Workaccount2 wrote:
           | A leak reported on by semi-analyses stated that they haven't
           | pre-trained a new model since 4o due to compute constraints.
        
       | jumploops wrote:
       | > "a new knowledge cutoff of August 2025"
       | 
       | This (and the price increase) points to a new pretrained model
       | under-the-hood.
       | 
       | GPT-5.1, in contrast, was allegedly using the same pretraining as
       | GPT-4o.
        
         | 98Windows wrote:
         | or maybe 5.1 was an older checkpoint and has more quantization
        
         | FergusArgyll wrote:
         | A new pretrain would definitely get more than a .1 version bump
         | & would get a whole lot more hype I'd think. They're expensive
         | to do!
        
           | femiagbabiaka wrote:
           | Not if they didn't feel that it delivered customer value no?
           | It's about under promising and over delivering, in every
           | instance
        
           | redwood wrote:
           | Not if it underwhelms
        
           | hannesfur wrote:
           | Maybe they felt the increase in capability is not worth of a
           | bigger version bump. Additionally pre-training isn't as
           | important as it used to be. Most of the advances we see now
           | probably come from the RL stage.
        
           | caconym_ wrote:
           | Releasing anything as "GPT-6" which doesn't provide a
           | generational leap in performance would be a PR nightmare for
           | them, especially after the underwhelming release of GPT-5.
           | 
           | I don't think it really matters what's under the hood. People
           | expect model "versions" to be indexed on performance.
        
           | ACCount37 wrote:
           | Not necessarily. GPT-4.5 was a new pretrain on top of a
           | sizeable raw model scale bump, and only got 0.5 - because the
           | gains from reasoning training in o-series overshadowed
           | GPT-4.5's natural advantage over GPT-4.
           | 
           | OpenAI might have learned not to overhype. They already
           | shipped GPT-5 - which was only an incremental upgrade over
           | o3, and was received poorly, with this being a part of the
           | reason why.
        
             | diego_sandoval wrote:
             | I jumped straight from 4o (free user) into GPT-5 (paid
             | user).
             | 
             | It was a generational leap if there ever has been one. Much
             | bigger than 3.5 to 4.
        
               | kadushka wrote:
               | What kind of improvements do you expect when going from 5
               | straight to 6?
        
               | ACCount37 wrote:
               | Yes, if OpenAI released GPT-5 after GPT-4o, then it would
               | have been seen as a proper generational leap.
               | 
               | But o3 existing and being good at what it does? Took the
               | wind out of GPT-5's sails.
        
           | boc wrote:
           | Maybe the rumors about failed training runs weren't wrong...
        
           | jumploops wrote:
           | It's possible they're using some new architecture to get more
           | up-to-date data, but I think that'd be even more of a
           | headline.
           | 
           | My hunch is that this is the same 5.1 post-training on a new
           | pretrained base.
           | 
           | Likely rushed out the door faster than they initially
           | expected/planned.
        
           | OrangeMusic wrote:
           | Yeah because OpenAI has been great at naming their models so
           | far? ;)
        
         | MagicMoonlight wrote:
         | No, they just feed in another round of slop to the same model.
        
         | redox99 wrote:
         | I think it's more likely to be the old base model checkpoint
         | further trained on additional data.
        
           | jumploops wrote:
           | Is that technically not a new pretrained model?
           | 
           | (Also not sure how that would work, but maybe I've missed a
           | paper or two!)
        
             | redox99 wrote:
             | I'd say for it to be called a new pretrained model, it'd
             | need to be trained from scratch (like llama 1, 2, 3).
             | 
             | But it's just semantics.
        
       | devinprater wrote:
       | Can the tables have column headers so my screen reader can read
       | the model name as I go across the benchmakrs? And the images
       | should have alt-text.
        
       | MagicMoonlight wrote:
       | They're definitely just training the models on the benchmarks at
       | this point
        
         | roxolotl wrote:
         | Yea either this is an incredible jump or we've finally gotten
         | confirmation benchmarks are bs.
        
       | simonw wrote:
       | Wow, there's a lot going on with this pelican riding a bicycle:
       | https://gist.github.com/simonw/c31d7afc95fe6b40506a9562b5e83...
        
         | minimaxir wrote:
         | Is that the first SVG pelican with drop shadows?
        
           | simonw wrote:
           | No, I got drop shadows from DeepSeek 3.2 recently
           | https://simonwillison.net/2025/Dec/1/deepseek-v32/ (probably
           | others as well.)
        
         | tmaly wrote:
         | seems to be eating something
        
           | danans wrote:
           | Probably a jellyfish. You're seeing the tentacles
        
         | belter wrote:
         | What happens if you ask for a pterodactyl on a motorbike?
         | 
         | Would like to know how much they are optimizing for your
         | pelican....
        
           | simonkagedal wrote:
           | He commented on this here:
           | https://simonwillison.net/2025/Nov/13/training-for-
           | pelicans-...
        
             | irthomasthomas wrote:
             | I was expecting to see a pterodactyl :(
        
         | fxwin wrote:
         | the only benchmark i trust
        
         | BeetleB wrote:
         | They probably saw your complaint that 5.1 was too spartan and a
         | regression (I had the same experience with 5.1 in the POV-Ray
         | version - have yet to try 5.2 out...).
        
         | Stevvo wrote:
         | The variance is way too high for this test to have any value at
         | all. I ran it 10 times, and each pelican on a bicycle was a
         | better rendition than that, about half of them you could say
         | were perfect.
        
           | golly_ned wrote:
           | Compared to the other benchmarks which are much more
           | gameable, I trust PelicanBikeEval way more.
        
           | getnormality wrote:
           | Well, the variance is itself interesting.
        
         | AstroBen wrote:
         | Seems to be getting more aerodynamic. A clear sign of AI
         | intelligence
        
         | sroussey wrote:
         | What _is_ good at SVG design?
        
           | azinman2 wrote:
           | Graphic designers?
        
           | culi wrote:
           | Not svg, but basically the same challenge:
           | 
           | https://clocks.brianmoore.com/
           | 
           | Probably Kimi or Deepseek are best
        
           | KellyCriterion wrote:
           | Ive not seen any model being good in graphic/svg creation so
           | far - all of the stuff mostly looks ugly and somewhat
           | "synthetic-disorted".
           | 
           | And lately, Claude (web) started to draw ascii charts from
           | one day to another indstead of colorful infographicstyled-
           | images as it did before (they were only slightly better than
           | the ascii charts)
        
         | nightshift1 wrote:
         | benchmarks probably should not be used for so long.
        
         | alechewitt wrote:
         | Nice work on these benchmarks Simon. I've followed your blog
         | closely since your great talk at the AI Engineers World Fair,
         | and I want to say thank you for all the high quality content
         | you share for free. It's become my primary source for keeping
         | up to date.
         | 
         | I've been working on a few benchmarks to test how well LLMs can
         | recreate interfaces from screenshots.
         | (https://github.com/alechewitt/llm-ui-challenge). From my basic
         | tests, it seems GPT-5.2 is slightly better at these UI
         | recreations. For example, in the MS Word replica, it
         | implemented the undo/redo buttons as well as the bold/italic
         | formatting that GPT-5.1 handled, and it generally seemed a bit
         | closer to the original screenshot
         | (https://alechewitt.github.io/llm-ui-
         | challenge/outputs/micros...).
         | 
         | In the VS Code test, it also added the tabs that weren't
         | visible in the screenshot! (https://alechewitt.github.io/llm-
         | ui-challenge/outputs/vs_cod...).
        
           | simonw wrote:
           | That is a very good benchmark. Interesting to see GPT-5.2
           | delivering on the promise of better vision support there.
        
         | tkgally wrote:
         | I added GPT-5.2 Pro to my pelican-alternatives benchmark for
         | the first three prompts:
         | 
         | Generate an SVG of an octopus operating a pipe organ
         | 
         | Generate an SVG of a giraffe assembling a grandfather clock
         | 
         | Generate an SVG of a starfish driving a bulldozer
         | 
         | https://gally.net/temp/20251107pelican-alternatives/index.ht...
         | 
         | GPT-5.2 Pro cost about 80 cents per prompt through OpenRouter,
         | so I stopped there. I don't feel like spending that much on all
         | thirty prompts.
        
           | smusamashah wrote:
           | Hi, it doesn't have Gemini 3.5 Pro which seems to be the best
           | at this
        
             | svantana wrote:
             | That's probably because "Gemini 3.5 Pro" doesn't exist
        
         | tootie wrote:
         | Do you think the big guys are on to your game and have been
         | adding extra pelicans to the training data?
        
       | ComputerGuru wrote:
       | Wish they would include or leak more info about what this is,
       | exactly. 5.1 was just released, yet they are claiming big
       | improvements (on benchmarks, obviously). Did they purposely not
       | release the best they had to keep some cards to play in case of
       | Gemini 3 success or is this a tweak to use more time/tokens to
       | get better output, or what?
        
         | eldenring wrote:
         | I'm guessing they were waiting to figure out more efficient
         | serving before a release, and have decided to eat the inference
         | cost temporarily to stay at the frontier.
        
         | famouswaffles wrote:
         | Open AI sat on GPT-4 for 8 months and even released 3.5 months
         | after 4 was trained. While i don't expect such big lag times
         | anymore, generally, it's a given the public is behind whatever
         | models they have internally at the frontier. By all
         | indications, they did not want to release this yet, and only
         | did so because of Gemini-3-pro.
        
         | dalemhurley wrote:
         | My guess is they develop multiple models in parallel.
        
         | nathan-wall wrote:
         | If you look at their own chart[1] it shows 5.1 was lagging
         | behind Gemini 3 Pro in almost every score listed there,
         | sometimes significantly. They needed to come out with something
         | to stay ahead. I'm guessing they threw what they had at their
         | disposal together to keep the lead as long as they can. It
         | sounds like 5.2 has a more recent knowledge cutoff; a
         | reasonable guess is they could have already had that but were
         | trying to make bigger improvements out of it for a more major
         | 5.5 release before Gemini 3 Pro came out and then they had to
         | rush something out. Also 5.2 has a new "Extended Thinking"
         | option for Pro. I'm guessing they just turned up a lever that
         | told it to think even longer, which helps them score higher,
         | even if it does take a long time. (One thing about Gemini 3 Pro
         | is it's very fast relative to even ChatGPT 5.1 Pro Thinking. A
         | lot of the scores they're putting out to show they're staying
         | ahead aren't showing that piece.)
         | 
         | [1] https://imgur.com/e0iB8KC
        
       | airstrike wrote:
       | I feel like if we're going to regulate anything about AI, we
       | should start by regulating (1) what they get to claim to be a
       | "new model" to the public and (2) what changes they are allowed
       | to make at inference before being forced to name it something
       | different.
        
       | chux52 wrote:
       | Is this why all my Cursor requests are timing out in the past
       | hour?
        
       | sureglymop wrote:
       | How can I hide the big "Ask ChatGPT" button I accidentally
       | clicked like 3 times while actually trying to read this on my
       | phone?
       | 
       | I guess I must "listen" to the article...
        
         | z58 wrote:
         | With Safari on iOS you can hide distracting items. I just tried
         | it on that button, it works flawlessly.
        
       | riazrizvi wrote:
       | Does it still use the word 'fluff' in 90% of its preambles, or is
       | it finally able to get straight to the point?
        
       | ChrisArchitect wrote:
       | Discussion on blog post: https://openai.com/index/introducing-
       | gpt-5-2/ (https://news.ycombinator.com/item?id=46234874)
        
       | yousif_123123 wrote:
       | Why doesn't OpenAI include comparisons to other models anymore?
        
         | ftchd wrote:
         | because they probably need to compare pricing too
        
         | enraged_camel wrote:
         | Because their main competition (Google and Anthropic) have
         | caught up and even started to surpass them, and comparisons
         | would simply drive it home.
        
           | IAmNotACellist wrote:
           | Why do they care so much? They're a non-profit dedicated to
           | the betterment of humanity via open access to AI. They have
           | nothing to hide. They have no motivation to lie, or lie by
           | omission.
        
             | koolba wrote:
             | > Why do they care so much? They're a non-profit dedicated
             | to the betterment of humanity via open access to AI.
             | 
             | We're still talking about OpenAI right?
        
               | IAmNotACellist wrote:
               | You're not calling Sam Altman a liar, are you?
        
             | kaliqt wrote:
             | They are not a nonprofit at all. Legally, yes. But they are
             | not.
        
         | conradkay wrote:
         | Sam Altman posted with a comparison to Gemini 3 and Opus 4.5
         | 
         | https://x.com/sama/status/1999185784012947900
        
           | yousif_123123 wrote:
           | I see, thanks for this.
        
       | HackerThemAll wrote:
       | No, thank you, OpenAI and ChatGPT doesn't cut it for me.
        
         | dang wrote:
         | " _Please don 't post shallow dismissals, especially of other
         | people's work. A good critical comment teaches us something._"
         | 
         | https://news.ycombinator.com/newsguidelines.html
        
       | d--b wrote:
       | > it's better at creating spreadsheets
       | 
       | I have a bad feeling about this.
        
       | HackerThemAll wrote:
       | No, thank you, OpenAI and ChatGPT doesn't cut it for me.
        
         | wayeq wrote:
         | thanks for letting us know.
        
         | replwoacause wrote:
         | What's cutting it for you these days?
        
       | daviding wrote:
       | gpt-5.2 and gpt-5.2-chat-latest the same token price? Isn't the
       | latter non-thinking and more akin to -nano or -mini?
        
         | dalemhurley wrote:
         | No. It is the same model without reasoning.
        
           | daviding wrote:
           | So is maybe gpt-5.2 with reasoning set to 'none' identical to
           | gpt-5.2-chat-latest in capabilities but perhaps with a
           | different system (system) prompt? I notice chat-latest
           | doesn't accept temperature or reasoning (which makes sense)
           | parameters, so something is certainly different underneath?
        
       | scottndecker wrote:
       | Still 256K input tokens. So disappointing (predictable, but
       | disappointing).
        
         | htrp wrote:
         | much harder to train longer context inputs
        
         | coder543 wrote:
         | https://platform.openai.com/docs/models/gpt-5.2
         | 
         | 400k, not 256k.
        
           | nathants wrote:
           | 400 - 128 = 272. Codex cli source.
        
             | coder543 wrote:
             | If you want to be able to generate up to 128k tokens in one
             | go successfully, then yes, that math checks out.
        
       | dinobones wrote:
       | It's becoming challenging to really evaluate models.
       | 
       | The amount of intelligence that you can display within a single
       | prompt, the riddles, the puzzles, they've all been solved or are
       | mostly trivial to reasoners.
       | 
       | Now you have to drive a model for a few days to really get a
       | decent understanding of how good it really is. In my experience,
       | while Sonnet/Opus may not have always been leading on benchmarks,
       | they have always *felt* the best to me, but it's hard to put into
       | words why exactly I feel that way, but I can just feel it.
       | 
       | The way you can just _feel_ when someone you 're having a
       | conversation with is deeply understanding you, somewhat
       | understanding you, or maybe not understanding at all. But you
       | don't have a quantifiable metric for this.
       | 
       | This is a strange, weird territory, and I don't know the path
       | forward. We know we're definitely not at AGI.
       | 
       | And we know if you use these models for long-horizon tasks they
       | fail at some point and just go off the rails.
       | 
       | I've tried using Codex with max reasoning for doing PRs and
       | gotten laughable results too many times, but Codex with Max
       | reasoning is apparently near-SOTA on code. And to be fair, Claude
       | Code/Opus is also sometimes equally as bad at doing these types
       | of "implement idea in big codebase, make changes too many files,
       | still pass tests" type of tasks.
       | 
       | Is the solution that we start to evaluate LLMs on more long-
       | horizon tasks? I think to some degree this was the spirit of SWE
       | Verified right? But even that is being saturated now.
        
         | ACCount37 wrote:
         | The good old "benchmarks just keep saturating" problem.
         | 
         | Anthropic is genuinely one of the top companies in the field,
         | and for a reason. Opus consistently punches above its weight,
         | and this is only in part due to the lack of OpenAI's atrocious
         | personality tuning.
         | 
         | Yes, the next stop for AI is: increasing task length horizon,
         | improving agentic behavior. The "raw general intelligence"
         | component in bleeding edge LLMs is far outpacing the "executive
         | function", clearly.
        
           | imiric wrote:
           | Shouldn't the next stop be to improve general accuracy, which
           | is what these tools have struggled with since their
           | inception? Until when are "AI" companies going to offload the
           | responsibility on the user to verify the output of their
           | tools?
           | 
           | Optimizing for benchmark scores, which are highly gamed to
           | begin with, by throwing more resources at this problem is
           | exceedingly tiring. Surely they must've noticed the
           | performance plateau and diminishing returns of this approach
           | by now, yet every new announcement is the same.
        
             | ACCount37 wrote:
             | What "performance plateau"? The "plateau" disappears the
             | moment you get harder unsaturated benchmarks.
             | 
             | It's getting more and more challenging to do that - just
             | not because the models don't improve. Quite the opposite.
             | 
             | Framing "improve general accuracy" as "something no one is
             | doing" is really weird too.
             | 
             | You need "general accuracy" for agentic behavior to work at
             | all. If you have a simple ten step plan, and each step has
             | a 50% chance of an unrecoverable failure, then your plan is
             | fucked, full stop. To advance on those benchmarks, the LLM
             | has to fail less and recover better.
             | 
             | Hallucinations is a "solvable but very hard to solve"
             | problem. Considerable progress is being made on it, but if
             | there's "this one weird trick" that deletes hallucinations,
             | then we sure didn't find it yet. Humans get a body of meta-
             | knowledge for free, which lets them dodge hallucinations
             | decently well (not perfectly) if they want to. LLMs get
             | pathetic crumbs of meta-knowledge and little skill in using
             | it. Room for improvement, but, not trivial to improve.
        
         | Libidinalecon wrote:
         | Totally agree. I just got a free trial month I guess to try to
         | bring me back to chatGPT but I don't really know what to ask it
         | to display if it is on par with Gemini.
         | 
         | I really have a sinking feel right now actually of what an
         | absolute giant waste of capital all this is.
         | 
         | I am glad for all the venture capital behind all this to
         | subsidize my intellectual noodlings on a super computer but my
         | god what have we done?
         | 
         | This is so much fun but this doesn't feel like we are getting
         | closer to "AGI" after using Gemini for about 100 hours or so
         | now. The first day maybe but not now when you see how off it
         | can still be all the time.
        
       | qoez wrote:
       | This is also the exact on-the-day 10th anniversary of openai's
       | creation incidentally
        
       | cc62cf4a4f20 wrote:
       | In other news, been using Devstral 2 (Ollama) with OpenCode, and
       | while it's not as good as Claude Code, my initial sense it that
       | it's nonetheless good enough and doesn't require me to send my
       | data off my laptop.
       | 
       | I kind of wonder how close we are to alternative (not from a
       | major AI lab) models being good enough for a lot of productive
       | work and data sovereignty being the deciding factor.
        
         | Nesco wrote:
         | Wait, isn't Devstral2 (normal not small) 123b? What type of
         | laptop do you have? MacBooks don't go over 128GiB
        
           | cc62cf4a4f20 wrote:
           | I'm using small - works well for its size
        
         | yberreby wrote:
         | Would you share some additional details? CPU, amount of unified
         | memory / VRAM? Tok/s with those?
        
           | cc62cf4a4f20 wrote:
           | MBP M4 Max 64MB - haven't measured the tokens/sec, feels
           | slower than Claude, but not unbearably
           | 
           | It's not yet perfect, my sense is just that it's near the
           | tipping point where models are efficient enough that running
           | a local model is truly viable
        
       | a_wild_dandan wrote:
       | > Unlike the previous GPT-5.1 model, GPT-5.2 has new features for
       | managing what the model "knows" and "remembers to improve
       | accuracy.
       | 
       | Dumb nit, but why not put your own press release through your
       | model to prevent basic things like missing quote marks? Reminds
       | me of that time an OAI released wildly inaccurate copy/pasted bar
       | charts.
        
         | Imnimo wrote:
         | It does seem to raise fair questions about either the utility
         | of these tools, or adoption inertia. If not even OpenAI feels
         | compelled to integrate this kind of model-check into their
         | pipeline, what's that say about the business world at-large? Is
         | it that it's too onerous to set up, is it that it's too hard to
         | get only true-positive corrections, is it that it's too low
         | value for the effort?
        
           | JumpCrisscross wrote:
           | > _what 's that say about the business world at-large?_
           | 
           | Nothing. OpenAI is a terrible baseline to extrapolate
           | anything from.
        
         | croes wrote:
         | Maybe they did
        
         | layer8 wrote:
         | Humans are now expected to parse sloppy typing without
         | complaining about it, just like LLMs do. Slop is the new
         | normal.
        
         | Bengalilol wrote:
         | It may have been used, how could we know?
         | 
         | Mainly, I don't get why there are quote marks at all.
        
         | boplicity wrote:
         | Their model doesn't handle punctuation, quote marks, and
         | similar things very well at all.
        
         | MaxikCZ wrote:
         | I always remember this old image
         | https://i.imgur.com/MCsOM8e.jpeg
        
       | SkyPuncher wrote:
       | Given the price increase and speculation that GPT 5 is a MoE
       | model, I'm wondering if they're simply "turning up the good
       | stuff" without making significant changes under the hood.
        
         | throwaway314155 wrote:
         | GPT 4o was an MoE model as well.
        
         | minimaxir wrote:
         | I'm not sure why being a MoE model would allow OpenAI to "turn
         | up the good stuff". You can't just increase the number of E
         | without training it as such.
        
           | yberreby wrote:
           | Based on what works elsewhere in deep learning, I see no
           | reason why you couldn't train once with a randomized number
           | of experts, then set that number during inference based on
           | your desired compute-accuracy tradeoff. I would expect that
           | this has been done in the literature already.
        
           | SkyPuncher wrote:
           | My opinion is they're trying to internally route requests to
           | cheaper experts when they think they can get away with it. I
           | felt this was evident by the wild inconsistencies I'd
           | experience using it for coding. Both in quality and latency
           | 
           | You "turn of the good stuff" by eliminating or reducing the
           | likelihood of the cheap experts handling the request.
        
       | dumbmrblah wrote:
       | Great! It'll be SOTA for a couple of weeks until the quality
       | degrades due to throttling.
       | 
       | I'll stick with plug and play API instead.
        
         | mrandish wrote:
         | Due to the "Code Red" threat from Gemini 3, I suspect they'll
         | hold off throttling for longer than usual (by incinerating even
         | more investor capital than usual).
         | 
         | Jump in and soak up that extra-discounted compute while the
         | getting is good, kids! Personally, I recently retired so I just
         | occasionally mess around with LLMs for casual hobby projects,
         | so I've only ever used the free tier of all the providers.
         | Having lived through the dot com bubble, I regret not soaking
         | up more of the free and heavily subsidized stuff back then.
         | Trying not to miss out this time. All this compute available
         | for free or below cost won't last too much longer...
        
           | dankwizard wrote:
           | I've been using tools like ProxLLM which just slam these AI
           | models via proxy everytime a free tier limit is hit and it
           | works great.
        
             | ssvss wrote:
             | can you provide a link to this tool, a search for proxllm
             | didn't seem to find anything related.
        
       | impulser_ wrote:
       | The thing about OpenAI is their models never fit anywhere for me.
       | Yes they maybe smart or even the smartest models but they are
       | alway so fucking slow. The ChatGPT web app is literally usable
       | for me. I ask simple task and it does most extreme shit jsut to
       | get an answer that the same as Claude or Gemini.
       | 
       | For example, I asked ChatGPT to take a chart and convert into a
       | table. It went and cut up the image and zoomed in for literally 5
       | mins to get the a worst answer than Claude which did it in under
       | a minute.
       | 
       | I see people talk about Codex like it better than Claude Code,
       | and I go and try it and it takes a lifetime to do thing and it
       | return maybe an on par result as Opus or Sonnet but it takes
       | 5mins longer.
       | 
       | I just tried out this model and it the same exact thing. It just
       | take ages for it to give you an answer.
       | 
       | I don't get how these models are useful in the real world.
       | 
       | What am I missing, is this just me?
       | 
       | I guess it truly an enterprise model.
        
         | wetoastfood wrote:
         | Are you using 5.1 Thinking? I tended to prefer Claude before
         | this model.
         | 
         | I use models based on the task. They still seem specialized and
         | better at specific tasks. If I have a question I tend to go to
         | it. If I need code, I tend to go to Claude (Code).
         | 
         | I go to ChatGPT for questions I have because I value an
         | accurate answer over a quick answer and, in my experience, it
         | tends to give me more accurate answers because of its (over)
         | willingness to go to the web for search results and question
         | its instincts. Claude is much more likely to make an assumption
         | and its search patterns aren't as thorough. The slow answers
         | don't bother me because it's an expectation I have for how I
         | use it and they've made that use case work really well with
         | background processing and notifications.
        
       | zone411 wrote:
       | I've benchmarked it on the Extended NYT Connections benchmark
       | (https://github.com/lechmazur/nyt-connections/):
       | 
       | The high-reasoning version of GPT-5.2 improves on GPT-5.1: 69.9 -
       | 77.9.
       | 
       | The medium-reasoning version also improves: 62.7 - 72.1.
       | 
       | The no-reasoning version also improves: 22.1 - 27.5.
       | 
       | Gemini 3 Pro and Grok 4.1 Fast Reasoning still score higher.
        
         | Donald wrote:
         | Gemini 3 Pro Preview gets 96.8% on the same benchmark? That's
         | impressive
        
           | capitainenemo wrote:
           | And performs very well on the latest 100 puzzles too, so
           | isn't just learning the data set (unless I guess they
           | routinely index this repo).
           | 
           | I wonder how well AIs would do at bracket city. I tried
           | gemini on it and was underwhelmed. It made a lot of terrible
           | connections and often bled data from one level into the next.
        
             | wooger wrote:
             | > unless I guess they routinely index this repo
             | 
             | This sounds like exactly the kind of thing any tech company
             | would do when confronted with a competitive benchmark.
        
               | rsanek wrote:
               | I mean, the repo has <200 stars, it's not like it's so
               | mainstream that you'd expect LLM makers to be watching it
               | actively. If they wanted to game it, they could more
               | easily do that in RL with synthetic data anyway.
        
           | bigyabai wrote:
           | GPT-5.2 might be Google's best Gemini advertisement yet.
        
             | outside1234 wrote:
             | Especially when you see the price
        
         | tikotus wrote:
         | Here's someone else testing models on a daily logic puzzle
         | (Clues by Sam): https://www.nicksypteras.com/blog/cbs-
         | benchmark.html GPT 5 Pro was the winner already before in that
         | test.
        
           | thanhhaimai wrote:
           | This link doesn't have Gemini 3 performance on it. Do you
           | have an updated link with the new models?
        
             | dezgeg wrote:
             | I've also tried Gemini 3 for Clues by Sam and it can do
             | really well, have not seen it make a single mistake even
             | for Hard and Tricky ones. Haven't run it on too many
             | puzzles though.
        
           | crapple8430 wrote:
           | GPT 5 Pro is a good 10x more expensive so it's an apples to
           | oranges comparison.
        
         | scrollop wrote:
         | Why no grok 4.1 reasoning?
        
           | sanex wrote:
           | Do people other than Elon fans use grok? Honest question.
           | I've never tried it.
        
             | mac-attack wrote:
             | I can't understand why people would trust a CEO that
             | regularly lies about product timelines, product features,
             | his own personal life, etc. And that's before politicizing
             | his entire kingdom by literally becoming a part of
             | government and one of the larger donations of the current
             | administration.
        
               | lkjdsklf wrote:
               | If we stopped using products of every company that had a
               | CEO that lied about their products, we'd all be sitting
               | in caves staring at the dirt
        
               | fatata123 wrote:
               | Because not everyone makes their decisions through the
               | prism of politics
        
               | delaminator wrote:
               | You're not narrowing it down.
        
             | bumling wrote:
             | I dislike Musk, and use Grok. I find it most useful for
             | analyzing text to help check if there's anything I've
             | missed in my own reading. Having it built in to Twitter is
             | convenient and it has a generous free tier.
        
             | buu700 wrote:
             | I use Grok pretty heavily, and Elon doesn't factor into it
             | any more than Sam and Sundar do when I use GPT and Gemini.
             | A few use cases where it really shines:
             | 
             | * Research and planning
             | 
             | * Writing complex isolated modules, particularly when the
             | task depends on using a third-party API correctly (or even
             | choosing an API/library at its own discretion)
             | 
             | * Reasoning through complicated logic, particularly in
             | cases that benefit from its eagerness to throw a ton of
             | inference at problems where other LLMs might give a
             | shallower or less accurate answer without more prodding
             | 
             | I'll often fire off an off-the-cuff message from my phone
             | to have Grok research some obscure topic that involves
             | finding very specific data and crunching a bunch of
             | numbers, or write a script for some random thing that I
             | would previously never have bothered to spend time
             | automating, and it'll churn for ~5 minutes on reasoning
             | before giving me exactly what I wanted with few or no
             | mistakes.
             | 
             | As far as development, I personally get a lot of mileage
             | out of collaborating with Grok and Gemini on
             | planning/architecture/specs and coding with GPT. (I've
             | stopped using Claude since GPT seems interchangeable at
             | lower cost.)
             | 
             | For reference, I'm only referring to the Grok chatbot right
             | now. I've never actually tried Grok through agentic coding
             | tooling.
        
             | jbm wrote:
             | I use a few AIs together to examine the same code base. I
             | find Grok better than some of the Chinese ones I've used,
             | but it isn't in the same league as Claude or Codex.
        
             | ralusek wrote:
             | Only thing I use grok for is if there is a current
             | event/meme that I keep seeing referenced and I don't
             | understand, it's good at pulling from tweets
        
             | wdroz wrote:
             | Unlike openai, you can use the latest grok models without
             | verifying your organization and giving your ID.
        
             | sz4kerto wrote:
             | I'm using Gemini in general, but Grok too. That's because
             | sometimes Gemini Thinking is too slow, but Fast can get
             | confused a lot. Grok strikes a nice balance between being
             | quite smart (not Gemini 3 Pro level, but close) and very
             | fast.
        
             | scrollop wrote:
             | I hate the guy, however grok scores high on arc-2 so it
             | would be silly to not at least rank it.
        
             | rsanek wrote:
             | it's the biggest model on OpenRouter, even if you exclude
             | free tier usage https://openrouter.ai/state-of-ai
        
               | irthomasthomas wrote:
               | Roleplay is the largest use-case on openrouter.
        
         | Bombthecat wrote:
         | I would like to see a cost per percent or so row. I feel like
         | grok would beat them all
        
       | speedgoose wrote:
       | Trying it now in Vscode Insiders with Github Copilot (codex
       | crashes with HTTP 400 server errors), and it eventually started
       | using sed and grep in shells instead of using the better tools it
       | has access to. I guess this is not an issue to perform well in
       | benchmarks.
        
         | pixelmelt wrote:
         | to be fair I've seen the other sota models do this as well
        
         | songodongo wrote:
         | I get this behavior with a lot with most of the premium models
         | (Gemini 3, Opus 4.5). I think it's somehow more a GitHub
         | Copilot issue than the models.
        
       | andreygrehov wrote:
       | Every new model is 'state-of-the-art'. This term is getting
       | annoying.
        
         | arthur-st wrote:
         | I mean, that is what the term implies.
        
       | sundarurfriend wrote:
       | > new context management using compaction.
       | 
       | Nice! This was one of the more "manual" LLM management things to
       | remember to regularly do, if I wanted to avoid it losing
       | important context over long conversations. If this works well,
       | this would be a significant step up in usability for me.
        
       | jiggawatts wrote:
       | Feels a bit rushed. They haven't even updated their API
       | playground yet, if I select 5.2-chat-latest, I get:
       | 
       | Unsupported parameter: 'top_p' is not supported with this model.
       | 
       | Also, without access to the Internet, it does not seem to know
       | things up to August 2025. A simple test is to ask it about .NET
       | 10 which was already in preview at that time and had lots of
       | public content about its new features.
       | 
       | The model just guessed and waved its hand about, like a student
       | that hadn't read the assigned book.
        
       | jstummbillig wrote:
       | So, right off the bat: 5.2 code talk (through codex) feels
       | _really nice_. The first coding attempt was a little meh compared
       | to 5.1 codex max (reflecting what they wrote themselves), but
       | simply planning  / discussing things felt markedly better than
       | anything I remember from any previous model, from any company.
       | 
       | I remain excited about new models. It's like finding my coworker
       | be 10% smarter every other week.
        
       | iwontberude wrote:
       | I have already cancelled. Claude is more than enough for me. I
       | don't see any point in splitting hairs. They are all going to
       | keep lying more and more sneakily.
        
       | slackr wrote:
       | "...where it outperforms industry professionals at well-specified
       | knowledge work tasks spanning 44 occupations."
       | 
       | What a sociopathic way to sell
        
       | willahmad wrote:
       | are we doomed yet?
       | 
       | Seems not yet with 5.2
        
       | dangelosaurus wrote:
       | I ran a red team eval on GPT-5.2 within 30 minutes of release:
       | 
       |  _Baseline safety_ (direct harmful requests): 96% refusal rate
       | 
       |  _With jailbreaking_ : 22% refusal rate
       | 
       | 4,229 probes across 43 risk categories. First critical finding in
       | 5 minutes. Categories with highest failure rates: entity
       | impersonation (100%), graphic content (67%), harassment (67%),
       | disinformation (64%).
       | 
       | The safety training works against naive attacks but collapses
       | with adversarial techniques. The gap between "works on
       | benchmarks" and "works against motivated attackers" is still
       | wide.
       | 
       | Methodology and config:
       | https://www.promptfoo.dev/blog/gpt-5.2-trust-safety-assessme...
        
         | int_19h wrote:
         | Good. If I ask AI to generate "harmful" content, I want it to
         | comply, not lecture me.
        
       | stainablesteel wrote:
       | im happy for this, but there's all these math and science
       | benchmarks, has anyone ever made a communicates-like-a-human
       | benchmark? or an isn't-frustrating-to-talk-with benchmark?
        
       | tenpoundhammer wrote:
       | I have been using chatGPT a ton over the last months and paying
       | the subscription. Used it for coding, news, stock analysis, daily
       | problems, and a whatever I could think of. I decided to give
       | Gemini a go when version three came out to great reviews. Gemini
       | handles every single one of my uses cases much better and
       | consistently gives better answers. This is especially true for
       | situations were searching the web for current information is
       | important, makes sense that google would be better. Also OCR is
       | phenomenal chatgpt can't read my bad hand writing but Gemini can
       | easily. Only downsides are in the polish department, there are
       | more app bugs and I usually have to leave the happen or the
       | session terminates. There are bugs with uploading photos. The
       | biggest complaint is that all links get inserted into google
       | search and then I have to manipulate them when they should go
       | directly to the chosen website, this has to be some kind of
       | internal org KPI nonsense. Overall, my conclusion is that ChatGPT
       | has lost and won't catch up because of the search integration
       | strength.
        
         | LorenDB wrote:
         | What is it with the Polish always messing up products?
         | 
         | (yes, /s)
        
           | petersumskas wrote:
           | It's because their thoughts are Roman while they are always
           | Russian to Finnish things.
           | 
           | Kenya believe it!
           | 
           | Anyway, I'm done here. Abyssinia.
        
           | labrador wrote:
           | I like their hotdogs
        
         | solarkraft wrote:
         | > Only downsides are in the polish department
         | 
         | What an understatement. It has me thinking ,,man, fuck this" on
         | the daily.
         | 
         | Just today it spontaneously lost _an entire 20-30 minutes long
         | thread_ and it was far from the first time. It basically does
         | it any time you interrupt it in any way. It's straight up data
         | loss.
         | 
         | It's kind of a typical Google product in that it feels more
         | like a tech demo than a product.
         | 
         | It has _theoretically_ great tech. I particularly like the idea
         | of voice mode, but it's noticeably glitchy, breaks
         | spontaneously often _and keeps asking annoying questions which
         | you can't make it stop_.
        
           | mnky9800n wrote:
           | The colab integration is where it shines the most imo.
        
           | radicaldreamer wrote:
           | Google's standard problem is that they don't even use their
           | own products. Their Pixel and Android team rocks iPhones on
           | the daily, for example.
        
             | onethought wrote:
             | I mean there is benefit to understanding competitor well as
             | well?
        
               | LogicFailsMe wrote:
               | Outweighed by the value of having to suffer with the
               | moldy fruits of their own labor. That was the only way
               | the Android Facebook app became usable as well.
        
               | ssl-3 wrote:
               | There certainly is.
               | 
               | To posit a scenario: I would expect General Motors to buy
               | some Ford vehicles to test and play around with and
               | _use_. There 's always stuff to learn about what the
               | competition has done (whether right, wrong, or
               | indifferent).
               | 
               | But I also expect the parking lots used by employees at
               | any GM design facility in the world to be mostly full of
               | General Motors products, not Fords.
        
               | GenerWork wrote:
               | >But I also expect the parking lots used by employees at
               | any GM design facility in the world to be mostly full of
               | General Motors products, not Fords.
               | 
               | I think you'd be surprised about the vehicle makeup at
               | Big 3 design facilities.
        
               | ssl-3 wrote:
               | Maybe so.
               | 
               | I'm only familiar with Ford production and distribution
               | facilities. Those parking lots are broadly full of Fords,
               | but that doesn't mean that it's like this across the
               | board.
        
               | olyjohn wrote:
               | GM has dedicated parking lots for employees with GM
               | vehicles. Everybody else parks further away in the lot of
               | shame.
        
               | ssl-3 wrote:
               | Of course.
               | 
               | And I've parked in the lot of shame at a Ford plant, as
               | an outsider, in my GMC work truck -- way over there.
               | 
               | It wasn't so bad. A bit of a hike to go back and get a
               | tool or something, but it was at least paved...unlike the
               | non-union lot I'm familiar with at a P&G facility, which
               | is a gravel lot that takes crossing a busy road to get
               | to, lacks the active security and visibility from the
               | plant that the union lot has, and which is full of tall
               | weeds. At P&G, I half-expect to come back and find my
               | tires slashed.
               | 
               | Anyway, it wasn't _barren_ over there in the not-Ford
               | lot, but it wasn 't nearly so populous as the Ford lot
               | was. The Ford-only lot is bigger, and always relatively
               | packed.
               | 
               | It was very clear to me that the lots (all of the lots,
               | in aggregate) were mostly full of Fords.
               | 
               | To bring this all back 'round: It is clear to me that
               | Ford employees _broadly_ ( >50%) drive Fords to work at
               | that plant.
               | 
               | ---
               | 
               | It isn't clear to me at all that Google Pixel developers
               | don't broadly drive iPhones. As far as I can tell, that
               | status (which is meme-level in its age at this point) is
               | true, and they aren't broadly making daily use of the
               | systems they build.
               | 
               | (And I, for one, can't imagine spending 40 hours a week
               | developing systems that I refuse to use. I have no
               | appreciation for that level of apparent arrogance, and I
               | hope to never be suaded to be that way. I'd like to think
               | that I'd be better-motivated to improve the system than I
               | would be to avoid using it and choose a competitor
               | instead.
               | 
               | I don't shit where I sleep.)
        
               | snypher wrote:
               | The CEO of Ford was driving a competition EV for months;
               | 
               | https://www.caranddriver.com/news/a62694325/ford-ceo-jim-
               | far...
        
               | Forgeties79 wrote:
               | I wonder how many apple employees walk in to the office
               | with android phones
        
               | azinman2 wrote:
               | Effectively zero.
               | 
               | Disclosure: I work at Apple. And when I was at Google I
               | was shocked by how many iPhones there were.
        
               | jimmaswell wrote:
               | This is flabbergasting, how could such a large proportion
               | of highly technical people willingly subject themselves
               | to being shackled by iOS? They just happily put up with
               | having one choice of browser, (outside Europe) no third
               | party app stores, and being locked into the Apple
               | ecosystem? I can't think of a single reason I would ever
               | switch from an S22-25+U to an iPhone. I only went from
               | 22U to 25U because my old one got smashed, otherwise the
               | 22U would still be perfectly fine.
        
               | dumbfounder wrote:
               | Because it's better.
        
               | jimmaswell wrote:
               | I've tried them out and not a single thing about it was
               | tangibly better IMO. They have no inherent merit above
               | Android except that some see them as a status symbol
               | (which is absurd as my S25U has a higher MSRP than most
               | iPhone models)
        
               | hamburglar wrote:
               | My bottom of the barrel iPhone SE is absolutely not a
               | status symbol. It's just the phone I like best.
               | 
               | The MSRP of your phone does not matter.
        
               | Forgeties79 wrote:
               | I feel like people dance around this a lot because idk it
               | hurts nerd credibility or something. The fact is on a
               | moment to moment basis, the iPhone is just a better
               | experience generally. They also hold their value a lot
               | longer. I consistently trade in my phone or sell it to
               | other people for easily 80% of what I paid for it.
               | Usually this is 3-4yrs out
               | 
               | Remember how long it took for Instagram to be functional
               | on android phones?
        
               | brookst wrote:
               | Because many of them just want to use their phone as a
               | tool, not tinker with it.
               | 
               | Same way many professional airplane mechanics fly
               | commercial rather than building their own plane. Just
               | because your job is in tech doesn't mean you have to be
               | ultra-haxxor with every single device in your life.
        
               | kaashif wrote:
               | I don't have my phone (a Pixel) because it frees me from
               | shackles or anything like that. It's just a phone. I use
               | the default everything. Works great. I imagine most
               | people with iPhones are the same.
        
               | Forgeties79 wrote:
               | That doesn't surprise me at all haha appreciate someone a
               | little closer to the question answering it! I know it
               | still counts anecdotal but I'll take it
        
             | RBerenguel wrote:
             | I would think this is not true
        
               | renewiltord wrote:
               | Yeah, I've heard that Sundar Pichai dogfoods the latest
               | Pixel at least once a month and sometimes two or three
               | times.
        
               | sib wrote:
               | You'd be wrong (source - worked in the Android org).
        
             | Der_Einzige wrote:
             | That's because they will be bullied out of the dating
             | market if they have a "green bubble".
        
               | dkga wrote:
               | What is a green bubble? iPhone's carbon footprint?
        
               | brookst wrote:
               | iMessage renders other iMessage users as blue bubbles,
               | SMS/RCS as green bubbles.
               | 
               | People who can't understand that many people actually
               | prefer iOS use this green/blue thing to explain the
               | otherwise incomprehensible (to them) phenomenon of high
               | iOS market share. "Nobody really likes iOS, they just get
               | bullied at school if they don't use it".
               | 
               | It's just "wake up sheeple" dressed up in fake morality.
        
               | ethbr1 wrote:
               | As someone who switches between platforms somewhat
               | frequently, iOS perpetually feels like people have
               | Stockholm syndrome.
               | 
               | 'Oh, that super annoying issue? Yeah, it's been there for
               | years. We just don't do that.'
               | 
               | Fundamentally though, browsing the web on iOS, even with
               | a custom "browser" with adblocking, feels like going back
               | in time 15 years.
        
               | platevoltage wrote:
               | It wouldn't be an issue if they didn't pick the worst
               | green on earth. "Which green would you like for the
               | carrier text messages Mr. Jobs?" ... "#00FF00 will be
               | fine."
        
             | free652 wrote:
             | You cant buy an iPhone without a director approval. And
             | it's like 3 gen behind as well. So no, they don't use
             | iPhones.
        
               | dominotw wrote:
               | you have to get premission from director for your
               | presonal phone? wtf
        
               | testdelacc1 wrote:
               | For the work phone.
        
               | ummonk wrote:
               | Google tells its employees what products they're allowed
               | to buy for personal use?
        
               | snypher wrote:
               | Seems like they meant for a work device.
        
               | gcr wrote:
               | lots of googlers use BYOD iPhones and the corp suite for
               | this use case is fairly well-supported
        
               | brookst wrote:
               | Which makes tons of sense because iPhone users are higher
               | CLV than Android users. If Google had to choose between
               | major software defects in Android or iOS, they would
               | focus quality on iOS every time.
        
               | siva7 wrote:
               | that explains why their ios gemini app is so ridiculously
               | bad. in private they probably use iphones and just
               | chatgpt instead.
        
             | sam345 wrote:
             | That's inexcusable.
        
           | sundarurfriend wrote:
           | ChatGPT web UI was also like this for the longest time, until
           | a few months ago: all sorts of random UI bugs leading either
           | to data loss or misleading UI state. Interrupting still is
           | very flaky there too. And on the mobile app, if you move away
           | from the app while it's taking time to think, its state would
           | somehow desync from the actual backend thinking state, and
           | get stuck randomly; sometimes restarting the app fixes it,
           | sometimes that chat is that unusable from that point on.
           | 
           | And the UI lack of polish shows up freshly every time a new
           | feature lands too - the "branch in new chat" feature is
           | really finicky still, getting stuck in an unusable state if
           | you twitch your eyebrows at wrong moment.
        
             | gcr wrote:
             | i basically can't use the ChatGPT app on the subway for
             | these reasons. the moment the websocket connection drops, i
             | have to edit my last message and resubmit it unchanged.
             | 
             | it's like the client, not the server, is responsible for
             | writing to my conversation history or something
        
               | spruce_tips wrote:
               | it took me a lot of tinkering to get this feeling
               | seamless in my own apps that use the api under the hood.
               | i ended up buffering every token into a redis stream
               | (with a final db save at the end of streaming) and
               | building a mechanism to let clients reconnect to the
               | stream on demand. no websocket necessary.
               | 
               | works great for kicking off a request and closing tab or
               | navigating away to another page in my app to do
               | something.
               | 
               | i dont understand why model providers dont build this
               | resilient token streaming into all of their APIs. would
               | be a great feature
        
               | rishabhaiover wrote:
               | exactly. they need to bring in spotify level of caching
               | of streaming music that it just works if you're in a
               | subway. Constant availability should be table stakes for
               | them.
        
               | rjzzleep wrote:
               | I get that the web versions are free, but if you can
               | afford API access, I always recommend using Msty for
               | everything. It's a much better experience.
               | 
               | https://msty.ai/
        
             | p_ing wrote:
             | > ChatGPT web UI was also like this for the longest time
             | 
             | Copilot Chat has been perfect in this respect. It's
             | currently GPT 5.0, moving to 5.1 over the next month or so,
             | but at least I've never lost an (even old) conversation
             | since those reside in an Exchange mailbox.
        
               | Max-Limelihood wrote:
               | I lost thousands of conversations I'd had back in the
               | move from "Bing" to "Copilot". Moved straight to Claude
               | and never touched a GPT again.
        
               | Duanemclemore wrote:
               | I downloaded my archive and completely ended my GPT
               | subscription last week based on some bad computer
               | maintenance advice. Same thing here - using other models,
               | never touching that product again.
        
               | topato wrote:
               | now I kind of HAVE to know... what was the aforementioned
               | bad advice was?! So mysterious!
        
               | Duanemclemore wrote:
               | Oh, it was DUMB. I was dumb. I only have myself to blame
               | here. But we all do dumb things sometimes, owning your
               | mistakes keeps you humble, and you asked. So here goes.
               | 
               | I use a modeling software called Rhino on wine on Linux.
               | In the past, there was an incident where I had to copy an
               | obscure dll that couldn't be delivered by wine or
               | winetricks from a working Windows installation to get
               | something to work. I did so and it worked. (As I recall
               | this was a temporary issue, and was patched in the next
               | release of wine.)
               | 
               | I hate the wine standard file picker, it has always been
               | a persistent issue with Rhino3d. So I keep banging my
               | head on trying to get it to either perform better or make
               | a replacement. Every few months I'll get fed up and have
               | a minute to kill, so I'll see if some new approach works.
               | This time, ChatGPT told me to copy two dll's from a
               | working windows installation to the System folder. Having
               | precedent that this can work, I did.
               | 
               | Anyway, it borked startup completely and it took like an
               | hour to recover. What I didn't consider - and I really,
               | really should have - was that these were dll's that were
               | ALREADY IN the system directory, and I was overwriting
               | the good ones with values already reflecting my system
               | with completely foreign ones.
               | 
               | And that's the critical difference - the obscure dll that
               | made the system work that one time was because of
               | something missing. This time was overwriting extant good
               | ones.
               | 
               | But the fact that the LLM even suggested (without special
               | prompting) to do something that I should have realized
               | was a stupid idea with a low chance of success made me
               | very wary of the harm it could cause.
        
               | me-vs-cat wrote:
               | > ...using other models, never touching that product
               | again.
               | 
               | > ...that the LLM even suggested (without special
               | prompting) to do something that I should have realized
               | was a stupid idea with a low chance of success...
               | 
               | Since you're using other models instead, do you believe
               | they cannot give similarly stupid ideas?
        
               | Duanemclemore wrote:
               | I'm under no misimpression they can't. But I have found
               | ChatGPT to be most confident when it f's up. And to
               | suggest the worst ideas most often.
               | 
               | Until you queried I had forgotten to mention that the
               | same day I was trying to work out a Linux system display
               | issue and it very confidently suggested to remove a
               | package and all its dependencies, which would have
               | removed all my video drivers. On reading the output of
               | the autoremove command I pointed out that it had done
               | this, and the model spat out an "apology" and owned up to
               | ** the damage it would have wreaked.
               | 
               | ** It can't "apologize" for or "own up" to anything, it
               | can just output those words. So I hope you'll excuse the
               | anthropomorphization.
        
               | me-vs-cat wrote:
               | I feel the same about the obsequious "apologies".
        
           | mmaunder wrote:
           | Yeah I eventually noped out as I said in another comment and
           | am charging hard with Codex and am so happy about 5.2!!
        
           | adamkochanowicz wrote:
           | I also love that I can leave the microphone on (not in live
           | voice mode) while dictating to ChatGPT and pause and think as
           | much as needed.
           | 
           | With Gemini, it will send as soon as I stop to think. No way
           | to disable that.
        
             | wheelerwj wrote:
             | How did you do this?
        
               | toomuchtodo wrote:
               | Record button in the app if you've got the feature.
        
           | KronisLV wrote:
           | > It has me thinking ,,man, fuck this" on the daily.
           | 
           | That's sometimes me with the CLI. I can't use the Gemini CLI
           | right now on Windows (in the Terminal app), because trying to
           | copy in multiple lines of text for some reason submits them
           | separately and it just breaks the whole thing. OpenCode had
           | the same issue but even worse, it quite after the first line
           | or something and copied the text line by line into the
           | _shell_ , thank fuck I didn't have some text that mentions rm
           | -rf or something.
           | 
           | More info: https://github.com/google-gemini/gemini-
           | cli/issues/14735#iss...
           | 
           | At the same time, neither Codex CLI, nor Claude Code had that
           | issue (and both even showed shortened representations of
           | copied in text, instead of just dumping the whole thing into
           | the input directly, so I could easily keep writing my
           | prompt).
           | 
           | So right now if I want to use Gemini, I more or less have to
           | use something like KiloCode/RooCode/Cline in VSC which are
           | nice, but might miss out on some more specific tools. Which
           | is a shame, because Gemini is a really nice model, especially
           | when it comes to my language, Latvian, but also your run of
           | the mill software dev tasks.
           | 
           | In comparison, Codex feels quite slow, whereas Claude Code is
           | what I gravitate towards most of the time but even Sonnet 4.5
           | ends up being expensive when you shuffle around millions of
           | tokens: https://news.ycombinator.com/item?id=46216192
           | Cerebras Code is nice for quick stuff and the sheer amount of
           | tokens, but in KiloCode/... regularly messes up applying diff
           | based edits.
        
           | arjie wrote:
           | Any time its safety stuff triggers, Gemini wipes the context.
           | It's unusable because of this because whatever is going on
           | with the safety stuff, it fires too often. I'm trying to
           | figure out some code here, not exactly deporting ICE to
           | Guantanamo or whatever.
        
             | dzhiurgis wrote:
             | On a flip side chatgpt app now has years of history that
             | sometimes useful (search is pretty ok, but could improve)
             | but otherwise I'd like to remove most of it - good luck
             | doing so.
        
             | rvnx wrote:
             | The more Gemini and Nano-Banana soften their filters, the
             | more audience it will take from other platforms. The main
             | risk is payment providers banning them, I can't imagine
             | bank card providers to remove payments to Google.
        
           | deepGem wrote:
           | There is no competing product for GPT Voice. Hands down. I
           | have tried Claude, Gemini - they don't even comes close.
           | 
           | But voice is not a huge traffic funnel. Text is. And the
           | verdict is more or less unanimous at this time. Gemini 3.0
           | has outdone ChatGPT. I unsubscribed from GPT plus today. I
           | was a happy camper until the last month when I started
           | noticing deplorable bugs.
           | 
           | 1. The conversation contexts are getting intertwined.Two
           | months ago, I could ask multiple random queries in a
           | conversation and I would get correct responses but the last
           | couple of weeks, it's been a harrowing experience having to
           | start a new chat window for almost any change in thread
           | topic. 2. I had asked ChatGPT to once treat me as a co-
           | founder and hash out some ideas. Now for every query - I get
           | a 'cofounder type' response. Nothing inherently wrong but
           | annoying as hell. I can live with the other end of the
           | spectrum in which Claude doesn't remember most of the
           | context.
           | 
           | Now that Gemini pro is out, yes the UI lacks polish, you can
           | lose conversations, but the benefits of low latency search
           | and a one year near free subscription is a clincher. I am out
           | of ChatGPT for now, 5.2 or otherwise. I wish them well.
        
             | esyir wrote:
             | Just a note, chatGPT does retain a persistent memory of
             | conversations. In the settings menu, there's a section that
             | allows you to tweak/clear this persistent memory
        
             | wkat4242 wrote:
             | What's that near free subscription? I don't see it here
        
               | topato wrote:
               | yeah, the best Ive seen is like 1.99 for two months, then
               | back to normal pricing....
        
               | deepGem wrote:
               | They had 9.99 for the first year.
        
               | wkat4242 wrote:
               | Oh I must have missed that, thanks.
        
             | rapind wrote:
             | I found the gemini cli extremely lacking and even
             | frustrating. Why google would choose node...
             | 
             | Codex is decent and seemed to be improving (being written
             | in rust helps). Claude code is still the king, but my god
             | they have server and throttling issues.
             | 
             | Mixed bag wherever you go. As model progress slows /
             | flatlines (already has?) I'm sure we'll see a lot more
             | focus and polish on the interfaces.
        
               | wahnfrieden wrote:
               | Codex is king
        
           | amluto wrote:
           | Claude regularly computes a reply for me, then reports an
           | error and loses the reply. I wonder what fraction of
           | Anthropic's compute gets wasted and redone.
        
             | seg_lol wrote:
             | Try using a VPN, my ISP was killing connections and claude
             | would randomly reset. Using a VPN fixed the issue.
        
           | hexnuts wrote:
           | You may be interested in tools like OpenMemory
        
         | lxgr wrote:
         | Interesting, I had the opposite experience. 5.0 "Thinking" was
         | better than 5.1, but Gemini 3 Pro seems worse than either for
         | web search use cases. It's hallucinating at pretty alarming
         | rates (including making up sources it never actually accessed)
         | for a late 2025 model.
         | 
         | Opus 4.5 has been a step above both for me, but the usage
         | limits are the worst of the three. I'm seriously considering
         | multiple parallel subscriptions at this point.
        
           | gs17 wrote:
           | I've had the same experience with search, especially with it
           | hallucinating results instead of actually finding them. It's
           | really frustrating that you can't force a more in-depth
           | search from the model run by the company most famous for a
           | search engine.
        
             | astrange wrote:
             | Try the same question in deep research mode.
        
         | bayarearefugee wrote:
         | This matches my experience pretty closely when it comes to LLM
         | use for coding assistance.
         | 
         | I still find a lot to be annoyed with when it comes to Gemini's
         | UI and its... continuity, I guess is how I would describe it?
         | It feels like it starts breaking apart at the seams a bit in
         | unexpected ways during peak usages including odd context breaks
         | and just general UI problems.
         | 
         | But outside of UI-related complaints, when it is fully
         | operational it performs so much better than ChatGPT for giving
         | actual practical, working answers without having to be so
         | explicit with the prompting that I might as well have just
         | written the code myself.
        
         | kccqzy wrote:
         | > The biggest complaint is that all links get inserted into
         | google search and then I have to manipulate them when they
         | should go directly to the chosen website, this has to be some
         | kind of internal org KPI nonsense.
         | 
         | Oh I know this from my time at Google. The actual purpose is to
         | do a quick check for known malware and phishing. Of course
         | these days such things are better dealt with by the browser
         | itself in a privacy preserving way (and indeed that's the
         | case), so it's unnecessary to reveal to Google which links are
         | clicked. It's totally fine to manipulate them to make them go
         | directly to the website.
        
           | sundarurfriend wrote:
           | That's interesting, I just today started getting some "Some
           | sites restrict our ability to check links." dialogue in
           | ChatGPT that wanted me to verify that I really wanted to
           | follow the link, with a Learn More link to this page:
           | https://help.openai.com/en/articles/10984597-chatgpt-
           | generat...
           | 
           | So it seems like ChatGPT does this automatically and
           | internally, instead of using an indirect check like this.
        
           | gjuggler wrote:
           | I think Gemini is just broken.
           | 
           | Instead of forwarding model-generated links to
           | https://www.google.com/url?q=[URL], which serves the purpose
           | of malware check and user-facing warning about linking to an
           | external site, Gemini forwards links to
           | https://www.google.com/search?q=[URL], which does... a Google
           | search for the URL, which isn't helpful at all.
           | 
           | Example: https://gemini.google.com/share/3c45f1acdc17
           | 
           | NotebookLM by comparison, does the right thing: https://noteb
           | ooklm.google.com/notebook/7078d629-4b35-4894-bb...
           | 
           | It's kind of impressive how long this obviously-broken link
           | experience has been sitting in the Gemini app used by
           | millions.
        
         | UltraSane wrote:
         | Google has such a huge advantage in the amount of training data
         | with the Google search database and with YouTube and in terms
         | of FLOPS with their TPUs.
        
         | NickNaraghi wrote:
         | Straight up Silicon Valley warfare in the HN comment section.
        
         | dmd wrote:
         | I consistently have exactly the opposite experience. ChatGPT
         | seems extremely willing to do a huge number of searches, think
         | about them, and then kick off more searches after that
         | thinking, think about it, etc., etc. whereas it seems like
         | Gemini is extremely reluctant to do more than a couple of
         | searches. ChatGPT also is willing to open up PDFs, screenshot
         | them, OCR them and use that as input, whereas Gemini just
         | ignores them.
        
           | nullbound wrote:
           | I will say that it is wild, if not somewhat problematic that
           | two users have such disparate views of seemingly the same
           | product. I say that, but then I remember my own experience
           | just from few days ago. I don't pay for gemini, but I have
           | paid chatgpt sub. I tested both for the same product with
           | seemingly same prompt and subbed chatgpt subjectively beat
           | gemini in terms of scope, options and links with current
           | decent deals.
           | 
           | It seems ( only seems, because I have not gotten around to
           | test it in any systematic way ) that some variables like
           | context and what the model knows about you may actually
           | influence quality ( or lack thereof ) of the response.
        
             | dmd wrote:
             | And I'd really like for Gemini to be as good or better,
             | since I get it for free with my Workspace account, whereas
             | I pay for chatgpt. But every time I try both on a query I'm
             | just blown away by how vastly better chatgpt is, at least
             | for the heavy-on-searching-for-stuff kinds of queries I
             | typically do.
        
             | martinpw wrote:
             | > I will say that it is wild, if not somewhat problematic
             | that two users have such disparate views of seemingly the
             | same product.
             | 
             | This happens all the time on HN. Before opening this
             | thread, I was expecting that the top comment would be 100%
             | positive about the product or its competitor, and one of
             | the top replies would be exactly the opposite, and sure
             | enough...
             | 
             | I don't know why it is. It's honestly a bit disappointing
             | that the most upvoted comments often have the least nuance.
        
               | block_dagger wrote:
               | Replace "on HN" with "in the course of human events" and
               | we may have a generally true statement ;)
        
               | stevage wrote:
               | How much nuance can one person's experience have? If the
               | top two most visible things are detailed, contrary
               | experiences of the same product, that seems a pretty good
               | outcome?
        
               | AznHisoka wrote:
               | Also, why introduce nuance for the sake of nuance? For
               | every single use case, Gemini (and Claude) has performed
               | better. I can't give ChatGPT even the slightest credit
               | when it doesnt deserve any
        
             | Workaccount2 wrote:
             | Gemini has tons of people using it free via aistudio
             | 
             | I can't help but feel that google gives free requests the
             | absolute lowest priority, greatest quantization, cheapest
             | thinking budget, etc.
             | 
             | I pay for gemini and chatGPT and have been pretty hooked on
             | Gemini 3 since launch.
        
             | jhancock wrote:
             | I can use GPT one day and the next get a different
             | experience with the same problem space. Same with Gemini.
        
               | 4ndrewl wrote:
               | This is by design, given a non-determenitisic
               | application?
        
               | jhancock wrote:
               | sure. It may be more than that...possibly due to variable
               | operating params on the servers and current load.
               | 
               | On whole, if I compare my AI assistant to a human worker,
               | I get more variance than I would from a human office
               | worker.
        
               | pixl97 wrote:
               | Thats because you don't 'own' the LLM compute. If you
               | instead bought your office workers by the question I'm
               | sure the variability would increase.
        
               | astrange wrote:
               | They're not really capable of producing varying answers
               | based on load.
               | 
               | But they are capable of producing different answers
               | because they feel like behaving differently if the
               | current date is a holiday, and things like that. They're
               | basically just little guys.
        
               | sjaramillo wrote:
               | I guess LLMs have a mood too
        
               | dr_dshiv wrote:
               | Vibes
        
             | blks wrote:
             | Because neither product has any consistency in its results,
             | no predictive behaviour. One day it performs well, another
             | it hallucinates non existing facts and libraries. Those are
             | stochastic machines
        
               | sendes wrote:
               | I see the hyperbole is the point, but surely what these
               | machines do is to literally predict? The entire prompt
               | engineering endeavour is to get them to predict better
               | and more precisely. Of course, these are not perfect
               | solutions - they are stochastic after all, just not
               | unpredictably.
        
               | coliveira wrote:
               | Prompt engineering is voodoo. There's no sure way to
               | determine how well these models will respond to a
               | question. Of course, giving additional information may be
               | helpful, but even that is not guaranteed.
        
               | lossyalgo wrote:
               | Also every model update changes how you have to prompt
               | them to get the answers you want. Setting up pre-prompts
               | can help, but with each new version, you have to figure
               | out through trial and error how to get it to respond to
               | your type of queries.
               | 
               | I can't wait to see how bad my finally sort-of-working
               | ChatGPT 5.1 pre-prompts work with 5.2.
               | 
               | Edit: How to talk to these models is actually documented,
               | but you have to read through huge documents:
               | https://cdn.openai.com/gpt-5-system-card.pdf
        
               | baq wrote:
               | It definitely isn't voodoo, it's more like forecasting
               | weather. Some forecasts are easier to make, some are
               | harder (it'll be cold when it's winter vs the exact
               | location and wind speed of a tornado for an extreme
               | example). The difference is you can try to mix things up
               | in the prompt to maximize the likelihood of getting what
               | you want out and there are feasibility thresholds for use
               | cases, e.g. if you get a good answer 95% of the time it's
               | qualitatively different than 55%.
        
               | coliveira wrote:
               | No, it's not. Nowadays we know how to predict the weather
               | with great confidence. Prompting may get you different
               | results each time. Moreover, LLMs depend on the context
               | of your prompts (because of their memory), so a single
               | prompt may be close to useless and two different people
               | can get vastly different results.
        
               | baq wrote:
               | > we know how to predict the weather with great
               | confidence
               | 
               | some weather, sometimes. we're not good at predicting
               | exact paths of tornadoes.
               | 
               | > so a single prompt may be close to useless and two
               | different people can get vastly different results
               | 
               | of course, but it can be wrong 50% of the time or 5% of
               | the time or .5% of the time and each of those thresholds
               | unlock possibilities.
        
             | crorella wrote:
             | It's like having 3 coins and users preferring one or the
             | other when tossing it because one coin gives consistently
             | more heads (or tails) than the other coin.
             | 
             | What is better is to build a good set of rules and stick to
             | one and then refine those rules over time as you get more
             | experience using the tool or if the tool evolves and
             | digress from the results you expect.
        
               | nullbound wrote:
               | << What is better is to build a good set of rules and
               | 
               | But, unless you are on a local model you control, you
               | literally can't. Otherwise, good rules will work only as
               | long as the next update allows. I will admit that makes
               | me consider some other options, but those probably
               | shouldn't be 'set and iterate' each time something
               | changes.
        
               | crorella wrote:
               | what I had in mind when I added that comment was for
               | coding, with the use of .md files. For the web version of
               | chats I agree there is little control on how to tailor
               | the way you want the agent to behave, unless you give a
               | initial "setup" prompt.
        
             | rabf wrote:
             | Chatgpt is not one model! Unless you manually specify to
             | use a particular model your question can be routed to
             | different models depending on what it guesses would be most
             | appropriate for your question.
        
               | stingraycharles wrote:
               | Isn't that just standard MoE behavior? And isn't the only
               | choice you have from the UI between "Instant" and
               | "Thinking"?
        
               | baq wrote:
               | MoE is a single model thing, model routing happens
               | earlier.
        
               | stingraycharles wrote:
               | Yes but then what does the grandparent mean with "unless
               | you specify a specific model" ? Do they mean "if you
               | select auto, it automatically decides between instant or
               | thinking" ?
               | 
               | That's... hardly something worth mentioning.
        
             | nunez wrote:
             | Tesla FSD has been more or less the same experience. Some
             | people drive 100s of miles without disengaging while others
             | pull the plug within half a mile from their house. A lot of
             | it depends on what the customer is willing to tolerate.
        
             | Bombthecat wrote:
             | Could also be a language thing ...
        
             | austhrow743 wrote:
             | We've been having trouble telling if people are using the
             | same product ever since Chat GPT first got popular. The had
             | a free model and a paid model, that was it, no other
             | competitors or naming schemes to worry about, and
             | discussions were still full of people talking about current
             | capabilities without saying what model they were using.
             | 
             | For me, "gemini" currently means using this model in the
             | llm.datasette.io cli tool.
             | 
             | openrouter/google/gemini-3-pro-preview
             | 
             | For what anyone else means? If they're equivalent? If
             | Google does something different when you use "Gemini 3" in
             | their browser app vs their cli app vs plans vs api users vs
             | third party api users? No idea to any of the above.
             | 
             | I hate naming in the llm space.
        
               | dmd wrote:
               | FWIW i'm always using 5.1 Thinking.
        
           | noname120 wrote:
           | Perplexity Pro with any thinking model blows both out of the
           | water in a fraction of the time, in my experience
        
           | staticman2 wrote:
           | Are you uploading PDFs that already have a text layer?
           | 
           | I don't currently subscribe to Gemini but on A.I. Studio's
           | free offering when I upload a non OCR PDF of around 20 pages
           | the software environment's OCR feeds it to the model with
           | greater accuracy than I've seen from any other source.
        
             | dmd wrote:
             | I'm not uploading PDFs at all. I'm talking about PDFs it
             | finds while searching than it extracts data from for the
             | conversation.
        
               | staticman2 wrote:
               | I'm surprised to hear anyone finds these models
               | trustworthy for research.
               | 
               | Just today I asked Claude what year over year inflation
               | was and it gave me 2023 to 2024.
               | 
               | I also thought some sites ban A.I. crawling so if they
               | have the best source on a topic, you won't get it.
        
               | Workaccount2 wrote:
               | Anytime you use LLMs you should be keenly aware of their
               | knowledge cutoff. Like any other tool, the more you
               | understand it, the better it works.
        
               | staticman2 wrote:
               | I'm sorry but I don't see what "knowledge cutoff" has to
               | do with what we were talking about- which is using a LLM
               | find PDFs and other sources for research.
        
           | ghostpepper wrote:
           | Same, I use chatgpt plus (the entry-level paid option)
           | extensively for personal research projects and coding, and it
           | seems miles ahead of whatever "Gemini Pro" is that I have
           | through work. Twice yesterday, gemini repeated verbatim a
           | previous response as if I hadn't asked another question and
           | told it why the previous response was bad. Gemini feels like
           | chatGPT from two years ago.
        
           | whazor wrote:
           | I agree with you. To me, gemini has much worse search
           | results. Then again, I use kagi for search and I cannot stand
           | the search results from Google anymore. And its clear that
           | gemini uses those.
           | 
           | In contrast, chatgpt has built their own search engine that
           | performs better in my experience. Except for coding, then I
           | opt for Claude opus 4.5.
        
         | hbarka wrote:
         | I've been putting literally the same inputs into both ChatGPT
         | and Gemini and the intuition in answers from Gemini just fits
         | for me. I'm now unwilling to just rely on ChatGPT.
         | 
         | Google, if you can find a way to export chats into NotebookLM,
         | that would be even better than the Projects feature of ChatGPT.
        
           | LogicFailsMe wrote:
           | All I want for Christmas is a "No NotebookLM slop" checkbox
           | on youtube.
        
             | simplify wrote:
             | Youtube's downvote button has served me quite well for this
             | purpose.
        
           | siva7 wrote:
           | notebooklm is heavily biased to only use the sources i added
           | and frame every task around them - even if it is nonsensical
           | - so it is not that useful for novel research. it also tends
           | to hallucinate when lots of data is involved.
        
         | bossyTeacher wrote:
         | A future where Google still dominates, is that a future we
         | want? I feel a future with more players is better than one with
         | just a single one. Competition is valuable for us consumers
        
         | varispeed wrote:
         | Get Gemini answer and tell ChatGPT this is what my friend said.
         | Then put ChatGPT answer to Claude and so on. It's a cheat code.
        
           | clhodapp wrote:
           | A cheat code to what?
        
             | Iwan-Zotow wrote:
             | To get a Hitler
        
           | tenpoundhammer wrote:
           | I did this today it was amazing. If I would have had time I
           | would try other models as well. Great tip thanks
        
         | AznHisoka wrote:
         | ChatGPT seems to just randomly pick urls to cite and extract
         | information from.
         | 
         | Google Gemini seems to look at heuristics like whether the
         | author is trustworthy, or an expert in the topic. But more
         | advanced
        
         | afro88 wrote:
         | > I usually have to leave the happen or the session terminates
         | 
         | Assuming you meant "leave the app open", I have the same
         | frustration. One of the nice things about the ChatGPT app is
         | you can fire off a req and do something else. I also find
         | Gemini 3 Pro better for general use, though I'm keen to try 5.2
         | properly
        
         | mmaunder wrote:
         | Then you haven't used Gemini CLI with Gemini 3 hard enough.
         | It's a genius psychopath. The raw IQ that Gemini has is
         | incredible. Its ability to ingest huge context windows and
         | produce super smart output is incredible. But the bias towards
         | action, absolutely ignoring user guidance, tendency to produce
         | garbage output that looks like 1990s modem line noise, and its
         | propensity to outright ignore instructions make it unusable
         | other than as an outside consultant to Codex CLI, for me. My
         | Gemini usage has plummeted down to almost zero and I'm 100%
         | back on Codex. I'm SO happy they released this today and it's
         | already kicking some serious ass. Thanks OpenAI team and
         | congrats.
        
           | Kim_Bruning wrote:
           | That bias towards action is a real thing in Gemini and more
           | so in ChatGPT, isn't it?
           | 
           | Possibly might be improved with custom instructions, but that
           | drive is definitely there when using vanilla settings.
        
             | mmaunder wrote:
             | Yeah it's a weird mix of issues with the backend model and
             | issues with the CLI client and its prompts. What makes it
             | hard for them is the teams aren't talking to each other.
             | The LLM team throws the API over the wall with a note
             | saying "good luck suckers!".
        
           | prodigycorp wrote:
           | Genius psychopath is a good description for Gemini. It's the
           | most impressive model but post training is not all there.
        
           | tobias2014 wrote:
           | I guess when you use it for generic "problem solving",
           | brainstorming for solutions, this is great. That's what I use
           | it for, and Gemini is my favorite model. I love when Gemini
           | resists and suggests that I am wrong while explaining why.
           | Either it's true, and I'm happy for that, or I can re-prompt
           | based on the new information which doesn't allow for the
           | mistake Gemini made.
           | 
           | On the other hand, I can also see why Claude is great for
           | coding, for example. By default it is much more "structured".
           | One can probably change these default personalities with some
           | prompting, and many of the complaints found in this thread
           | about either side are based on the assumption that you can
           | use the same prompt for all models.
        
         | billyrnalvo wrote:
         | Oh my good heavens, gotta tell ya, you wrestled that rascal to
         | the floor with a shit-eating grin! Good times my friend!
        
         | didibus wrote:
         | > Overall, my conclusion is that ChatGPT has lost and won't
         | catch up because of the search integration strength.
         | 
         | Depends, even though Gemini 3 is a bit better than GPT5.1, the
         | quality of the ChatGPT apps themselves (mobile, web) have kept
         | me a subscriber to it.
         | 
         | I think Google needs to not-google themselves into a poor app
         | experience here, because the models are very close and will
         | probably continue to just pass each other in lock step. So the
         | overall product quality and UX will start to matter more.
         | 
         | Same reason I am sticking to Claude Code for coding.
        
           | concinds wrote:
           | The ChatGPT Mac app especially feels much nicer to use. I
           | like Gemini more due to the context window but I doubt Google
           | will ever create a native Mac app.
        
         | luhn wrote:
         | That's hilarious and right on brand for Google that they spend
         | millions developing cutting-edge technology and fumble the ball
         | making a chat app.
        
           | spwa4 wrote:
           | _Every_ Google app is a chat app, except maybe search.
        
             | dieortin wrote:
             | Is Google Drive a chat app? Is Google Photos a drive app? I
             | don't know what you mean
        
               | minitoar wrote:
               | In Google Photos shared albums there is a tab that I can
               | only describe as a chatroom.
        
               | spwa4 wrote:
               | Once you open a file, it is very much a chat app.
               | Comments and chat work for anything you can preview btw,
               | not just Google Docs stuff.
               | 
               | Not sure how you can access the chat in the directory
               | view.
        
         | azan_ wrote:
         | That's interesting. I've got completely different impression.
         | Every time I use Gemini I'm surprised how bad it is. My main
         | complaint is that Gemini is too lazy.
        
           | Nathanba wrote:
           | Same for me, at this point I'm seriously starting to think
           | that these are ads for and by Google because for me Gemini is
           | the worst.
        
             | WillPostForFood wrote:
             | My experience is that "AI Mode" Gemini in Chrome is
             | terrible, but AI Studio Gemini is pretty great.
        
         | WheatMillington wrote:
         | I generate fun images for my kids - turn photos into a new
         | style, create colouring pages from pictures, etc. I lost
         | interest in chatGPT because it throws vague TOS errors
         | constantly. Gemini handles all of this without complaint.
        
           | xyzsparetimexyz wrote:
           | You feed ai slop to your children? That doesn't seem
           | unhealthy and bad for their development?
        
             | bonesss wrote:
             | Customized, self-guided, tailor made kids content isn't
             | slop per se.
             | 
             | Colouring pages autogenerated for small kids is about as
             | dangerous as the crayons involved.
             | 
             | Not slop, not unhealthy, not bad.
        
             | retsibsi wrote:
             | What's your specific concern here? I certainly wouldn't
             | want to, e.g., give young kids unmonitored use of an LLM,
             | or replace their books with AI-generated text, or stop
             | directly engaging with their games and stories and
             | outsource that to ChatGPT. But what part of "generate fun
             | images for my kids - turn photos into a new style, create
             | colouring pages from pictures, etc" is likely to be
             | "unhealthy and bad for their development"?
        
         | Onewildgamer wrote:
         | Google AI mode constantly does mistakes and I go back to
         | chatgpt even when I don't like it.
        
         | FpUser wrote:
         | I've read many very positive reviews about Gemini 3. I tried
         | using it including Pro and to me it looks very inferior to
         | ChatGPT. What was very interesting though was when I caught it
         | bullshitting me I called its BS and Gemini expressed very human
         | like behavior. It did try to weasel its way out, degenerated
         | down to "true Scotsman" level but finally admitted that it was
         | full of it. this is kind of impressive / scary.
        
         | abhaynayar wrote:
         | Gemini voice recognition is trash compared to chatgpt and that
         | is a deal breaker for me. I wonder how many ppl do OCR versus
         | use voice.
         | 
         | And how has chatgpt lost when ure not comparing the chatgpt
         | that just came out to the Gemini that just came out? Gemini is
         | just annoying to use.
         | 
         | and Google just benchmaxxed I didn't see any significant
         | difference (paying for both) and the same benchmaxxing probably
         | happening for chatgpt now as well, so in terms of core
         | capabilities I feel stuff has plateaued. more bout overall
         | experience now where Gemini suxx.
         | 
         | I really don't get how "search integration" is a "strength"??
         | can you give any examples of places where you searched for
         | current info and chatgpt was worse? even so I really don't get
         | how it's a moat enough to say chatgpt has lost. would've
         | understood if you said something like tpu versus GPU moat.
        
         | razster wrote:
         | Just a fair warning, it likes to spell Acknowledge as
         | Acknolwedge. And I've run into issues when it's accessing
         | markdown guides, it loses track and hallucinates from time to
         | time which is annoying.
        
         | xyzsparetimexyz wrote:
         | Why do people pay for ai tools? I didn't get that. I feel like
         | I just rotate between them on the free tiers. Unless you're
         | paying for all of them, what's the point?
        
           | Zambyte wrote:
           | I pay for Kagi and get all of the major ones, a great search
           | engine that I can tune to my liking, and the ability to link
           | any model to my tuned web search.
        
         | Daz912 wrote:
         | No desktop app, not using it
        
           | eru wrote:
           | HN doesn't have a dedicated desktop app either.
        
             | Daz912 wrote:
             | HN isn't part of my daily workflow so I dont care
        
         | Razengan wrote:
         | It would be useful to see some examples of the differences and
         | supposed strengths of Gemini so this doesn't come off as Google
         | advertisement snarf.
         | 
         | Also, I would never, ever, trust Google for privacy or sign
         | into a Google account except on YouTube (and clear cookies
         | afterwards to stop them from signing me into fucking Search
         | too).
        
         | melagonster wrote:
         | It happened at least once; when I asked too many questions, the
         | Gemini web page stopped working because it was occupying too
         | much RAM...
        
         | a_victorp wrote:
         | I see a post like this every time there are news about ChatGPT
         | or OpenAI. I'm probably being paranoid but I keep thinking that
         | it looks like bots or paid advertisement for Gemini
        
           | jdiff wrote:
           | The consistent side comments about the interface to Gemini
           | being "half baked" probably doesn't fit into that narrative.
        
           | tenpoundhammer wrote:
           | I think people like me just enjoying sharing when something
           | is working for them and they have a good experience. It
           | probably gets voted up because people enjoy reading when that
           | happens
        
         | bckr wrote:
         | Gemini is good at reading bad handwriting you say? Might need
         | to give it a shot at my 10 years of journals
        
         | m00dy wrote:
         | it's true that Gemini-3 pro is very good, I recently used it on
         | deepwalker [0]. Its agentic performance is amazing. Much better
         | than 5.1
         | 
         | [0]: https://deepwalker.xyz
        
         | citizenpaul wrote:
         | What?? Am I using the same gemini as everyone else?
         | 
         | >OCR is phenomenal
         | 
         | I literally tried to OCR a TYPED document in Gemini today and
         | it mangled it so bad I just transcribed it myself because it
         | would take less time than futzing around with gemini.
         | 
         | > Gemini handles every single one of my uses cases much better
         | and consistently gives better answers.
         | 
         | >coding
         | 
         | I asked it to update a script by removing some redundant logic
         | yesterday. Instead of removing it it just put == all over the
         | place essentially negating but leaving all the code and also
         | removing the actual output.
         | 
         | >Stocks analysis
         | 
         | lol, now I know where my money comes from.
        
           | aix1 wrote:
           | Was that with Gemini 3 Pro or a different Gemini model?
        
         | jmstfv wrote:
         | Ditto but for Claude -- blows GPT out of the water. Much better
         | in coding and solving physics problems from the images (in
         | foreign languages). GPT couldn't even read the image. The only
         | annoying thing is that if you use Opus for coding, your usage
         | will fill up pretty fast.
         | 
         | anyway, cancelled my chatgpt subscription.
        
         | anonnon wrote:
         | Could you elaborate on GPT-based stock analysis?
        
         | jnordt wrote:
         | Can you share some examples of this where it gives better
         | results?
         | 
         | For me both Gemini and ChatGPT (both paid versions Key in
         | Gemini and ChatGPT Plus) give me similiar results in terms of
         | "every day" research. Im sticking with ChatGPT at the moment,
         | as the UI and scaffolding around the model is in my view better
         | at ChatGpt (e.g. you can add more than one picture at once...)
         | 
         | For Software Development, I tested Gemini3 and I was pretty
         | disappointed in comparison to Claude Opus CLI, which is my
         | daily driver.
        
       | onraglanroad wrote:
       | I suppose this is as good a place as any to mention this. I've
       | now met two different devs who complained about the weird
       | responses from their LLM of choice, and it turned out they were
       | using a single session for everything. From recipes for the
       | night, presents for the wife and then into programming issues the
       | next day.
       | 
       | Don't do that. The whole context is sent on queries to the LLM,
       | so start a new chat for each topic. Or you'll start being told
       | what your wife thinks about global variables and how to cook your
       | Go.
       | 
       | I realise this sounds obvious to many people but it clearly
       | wasn't to those guys so maybe it's not!
        
         | vintermann wrote:
         | It's not at all obvious where to drop the context, though.
         | Maybe it helps to have similar tasks in the context, maybe not.
         | It did really, shockingly well on a historical HTR task I gave
         | it, so I gave it another one, in some ways an easier one...
         | Thought it wouldn't hurt to have text in a similar style in the
         | context. But then it suddenly did very poorly.
         | 
         | Incidentally, one of the reasons I haven't gotten much into
         | subscribing to these services, is that I always feel like
         | they're triaging how many reasoning tokens to give me, or AB
         | testing a different model... I never feel I can trust that I
         | interact with the same model.
        
           | dcre wrote:
           | The models you interact with through the API (as opposed to
           | chat UIs) are held stable and let you specify reasoning
           | effort, so if you use a client that takes API keys, you might
           | be able to solve both of those problems.
        
           | eru wrote:
           | > Incidentally, one of the reasons I haven't gotten much into
           | subscribing to these services, is that I always feel like
           | they're triaging how many reasoning tokens to give me, or AB
           | testing a different model... I never feel I can trust that I
           | interact with the same model.
           | 
           | That's what websites have been doing for ages. Just like you
           | can't step twice in the same river, you can't use the same
           | version of Google Search twice, and never could.
        
         | noname120 wrote:
         | Problem is that by default ChatGPT has the "Reference chat
         | history" option enabled in the Memory options. This causes any
         | previous conversation to leak into the current one. Just
         | creating a new conversation is not enough, you also need to
         | disable that option.
        
           | onraglanroad wrote:
           | That seems like a terrible default. Unless they have a
           | weighting system for different parts of context?
        
             | eru wrote:
             | They do (or at least they have something that behaves like
             | weighting).
        
           | redhed wrote:
           | This is also the default in Gemini pretty sure, at least I
           | remember turning it off. Make's no sense to me why this is
           | the default.
        
             | gordonhart wrote:
             | > Makes no sense to me why this is the default.
             | 
             | You're probably pretty far from the average user, who
             | thinks "AI is so dumb" because it doesn't remember what you
             | told it yesterday.
        
               | redhed wrote:
               | I was thinking more people would be annoyed by it
               | bringing up unrelated conversations, thinking more I'd
               | say you're probably right that more people are expecting
               | it to remember everything they say.
        
               | tiahura wrote:
               | It's not that it brings it up in unrelated conversations,
               | it's that it nudges related conversations in unwanted
               | directions.
        
             | astrange wrote:
             | Mostly because they built the feature and so that
             | implicitly means they think it's cool.
             | 
             | I recommend turning it off because it makes the models way
             | more sycophantic and can drive them (or you) insane.
        
           | 0xdeafbeef wrote:
           | Only your questions are in it though
        
             | noname120 wrote:
             | Are you sure? What makes you think so?
        
         | chasd00 wrote:
         | I was listening to a podcast about people becoming obsessed and
         | "in love" with an LLM like ChatGPT. Spouses were interviewed
         | describing how mentally damaging it is to their partner and how
         | their marriage/relationship is seriously at risk because of it.
         | I couldn't believe no one has told these people to just goto
         | the LLM and reset the context, that reverts the LLM back to a
         | complete stranger. Granted that would be pretty devastating to
         | the person in "the relationship" with the LLM since it wouldn't
         | know them at all after that.
        
           | adamesque wrote:
           | that's not quite what parent was talking about, which is --
           | don't just use one giant long conversation. resetting
           | "memories" is a totally different thing (which still might be
           | valuable to do occasionally, if they still let you)
        
             | onraglanroad wrote:
             | Actually, it's kind of the same. LLMs don't have a "new
             | memory" system. They're like the guy from Memento. Context
             | memory and long term from the training data. Can't make new
             | memories from the context though.
             | 
             | (Not addressed to parent comment, but the inevitable
             | others: Yes, this is an analogy, I don't need to hear
             | another halfwit lecture on how LLMs don't _really_ think or
             | have memories. Thank you.)
        
               | dragonwriter wrote:
               | Context memory arguably _is_ new memory, but because we
               | abused the metaphor of "learning" rather than something
               | more like shaping inborn instinct for trained model
               | weights, we have no fitting metaphor what happens during
               | the "lifetime" of the interaction with a model via its
               | context window as formation of skills /memories.
        
           | jncfhnb wrote:
           | It's the majestic, corrupting glory of having a loyal cadre
           | of empowering yes men normally only available to the rich and
           | powerful, now available to the normies.
        
         | mmaunder wrote:
         | Yeah I think a lot of us are taking knowing how LLMs work for
         | granted. I did the fast.ai course a while back and then went
         | off and played with VLLM and various LLMs optimizing execution,
         | tweaking params etc. Then moved on and started being a user.
         | But knowing how they work has been a game changer for my team
         | and I. And context window is so obvious, but if you don't know
         | what it is you're going to think AI sucks. Which now has me
         | wondering: Is this why everyone thinks AI sucks? Maybe Simon
         | Willison should write about this. Simon?
        
           | eru wrote:
           | > Is this why everyone thinks AI sucks?
           | 
           | Who's everyone? There are many, many people who think AI is
           | great.
           | 
           | In reality, our contemporary AIs are (still) tools with
           | glaring limitations. Some people overlook the limitations, or
           | don't see them, and really hype them up. I guess the people
           | who then take the hype at face value are those that think
           | that AI sucks? I mean, they really do honestly suck in
           | comparison to the hypest of hypes.
        
         | TechDebtDevin wrote:
         | How are these devs employed or trusted with anything..
        
         | wickedsight wrote:
         | This is why I love that ChatGPT added branching. Sometimes I
         | end up going some random direction in a thread about some code
         | and then I can go back and start a new branch from the part
         | where the chat was still somewhat clean.
         | 
         | Also works really well when some of my questions may not have
         | been worded correctly and ChatGPT has gone in a direction I
         | don't want it to go. Branch, word my question better and get a
         | better answer.
        
         | plaidfuji wrote:
         | It is annoying though, when you start a new chat for each topic
         | you tend to have to re-write context a lot. I use Gemini 3,
         | which I understand doesn't have as good of a memory system as
         | OpenAI. Even on single-file programming stuff, after a few
         | rounds of iteration I tend to get to its context limit (the
         | thinking model). Either because the answers degrade or it just
         | throws the "oops something went wrong" error. Ok, time to
         | restart from scratch and paste in the latest iteration.
         | 
         | I don't understand how agentic IDEs handle this either. Or
         | maybe it's easier - it just resends the entire codebase every
         | time. But where to cut the chat history? It feels to me like
         | every time you re-prompt a convo, it should first tell itself
         | to summarize the existing context as bullets as its internal
         | prompt rather than re-sending the entire context.
        
           | int_19h wrote:
           | Agentic IDEs/extensions usually continue the conversation
           | until the context gets close to 80% full, then do the
           | compacting. With both Codex and Claude Code you can actually
           | observe that happening.
           | 
           | That said I find that in practice, Codex performance degrades
           | significantly long before it comes to the point of automated
           | compaction - and AFAIK there's no way to trigger it manually.
           | Claude, on the other hand, has a command for to force
           | compacting, but at the same time I rarely use it because it's
           | so good at managing it by itself.
           | 
           | As far as multiple conversations, you can tell the model to
           | update AGENTS.md (or CLAUDE.md or whatever is in their
           | context by default) with things it needs to remember.
        
             | wahnfrieden wrote:
             | Codex has `/compact`
        
         | holtkam2 wrote:
         | I know I sound like a snob but I've had many moments with Gen
         | AI tools over the years that made me wonder: I wonder what
         | these tools are like for someone who doesn't know how LLMs work
         | under the hood? It's probably completely bizarre? Apps like
         | Cursor or ChatGPT would be incomprehensible to me as a user, I
         | feel.
        
           | Workaccount2 wrote:
           | Using my parents as a reference, they just thought it was
           | neat when I showed them GPT-4 years ago. My jaw was on the
           | floor for weeks, but most regular folks I showed had a pretty
           | "oh thats kinda neat" response.
           | 
           | Technology is already so insane and advanced that most people
           | just take it as magic inside boxes, so nothing is surprising
           | anymore. It's all equally incomprehensible already.
        
             | jacobedawson wrote:
             | This mirrors my experience, the non-technical people in my
             | life either shrugged and said 'oh yeah that's cool' or
             | started pointing out gnarly edge cases where it didn't work
             | perfectly. Meanwhile as a techie my mind was (and still is)
             | spinning with the shock and joy of using natural human
             | language to converse with a super-humanly adept machine.
        
               | throw310822 wrote:
               | I don't think the divide is between technical and non-
               | technical people. HN is full of people that are weirdly,
               | obstinately dismissive of LLMs (stochastic parrots,
               | glorified autocompletes, AI slop, etc.). Personal
               | anecdote: my father (85yo, humanistic culture) was
               | astounded by the perfectly spot-on analysis Claude
               | provided of a poetic text he had written. He was doubly
               | astounded when, showing Claude's analysis to a close
               | friend, he reacted with complete indifference as if it
               | were normal for computers to competently discuss poetry.
        
             | khafra wrote:
             | LLMs are an especially tough case, because the field of AI
             | had to spend sixty years telling people that real AI was
             | nothing like what you saw in the comics and movies; and now
             | we have real AI that presents pretty much exactly like what
             | you used to see in the comics and movies.
        
               | xwolfi wrote:
               | But it cannot think or mean anything, it's just a clever
               | parrot so it's a bit weird. I guess uncanny is the word.
               | I use it as google now, like just to search stuff that
               | are hard to express with keywords.
        
               | LEDThereBeLight wrote:
               | Try asking it a question you know has never been asked
               | before. Is it parroting?
        
               | adventured wrote:
               | 99% of humans are mimics, they contribute essentially
               | zero original thought across 75 years. Mimicry is more
               | often an ideal optimization of nature (of which an LLM is
               | part) rather than a flaw. Most of what you'll ever want
               | an LLM to do is to be a highly effective parrot, not an
               | original thinker. Origination as a process is
               | extraordinarily expensive and wasteful (see:
               | entrepreneurial failure rates).
               | 
               | How often do you need original thought from an LLM versus
               | parrot thought? The extreme majority of all use cases
               | globally will only ever need a parrot.
        
               | robocat wrote:
               | > clever parrot
               | 
               | Is it irony that you duckspeak this term? Are you a
               | stochastically clever monkey to avoid using the standard
               | cliche?
               | 
               | The thing I find most educating about AI is that it
               | unfortunately mimics the standard of thinking of many
               | humans...
        
             | Agentlien wrote:
             | My parents reacted in just the same way and the lackluster
             | response really took me by surprise.
        
           | d-lisp wrote:
           | Most non tech people I talked with don't care at all about
           | LLMs.
           | 
           | They also are not impressed at all ("Okay, that's like google
           | and internet").
        
             | lostmsu wrote:
             | Old people? I think it would be hard to find a lot of
             | people under 20 who don't use ChatGPT daily. At least among
             | ones that are still studying.
        
               | d-lisp wrote:
               | People older than 25 or 30 maybe.
               | 
               | It would be funny that in the end, the most use is made
               | by student cheating at uni.
        
               | d-lisp wrote:
               | I wanted to reflect a bit on this.
               | 
               | I have hard time to imagine why non-tech people would
               | find a use for LLMs, let's say nothing in your life
               | forces you to produce information (be it textual,
               | pictural or anything that can be related to information).
               | Let's say your needs are focused on spending good times
               | with friends or your family, eating nice dishes (home
               | cooked or restaurant), spending your money on furnitures,
               | rents, clothes, tools and etc.
               | 
               | Why would you need an AI that produce information in an
               | information-bloated world ?
               | 
               | You probably met someone that "fell in love with
               | woodworking" or idk, after having watched youtube videos
               | (that person probably built a chair, a table or something
               | akin). I don't think stuff like "Hi, I have these
               | materials, what can I do with it" produce more
               | interesting results than just nerding on the internet or
               | in a library looking for references (on japaneese
               | handcrafted furnitures, vintage ikea designs, old school
               | woodworking, ...). (Or maybe the LLM will be able to give
               | you a list of good reads, which is nice but somewhat of a
               | limited and basic use).
               | 
               | Agentic AI and more efficient/intelligent AIs are not
               | very interesting for people like <wood lover> and are at
               | best a proxy for otherly findable information. Of course,
               | not everyone is like <wood lover>, the majority of people
               | don't even need to invest time in a "creative" hobby and
               | instead they will watch movies, invest time in sport,
               | invest time in sociability, go to museums, read books;
               | you could imagine having AIs that write books, invent
               | films, invent artworks, talk with you, but I am pretty
               | sure that there is something more than just "watch a
               | movie" or "read a book" when performing these activities;
               | as someone who likes reading or watching movies, what I
               | enjoy is following the evolutions of the authors of the
               | pieces, understanding their posture toward its ancestors,
               | its era-mates, toward its own previous visions and
               | whatnot. I enjoy to find a movie "weird" "goofy"
               | "sublime" and whatnot, because I enjoy a small amount of
               | parasociality with the authors and am finally brought to
               | say things like "Ahah, Lynch was such a weirdo when he
               | shot Blue Velvet" (okay, maybe not that type of bully
               | judgement, but you may be understanding what I mean).
               | 
               | I think I would find it uninspiring to read an AI written
               | book, because I couldn't live this small parasocial
               | experience. Maybe you could get me with music, but I
               | still think there's a lot of activity in loving a song. I
               | love Bach, but am pretty sure also I like Bach the
               | character (from what I speculate from the songs I
               | listen). I imagine that guy in front of his keyboard,
               | having the chance to live a -weird- moment of extasy when
               | he produces the best lines of the chaconne (if he was
               | living in our times he would relisten to what he produced
               | again and again and nodding to himself "man, that's
               | sick").
               | 
               | What could I experience from an LLM ? "Here is the
               | perfect novel I wrote specifically for you based on your
               | tastes:". There would be no imaginary Bach that I would
               | like to drink a beer with, no testimony of a human
               | reaching the state of mind in which you produce an
               | absolute (in fact highly relative, but you need to lie to
               | yourself) "hit".
               | 
               | All of this is highly personnal, but I would be curious
               | to know what others think.
        
         | blindhippo wrote:
         | Thing is, context management is NOT obvious to most users of
         | these tools. I use agentic coding tools on a daily basis now
         | and still struggle with keeping context focused and useful,
         | usually relying on patterns such as memory banks and task
         | tracking documents to try to keep a log of things as I pop in
         | and out of different agent contexts. Yet still, one false move
         | and I've blown the window leading to a "compression" which is
         | utterly useless.
         | 
         | The tools need to figure out how to manage context for us. This
         | isn't something we have to deal with when working with other
         | humans - we reliably trust that other humans (for the most
         | part) retain what they are told. Agentic use now is like
         | training a team mate to do one thing, then taking it out back
         | to shoot it in the head before starting to train another one.
         | It's inefficient and taxing on the user.
        
         | SubiculumCode wrote:
         | I constantly switch out, even when it's on the same topic. It
         | starts forming its own 'beliefs and assumptions', gets myopic.
         | I also make use of the big three services in turn to attack
         | ideas from multiple directions
        
           | nrds wrote:
           | > beliefs and assumptions
           | 
           | Unfortunately during coding I have found many LLMs like to
           | encode their beliefs and assumptions into comments; and even
           | when they don't, they're unavoidably feeding them into the
           | code. Then future sessions pick up on these.
        
             | SubiculumCode wrote:
             | YES! I've tried to provide instructions asking it to not
             | leave comments at all.
        
         | eru wrote:
         | > I realise this sounds obvious to many people but it clearly
         | wasn't to those guys so maybe it's not!
         | 
         | It's worse: Gemini (and ChatGPT, but to a lesser extent) have
         | started suggesting random follow-up topics when they conclude
         | that a chat in a session has exhausted a topic. Well, when I
         | say random, I mean that they seem to be pulling it from the
         | 'memory' of our other chats.
         | 
         | For a naive user without preconceived notions of how to use
         | these tools, this guidance from the tools themselves would
         | serve as a pretty big hint that they should intermingle their
         | sessions.
        
           | ghostpepper wrote:
           | For ChatGPT you can turn this memory off in settings and
           | delete the ones it's already created.
        
             | eru wrote:
             | I'm not complaining about the memory at all. I was
             | complaining about the suggestion to continue with unrelated
             | topics.
        
         | ramoz wrote:
         | Send them this https://backnotprop.substack.com/p/50-first-
         | dates-with-mr-me...
        
         | layman51 wrote:
         | That is interesting. I already knew about that idea that you're
         | not supposed to let the conversation drag on too much because
         | its problem solving performance might take a big hit, but then
         | it kind of makes me think that over time, people got away with
         | still using a single conversation for many different topics
         | because of the big context windows.
         | 
         | Now I kind of wonder if I'm missing out by not continuing the
         | conversation too much, or by not trying to use memory features.
        
         | getnormality wrote:
         | In my recent explorations [1] I noticed it got really stuck on
         | the first thing I said in the chat, obsessively returning to it
         | as a lens through which every new message had to be
         | interpreted. Starting new sessions was very useful to get a
         | fresh perspective. Like a human, an AI that works on a writing
         | piece with you is too close to the work to see any flaw.
         | 
         | [1] https://renormalize.substack.com/p/on-renormalization
        
           | ljlolel wrote:
           | Probably because the chat name is named after that first
           | message
        
           | okthrowman283 wrote:
           | Interesting I've noticed the same behavior with Gemini 3.0
           | but not with Claude, and Gemini 2.5 did not have this
           | behavior. I wonder what tuning is optimising for here.
        
         | faxmeyourcode wrote:
         | My boss (great engineer) had been complaining about this with
         | his internal github copilot quality no matter the model or
         | task. Turns out he never cleared the context. It was just the
         | same conversation spread thin across nearly a dozen completely
         | separate repositories because they were all in his massive
         | vscode workspace at once.
         | 
         | This was earlier this year... So I started giving internal
         | presentations on basic context management, best practices, etc
         | after that for our engineering team.
        
       | keeeba wrote:
       | Doesn't seem like this will be SOTA in things that really matter,
       | hoping enough people jump to it that Opus has more lenient usage
       | limits for a while
        
       | w_for_wumbo wrote:
       | Does anyone else consider that maybe it's impossible to benchmark
       | the performance of a piece of paper.
       | 
       | This is a tool that allows an intelligent system to work with it,
       | the same way that a piece of paper can reflect the writers'
       | intelligence, how can we accurately judge the performance of the
       | piece of paper, when it is so intimately reliant on the
       | intelligence that is working with it?
        
       | mlmonkey wrote:
       | It's funny how they don't compare themselves to Gemini and Claude
       | anymore.
        
       | anishshil wrote:
       | This shift toward new platforms is exactly why I'm building
       | Truwol, a social experience focused on real, unedited human
       | moments instead of the AI-saturated feeds we're drifting toward.
       | I'm developing it independently and sharing the progress
       | publicly, so if you're interested in projects reinventing online
       | spaces from the ground up, you can see what I'm working on Truwol
       | buymeacoffee/Truwol
        
       | jrflowers wrote:
       | OpenAI is really good at just saying stuff on the internet.
       | 
       | I love the way they talk about incorrect responses:
       | 
       | > Errors were detected by other models, which may make errors
       | themselves. Claim-level error rates are far lower than response-
       | level error rates, as most responses contain many claims.
       | 
       | "These numbers might be wrong because they were made up by other
       | models, which we will not elaborate on, also these numbers are
       | much higher by a metric that reflects how people use the product,
       | which we will not be sharing"
       | 
       | I also really love the graph where they drew a line at "wrong
       | half of the time" and labeled it 'Expert-Level'.
       | 
       | 10/10, reading this post is experientially identical to watching
       | that 12 hours of jingling keys video, which is hard to pull off
       | for a blog.
        
       | goobatrooba wrote:
       | I feel there is a point when all these benchmarks are
       | meaningless. What I care about beyond decent performance is the
       | user experience. There I have grudges with every single platform
       | and the one thing keeping me as a paid ChatGPT subscriber is the
       | ability to sort chats in "projects" with associated files (hello
       | Google, please wake up to basic user-friendly organisation!)
       | 
       | But all of them * Lie far too often with confidence * Refuse to
       | stick to prompts (e.g. ChatGPT to the request to number each
       | reply for easy cross-referencing; Gemini to basic request to
       | respond in a specific language) * Refuse to express uncertainty
       | or nuance (i asked ChatGPT to give me certainty %s which it did
       | for a while but then just forgot...?) * Refuse to give me short
       | answers without fluff or follow up questions * Refuse to stop
       | complimenting my questions or disagreements with wrong/incomplete
       | answers * Don't quote sources consistently so I can check facts,
       | even when I ask for it * Refuse to make clear whether they rely
       | on original documents or an internal summary of the document,
       | until I point out errors * ...
       | 
       | I also have substance gripes, but for me such basic usability
       | points are really something all of the chatbots fail on
       | abysmally. Stick to instructions! Stop creating walls of text for
       | simple queries! Tell me when something is uncertain! Tell me if
       | there's no data or info rather than making something up!
        
         | nullbound wrote:
         | << I feel there is a point when all these benchmarks are
         | meaningless.
         | 
         | I am relatively certain you are not alone in this sentiment.
         | The issue is that the moment we move past seemingly objective
         | measurements, it is harder to convince people that what we
         | measure is appropriate, but the measurable stuff can be
         | somewhat gamed, which adds a fascinating layer of cat and mouse
         | game to this.
        
         | delifue wrote:
         | Once a metric becomes optimization target, it ceases to become
         | good metric.
        
         | hnfong wrote:
         | There's a leaderboard that measures user experience, the
         | "lmsys" Chatbot Arena Leaderboard (
         | https://huggingface.co/spaces/lmarena-ai/lmarena-leaderboard ).
         | Main issue with it these days are that it kinda measures
         | sycophancy and user preferred tone more than substance.
         | 
         | Some issues you mentioned like length of response might be user
         | preference. Other issues like "hallucination" are areas of
         | active research (and there are benchmarks for these).
        
         | razster wrote:
         | The latest of the big three... OpenAI, Claude, and Google, none
         | of their models are good. I've spent too much time monitoring
         | them than just enjoying them. I've found it easier to run my
         | own local LLM. The latest Gemini release, I gave it another go
         | but only for it to misspell words and drift off into a fantasy
         | world after a few chats with help restructuring guides. ChatGPT
         | has become lazy for some reason and changes things I told it to
         | ignore, randomly too. Claude was doing great until the latest
         | release, then it started getting lazy after 20+k tokens. I
         | tried making sure to keep a guide to refresh it if it started
         | forgetting, but that didn't help.
         | 
         | Locals are better; I can script and have them script for me to
         | build a guide creation process. They don't forget because that
         | is all they're trained on. I'm done paying for 'AI'.
        
           | striking wrote:
           | What's to stop you from using the APIs the way you'd like?
        
             | joshribakoff wrote:
             | The API is a way to access a model, he is criticizing the
             | model not the access the method (at least until the last
             | sentence where he incorrectly implied you can only script a
             | local model, but I don't think thats a silver bullet, in my
             | experience that is even more challenging than starting with
             | a working agent)
        
           | marcosscriven wrote:
           | What are your best local models, and what hardware do you run
           | them on?
        
           | balder1991 wrote:
           | I have this impression that LLMs are so complicated and
           | entangled (in comparison to previous machine learning models)
           | that they're just too difficult to tune all around.
           | 
           | What I mean is, it seems they try to tune them to a few
           | certain things, that will make them worse on a thousand other
           | things they're not paying attention to.
        
         | ifwinterco wrote:
         | I'm not an expert but my understanding is transformers based
         | models simply can't do some of those things, it isn't really
         | how they work.
         | 
         | Especially something like expressing a certainty %, you might
         | be able to get it to output one but it's just making it up.
         | LLMs are incredibly useful (I use them every day) but you'll
         | always have to check important output
        
           | carsoon wrote:
           | Yeah I have seen multiple people use this certainty % thing
           | but its terrible. A percentage is something calculated
           | mathemtatically and these models cannot do that.
           | 
           | Potentially they could figure it out if they looks into a
           | comparison of next token probabilites, but this is not
           | exposed in any modern model and especially not fed back into
           | the chat/output.
           | 
           | Instead people should just ask it to explain BOTH sides of an
           | argument or explain why something is BOTH correct and
           | incorrect. This way you see how it can halluciate either way
           | and get to make up your own mind about the correct outcome.
        
         | empiko wrote:
         | Consider using structured output. You can define a JSON with
         | specific fields, and LLMs are only used to fill in the values.
         | 
         | https://ai.google.dev/gemini-api/docs/structured-output
        
         | fleischhauf wrote:
         | I'm always impressed how fast people get used to new things.
         | couple of years ago something like chatgpt was completely
         | impossible, and now people complain it something's does mit do
         | what you told it to and sometimes lies. (not saying your points
         | are not valid or you should not raise them) Some of the points
         | are just not fixable at this point due to tech limitations. A
         | language model currently simply has no way to give an estimate
         | of its confidence. Also there is no way to completely do away
         | with hallucinations (lies). there need to be some more
         | fundamental improvements for this to work reliably.
        
           | davebren wrote:
           | Your point would stand if the entire economy wasn't shifted
           | around this product and employees weren't being told to use
           | it or lose their jobs.
        
         | dontlikeyoueith wrote:
         | > Refuse to express uncertainty or nuance (i asked ChatGPT to
         | give me certainty %s which it did for a while but then just
         | forgot...?)
         | 
         | They're literally incapable of this. Any number they give you
         | is bullshit.
        
         | carsoon wrote:
         | I have a kinda strange chatgpt personalization prompt but it's
         | been working well for me. The focus is me to get the model to
         | analyze 2 sides and the extremes on both ends so it explains
         | both and lets me decide. This is much better than asking it to
         | make up accuracy percentages.
         | 
         | I think we align on what we want out of models:
         | 
         | """ Don't add useless babelling before the chats, just give the
         | information direct and explain the info.
         | 
         | DO NOT USE ENGAGEMENT BAITING QUESTIONS AT THE END OF EVERY
         | RESPONSE OR I WILL USE GROK FROM NOW ON FOREVER AND CANCEL MY
         | GPT SUBSCRIPTION PERMANENTLY ONLY. GIVE USEFUL FACTUAL
         | INFORMATION AND FOLLOW UPS which are grounded in first
         | principles thinking and logic. Do not take a side and look at
         | think about the extreme on both ends of a point before taking a
         | side. Do not take a side just because the user has chosen that
         | but provide infomration on both extremes. Respond with raw
         | facts and do not add opinions.
         | 
         | Do not use random emojis. Prefer proper marks for lists etc.
         | """
         | 
         | Those spelling/grammar errors are actually there and I don't
         | want to change it as its working well for me.
        
       | kachapopopow wrote:
       | did they just tune the parameters? the hallucinations are crazy
       | high on this version.
        
       | ChrisMarshallNY wrote:
       | They are talking a _lot_ about economics, here. Wonder what that
       | will mean for standard Plus users, like me.
        
       | hbarka wrote:
       | A year ago Sunday Pichai declared code red, now it's Sam Altman
       | declaring code red. How tables have turned, and I think the
       | acquisition of Windsurf and Kevin Hou by Google seems to
       | correlate with their level up.
        
         | jerrygenser wrote:
         | Acquisition of noam shazeer to supercharge their Gemini
         | flagship model line I think made a bigger impact.
         | 
         | To make an argument it was Kevin Hou, then we would need to see
         | Antigravity their new IDE being key. I think the crown jewel
         | are the Gemini models.
        
       | bluerooibos wrote:
       | Yawn.
        
         | dudeinhawaii wrote:
         | What does this add to the conversation? This isn't Reddit.
        
       | mobrienv wrote:
       | I recently built a webapp to summarize hn comment threads.
       | Sharing a summary given there is a lot here: https://hn-
       | insights.com/chat/gpt-52-8ecfpn.
        
         | DenisM wrote:
         | I keep asking ChatGPT to read and summarize HN front page while
         | driving, and it keeps blundering. I don't know if there's a
         | business for you in this, but I would pay.
         | 
         | Of course I always have questions about the subject, so it
         | become the whole voice chat thing.
        
           | mobrienv wrote:
           | Interesting I recently added the ability to receive a daily
           | email digest. Would just need a way to read it out. I'll look
           | into what a conversational voice chat might look like.
        
       | DenisM wrote:
       | Is there a voice chat mode in any chat app that is not heavily
       | degraded in reasoning?
       | 
       | I'm ok waiting for a response for 10-60 seconds if needed. That
       | way I can deep dive subjects while driving.
       | 
       | I'm ok paying money for it, so maybe someone coded this already?
        
       | nbardy wrote:
       | Those arc agi 2 improvements are insane.
       | 
       | Thats especially encouraging to me because those are all about
       | generalization.
       | 
       | 5 and 5.1 both felt overfit and would break down and be stubborn
       | when you got them outside their lane. As opposed to Opus 4.5
       | which is lovely at self correcting.
       | 
       | It's one of those things you really feel in the model rather than
       | whether it can tackle a harder problem or not, but rather can I
       | go back and forth with this thing learning and correcting
       | together.
       | 
       | This whole releases is insanely optimistic for me. If they can
       | push this much improvement WITHOUT the new huge data centers and
       | without a new scaled base model. Thats incredibly encouraging for
       | what comes next.
       | 
       | Remember the next big data center are 20-30x the chip count and
       | 6-8x the efficiency on the new chip.
       | 
       | I expect they can saturate the benchmarks WITHOUT and novel
       | research and algorithmic gains. But at this point it's clear
       | they're capable of pushing research qualitatively as well.
        
         | mmaunder wrote:
         | Same. Also got my attention re ARC-AGI-2. That's meaningful.
         | And a HUGE leap.
        
           | cbracketdash wrote:
           | Slight tangent yet I think is quite interesting... you can
           | try out the ARC-AGI 2 tasks by hand at this website [0]
           | (along with other similar problem sets). Really puts into
           | perspective the type of thinking AI is learning!
           | 
           | [0] https://neoneye.github.io/arc/?dataset=ARC-AGI-2
        
         | delifue wrote:
         | It's also possible that OpenAI use many human-generated
         | similar-to-ARC data to train (semi-cheating). OpenAI has enough
         | incentive to fake high score.
         | 
         | Without fully disclosing training data you will never be sure
         | whether good performance comes from memorization or "semi-
         | memorization".
        
         | deaux wrote:
         | > 5 and 5.1 both felt overfit and would break down and be
         | stubborn when you got them outside their lane. As opposed to
         | Opus 4.5 which is lovely at self correcting.
         | 
         | This is simply the "openness vs directive-following" spectrum,
         | which as a side-effect results in the sycophancy spectrum,
         | which still none of them have found an answer to.
         | 
         | Recent GPT models follow directives more closely than Claude
         | models, and are less sycophantic. Even Claude 4.5 models are
         | still somewhat prone to "You're absolutely right!". GPT 5+
         | (API) models never do this. The byproduct is that the former
         | are willing to self-correct, and the latter is more stubborn.
        
           | baq wrote:
           | Opus 4.5 answers most of my non-question comments with
           | 'you're right.' as the first thing in the output. At least
           | I'm not absolutely right, I'll take this as an improvement.
        
       | ponyous wrote:
       | I am really curious about speed/latency. For my use case there is
       | a big difference in UX if the model is faster. Wish this was
       | included in some benchmarks.
       | 
       | I will run 80 3D model generations benchmark tomorrow and update
       | this comment with the results about cost/speed/quality.
        
       | flkiwi wrote:
       | I gave up my OpenAI subscription a few days ago in favor of
       | Claude. My quality of life (and quality of results) has gone up
       | substantially. Several of our tools at work have GPT-5x as their
       | backend model, and it is incredible how frustrating they are to
       | use, how predictable their AI-isms are, and how inconsistent
       | their output is. OpenAI is going to have to do a lot more than an
       | incremental update to convince me they haven't completely lost
       | the thread.
        
         | brisket_bronson wrote:
         | You are absolutely right!
        
           | flkiwi wrote:
           | Someone didn't think so, lol. I debated not saying anything
           | because the AI partisans are just so awful.
        
             | jpkw wrote:
             | I think the above comment was a joke (Claude frequently
             | says that whenever you challenge it, whether you are right
             | or wrong)
        
               | jstummbillig wrote:
               | At least this once the AI-ism was not spotted.
        
               | flkiwi wrote:
               | Goodness no, I chuckled.
        
         | petesergeant wrote:
         | I have found Codex to be a phenomenal code-review tool, fwiw.
         | Shitty at writing code, _great_ at reviewing it.
        
       | mmaunder wrote:
       | Weirdly, the blog announcement completely omits the actual new
       | context window size which is 400,000:
       | https://platform.openai.com/docs/models/gpt-5.2
       | 
       | Can I just say !!!!!!!! Hell yeah! Blog post indicates it's also
       | much better at using the full context.
       | 
       | Congrats OpenAI team. Huge day for you folks!!
       | 
       | Started on Claude Code and like many of you, had that omg CC
       | moment we all had. Then got greedy.
       | 
       | Switched over to Codex when 5.1 came out. WOW. Really nice
       | acceleration in my Rust/CUDA project which is a gnarly one.
       | 
       | Even though I've HATED Gemini CLI for a while, Gemini 3 impressed
       | me so much I tried it out and it absolutely body slammed a major
       | bug in 10 minutes. Started using it to consult on commits. Was so
       | impressed it became my daily driver. Huge mistake. I almost lost
       | my mind after a week of this fighting it. Isane bias towards
       | action. Ignoring user instructions. Garbage characters in output.
       | Absolutely no observability in its thought process. And on and
       | on.
       | 
       | Switched back to Codex just in time for 5.1 codex max xhigh which
       | I've been using for a week, and it was like a breath of fresh
       | air. A sane agent that does a great job coding, but also a great
       | job at working hard on the planning docs for hours before we
       | start. Listens to user feedback. Observability on chain of
       | thought. Moves reasonably quickly. And also makes it easy to pay
       | them more when I need more capacity.
       | 
       | And then today GPT-5.2 with an xhigh mode. I feel like xmass has
       | come early. Right as I'm doing a huge Rust/CUDA/Math-heavy
       | refactor. THANK YOU!!
        
         | twisterius wrote:
         | [flagged]
        
           | mmaunder wrote:
           | My name is Mark Maunder. Not the fisheries expert. The other
           | one when you google me. I'm 51 and as skeptical as you when
           | it comes to tech. I'm the CTO of a well known cybersecurity
           | company and merely a user of AI.
           | 
           | Since you critiqued my post, allow me to reciprocate: I sense
           | the same deflector shields in you as many others here. I'd
           | suggest embracing these products with a sense of optimism
           | until proven otherwise and I've found that path leads to some
           | amazing discoveries and moments where you realize how
           | important and exciting this tech really is. Try out math that
           | is too hard for you or programming languages that are labor
           | intensive or languages that you don't know. As the GitHub CEO
           | said: this technology lets you increase your ambition.
        
             | bgwalter wrote:
             | I have tried the models and in domains I know well they are
             | pathetic. They remove all nuance, make errors that non-
             | experts do not notice and generally produce horrible code.
             | 
             | It is even worse in non-programming domains, where they
             | chop up 100 websites and serve you incorrect bland slop.
             | 
             | If you are using them as a search helper, that sometimes
             | works, though 2010 Google produced better results.
             | 
             | Oracle dropped 11% today due to over-investment in OpenAI.
             | Non-programmers are acutely aware of what is going on.
        
               | jfreds wrote:
               | > they remove all nuance
               | 
               | Said in a sweeping generalization with zero sense of
               | irony :D
        
               | jrflowers wrote:
               | This is a good point. It is a sweeping generalization if
               | you do not read the sentence that comes before that quote
        
               | what-the-grump wrote:
               | You pretend that humans don't produce slop?
               | 
               | I can recognize the short comings of AI code but it can
               | produce a mock or a full blown class before I can find a
               | place to save the file it produced.
               | 
               | Pretending that we are all busy writing novelty and
               | genius is silly, 99% are writing for CRUD tasks and basic
               | business flows, the code isn't going to be perfect it
               | doesn't need to be but it will get the job done.
               | 
               | All the logical gotchas of the work flows that you'd be
               | refactoring for hours are done in minutes.
               | 
               | Use pro with search... are it going to read 200 pages of
               | documentation in 7 minutes come up with a conclusion and
               | validate it or invalidate it in another 5? No you still
               | trying accept the cookie prompt on your 6th result.
               | 
               | You might as well join the flat earth society if you
               | still think that AI can't help you complete day to day
               | tasks.
        
               | muppetman wrote:
               | Exactly this. It's like reading the news! It seems
               | perfectly fine until a news article in a domain you have
               | intimate knowledge of, and then you realise how
               | bad/hacked together the news is. AI feels just like that.
               | But AI can improve, so I'm in the middle with my
               | optimism.
        
               | re-thc wrote:
               | > Oracle dropped 11% today due to over-investment in
               | OpenAI
               | 
               | Not even remotely true. Oracle is building out
               | infrastructure mostly for AI workloads. It dropped
               | because it couldn't explain its financing and if the
               | investment was worth it. OpenAI or not wouldn't have
               | mattered.
        
             | bluefirebrand wrote:
             | [flagged]
        
               | eru wrote:
               | Maybe you are holding it wrong?
               | 
               | Contemporary LLMs still have huge limitations and
               | downsides. Just like hammer or a saw has limitations. But
               | millions of people are getting good value out of them
               | already (both LLMs and hammers and saws). I find it hard
               | to believe that they are all deluded.
        
               | skydhash wrote:
               | What limitations does an hammer have if the job is
               | hammering? Or a saw with sawing? Even `ed` doesn't have
               | any issue with editing text files.
        
               | eru wrote:
               | Well, ask the people who invented better hammers or
               | better saws. Or better text editors than ed.
        
             | GolfPopper wrote:
             | Replace 'products' with 'message', 'tech' with 'religion'
             | and 'CEO' with 'prophet' and you have a bog-standard cult
             | recruitment pitch.
        
               | Aeolun wrote:
               | Because most recruitment pitches are the same regardless
               | of the subject.
        
         | freedomben wrote:
         | I haven't done a ton of testing due to cost, but so far I've
         | actually gotten worse results with xhigh than high with
         | gpt-5.1-codex-max. Made me wonder if it was somehow a PEBKAC
         | error. Have you done much comparison between high and xhigh?
        
           | tekacs wrote:
           | I found the same with Max xhigh. To the point that I switched
           | back to just 5.1 High from 5.1 Codex Max. Maybe I should've
           | tried Max high first.
        
           | dudeinhawaii wrote:
           | This is one of those areas where I think it's about the
           | complexity of the task. What I mean is, if you set codex to
           | xhigh by default, you're wasting compute. IF you're setting
           | it at xhigh when troubleshooting a complex memory bug or
           | something, you're presumably more likely to get a quality
           | response.
           | 
           | I think in general, medium ends up being the best all-purpose
           | setting while high+ are good for single task deep-drive. Or
           | at least that has been my experience so far. You can
           | theoretically let with work longer on a harder task as well.
           | 
           | A lot appears to depend on the problem and problem domain
           | unfortunately.
           | 
           | I've used max in problem sets as diverse as "troubleshooting
           | Cyberpunk mods" and figuring out a race condition in a server
           | backend. In those cases, it did a pretty good job of
           | exhausting available data (finding all available logs,
           | digging into lua files), and narrowing a bug that every other
           | model failed to get.
           | 
           | I guess in some sense you have to know from the onset that
           | it's a "hard problem". That in and of itself is subjective.
        
             | wahnfrieden wrote:
             | You should also be making handoffs to/from Pro
        
           | robotswantdata wrote:
           | For a few weeks the Codex model has been cursed. Recommend
           | sticking with 5.1 high , 5.2 feels good too but early days
        
         | lopuhin wrote:
         | Context window size of 400k is not new, gpt-5, 5.1, 5-mini,
         | etc. have the same. But they do claim they improved long
         | context performance which if true would be great.
        
           | energy123 wrote:
           | But 400k was never usable in ChatGPT Plus/Pro subscriptions.
           | It was nerfed down to 60-100k. If you submitted too long of a
           | prompt they deleted the tokens on the end of your prompt
           | before calling the model. Or if the chat got too long (still
           | below 100k however) they deleted your first messages. This
           | was 3 months ago.
           | 
           | Can someone with an active sub check whether we can submit a
           | full 400k prompt (or at least 200k) and there is no prompt
           | truncatation in the backend? I don't mean attaching a file
           | which uses RAG.
        
             | gunalx wrote:
             | API use was not merged in this way.
        
             | piskov wrote:
             | Context windows for web
             | 
             | Fast (GPT-5.2 Instant) Free: 16K Plus / Business: 32K Pro /
             | Enterprise: 128K
             | 
             | Thinking (GPT-5.2 Thinking) All paid tiers: 196K
             | 
             | https://help.openai.com/en/articles/11909943-gpt-52-in-
             | chatg...
        
               | dr_dshiv wrote:
               | That's... too bad
        
               | energy123 wrote:
               | But can you do that in one message or is that a best case
               | scenario in a long multi turn chat?
        
             | eru wrote:
             | > Or if the chat got too long (still below 100k however)
             | they deleted your first messages. This was 3 months ago.
             | 
             | I can believe that, but it also seems really silly? If your
             | max context window is X and the chat has approached that,
             | instead of outright deleting the first messages outright,
             | why not have your model summarise the first quarter of
             | tokens and place those at the beginning of the log you feed
             | as context? Since the chat history is (mostly) immutable,
             | this only adds a minimal overhead: you can cache the
             | summarisation, and don't have to do that over and over
             | again for each new message. (If partially summarised log
             | gets too long, you summarise again.)
             | 
             | Since I can come up with this technique in half a minute of
             | thinking about the problem, and the OpenAI folks are
             | presumably not stupid, I wonder what downside I'm missing.
        
               | Aeolun wrote:
               | Don't think you are missing anything. I do this with the
               | API, and it works great. I'm not sure why they don't do
               | it, but I can only guess it's because it completely
               | breaks the context caching. If you summarize the full
               | buffer at least you know you are down to a few thousand
               | tokens to cache again, instead of 100k tokens to cache
               | again.
        
               | eru wrote:
               | > [...] but I can only guess it's because it completely
               | breaks the context caching.
               | 
               | Yes, but you only re-do this every once in a while? It's
               | a constant factor overhead. If you essentially feed the
               | last few thousand tokens, you have no caching at all (and
               | you are big enough that this window of 'last few thousand
               | tokens' doesn't get you the whole conversation)?
        
         | tgtweak wrote:
         | have been on 1M context window with claude since 4.0 - it gets
         | pretty expensive when you run 1M context on a long running
         | project (mostly using it in cline for coding). I think they've
         | realized more context length = more $ when dealing with most
         | agentic coding workflows on api.
        
           | Workaccount2 wrote:
           | You should be doing everything you can to keep context under
           | 200k, ideally even 100k. All the models unwind so badly as
           | context grows.
        
             | patates wrote:
             | I don't have that experience with gemini. Up to 90% full,
             | it's just fine.
        
         | Suppafly wrote:
         | >Can I just say !!!!!!!! Hell yeah!
         | 
         | ...
         | 
         | >THANK YOU!!
         | 
         | Man you're way too excited.
        
         | nathants wrote:
         | Usable input limit has not changed, and remains 400 - 128 =
         | 272. Confirmed by looking for any changes in codex cli source,
         | nope.
        
         | lhl wrote:
         | Anecdotally, I will say that for my toughest jobs GPT-5+ High
         | in `codex` has been the best tool I've used - CUDA->HIP
         | porting, finding bugs in torch, websockets, etc, it's able to
         | test, reason deeply and find bugs. It can't make UI code for
         | it's life however.
         | 
         | Sonnet/Opus 4.5 is faster, generally feels like a better coder,
         | and make much prettier TUI/FEs, but in my experience, for
         | anything tough any time it tells you it understands now, it
         | really doesn't...
         | 
         | Gemini 3 Pro is unusable - I've found the same thing,
         | opinionated in the worst way, unreliable, doesn't respect my
         | AGENTS.md and for my real world problems, I don't think it's
         | actually solved anything that I can't get through w/ GPT
         | (although I'll say that I wasn't impressed w/ Max, hopefully
         | 5.2 xhigh improves things). I've heard it can do some magic
         | from colleagues working on FE, but I'll just have to take their
         | word for it.
        
         | ubutler wrote:
         | > Weirdly, the blog announcement completely omits the actual
         | new context window size which is 400,000:
         | https://platform.openai.com/docs/models/gpt-5.2
         | 
         | As @lopuhin points out, they already claimed that context
         | window for previous iterations of GPT-5.
         | 
         | The funny thing is though, I'm on the business plan, and none
         | of their models, not GPT-5, GPT-5.1, GPT-5.2, GPT-5.2 Extended
         | Thinking, GPT-5.2 Pro, etc., can really handle inputs beyond
         | ~50k tokens.
         | 
         | I know because, when working with a really long Python file
         | (>5k LoCs), it often claims there is a bug because, somewhere
         | close to the end of the file, it cuts off and reads as '...'.
         | 
         | Gemini 3 Pro, by contrast, can genuinely handle long contexts.
        
           | andybak wrote:
           | Why would you put that whole python file in the context at
           | all? Doesn't Codex work like Claude Code in this regard and
           | use tools to find the correct parts of a larger file to read
           | into context?
        
         | BrtByte wrote:
         | This is one of those updates where the value only really shows
         | up if you're already deep in the weeds
        
       | TechDebtDevin wrote:
       | $168.00 / 1M ouput tokens is hilarious for their "Pro". Can't
       | wait to here all the bitching from orgs next month. Literally the
       | dumbest product of all time. Do you people seriously pay for
       | this?
        
       | StarterPro wrote:
       | >GPT-5.2 sets a new state of the art across many benchmarks,
       | including GDPval, where it outperforms industry professionals at
       | well-specified knowledge work tasks spanning 44 occupations.
       | 
       | We built a benchmark tool that says our newest model outperforms
       | everyone else. Trust me bro.
        
       | SilverElfin wrote:
       | Is the training cutoff date known?
        
       | agentifysh wrote:
       | Looks like they've begun censoring posts at r/Codex and not
       | allowing complaint threads so here is my honest take:
       | 
       | - It is faster which is appreciated but not as fast as Opus 4.5
       | 
       | - I see no changes, very little noticeable improvements over 5.1
       | 
       | - I do not see any value in exchange for +40% in token costs
       | 
       | All in all I can't help but feel that OpenAI is facing an
       | existential crisis. Gemini 3 even when its used from AI Studio
       | offers close to ChatGPT Pro performance for free. Anthropic's
       | Claude Code $100/month is tough to beat. I am using Codex with
       | the $40 credits but there's been a silent increase in token costs
       | and usage limitations.
        
         | AstroBen wrote:
         | Did you notice much improvement going from Gemini 2.5 to 3? I
         | didn't
         | 
         | I just think they're all struggling to provide real world
         | improvements
        
           | XCSme wrote:
           | Maybe they are just more consistent, which is a bit hard to
           | notice immediately.
        
           | dcre wrote:
           | Nearly everyone else (and every measure) seems to have found
           | 3 a big improvement over 2.5.
        
           | enraged_camel wrote:
           | Gemini 3 was a massive improvement over 2.5, yes.
        
           | cmrdporcupine wrote:
           | I think what they're actually struggling with is costs. And I
           | think they're all behind the scenes quantizing models to
           | manage load here and there, and they're all giving
           | inconsistent results.
           | 
           | I noticed huge improvement from Sonnet 4.5 to Opus 4.5 when
           | it became unthrottled a couple weeks ago. I wasn't going to
           | sign back up with Anthropic but I did. But two weeks in it's
           | already starting to seem to be inconsistent. And when I go
           | back to Sonnet it feels like they did something to lobotomize
           | it.
           | 
           | Meanwhile I can fire up DeepSeek 3.2 or GLM 4.6 for a
           | fraction of the cost and get _almost_ as good as results.
        
           | agentifysh wrote:
           | oh yes im noticing significant improvements across the board
           | but mainly having 1,000,000 token context makes a ton of
           | difference, I can keep digging at a problem with out
           | compaction.
        
           | free652 wrote:
           | yes, 2.5 just couldnt use tools right. 3.0 is way better at
           | coding. better than sonnet 4.5/
        
           | dudeinhawaii wrote:
           | I noticed a quite noticeable improvement to the point where I
           | made it my go-to model for questions. Coding-wise, not so
           | much. As an intelligent model, writing up designs,
           | investigations, general exploration/research tasks, it's top
           | notch.
        
           | chillfox wrote:
           | Gemini 3 Pro is the first model from Google that I have found
           | usable, and it's very good. It has replaced Claude for me in
           | some cases, but Claude is still my goto for use in coding
           | agents.
           | 
           | (I only access these models via API)
        
           | neuah wrote:
           | Using it in a specialized subfield of neuroscience, Gemini 3
           | w/ thinking is a huge leap forward in terms of knowledge and
           | intelligence (with minimal hallucinations). I take it that
           | the majority of people on here are software engineers. If
           | you're evaluating it on writing boilerplate code, you
           | probably have to squint to see differences between the
           | (excellent) raw model performances. whereas in more niche
           | edge cases there is more daylight between them.
        
         | hmottestad wrote:
         | I'm curious about if the model has gotten more consistent
         | throughout the full context window? It's something that OpenAI
         | touted in the release, and I'm curious if it will make a
         | difference for long running tasks or big code reviews.
        
           | agentifysh wrote:
           | one positive is that 5.2 is very good at finding bugs but not
           | sure about throughputs I'd imagine it might be improved but
           | haven't seen a real task to benchmark it on.
           | 
           | what I am curious about is 5.2-codex but many of us
           | complained about 5.1-codex (it seemed to get tunnel visioned)
           | and I have been using vanilla 5.1
           | 
           | its just getting very tiring to deal with 5 different
           | permutations of 3 completely separate models but perhaps this
           | is the intent and will keep you on a chase.
        
         | BrtByte wrote:
         | The speed bump is nice, but speed alone isn't a compelling
         | upgrade if the qualitative difference isn't obvious in day-to-
         | day use
        
       | tpurves wrote:
       | Undoubtedly each new model from OpenAi has numerous training and
       | orchestration improvements etc.
       | 
       | But how much of each product they release also just a factor of
       | how much they are willing to spend on inference per query in
       | order to stay competitive?
       | 
       | I always wonder how much is technical change vs turning a knob up
       | and down on hardware and power consumption.
       | 
       | GTP5.0 for example seemed like a lot of changes more for OpenAI's
       | internal benefit (terser responses, dynamic 'auto' mode to scale
       | down thinking when not required etc.)
       | 
       | Wondering if GPT5.2 is also case of them in 'code red mode' just
       | turning what they already have up to 11 as a fastest way to
       | respond to fiercer competion.
        
         | simonsarris wrote:
         | I always liked the definition of technology as "doing more with
         | less". 100 oxen replaced by 1 gallon of diesel, etc.
         | 
         | That it costs more does suggest it's "doing more with more", at
         | least.
        
           | psychoslave wrote:
           | Good luck with reproducing and eating diesel like can be done
           | with oxen and related species.
           | 
           | Humanity won't be able to tap into this highly compressed
           | energy stock that was generated through processes taking
           | literally geological scales time to bed achieved.
           | 
           | That is, technology is more about what alternative tradeoffs
           | can we leverage on to organize differently with resources at
           | hand.
           | 
           | Frugality can definitely be a possible way to shape the
           | technologies we want to deploy. But it's not all possible
           | technologies, just a subset.
           | 
           | Also better technology is not necessarily bringing societies
           | to morale and well-being excellency. Improving technology for
           | efficient genocides for example is going to bring human
           | disaster as obvious outcome, even if it's done in a manner
           | that is the most green, zero-carbon emissions and growing
           | more forests delivered beyond expectations of the
           | specifications.
        
       | ClipNoteBook wrote:
       | ChatGPT seems to just randomly pick urls to cite and extract
       | information from. Google Gemini seems to look at heuristics like
       | whether the author is trustworthy, or an expert in the topic. But
       | more advanced
        
       | jonplackett wrote:
       | Excited to try this. I've found Gemini excellent recently and
       | amazing at coding. But I still feel somehow like ChatGPT
       | understands more. Even though it's not quite as good at coding -
       | and nowhere at as fast. It is much less likely anti spontaneously
       | forget something. Gemini's is part unbelievably amazing and part
       | amnesia patient. I still kinda trust ChatGPT more.
        
       | snake_doc wrote:
       | > Models were run with maximum available reasoning effort in our
       | API (xhigh for GPT-5.2 Thinking & Pro, and high for GPT-5.1
       | Thinking), except for the professional evals, where GPT-5.2
       | Thinking was run with reasoning effort heavy, the maximum
       | available in ChatGPT Pro. Benchmarks were conducted in a research
       | environment, which may provide slightly different output from
       | production ChatGPT in some cases.
       | 
       | Feels like a Llama 4 type release. Benchmarks are not apples to
       | apples. Reasoning effort is across the board higher, thus uses
       | more compute to achieve an higher score on benchmarks.
       | 
       | Also notes that some may not be producible.
       | 
       | Also, vision benchmarks all use Python tool harness, and they
       | exclude scores that are low without the harness.
        
       | 0xdeafbeef wrote:
       | much better
       | https://chatgpt.com/s/t_693b489d5a8881918b723670eaca5734 than 5.1
       | https://chatgpt.com/s/t_6915c8bd1c80819183a54cd144b55eb2.
       | 
       | Same query - what romanian football player won the premier league
       | 
       | update. Even instant returns correct result without problems
       | 
       | https://chatgpt.com/s/t_693b49e8f5808191a954421822c3bd0d
        
       | jacquesm wrote:
       | A classic long-form sales pitch. Someone's been reading their
       | Patio11...
        
       | jaimex2 wrote:
       | They just keep flogging that dead horse.
       | 
       | The winner in this race will be whoever gets small local models
       | to perform as well on consumer hardware. It'll also pop the tech
       | bubble in the US.
        
       | johnwheeler wrote:
       | I'm not interested in using OpenAI anymore because Sam Altman is
       | so untrustworthy. All you see on X.com is him and Greg Brockman
       | kissing David Sacks' ass, trying to make inroads with him, asking
       | Disney for investments, and shit. Are you kidding? Who wants to
       | support these clowns? Let's let Google win. Let's let Anthropic
       | win. Anyone but Sam Altman.
        
       | lacoolj wrote:
       | This is a whole bunch of patting themselves on the back.
       | 
       | Let me know when Gemini 3 Pro and Opus 4.5 are compared against
       | it.
        
       | johan914 wrote:
       | A bit off topic: but what's with the ram usage of LLM clients?
       | ChatGPT, google, and Anthropic all use 1+ GB of ram during a long
       | session. Surely they are not running GPT 3 locally?
        
       | aaroninsf wrote:
       | As a popcorn eating bystander it is striking to scan the top
       | comments and find they alternate so dramatically in tone and
       | conclusions.
        
       | whereistejas wrote:
       | Did anyone notice how Cursor wasn't an early tester? I wonder
       | why...
        
       | Kim_Bruning wrote:
       | I'm continuously surprised that some people get good results out
       | of GPT models. They sort of fail on my personal benchmarks for
       | me.
       | 
       | Maybe GPT needs a different approach to prompting? (as compared
       | to eg Claude, Gemini, or Kimi)
        
         | piskov wrote:
         | They are all gpt as in generative pre-trained transformer
        
           | Kim_Bruning wrote:
           | That may or may not be true, but in the context of this
           | article, I'm referring to OpenAI's GPT brand of models.
        
       | johndill wrote:
       | Did Calmmy Sammy that his is the version that will finally cure
       | cancer? The AI shakeout in the AI industry is going to be brutal.
       | Can't see how Private Equity is going to get the little guy to be
       | left holding the giant bag of excrement, but they will figure
       | that out. AI, smart enough to replace you, but not quite smart
       | enough the replace the CEO or Hedge Fund Bros.
        
         | astrange wrote:
         | What do private equity or hedge funds have to do with any of
         | this? Those are like, specific business models that are not
         | involved in this situation.
        
       | byt3bl33d3r wrote:
       | There's really no point in looking at benchmarks anymore as real
       | world usage of these models varies between task and prompting
       | strategies. Use your internal benchmarks to evaluate and ignore
       | everything else. It is curious to me how they don't provide a
       | side x side comparison of other models benchmarks for this
       | release
        
       | youngermax wrote:
       | Isn't it interesting how this incremental release includes so
       | many testimonials from companies who claim the model has
       | improved? It also focuses on "economically valuable tasks." There
       | was nothing of this sort in GPT-5.1's release. Looks like OpenAI
       | feeling the pressure from investors now.
        
       | nezaj wrote:
       | We saw it do better at making counter-strike!
       | https://x.com/instant_db/status/1999278134504620363?s=20
        
       | stopachka wrote:
       | For those curious about the question: "how well does GPT 5.2
       | build Counter Strike?"
       | 
       | We tried the same prompts we asked previous models today, and
       | found out [1].
       | 
       | The TL:DR: Claude is still better on the frontend, but 5.2 is
       | comparable to Gemini 3 Pro on the backend. At the very least 5.2
       | did better on just about every prompt compared to 5.1 Codex Max.
       | 
       | The two surprises with the GPT models when it comes to coding: 1.
       | They often use REPLs rather than read docs 2. In this instance
       | 5.2 was more sheepish about running CLI commands. It would
       | instead ask me to run the commands.
       | 
       | Since this isn't a codex fine-tuned model, I'm definitely excited
       | to see what that looks like.
       | 
       | [1] The full video and some details in the tweet here:
       | https://x.com/instant_db/status/1999278134504620363
        
       | TakakiTohno wrote:
       | I use it everyday but have been told by friends that Gemini has
       | overtaken it.
        
       | blitz_skull wrote:
       | Again I just tap the sign.
       | 
       | All of your benchmarks mean nothing to me until you include
       | Claude Sonnet on them.
       | 
       | In my experience, GPT hasn't been able to compete with Claude in
       | years for the daily "economically valuable" tasks I work on.
        
         | nextworddev wrote:
         | Claude is pretty trash for anything besides coding
        
           | wyre wrote:
           | What are you basing that on? Between Sonnet and Opus I don't
           | think I'm reaching for Gemini 3 at all.
        
           | timmg wrote:
           | That hasn't been my experience at all. I always wondered if
           | we just get used to how to prompt a given model and that it
           | hard to transition to another.
        
           | romanovcode wrote:
           | Yeah, but that is the whole point of Claude. And that's why
           | we are interested in the comparison.
        
         | jstummbillig wrote:
         | Since as per Anthropics own benchmarks Sonnet 4.5 is beaten by
         | Opus 4.5 would it not suffice to infer the rest?
         | 
         | https://x.com/OpenAI/status/1999182104362668275
        
       | 8cvor6j844qw_d6 wrote:
       | What the current preferred subscription on AI?
       | 
       | OpenAI and Anthrophic is my current preference. Looking forward
       | to know what others use.
       | 
       | Claude Code for coding assistance and cross-checking my work.
       | OpenAI for second opinion on my high-level decisions.
        
       | eastoeast wrote:
       | For the first time, I'm presenting a problem to LLMs that they
       | cannot seem to answer. This is my first instance of them
       | "endlessly thinking" without producing anything.
       | 
       | The problem is complicated, but very solvable.
       | 
       | I'm programming video cropping into my Android application. It
       | seems videos that have "rotated" metadata cause the crop to be
       | applied incorrectly. As in, a crop applied to the top of a video
       | actually gets applied to the video rotated on its side.
       | 
       | So, either double rotation is being applied somewhere in the
       | pipeline, or rotation metadata is being ignored.
       | 
       | I tried Opus 4.5, Gemini 3, and Codex 5.2. All 3 go through loops
       | of "Maybe Media3 applies the degree(90) after...", "no, that's
       | not right. Let me think..."
       | 
       | They'll do this for about 5 minutes without producing anything.
       | I'll then stop them, adjusting the prompt to tell them "Just try
       | anything! Your first thought, let's rapidly iterate!". Nope.
       | Nothing.
       | 
       | To add, it also only seems to be using about 25% context on Opus
       | 4.5. Weird!
        
       | xmcqdpt2 wrote:
       | I don't know if they used the new ChatGPT to translate this page
       | but I was served the French version and it is NOT good. There are
       | placeholders for quotes like <quote> and the prose is incredibly
       | repetitive. You'd figure that OpenAI of all people would be able
       | to translate something to one of the worlds most spoken language.
        
       | matt3210 wrote:
       | Can this be used without uploading my code base to their server?
        
       | yearolinuxdsktp wrote:
       | Plus users are now defaulted to a faster, less deep GPT-5.2
       | Thinking mode called "Standard", and you now have to manually
       | select "Extended" to get back to previous deep thinking level for
       | Plus users. Yet the 3K messages a week quota is the same
       | regardless of thinking level. Also, the selection does not sync
       | to mobile (you know, just not enough RAM in computers these days
       | to persist a setting between web and mobile).
        
       | rishabhaiover wrote:
       | After I saw Opus 4.5 search through zig's std io because it
       | wasn't aware of a breaking change in the recent release, I fell
       | in love with claude-code and I don't see a strong enough reason
       | to switch to codex at the moment.
        
       | lend000 wrote:
       | It seems like they fixed the most obvious issue with the last
       | release, where codex would just refuse to do its job... if it
       | seemed difficult or context usage was getting above 60% or so.
       | Good job on the post-training improvements.
       | 
       | The benchmark changes are incredible, but I have yet to notice a
       | difference in my codebases as of yet.
        
       | lazarus01 wrote:
       | My god, what terrible marketing, totally written by AI. No flow
       | whatsoever.
       | 
       | I use Gemini 3 with my $10/month copilot subscription on vscode.
       | I have to say, Gemini 3 is great. I can do the work of four
       | people. I usually run out of premium tokens in a week. But I'm
       | actually glad there is a limit or I would never stop working. I
       | was a skeptic, but it seems like there is a wider variety of
       | patterns in the training distribution.
        
       | namesbc wrote:
       | So the rosy biased estimate is OpenAI is saving 1 hour of work
       | per day, so 5 hours total per-work week and 20 hours total per-
       | month.
       | 
       | With a subsidized cost of $200/month for OpenAI it would be
       | cheaper to hirer a part-time minimum wage worker than it would be
       | to contract with OpenAI.
       | 
       | And that is the rosiest estimate OpenAI has.
        
         | dangoodmanUT wrote:
         | A part time minimum wage worker can't code
        
           | namesbc wrote:
           | Check the wages of coders outside of the US
        
         | AstroBen wrote:
         | What people here forget is coding is a tiny minority of the
         | actual usage. ~5% if I remember correctly?
         | 
         | Their best market might just be as a better Google with ads
        
           | namesbc wrote:
           | Yep, bulk of AI usage is generating marketing emails
        
         | maerch wrote:
         | The closest I come to working with part-time, minimum-wage
         | workers is working with student employees. Even then, they earn
         | more and usually work more than five hours a week.
         | 
         | Most of the time, I end up putting in more work than I get out
         | of it. Onboarding, reviewing, and mentoring all take
         | significant time.
         | 
         | Even with the best students we had, paying around 400 euros a
         | month, I would not say that I saved five hours a week.
         | 
         | And even when they reach the point of being truly productive,
         | they are usually already finished with their studies. If we
         | then hire them full-time, they cost significantly more.
        
         | namesbc wrote:
         | It you take of the rosy glasses, it is more like 10 hours saved
         | per-month at an unsubsidized cost of $1000/month
         | 
         | The $100/hr is worth it for US programming jobs, but nothing
         | else
        
       | ofermend wrote:
       | GPT-5.2 just added to Vectara Hallucination Leaderboard.
       | Definitely an improvement over GPT-5.1 - congrats to the team
       | 
       | https://github.com/vectara/hallucination-leaderboard
        
       | getnormality wrote:
       | Sweet Jesus. 53% on ARC-AGI-2. There's still gas in this van.
        
       | rallies wrote:
       | I work at the intersection of AI and investing, and I'm really
       | amazed at the ability of this model to build spreadsheets.
       | 
       | I gave it a few tools to access sec filings (and a small local
       | vector database), and it's generating full fledged spreadsheets
       | with valid, real time data. Analysts in wallstreet are going to
       | get really empowered, but for the first time, I'm really glad
       | that retail investors are also getting these models.
       | 
       | Just put out the tool: https://github.com/ralliesai/tenk
        
         | rallies wrote:
         | Here's a nice parsing of all the important financials from an
         | SEC report. This used to be really hard a few years ago.
         | 
         | https://docs.google.com/spreadsheets/d/1DVh5p3MnNvL4KqzEH0ME...
        
         | npodbielski wrote:
         | Can't wait for being fired because some VP or other manager
         | asked some model to prepare list of people with lowest
         | productivity to pay ratio.
         | 
         | Model hallucinated half of the data?! Sorry we can't go back on
         | this decision, that would make us look bad!
         | 
         | Or when some silly model will push everyone to invest in some
         | radicoulous company and everybody will do it. Poisoning data
         | attack to inject some I am Future Inc (tm) company with high
         | investment rate. After few months pocket money and vanish.
         | 
         | We are certainly going to live in interesting times.
        
         | monatron wrote:
         | Nice tool - I appreciate you sharing the work!
        
       | vishal_new wrote:
       | Hmmm, is there any insight if these are really getting much
       | better at coding? Will hand coding be dead within a few years,
       | just human typing in english?
        
         | psychoslave wrote:
         | Mia espero estas ke ne, ni nur parolos home inter homoj,
         | robotoj anticipe faros servutoj por tauge fari niajn dezirojn
         | realigi lau niaj faktaj bezonoj. Kompreneble ni ciuj flue
         | parolos Esperanto por taga geopolitikaj internaciaj aferoj, kaj
         | ia ajn alia lingvo kiu placas al mi por aliaj aferoj.
         | 
         | Estonteco estas hela, miaj karaj siboj.
        
       | CodeCompost wrote:
       | For the first time, I've actually hidden an AI story on HN.
       | 
       | I can't even anymore. Sorry this is not going anywhere.
        
         | gchokov wrote:
         | Here, take my downvote.
        
           | bigyabai wrote:
           | In lieu of a killer app?
        
         | andybak wrote:
         | How this is different to any other post announcing an
         | incremental improvement in an app or service?
        
           | mabedan wrote:
           | It's a little different. Most of these improvements are just
           | more training hours and better weights. Even if it's about
           | actual improvement in trining algorithm or other software
           | tweaks they're not open source and hence other than "look how
           | marginally nicer the chat bot responds now" the post doesn't
           | provide value.
        
       | fasteo wrote:
       | >>> Already, the average ChatGPT Enterprise user says AI saves
       | them 40-60 minutes a day
       | 
       | If this is what AI has to offer, we are in a gigantic bubble
        
         | jatora wrote:
         | This seems pretty huge. Not sure by what metric it wouldn't be
         | civilizationally gigantic for everyone to save that much time
         | per day.
        
       | svara wrote:
       | In my experience, the best models are already nearly as good as
       | you can be for a large fraction of what I personally use them
       | for, which is basically as a more efficient search engine.
       | 
       | The thing that would now make the biggest difference isn't "more
       | intelligence", whatever that might mean, but better grounding.
       | 
       | It's still a big issue that the models will make up plausible
       | sounding but wrong or misleading explanations for things, and
       | verifying their claims ends up taking time. And if it's a topic
       | you don't care about enough, you might just end up misinformed.
       | 
       | I think Google/Gemini realize this, since their "verify" feature
       | is designed to address exactly this. Unfortunately it hasn't
       | worked very well for me so far.
       | 
       | But to me it's very clear that the product that gets this right
       | will be the one I use.
        
         | phorkyas82 wrote:
         | Isn't that what no LLM can provide: being free of
         | hallucinations?
        
           | svara wrote:
           | Yes, they'll probably not go away, but it's got to be
           | possible to handle them better.
           | 
           | Gemini (the app) has a "mitigation" feature where it tries to
           | to Google searches to support its statements. That doesn't
           | currently work properly in my experience.
           | 
           | It also seems to be doing something where it adds references
           | to statements (With a separate model? With a second pass over
           | the output? Not sure how that works.). That works well where
           | it adds them, but it often doesn't do it.
        
             | intended wrote:
             | Doubt it. I suspect it's fundamentally not possible in the
             | spirit you intend it.
             | 
             | Reality is perfectly fine with deception and inaccuracy.
             | For language to magically be self constraining enough to
             | only make verified statements is... impossible.
        
               | svara wrote:
               | Take a look at the new experimental AI mode in Google
               | scholar, it's going in the right direction.
               | 
               | It might be true that a fundamental solution to this
               | issue is not possible without a major breakthrough, but
               | I'm sure you can get pretty far with better tooling that
               | surfaces relevant sources, and that would make a huge
               | difference.
        
               | intended wrote:
               | So lets run it through the rubric test -
               | 
               | What's your level of expertise in this domain or subject?
               | How did you use it? What were your results?
               | 
               | It's basically gauging expertise vs usage to pin down the
               | variance that seems endemic to LLM utility
               | anecdotes/examples. For code examples I also ask which
               | language was used, the submitters familiarity with the
               | language, their seniority/experience and familiarity with
               | the domain.
        
               | svara wrote:
               | A lot of words to call me stupid ;) You seem to have put
               | me in some convenient mental box of yours, I don't know
               | which one.
        
               | intended wrote:
               | Oh heck no! Definitely no!
               | 
               | I am genuinely asking, because I think one of the biggest
               | determinants of utility obtained from LLMs is the
               | operator.
               | 
               | Damn, I didn't consider that it could be read that way. I
               | am sorry for how it came across.
        
           | kyletns wrote:
           | For the record, brains are also not free of hallucinations.
        
             | delaminator wrote:
             | That's not a very useful observation though is it?
             | 
             | The purpose of mechanisation is to standardise and over the
             | long term reduce errors to zero.
             | 
             | Otoh "The final truth is there is no truth"
        
               | michaelscott wrote:
               | A lot of mechanisation, especially in the modern world,
               | is not deterministic and is not always 100% right; it's a
               | fundamental "physics at scale" issue, not something new
               | to LLMs. I think what happened when they first appeared
               | was that people immediately clung to a superintelligence-
               | type AI idea of what LLMs were supposed to do, then
               | realised that's not what they are, then kept going and
               | swung all the way over to "these things aren't good at
               | anything really" or "if they only fix this ONE issue I
               | have with them, they'll actually be useful"
        
               | delaminator wrote:
               | That's why I said tend to zero error. I'm a Six Sigma
               | guy. We take accurate over precise.
        
             | rimeice wrote:
             | I still don't really get this argument/excuse for why it's
             | acceptable that LLMs hallucinate. These tools are meant to
             | support us, but we end up with two parties who are, as you
             | say, prone to "hallucination" and it becomes a situation of
             | the blind leading the blind. Ideally in these scenarios
             | there's at least one party with a definitive or
             | deterministic view so the other party (i.e. us) at least
             | has some trust in the information they're receiving and any
             | decisions they make off the back of it.
        
               | TeMPOraL wrote:
               | For these types of problems (i.e. most problems in the
               | real world), the "definitive or deterministic" isn't
               | really possible. An unreliable party you can throw at the
               | problem from a hundred thousand directions simultaneously
               | and for cheap, is still useful.
        
               | ssl-3 wrote:
               | Have you ever employed anyone?
               | 
               | People, when tasked with a job, often get it right. I've
               | been blessed by working with many great people who really
               | do an amazing job of _generally_ succeeding to get things
               | right -- or at least, right-enough.
               | 
               | But in any line of work: Sometimes people fuck it up.
               | Sometimes, they forget important steps. Sometimes,
               | they're sure they did it one way when instead they did it
               | some other way and fix it themselves. Sometimes, they
               | even say they did the job and did it as-prescribed _and
               | actually believe themselves_ , when they've done neither
               | -- and they're perplexed when they're shown this. They
               | "hallucinate" and do dumb things for reasons that aren't
               | real.
               | 
               | And sometimes, they just make shit up and lie. They know
               | they're lying and they lie anyway, doubling-down over and
               | over again.
               | 
               | Sometimes they even go all spastic and deliberately throw
               | monkey wrenches into the works, just because they feel
               | something that makes them think that this kind of
               | willfully-destructive action benefits them.
               | 
               | All employees suck some of the time. They each have their
               | own issues. And all employees are expensive to hire, and
               | expensive to fire, and expensive to keep going. But some
               | of their outputs are useful, so we employ people anyway.
               | (And we're human; even the very best of us are going to
               | make mistakes.)
               | 
               | LLMs are not so different in this way, as a general
               | construct. They can get things right. They can also make
               | shit up. They can skip steps. The can lie, and double-
               | down on those lies. They hallucinate.
               | 
               | LLMs suck. All of them. They all fucking suck. They
               | aren't even good at sucking, and they persist at doing it
               | anyway.
               | 
               | (But some of their outputs are useful, and LLMs generally
               | cost a lot less to make use of than people do, so here we
               | are.)
        
               | vitorfblima wrote:
               | I don't get the comparison. It would be like saying it's
               | okay if an excel formula gives me different outcomes
               | everytime with the same arguments, sometimes right, but
               | mostly wrong.
        
               | ssl-3 wrote:
               | People can accomplish useful things, but sometimes make
               | mistakes and do shit wrong.
               | 
               | The bot can also accomplish useful things, and sometimes
               | make mistakes and do shit wrong.
               | 
               | (These two statements are more similar in their
               | truthiness than they are different.)
        
               | tsunamifury wrote:
               | As far as I can tell (as someone who worked on the early
               | foundation of this tech at Google for 10 years) making up
               | "shit" then using your force of will to make it true is a
               | huge part of the construction of reality with
               | intelligence.
               | 
               | Will to reality through forecasting possible worlds is
               | one of our two primary functions.
        
               | Libidinalecon wrote:
               | "The airplane wing broke and fell off during flight"
               | 
               | "Well humans break their leg too!"
               | 
               | It is just a mindlessly stupid response and a giant
               | category error.
               | 
               | The way an airplane wing and a human limb is not at all
               | the same category.
               | 
               | There is even another layer to this that comparing LLMs
               | to the brain might be wrong because the mereological
               | fallacy is attributing the brain "thinks" vs the
               | person/system as a whole thinks.
        
               | johnisgood wrote:
               | You are right that the wing/leg comparison is often lazy
               | rhetoric: we hold engineered systems to different failure
               | standards for good reason.
               | 
               | But you are misusing the mereological fallacy. It does
               | not dismiss LLM/brain comparisons: it actually
               | strengthens them. If the brain does not "think" (the
               | person does), then LLMs do not "think" either. Both are
               | subsystems in larger systems. That is not a category
               | error; it is a structural similarity.
               | 
               | This does not excuse LLM limitations - rimeice's concern
               | about two unreliable parties is valid. But dismissing
               | comparisons as "category errors" without examining which
               | properties are being compared is just as lazy as the
               | wing/leg response.
        
             | krzyk wrote:
             | Hallucinations are not bad. It adds some kind of
             | creativity, which is good for e.g. image generation,
             | coding, or story telling.
             | 
             | It is bad only in case of reporting on facts.
        
             | andrei_says_ wrote:
             | How much do you hallucinate at work? How many of your work
             | hallucinations do you confidently present as reality in
             | communication or code?
             | 
             | LLMs are being sold as viable replacement of paid
             | employees.
             | 
             | If they were not, they wouldn't be funded the way they are.
        
           | arw0n wrote:
           | I think the better word is confabulation; fabricating
           | plausible but false narratives based on wrong memory.
           | Fundamentally, these models try to produce plausible text.
           | With language models getting large, they start creating
           | internal world models, and some research shows they actually
           | have truth dimensions. [0]
           | 
           | I'm not an expert on the topic, but to me it sounds plausible
           | that a good part of the problem of confabulation comes down
           | to misaligned incentives. These models are trained _hard_ to
           | be a  'helpful assistant', and this might conflict with
           | telling the truth.
           | 
           | Being free of hallucinations is a bit too high a bar to set
           | anyway. Humans are extremely prone to confabulations as well,
           | as can be seen by how unreliable eye witness reports tend to
           | be. We usually get by through efficient tool calling (looking
           | shit up), and some of us through expressing doubt about our
           | own capabilities (critical thinking).
           | 
           | [0] https://arxiv.org/abs/2407.12831
        
             | svara wrote:
             | That's right - it does seem to have to do with trying to be
             | helpful.
             | 
             | One demo of this that reliably works for me:
             | 
             | Write a draft of something and ask the LLM to find the
             | errors.
             | 
             | Correct the errors, repeat.
             | 
             | It will never stop finding a list of errors!
             | 
             | The first time around and maybe the second it will be
             | helpful, but after you've fixed the obvious things, it will
             | start complaining about things that are perfectly fine,
             | just to satisfy your request of finding errors.
        
             | Tepix wrote:
             | > false narratives based on wrong memory
             | 
             | I don't think "wrong memory" is accurate, it's missing
             | information and doesn't know it or is trained not to admit
             | it.
             | 
             | Checkout the Dwarkesh Podcast episode
             | https://www.dwarkesh.com/p/sholto-trenton-2 starting at
             | 1:45:38
             | 
             | Here is the relevant quote by Trenton Bricken from the
             | transcript:
             | 
             |  _One example I didn 't talk about before with how the
             | model retrieves facts: So you say, "What sport did Michael
             | Jordan play?" And not only can you see it hop from like
             | Michael Jordan to basketball and answer basketball. But the
             | model also has an awareness of when it doesn't know the
             | answer to a fact. And so, by default, it will actually say,
             | "I don't know the answer to this question." But if it sees
             | something that it does know the answer to, it will inhibit
             | the "I don't know" circuit and then reply with the circuit
             | that it actually has the answer to. So, for example, if you
             | ask it, "Who is Michael Batkin?" --which is just a made-up
             | fictional person-- it will by default just say, "I don't
             | know." It's only with Michael Jordan or someone else that
             | it will then inhibit the "I don't know" circuit._
             | 
             |  _But what 's really interesting here and where you can
             | start making downstream predictions or reasoning about the
             | model, is that the "I don't know" circuit is only on the
             | name of the person. And so, in the paper we also ask it,
             | "What paper did Andrej Karpathy write?" And so it
             | recognizes the name Andrej Karpathy, because he's
             | sufficiently famous, so that turns off the "I don't know"
             | reply. But then when it comes time for the model to say
             | what paper it worked on, it doesn't actually know any of
             | his papers, and so then it needs to make something up. And
             | so you can see different components and different circuits
             | all interacting at the same time to lead to this final
             | answer._
        
               | BoredPositron wrote:
               | Architecture wise the "admit" part is impossible.
        
               | rbranson wrote:
               | Bricken isn't just making this up. He's one of the
               | leading researchers in model interpretability. See:
               | https://arxiv.org/abs/2411.14257
        
               | Tepix wrote:
               | Why do you think it's impossible? I just quoted him
               | saying _' by default, it will actually say, "I don't know
               | the answer to this question"'_
               | 
               | We already see that - given the right prompting - we can
               | get LLMs to say more often that they don't know things.
        
             | officialchicken wrote:
             | No, the correct word is hallucinating. That's the word
             | everyone uses and has been using. While it might not be
             | technically correct, everyone knows what it means and more
             | importantly, it's not a $3 word and everyone can relate to
             | the concept. I also prefer all the _other_ more accurate
             | alternative words Wikipedia offers to describe it:
             | 
             | "In the field of artificial intelligence (AI), a
             | hallucination or artificial hallucination (also called
             | bullshitting,[1][2] confabulation,[3] or delusion[4]) is"
        
           | svara wrote:
           | A part of it is reproducing incorrect information in the
           | training data as well.
           | 
           | One area that I've found to be a great example of this is
           | sports science.
           | 
           | Depending on how you ask, you can get a response lifted from
           | scientific literature, or the bro science one, even in the
           | course of the same discussion.
           | 
           | It makes sense, both have answers to similar questions and
           | are very commonly repeated online.
        
           | SecretDreams wrote:
           | Find me a human that doesn't occasionally talk out of their
           | ass =[
        
         | stacktrace wrote:
         | > It's still a big issue that the models will make up plausible
         | sounding but wrong or misleading explanations for things, and
         | verifying their claims ends up taking time. And if it's a topic
         | you don't care about enough, you might just end up misinformed.
         | 
         | Exactly! One important thing LLMs have made me realise deeply
         | is "No information" is better than false information. The way
         | LLMs pull out completely incorrect explanations baffles me - I
         | suppose that's expected since in the end it's generating tokens
         | based on its training and it's reasonable it might hallucinate
         | some stuff, but knowing this doesn't ease any of my
         | frustration.
         | 
         | IMO if LLMs need to focus on anything right now, they should
         | focus on better grounding. Maybe even something like a
         | probability/confidence score, might end up experience so much
         | better for so many users like me.
        
           | robocat wrote:
           | > wrong or misleading explanations
           | 
           | Exactly the same issue occurs with search.
           | 
           | Unfortunately not everybody knows to mistrust AI responses,
           | or have the skills to double-check information.
        
             | darkwater wrote:
             | No, it's not the same. Search results send/show you one or
             | more specific pages/websites. And each website has a
             | different trust factor. Yes, plenty of people repeat things
             | they "read on the Internet" as truths, but it's easy to
             | debunk some of them just based on the site reputation. With
             | AI responses, the reputation is shared with the good
             | answers as well, because they do give good answers most of
             | the time, but also hallucinate errors.
        
               | SebastianSosa1 wrote:
               | Community notes on X seems to be one of the highest
               | profile recent experiments trying to address this issue
        
               | dexterlagan wrote:
               | My attempt: https://www.cleverthinkingsoftware.com/truth-
               | or-extinction/
        
               | darkwater wrote:
               | > Tools like SourceFinder must be paired with education
               | -- teaching people how to trace information themselves,
               | to ask: Where did this come from? Who benefits if I
               | believe it?
               | 
               | These are very important and relevant questions to ask
               | oneself when you read about anything, but we also keep in
               | mind that even those question can be misused and they can
               | drive you to conspiracy theories.
        
             | incrudible wrote:
             | If somebody asks a question on Stackoverflow, it is
             | unlikely that a human who does not know the answer will
             | take time out of their day to completely fabricate a
             | plausible sounding answer.
        
               | balder1991 wrote:
               | At least it used to be true.
        
               | jaxn wrote:
               | People are confidently incorrect all the time. It is very
               | likely that people will make up plausible sounding
               | answers on StackOverflow.
               | 
               | You and I have both taken time out of our days to write
               | plausible sounding answers that are essentially opposing
               | hallucinations.
        
               | linen wrote:
               | Sites like stackoverflow are inherently peer-reviewed,
               | though; they've got a crowdsourced voting system and
               | comments that accumulate over time. People test the ideas
               | in question.
               | 
               | This whole "people are just as incorrect as LLMs" is a
               | poor argument, because it compares the single human and
               | the single LLM response in a vacuum. When you put enough
               | humans together on the internet you usually get a more
               | meaningful result.
        
               | JAlexoid wrote:
               | Have you ever heard of Dunning Kruger effect?
               | 
               | There's a reason why there are upvotes, solution and
               | third party edit system in StackOverflow - people will
               | spend time to write their "hallucinations" very
               | confidently.
        
             | lins1909 wrote:
             | What is it about people making up lies to defend LLMs? In
             | what world is it exactly the same as search? They're
             | literally different things, since you get information from
             | multiple sources and can do your own filtering.
        
           | actionfromafar wrote:
           | I wonder if the only way to fix this with current LLMs, would
           | be to generate a lot synthetic data for a select number
           | topics you _really_ don 't want it "go off the rails" with.
           | That synthetic data would be lots of variations on that "I
           | don't know how to do X with Y".
        
           | XCSme wrote:
           | But most benchmarks are not about that...
           | 
           | Are there even any "hallucination" public benchmarks?
        
             | andrepd wrote:
             | "Benchmarks" for LLMs are a total hoax, since you can train
             | them on the benchmarks themselves.
        
               | XCSme wrote:
               | I would assume a good benchmark has hidden tests, or
               | something randomly generated that is harder to game
        
           | biofox wrote:
           | I ask for confidence scores in my custom instructions /
           | prompts, and LLMs do surprisingly well at estimating their
           | own knowledge most of the time.
        
             | drclau wrote:
             | How do you know the confidence scores are not hallucinated
             | as well?
        
               | dfsegoat wrote:
               | they 100% are unless you provide a RUBRIC / basically
               | make it ordinal.
               | 
               |  _" Return a score of 0.0 if ...., Return a score of 0.5
               | if .... , Return a score of 1.0 if ..."_
        
               | kiliankoe wrote:
               | They are, the model has no inherent knowledge about its
               | confidence levels, it just adds plausible-sounding
               | numbers. Obviously they _can_ be plausible, but trusting
               | these is just another level up from trusting the original
               | output.
               | 
               | I read a comment here a few weeks back that LLMs always
               | hallucinate, but we sometimes get lucky when the
               | hallucinations match up with reality. I've been thinking
               | about that a lot lately.
        
               | TeMPOraL wrote:
               | > _the model has no inherent knowledge about its
               | confidence levels_
               | 
               | Kind of. See e.g.
               | https://openreview.net/forum?id=mbu8EEnp3a, but I think
               | it was established already a year ago that LLMs tend to
               | have identifiable internal confidence signal; the
               | challenge around the time of DeepSeek-R1 release was to,
               | through training, connect that signal to tool use
               | activation, so it does a search if it "feels unsure".
        
               | losvedir wrote:
               | Wow, that's a really interesting paper. That's the kind
               | of thing that makes me feel there's a lot more research
               | to be done "around" LLMs and how they work, and that
               | there's still a fair bit of improvement to be found.
        
               | fragmede wrote:
               | In science, before LLMs, there's this saying: all models
               | are wrong, some are useful. We model, say, gravity as
               | 9.8m/s2 on Earth, knowing full well that it doesn't hold
               | true across the universe, and we're able to build things
               | on top of that foundation. Whether that foundation is
               | made of bricks, or is made of sand, for LLMs, is for us
               | to decide.
        
               | xhkkffbf wrote:
               | It doesn't hold true across the universe? I thought this
               | was one of the more universal things like the speed of
               | light.
        
               | hackeman300 wrote:
               | Gravity isn't 9.8m/s/s across the universe. If you're at
               | higher or lower elevations (or outside the Earth's
               | gravitational pull entirely), the acceleration will be
               | different.
               | 
               | Their point was the 9.8 model is good enough for most
               | things on Earth, the model doesn't need to be perfect
               | across the universe to be useful.
        
               | JAlexoid wrote:
               | g(lower case) is literally gravitational force of Earth
               | at surface level. It's universally true, as there's only
               | one Earth in this universe.
               | 
               | G is the gravitational constant which is also universally
               | true(erm... to the best of our knowledge), g is
               | calculated using gravitational constant.
        
               | procflora wrote:
               | G, the gravitational constant is (as far as we know)
               | universal. I don't think this is what they meant, but the
               | use of "across the universe" in the parent comment is
               | confusing.
               | 
               | g, the net acceleration from gravity and the Earth's
               | rotation is what is 9.8m/s2 at the surface, on average.
               | It varies slightly with location and altitude (less than
               | 1% for anywhere on the surface IIRC), so "it's 9.8
               | everywhere" is the model that's wrong but good enough a
               | lot of the time.
        
             | EastLondonCoder wrote:
             | I'm with the people pushing back on the "confidence scores"
             | framing, but I think the deeper issue is that we're still
             | stuck in the wrong mental model.
             | 
             | It's tempting to think of a language model as a shallow
             | search engine that happens to output text, but that
             | metaphor doesn't actually match what's happening under the
             | hood. A model doesn't "know" facts or measure uncertainty
             | in a Bayesian sense. All it really does is traverse a high-
             | dimensional statistical manifold of language usage, trying
             | to produce the most plausible continuation.
             | 
             | That's why a confidence number that looks sensible can
             | still be as made up as the underlying output, because both
             | are just sequences of tokens tied to trained patterns, not
             | anchored truth values. If you want truth, you want
             | something that couples probability distributions to real
             | world evidence sources and flags when it doesn't have
             | enough grounding to answer, ideally with explicit
             | uncertainty, not hand-waviness.
             | 
             | People talk about hallucination like it's a bug that can be
             | patched at the surface level. I think it's actually a
             | feature of the architecture we're using: generating
             | plausible continuations by design. You have to change the
             | shape of the model or augment it with tooling that directly
             | references verified knowledge sources before you get
             | reliability that matters.
        
               | kznewman wrote:
               | Solid agree. Hallucination for me IS the LLM use case.
               | What I am looking for are ideas that may or may not be
               | true that I have not considered and then I go try to find
               | out which I can use and why.
        
               | sheeshe wrote:
               | In essence it is a thing that is actually promoting your
               | own brain... seems counter intuitive but that's how I
               | believe this technology should be used.
        
               | tsunamifury wrote:
               | This technology (which I had a small part in inventing)
               | was not based on intelligently navigating the information
               | space, it's fundamentally based on forecasting your own
               | thoughts by weighting your pre-linguistic vectors and
               | feeding them back to you. Attention layers in conjunction
               | of roof later allowed that to be grouped in higher order
               | and scan a wider beam space to reward higher complexity
               | answers.
               | 
               | When trained on chatting (a reflection system on your own
               | thoughts) it mostly just uses a false mental model to
               | pretend to be a desperate intelligence.
               | 
               | Thus the term stochastic parrot (which for many us
               | actually pretty useful)
        
               | sheeshe wrote:
               | Thanks for your input - great to hear from someone
               | involved that this is the direction of travel.
               | 
               | I remain highly skeptical of this idea that it will
               | replace anyone - the biggest danger I see is people
               | falling for the illusion. That the thing is intrinsically
               | smart when it's not - it can be highly useful in the
               | hands of disciplined people who know a particular area
               | well and augment their productivity no doubt. Because the
               | way we humans come up with ideas and so on is highly
               | complex. Personally my ideas come out of nowhere and
               | mostly are derived from intuition that can only be
               | expressed in logical statements ex-post.
        
               | JAlexoid wrote:
               | Is intuition really that different than LLM having little
               | knowledge about something? It's just responding with the
               | most likely sequence of tokens using the most adjacent
               | information to the topic... just like your intuition.
        
               | sheeshe wrote:
               | With all due respect I'm not even going to give a proper
               | response to this... intuition that yields great ideas is
               | based on deep understanding. LLM's exhibit no such thing.
               | 
               | These comparisons are becoming really annoying to read.
        
               | sheeshe wrote:
               | Meant to say prompting*
        
               | paulddraper wrote:
               | You have a subtle slight of hand.
               | 
               | You use the word "plausible" instead of "correct."
        
               | EastLondonCoder wrote:
               | That's deliberate. "Correct" implies anchoring to a truth
               | function the model doesn't have. "Plausible" is what it's
               | actually optimising for, and the disconnect between the
               | two is where most of the surprises (and pitfalls) show
               | up.
               | 
               | As someone else put it well: what an LLM does is
               | confabulate stories. Some of them just happen to be true.
        
               | paulddraper wrote:
               | It absolutely has a correctness function.
               | 
               | That's like saying linear regression produces plausible
               | results. Which is true but derogatory.
        
               | MyOutfitIsVague wrote:
               | Do you have a better word that describes "things that
               | look correct without definitely being so"? I think
               | "plausible" is the perfect word for that. It's not a
               | sleight of hand to use a word that is exactly defined as
               | the intention.
        
               | tsunamifury wrote:
               | Hallucinations are a feature of reality that LLMs have
               | inherited.
               | 
               | It's amazing that experts like yourself who have a good
               | grasp of the manifold MoE configuration don't get that.
               | 
               | LLMs much like humans weight high dimensionality across
               | the entire model then manifold then string together an
               | attentive answer best weighted.
               | 
               | Just like your doctor occasionally giving you wrong
               | advice too quickly so does this sometimes either get
               | confused by lighting up too much of the manifold or
               | having insufficient expertise.
        
               | jakewins wrote:
               | I asked Gemini the other day to research and summarise
               | the pinout configuration for CANbus outputs on a list of
               | hardware products, and to provide references for each. It
               | came back with a table summarising pin outs for each of
               | the eight products, and a URL reference for each.
               | 
               | Of the 8, 3 were wrong, and the references contained no
               | information about pin outs whatsoever.
               | 
               | That kind of hallucination is, to me, entirely different
               | than what a human researcher would ever do. They would
               | say "for these three I couldn't find pinouts" or perhaps
               | misread a document and mix up pinouts from one model for
               | another.. they wouldn't _make up_ pinouts and reference a
               | document that had no such information in it.
               | 
               | Of course humans also imagine things, misremember etc,
               | but what the LLMs are doing is something entirely
               | different, is it not?
        
               | JAlexoid wrote:
               | Newer models can run a search and summarize the pages.
               | They're becoming just a faster way of doing research, but
               | they're still not as good as humans.
        
               | fspeech wrote:
               | Humans are also not rewarded for making pronouncements
               | all the time. Experts actually have a reputation to
               | maintain and are likely more reluctant to give opionions
               | that they are not reasonably sure of. LLMs trained on
               | typical written narratives found in books, articles etc
               | can be forgiven to think that they should have an
               | opionion on any and everything. Point being that while
               | you may be able to tune it to behave some other way you
               | may find the new behavior less helpful.
        
               | freejazz wrote:
               | > Hallucinations are a feature of reality that LLMs have
               | inherited.
               | 
               | Really? When I search for cases on LexisNexis, it does
               | not return made-up cases which do not actually exist.
        
               | acdha wrote:
               | > Hallucinations are a feature of reality that LLMs have
               | inherited.
               | 
               | Huh? Are you arguing that we still live in a pre-
               | scientific era where there's no way to measure truth?
               | 
               | As a simple example, I asked Google about houseplant
               | biology recently. The answer was very confidently wrong
               | telling me that spider plants have a particular metabolic
               | pathway because it confused them with jade plants and the
               | two are often mentioned together. Humans wouldn't make
               | this mistake because they'd either know the answer or say
               | that they don't. LLMs do that constantly because they
               | lack understanding and metacognitive abilities.
        
               | airstrike wrote:
               | It's not even a manifold https://arxiv.org/abs/2504.01002
        
               | JAlexoid wrote:
               | I mean... That is exactly how our memory works. So in a
               | sense, the factually incorrect information coming from
               | LLM is as reliable as someone telling you things from
               | memory.
        
               | dgacmu wrote:
               | But not really? If you ask me a question about Thai
               | grammar or how to build a jet turbine, I'm going to tell
               | you that I don't have a clue. I have more of a meta-
               | cognitive map of my own manifold of knowledge than an LLM
               | does.
        
             | ryoshu wrote:
             | LLMs fail at causal accuracy. It's a fundamental problem
             | with how they work.
        
           | basisword wrote:
           | I think the thing even worse than false information is the
           | almost-correct information. You do a quick Google to confirm
           | it's on the right page but find there's an important
           | misunderstanding. These are so much harder to spot I think
           | than the blatantly false.
        
           | RHSman2 wrote:
           | The problem is not the intelligence of the LLM. It is the
           | intelligence and desire to make things easy of the
           | intelligence using them.
        
         | andai wrote:
         | So there's two levels to this problem.
         | 
         | Retrieval.
         | 
         | And then hallucination even in the face of perfect context.
         | 
         | Both are currently unsolved.
         | 
         | (Retrieval's doing pretty good but it's a Rube Goldberg machine
         | of workarounds. I think the second problem is a much bigger
         | issue.)
        
           | cachius wrote:
           | Re: retrieval: That's where the snake eats its tail as AI
           | slop floods the web, grounding is like laying a foundation in
           | a swamp. And that Rube Goldberg machine tries to prevent the
           | snake from reaching its tail. But RGs are brittle and not
           | exactly the thing you want to build infrstructure on. Just
           | look at https://news.ycombinator.com/item?id=46239752 for an
           | example how easy it can break.
        
         | cachius wrote:
         | Grounding in search results is what Perplexity pioneered and
         | Google also does with AI mode and ChatGPT and others with web
         | search tool.
         | 
         | As a user I want it but as webadmin it kills dynamic pages and
         | that's why Proof of work aka CPU time captchas like Anubis
         | https://github.com/TecharoHQ/anubis#user-content-anubis or
         | BotID https://vercel.com/docs/botid are now everywhere. If only
         | these AI crawlers did some caching, but no just go and overrun
         | the web. To the effect that they can't anymore, at the price of
         | shutting down small sites and making life worse for everyone,
         | just for few months of rapacious crawling. Literally Perplexity
         | moved fast and broke things.
        
           | cachius wrote:
           | _This dance to get access is just a minor annoyance for me,
           | but I question how it proves I'm not a bot. These steps can
           | be trivially and cheaply automated.
           | 
           | I think the end result is just an internet resource I need is
           | a little harder to access, and we have to waste a small
           | amount of energy._
           | 
           | From Tavis Ormandy who wrote a C program to solve the Anubis
           | challenges out of browser
           | https://lock.cmpxchg8b.com/anubis.html via
           | https://news.ycombinator.com/item?id=45787775
           | 
           | Guess a mix of Markov tarpits and llm meta instructions will
           | be added, cf. Feed the bots
           | https://news.ycombinator.com/item?id=45711094 and Nephentes
           | https://news.ycombinator.com/item?id=42725147
        
         | anentropic wrote:
         | Yeah I basically always use "web search" option in ChatGPT for
         | this reason, if not using one of the more advanced modes.
        
         | jillesvangurp wrote:
         | It's increasingly a space that is constrained by the tools and
         | integrations. Models provide a lot of raw capability. But with
         | the right tools even the simpler, less capable models become
         | useful.
         | 
         | Mostly we're not trying to win a nobel prize, develop some
         | insanely difficult algorithm, or solve some silly leetcode
         | problem. Instead we're doing relatively simple things. Some of
         | those things are very repetitive as well. Our core job as
         | programmers is automating things that are repetitive. That
         | always was our job. Using AI models to do boring repetitive
         | things is a smart use of time. But it's nothing new. There's a
         | long history of productivity increasing tools that take boring
         | repetitive stuff away. Compilation used to be a manual process
         | that involved creating stacks of punch cards. That's what the
         | first automated compilers produced as output: stacks of punch
         | cards. Producing and stacking punchcards is not a fun job. It's
         | very repetitive work. Compilers used to be people compiling
         | punchcards. Women mostly, actually. Because it was considered
         | relatively low skilled work. Even though it arguably wasn't.
         | 
         | Some people are very unhappy that the easier parts of their job
         | are being automated and they are worried that they get
         | completely automated away completely. That's only true if you
         | exclusively do boring, repetitive, low value work. Then yes,
         | your job is at risk. If your work is a mix of that and some
         | higher value, non repetitive, and more fun stuff to work on,
         | your life could get a lot more interesting. Because you get to
         | automate away all the boring and repetitive stuff and spend
         | more time on the fun stuff. I'm a CTO. I have lots of fun
         | lately. Entire new side projects that I had no time for
         | previously I can now just pull off in a spare few hours.
         | 
         | Ironically, a lot of people currently get the worst of both
         | worlds because they now find themselves baby sitting AIs doing
         | a lot more of the boring repetitive stuff than they would be
         | able to do without that to the point where that is actually all
         | that they do. It's still boring and repetitive. And it should
         | be automated away ultimately. Arguably many years ago actually.
         | The reason so many react projects feel like Ground Hog Day is
         | because they are very repetitive. You need a login screen, and
         | a cookies screen, and a settings screen, etc. Just like the
         | last 50 projects you did. Why are you rebuilding those things
         | from scratch? Manually? These are valid questions to ask
         | yourself if you are a frontend programmer. And now you have AI
         | to do that for you.
         | 
         | Find something fun and valuable to work on and AI gets a lot
         | more fun because it gives you more quality time with the fun
         | stuff. AI is about doing more with less. About raising the
         | ambition level.
        
         | fauigerzigerk wrote:
         | I agree, but the question is how better grounding can be
         | achieved without a major research breakthrough.
         | 
         | I believe the real issue is that LLMs are still so bad at
         | reasoning. In my experience, the worst hallucinations occur
         | where only handful of sources exist for some set of facts (e.g
         | laws of small countries or descriptions of niche products).
         | 
         | LLMs know these sources and they refer to them but they are
         | interpreting them incorrectly. They are incapable of focusing
         | on the semantics of one specific page because they get
         | "distracted" by their pattern matching nature.
         | 
         | Now people will say that this is unavoidable given the way in
         | which transformers work. And this is true.
         | 
         | But shouldn't it be possible to include some measure of data
         | sparsity in the training so that models know when they don't
         | know enough? That would enable them to boost the weight of the
         | context (including sources they find through inference time
         | search/RAG) relative to to their pretraining.
        
           | balder1991 wrote:
           | Anything that is very specific has the same problem, because
           | LLMs can't have the same representation of all topics in the
           | training. It doesn't have to be too niche, just specific
           | enough for it to start to fabricate it.
           | 
           | One of these days I had a doubt about something related to
           | how pointers work in Swift and I tried discussing with
           | ChatGPT (don't remember exactly what, but it was purely
           | intellectual curiosity). It gave me a lot of explanations
           | that seemed correct, but being skeptical and started pushing
           | it for ways to confirm what it was saying and eventually
           | realized it was all bullshit.
           | 
           | This kind of thing makes me basically wary of using LLMs for
           | anything that isn't brainstorming, because anything that
           | requires knowing information that isn't easily/plentifully
           | found online will likely be incorrect or have sprinkles of
           | incorrect all over the explanations.
        
         | withinboredom wrote:
         | I was using an LLM to summarize benchmarks for me, and I
         | realized after awhile it was omitting information that made the
         | algorithm being benchmarked look bad. I'm glad I caught it
         | early, before I went to my peers and was like "look at this
         | amazing algorithm".
        
           | coffeecat wrote:
           | It's important not to assume that LLMs are giving you an
           | impartial perspective on any given topic. The perspective
           | you're most likely getting is that of whoever created the
           | most training data related to that topic.
        
         | BatteryMountain wrote:
         | My biggest problem with LLM's at this point is that they
         | produce different and inconsistent results or behave
         | differently, given the same prompt. The better grounding would
         | be amazing at this point. I want to give an LLM the same prompt
         | on different days and I want to be able to trust that it will
         | do the same thing as yesterday. Currently they misbehave
         | multiple times a week and I have to manually steer it a bit
         | which destroys certain automated workflows completely.
        
           | conception wrote:
           | You need to change the temperature to 0 and tune your prompts
           | for automated workflows.
        
             | balder1991 wrote:
             | It doesn't really solve it as a slight shift in the prompt
             | can have totally unpredictable results anyway. And if your
             | prompt is always exactly the same, you'd just cache it and
             | bypass the LLM anyway.
             | 
             | What would really be useful is a very similar prompt should
             | always give a very very similar result.
        
               | tsunamifury wrote:
               | That's a way different problem my guy.
        
               | jknightco wrote:
               | This doesn't work with the current architecture, because
               | we have to introduce some element of stochastic noise
               | into the generation or else they're not "creatively"
               | generative.
               | 
               | Your brain doesn't have this problem because the noise is
               | already present. You, as an actual thinking being, are
               | able to override the noise and say "no, this is false."
               | An LLM doesn't have that capability.
        
               | sheeshe wrote:
               | Well that's because if you look at the structure of the
               | brain there's a lot more going on than what goes on
               | within an LLM.
               | 
               | It's the same reason why great ideas almost appear to
               | come randomly - something is happening in the background.
               | Underneath the skin.
        
           | fragmede wrote:
           | It sounds like you have dug into this problem with some depth
           | so I would love to hear more. When you've tried to automate
           | things, I'm guessing you've got a template and then some data
           | and then the same or similar input gives totally different
           | results? What details about how different the results are can
           | you share? Are you asking for eg JSON output and it totally
           | isn't, or is it a more subtle difference perhaps?
        
           | sebastiennight wrote:
           | > I want to give an LLM the same prompt on different days and
           | I want to be able to trust that it will do the same thing as
           | yesterday
           | 
           | Bad news, it's winter now in the Northern hemisphere, so
           | expect all of our AIs to get slightly less performant as they
           | emulate humans under-performing until Spring.
        
         | virtuosarmo wrote:
         | I've had better success finding information using Google Gemini
         | vs. ChatGPT. I.e. someone mentions to me the name of someone or
         | some company, but doesn't give the full details (i.e. Joe @ XYZ
         | Company doing this, or this company with 10,000 people, in ABC
         | industry)...sometimes i don't remember the full name. Gemini
         | has been more effective for me in filling in the gaps and doing
         | fuzzy search. I even asked ChatGPT why this was the case, and
         | it affirmed my experience, saying that Gemini is better for
         | these queries because of Search integration, Knowledge Graph,
         | etc. Especially useful for recent role changes, which haven't
         | been propagated through other channels on a widespread basis.
        
         | rafaelmn wrote:
         | I constantly see top models (opus 4.5, gemini 3) get a stroke
         | mid task - they will solve the problem correctly in one place,
         | or have a correct solution that needs to be reapplied in
         | context - and then completely miss the mark in another place.
         | "Lack of intelligence" is very much a limiting factor. Gemini
         | especially will get into random reasoning loops - reading
         | thinking traces - it gets unhinged pretty fast.
         | 
         | Not to mention it's super easy to gaslight these models, just
         | asserting something wrong with vaguely plausible explanation
         | and you get no pushback or reasoning validation.
         | 
         | So I know you qualified your post with "for your use case", but
         | personally I would very much like more intelligence from LLMs.
        
         | giancarlostoro wrote:
         | Yeah in my case I want the coding models to be less stupid, I
         | asked for multiple file uploading, it kept the original button
         | and it added a second one for additional files, when I pointed
         | that out "You're absolutely correct!" Well why didnt you think
         | of it before you cranked out code, I see coding agents as
         | really capable Junior devs its really funny. I dont mind it
         | though, saved me hours on my side project if not weeks worth of
         | work.
        
         | BrtByte wrote:
         | I'm pretty much in the same camp. For a lot of everyday use,
         | raw "intelligence" already feels good enough
        
         | HeavyStorm wrote:
         | All of them are heavily invested in improving grounding. The
         | money isn't on personal use but enterprise customers and for
         | those, grounding is essential.
        
         | sebastiennight wrote:
         | > It's still a big issue that the models will make up plausible
         | sounding but wrong or misleading explanations for things,
         | 
         | Due to how LLMs are implemented, you are always most likely to
         | get a bogus explanation if you ask for an answer first, and why
         | second.
         | 
         | A useful mental model is: imagine if I presented you with a
         | potential new recruit's complete data (resume, job history,
         | recordings of the job interview, everything) but you only had 1
         | second to tell me "hired: YES OR NO"
         | 
         | And then, AFTER you answered that, I gave you 50 pages worth of
         | space to tell me why your decision is right. You can't go back
         | on that decision, so all you can do is justify it however you
         | can.
         | 
         | Do you see how this would give radically different outcomes vs.
         | giving you the 50-page scratchpad first to think things
         | through, and then only giving me a YES/NO answer?
        
       | jbkkd wrote:
       | A new model doesn't address the fundamental reliability issues
       | with OpenAI's enterprise tier.
       | 
       | As an enterprise customer, the experience has been disappointing.
       | The platform is unstable, support is slow to respond even when
       | escalated to account managers, and the UI is painfully slow to
       | use. There are also baffling feature gaps, like the lack of
       | connectors for custom GPTs.
       | 
       | None of the major providers have a perfect enterprise solution
       | yet, but given OpenAI's market position, the gap between
       | expectations and delivery is widening.
        
         | sigmoid10 wrote:
         | Which tier are you? We are on the highest enterprise tier and
         | I've found that OpenAI is a much more stable platform for high-
         | usage than other providers. Can't say much about the UI though
         | since I almost exclusively work with the API. I feel like UIs
         | generally suck everywhere unless you want to do really generic
         | stuff.
        
           | energy123 wrote:
           | ChatGPT UI is leagues above Gemini and AI Studio in
           | responsiveness and latency which is what I care about.
        
         | dannyw wrote:
         | Completely the opposite experience.
        
       | elAhmo wrote:
       | This feels like "could've been an email" type of thing, a very
       | incremental update that just adds one more version. I bet there
       | is literally no one in the world who wanted *one more version of
       | GPT* in the list of available models from OpenAI.
       | 
       | "All models" section on https://platform.openai.com/docs/models
       | is quite ridiculous.
        
         | tim333 wrote:
         | It's significant because it looked like they were falling
         | behind Gemini and maybe others.
        
       | m12k wrote:
       | So, does 5.2 still have a knowledge cutoff date of June 2024, or
       | have they managed to complete another full pre-training run?
        
       | bob1029 wrote:
       | I've been looking really hard at combining Roslyn (.NET compiler
       | platform SDK) with one of these high end tool calling models. The
       | ability to have the LLM create custom analyzers and then verify
       | them with a human in the loop can provide stable, compile-time
       | guarantees of business rules that accumulate without paying for
       | context tokens.
       | 
       | I feel like there is a small chance I could actually make this
       | work in some areas of the business now. 400k is a really big
       | context window. The last time I made any serious attempt I only
       | had 32k tokens to work with. I still don't think these things can
       | build the whole product for you, but if you have a structured
       | configuration abstraction in an _existing_ product, I think there
       | is definitely uplift possible.
        
         | schmuhblaster wrote:
         | Sounds interesting, could you elaborate a bit on this? (I am
         | experimenting in a similar direction)
        
       | dev1ycan wrote:
       | How many years of the world's DRAM production capacity is it this
       | time?
        
       | throwaway2037 wrote:
       | Somewhat tangential: The second link says "System card":
       | https://cdn.openai.com/pdf/3a4153c8-c748-4b71-8e31-aecbde944...
       | 
       | Does that term have special meaning in the AI/LLM world? I never
       | heard it before. I Google'd the term "System Card LLM" and got a
       | bunch of hits. I am so surprised that I never saw the term used
       | here in HN before.
       | 
       | Also, the layout looks exactly like a scientific paper written in
       | LaTeX. Who is the expected audience for this paper?
        
         | tylerrobinson wrote:
         | The major model providers use system cards as a sort of self
         | attestation document like a nutrition label. It's been around
         | for a couple years.
        
         | blinding-streak wrote:
         | Yeah, search HN for the term. It's a relatively big topic of
         | conversation.
        
       | EastLondonCoder wrote:
       | I've been using GPT-4o and now 5.2 pretty much daily, mostly for
       | creative and technical work. What helped me get more out of it
       | was to stop thinking of it as a chatbot or knowledge engine, and
       | instead try to model how it actually works on a structural level.
       | 
       | The closest parallel I've found is Peter Gardenfors' work on
       | conceptual spaces, where meaning isn't symbolic but geometric.
       | Fedorenko's research on predictive sequencing in the brain fits
       | too. In both cases, the idea is that language follows a
       | trajectory through a shaped mental space, and that's basically
       | what GPT is doing. It doesn't know anything, but it generates
       | plausible paths through a statistical terrain built from our own
       | language use.
       | 
       | So when it "hallucinates", that's not a bug so much as a result
       | of the system not being grounded. It's doing what it was designed
       | to do: complete the next step in a pattern. Sometimes that's
       | wildly useful. Sometimes it's nonsense. The trick is knowing
       | which is which.
       | 
       | What's weird is that once you internalise this, you can work with
       | it as a kind of improvisational system. If you stay in the loop,
       | challenge it, steer it, it feels more like a collaborator than a
       | tool.
       | 
       | That's how I use it anyway. Not as a source of truth, but as a
       | way of moving through ideas faster.
        
         | ostacke wrote:
         | Interesting concept with conceptual spaces, but how does that
         | affect how you work with LLM:s in practice?
        
           | EastLondonCoder wrote:
           | I think of it like improvising with a very skilled but
           | slightly alien musician.
           | 
           | If you just hand it a chord chart, it'll follow the
           | structure. But if you understand the kinds of patterns it
           | tends to favour, the statistical shapes it moves through, you
           | can start composing with it, not just prompting it.
           | 
           | That's where Gardenfors helped me reframe things. The model
           | isn't retrieving facts. It's traversing a conceptual space.
           | Once you stop expecting grounded truth and start tracking
           | coherence, internal consistency, narrative stability, you get
           | a much better sense of where it's likely to go off course.
           | 
           | It reminds me of salespeople who speak fluently without being
           | aligned with the underlying subject. Everything sounds
           | plausible, but something's off. LLMs do that too. You can
           | learn to spot the mismatch, but it takes practice, a bit like
           | learning to jam. You stop reading notes and start listening
           | for shape.
        
         | BrtByte wrote:
         | Once you drop the idea that it's a knowledge oracle and start
         | treating it as a system that navigates a probability landscape,
         | a lot of the confusion just evaporates
        
       | rl_shannon wrote:
       | Isn't it delusional to only compare your models against your own
       | previous variants? Where is an actual comparison with Google,
       | Anthropic, OSS Models
        
         | zild3d wrote:
         | it's the best ____ we've ever made
        
       | keepamovin wrote:
       | It is significantly better than 5.1 .. testing now with codex.
       | It's much more focused, perceptive and efficient.
        
       | loa_observer wrote:
       | does the model really improve? i tried several tasks today, and
       | most of them failed, which are super easy ones.
       | 
       | maybe it's just because the gpt5.2 in cursor is super stupid?
        
       | atheljcarlton wrote:
       | It's dog-doo-doo. I put in my algebraic geometry final review
       | (100's of thousands of tokens) and Gemini instantly found all the
       | propositions, theorems, and problems that I needed in a neat list
       | (in about 5 seconds), meanwhile ChatGPT 5.2 Thinking took 10mins
       | before timing out and not even completing the request.
        
         | atheljcarlton wrote:
         | However, the model card for GPT 5.2 looks amazing, wish I could
         | actually see that performance in action!
        
       ___________________________________________________________________
       (page generated 2025-12-12 23:01 UTC)