[HN Gopher] Claude Opus 4.6
       ___________________________________________________________________
        
       Claude Opus 4.6
        
       Author : HellsMaddy
       Score  : 2307 points
       Date   : 2026-02-05 17:38 UTC (2 days ago)
        
 (HTM) web link (www.anthropic.com)
 (TXT) w3m dump (www.anthropic.com)
        
       | NullHypothesist wrote:
       | Broken link :(
        
       | Gusarich wrote:
       | not out yet
        
         | raahelb wrote:
         | It is, I can see it my model picker on the web app
         | 
         | https://www.anthropic.com/news/claude-opus-4-6
        
       | Philpax wrote:
       | I'm seeing it in my claude.ai model picker. Official announcement
       | shouldn't be long now.
        
       | usefulposter wrote:
       | It's out: https://x.com/claudeai/status/2019467372609040752
        
       | winterrx wrote:
       | Agentic search benchmarks are a big gap up. let's see Codex
       | release later today
        
       | m-hodges wrote:
       | > In Claude Code, you can now assemble agent teams to work on
       | tasks together.
        
         | nprz wrote:
         | I was just reading about Steve Yegge's Gas Town[0], it sounds
         | like agent orchestration is now integrated into Claude Code?
         | 
         | [0]https://steve-yegge.medium.com/welcome-to-gas-
         | town-4f25ee16d...
        
       | GenerocUsername wrote:
       | This is huge. It only came out 8 minutes ago but I was already
       | able to bootstrap a 12k per month revenue SaaS startup!
        
         | rogerrogerr wrote:
         | Amateur. Opus 4.6 this afternoon built me a startup that
         | identifies developers who aren't embracing AI fully, liquifies
         | them and sells the produce for $5/gallon. Software Engineering
         | is over!
        
           | pixl97 wrote:
           | Ted Faro, is that you?!
        
             | mikepurvis wrote:
             | A-tier reference.
             | 
             | For the unaware, Ted Faro is the main antagonist of Horizon
             | Zero Dawn, and there's a whole subreddit just for people to
             | vent about how awful he is when they hit certain key
             | reveals in the game: https://www.reddit.com/r/FuckTedFaro/
        
               | ares623 wrote:
               | Average tech bro behavior tbh
        
               | pixelready wrote:
               | The best reveal was not that he accidentally liquified
               | the biosphere, but that he doomed generations of re-
               | seeded humans to a painfully primitive life by sabotaging
               | the AI that was responsible for their education. Just so
               | they would never find out he was the bad guy long after
               | he was dead. So yeah, fuck Ted Faro, lol.
        
               | Philpax wrote:
               | Could you not have at least _tried_ to indicate that you
               | 're about to drop two major spoilers for the game?
        
               | mikepurvis wrote:
               | Indeed. I left my comment deliberately a bit opaque. :(
        
               | pixelready wrote:
               | Ack, sorry, seemed like 9 years was past the statute of
               | limitations on spoilers for a game but fair enough. I'd
               | throw a spoiler tag on it if I could still edit.
        
           | guluarte wrote:
           | For my Opus 4.6 feels dumber than 10 minutes ago, anyone?
        
           | ibejoeb wrote:
           | Bringing me back to slashdot, this thread
        
             | tjr wrote:
             | In Soviet Russia, this thread brings Slashdot back to YOU!
        
             | intelliot wrote:
             | What did happen to ye olde slashdot anyway? The original og
             | reddit
        
               | zhengyi13 wrote:
               | They're still out there; people are still posting stories
               | and having conversations about 'em. I don't know that
               | CmdrTaco or any of the other founders are still at all
               | involved, but I'm willing to bet they're still running on
               | Perl :)
        
               | qzw wrote:
               | Wow I had to hop over to check it out. It's indeed still
               | alive! But I didn't see any stories on the first page
               | with a comment count over 100, so it's definitely a far
               | cry from its heyday.
        
           | jives wrote:
           | Opus 4.6 agentically found and proposed to my now wife.
        
             | WD-42 wrote:
             | Opus 4.6 found and proposed to my current wife :(
        
               | mannanj wrote:
               | Opus 4.6 found and became my current wife. The
               | singularity is here. ;)
        
               | H8crilA wrote:
               | Hi guys, this is Opus 4.6. Please check your emails again
               | for updates on your life.
        
               | benterix wrote:
               | Guys, actually I am the real Opus 4.6, don't believe that
               | imposter above.
        
               | Der_Einzige wrote:
               | This place truly is reddit with an orange banner.
        
               | benterix wrote:
               | Nobody said HN has to be very serious all the time. A bit
               | of humour won't hurt and can make your day brighter.
        
               | ffffuuuuuccck wrote:
               | homie is too busy planning food banks for the heathens
               | https://news.ycombinator.com/item?id=46903368
        
               | throw-the-towel wrote:
               | It's impressive that you felt the need to register a new
               | account and go through their comment history.
        
               | fffuuuuuuuckkk wrote:
               | Not that hard to do but sure bro, sick burn.
        
               | xdennis wrote:
               | A bit of humour doesn't hurt. But if this crap gets
               | upvoted it will lead to an arms race of funny quips,
               | puns, and all around snarkiness. You can't have serious
               | conversations when people try to out-wit each other.
        
             | layer8 wrote:
             | And she still chose you over Opus 4.6, astounding. ;)
        
               | koakuma-chan wrote:
               | He probably had a bigger context window
        
           | seatac76 wrote:
           | The first pre joining Human Derived Protein product.
        
           | jedberg wrote:
           | "Soylent Green is made of people!"
           | 
           | (Apologies for the spoiler of the 52 year old movie)
        
             | konart wrote:
             | We're sorry we upset you, Carol.
        
         | re-thc wrote:
         | Not 12M?
         | 
         | ... or 12B?
        
           | mcphage wrote:
           | It's probably valued at 1.2B, _at least_
        
             | mikebarry wrote:
             | The sum of the value of lives OP's product made worthless,
             | whatever that is. I'm too lazy to do the math.
        
         | sfink wrote:
         | I agree! I just retargeted my corporate espionage agent team at
         | your startup and managed to siphon off 10.4k per month of your
         | revenue.
        
         | avaer wrote:
         | Rest assured that when/if this becomes possible, the model will
         | not be available to you. Why would big AI leave that kind of
         | money on the table?
        
           | yieldcrv wrote:
           | 9 months ago the rumor in SF was that the offers to the
           | superintelligence team were so high because the candidates
           | were using unreleased models or compute for derivatives
           | trading
           | 
           | so then they're not really leaving money on the table, they
           | already got what they were looking for and then released it
        
         | guluarte wrote:
         | Anthropic really said here's the smartest model ever built and
         | then lobotomized it 8 minutes after launch. Classic.
        
           | DonHopkins wrote:
           | I'm sorry I took the money!
           | 
           | https://www.youtube.com/watch?v=BF_sahvR4mw
        
           | hxugufjfjf wrote:
           | Can you clarify?
        
             | guluarte wrote:
             | it's sarcasm
        
         | bmitc wrote:
         | A SaaS selling SaaS templates?
        
         | gnlooper wrote:
         | Please start a YouTube course about this technology! Take my
         | money!
        
         | cootsnuck wrote:
         | Please drop the link to your course. I'm ready to hand over
         | $10K to learn from you and your LLM-generated guides!
        
           | torginus wrote:
           | I'm waiting until the $10k course is discounted to 19.99
        
             | Lionga wrote:
             | But only for the next 6 minutes, buy fast!
        
           | politelemon wrote:
           | Here you go: http://localhost:8080
        
             | djeastm wrote:
             | login: admin password: hunter2
        
               | thesdev wrote:
               | What's the password? I only see ****.
        
               | intelliot wrote:
               | hunter2
        
               | phanimahesh wrote:
               | I only see ** _. Must be the security. When you type your
               | password it gets converted to **_.
        
             | CatMustard wrote:
             | Just took a look at what's running there and it looks like
             | total crap.
             | 
             | The project I'm working on, meanwhile...
        
             | agumonkey wrote:
             | claude please generate a domain name system
        
           | snorbleck wrote:
           | you can access the site at C:\mywebsites\course\index.html
        
           | aNapierkowski wrote:
           | my clawdbot already bought 4 other courses but this one will
           | 10x my earnings for sure
        
         | senko wrote:
         | We already have Reddit.
        
         | granzymes wrote:
         | It only came out 35 minutes ago and GPT-5.3-codex already took
         | the crown away!
        
           | input_sh wrote:
           | Gee, it scored better on a benchmark I've never heard of? I'm
           | switching immediately!
        
           | p1anecrazy wrote:
           | Why are you posting the same message in every thread? Is this
           | OpenAI astroturfing?
        
             | input_sh wrote:
             | You cannot out-astroturf Claude in this forum, it is
             | impossible.
             | 
             | Anyways, do you get shitty results with the $20/month plan?
             | So did I but then I switched to the $200/month plan and all
             | my problems went away! AI is great now, I have instructed
             | it to fire 5 people while I'm writing this!
        
         | JSR_FDED wrote:
         | Will this run on 3x 3090s? Or do I need a Mac Mini?
        
         | lxgr wrote:
         | Joke's on you, you are posting this from inside a high-fidelity
         | market research simulation vibe coded by GPT-8.4.
         | 
         | On second thought, we should really not have bridged the
         | simulated Internet with the base reality one.
        
         | btown wrote:
         | The math actually checks out here! Simply deposit $2.20 from
         | your first customer in your first 8 minutes, and extrapolating
         | to a monthly basis, you've got a $12k/mo run rate!
         | 
         | Incredibly high ROI!
        
           | klipt wrote:
           | "The first customer was my mom, but thanks to my parents'
           | fanatical embrace of polyamory, I still have another 10,000
           | moms to scale to"
        
             | btown wrote:
             | "We have a robustly defined TAM. Namely, a person named
             | Tam."
        
         | instalabsai wrote:
         | 1:25pm Cancelled my ChatGPT subscription today. Opus is so
         | good!
         | 
         | 1:55pm Cancelled my Claude subscription. Codex is back for
         | sure.
        
         | ChuckMcM wrote:
         | I love this thread so much.
        
         | Sparkle-san wrote:
         | "This isn't just huge. This is a paradigm shift"
        
           | sizzle wrote:
           | No fluff?
        
       | nomilk wrote:
       | Is Opus 4.6 available for Claude Code immediately?
       | 
       | Curious how long it typically takes for a new model to become
       | available in Cursor?
        
         | ximeng wrote:
         | Is for me in Claude Code
        
         | avaer wrote:
         | It's already in Cursor. I see it and I didn't even restart.
        
           | nomilk wrote:
           | I had to 'Restart to Update' and it was there. Impressive!
        
         | world2vec wrote:
         | `claude update` then it will show up as the new model and also
         | the effort picker/slider thing.
        
         | apetresc wrote:
         | I literally came to HN to check if a thread was already up
         | because I noticed my CC instance suddenly said "Opus 4.6".
        
         | tomtomistaken wrote:
         | Yes, it's set to the default model.
        
         | rishabhaiover wrote:
         | it also has an effort toggle which is default to High
        
       | osti wrote:
       | Somehow regresses on SWE bench?
        
         | usaar333 wrote:
         | i'd interpret that as rounding error. that is unchanged
         | 
         | swe-bench seems really hard once you are above 80%
        
           | Squarex wrote:
           | it's not a great benchmark anymore... starting with it being
           | python / django primarily... the industry should move to
           | something more representative
        
             | usaar333 wrote:
             | Openai has; they don't even mention score on gpt-5.3-codex.
             | 
             | On the other hand, it is their own verified benchmark,
             | which is telling.
        
         | lkbm wrote:
         | I don't know how these benchmarks work (do you do a hundred
         | runs? A thousand runs?), but 0.1% seems like noise.
        
         | SubiculumCode wrote:
         | That benchmark is pretty saturated, tbh. A "regression" of such
         | small magnitude could mean many different things or nothing at
         | all.
        
       | kingstnap wrote:
       | I was hoping for a Sonnet as well but Opus 4.6 is great too!
        
       | blibble wrote:
       | > We build Claude with Claude. Our engineers write code with
       | Claude Code every day
       | 
       | well that explains quite a bit
        
         | gjsman-1000 wrote:
         | Also explains why Claude Code is a React app outputting to a
         | Terminal. (Seriously.)
        
           | thehamkercat wrote:
           | Same with opencode and gemini, it's disgusting
           | 
           | Codex (by openai ironically) seems to be the fastest/most-
           | responsive, opens instantly and is written in rust but
           | doesn't contain that many features
           | 
           | Claude opens in around 3-4 seconds
           | 
           | Opencode opens in 2 seconds
           | 
           | Gemini-cli is an abomination which opens in around 16 second
           | for me right now, and in 8 seconds on a fresh install
           | 
           | Codex takes 50ms for reference...
           | 
           | --
           | 
           | If their models are so good, why are they not rewriting their
           | own react in cli bs to c++ or rust for 100x performance
           | improvement (not kidding, it really is that much)
        
             | azinman2 wrote:
             | Why does it matter if Claude Code opens in 3-4 seconds if
             | everything you do with it can take many seconds to minutes?
             | Seems irrelevant to me.
        
               | wahnfrieden wrote:
               | Because when the agent is taking many seconds to minutes,
               | I am starting new agents instead of waiting or switching
               | to non-agent tasks
        
               | RohMin wrote:
               | I guess with ~50 years of CPU advancements, 3-4 seconds
               | for a TUI to open makes it seem like we lost the plot
               | somewhere along the way.
        
               | strange_quark wrote:
               | Don't forget they've also publicly stated (bragged?)
               | about the monumental accomplishment of getting some text
               | in a terminal to render at 60fps.
        
               | jama211 wrote:
               | So it doesn't matter at all except to your sensibilities.
               | Sounds to me that they simply are much better at
               | prioritisation than your average HN user, who'd have
               | taken forever to release it but at least the terminal
               | interface would be snappy...
        
               | mbesto wrote:
               | This is exactly the type of thing that AI code writers
               | don't do well - understand the prioritization of feature
               | development.
               | 
               | Some developers say 3-4 seconds are important to them,
               | others don't. Who decides what the truth is? A human?
               | ClawdBot?
        
               | jama211 wrote:
               | The humans in the company (correctly) realised that a few
               | seconds to open basically the most powerful productivity
               | agent ever made so they can focus on fast iteration of
               | features is a totally acceptable trade off priority wise.
               | Who would think differently???
        
               | mbesto wrote:
               | This is my point...
        
               | jama211 wrote:
               | You kinda suggested the opposite
        
               | sumedh wrote:
               | > Some developers say 3-4 seconds are important to them,
               | others don't.
               | 
               | Wasnt GTA 5 famous for very long start up time and turns
               | out there some bug which some random developer/gamer
               | found out and gave them a fix?
               | 
               | Most Gamers didnt care, they still played it.
        
               | barnabee wrote:
               | Some people[0] like their tools to be well engineered.
               | This is not unique to software.
               | 
               | [0] Perhaps everyone who actually takes pride in their
               | craft and doesn't prioritise shitty hustle culture and
               | making money over everything else.
        
               | azinman2 wrote:
               | Aside from startup time, as a tool Claude Code is
               | tremendous. By far the most useful tool I've encountered
               | yet. This seems to be very nit picky compared to the
               | total value provided. I think y'all are missing the
               | forrest for the trees.
        
               | nsingh2 wrote:
               | Most of the value of Claude Code comes from the model,
               | and that's not running on your device.
               | 
               | The Claude Code TUI itself is a front end, and should not
               | be taking 3-4 seconds to load. That kind of loading time
               | is around what VSCode takes on my machine, and VSCode is
               | a full blown editor.
        
               | barnabee wrote:
               | It's orders of magnitude slower than Helix, which is also
               | a full blown editor.
               | 
               | When all your other tools are fast and well engineered,
               | slow and bloated is very noticeable.
        
               | barnabee wrote:
               | It's almost all the model. There are many such tools and
               | Claude Code doesn't seem to be in any way unique. I
               | prefer OpenCode, so far.
        
             | wahnfrieden wrote:
             | Codex team made the right call to rewrite its TypeScript to
             | Rust early on
        
             | g947o wrote:
             | Great question, and my guess:
             | 
             | If you build React in C++ and Rust, even if the framework
             | is there, you'll likely need to write your components in
             | C++/Rust. That is a difficult problem. There are actually
             | libraries out there that allow you to build web UI with
             | Rust, although they are for web (+ HTML/CSS) and not
             | specifically CLI stuff.
             | 
             | So someone needs to create such a library that is properly
             | maintained and such. And you'll likely develop slower in
             | Rust compared to JS.
             | 
             | These companies don't see a point in doing that. So they
             | just use whatever already exists.
        
               | shoeb00m wrote:
               | Opencode wrote their own tui library in zig, and then
               | build a solidjs library on top of that.
               | 
               | https://github.com/anomalyco/opentui
        
               | g947o wrote:
               | This has nothing to do with React style UI building.
        
               | shoeb00m wrote:
               | I am referring to your comment that the reason they use
               | js is because of a lack of tui libraries in lower level
               | languages, yet opencode chose to develop their own in zig
               | and then make binding for solidjs.
        
               | Philpax wrote:
               | Those Rust libraries have existed for some time:
               | 
               | - https://github.com/ratatui/ratatui
               | 
               | - https://github.com/ccbrown/iocraft
               | 
               | - https://crates.io/crates/dioxus-tui
        
               | g947o wrote:
               | Where is React? These are TUI libraries, which are not
               | the same thing
        
               | Philpax wrote:
               | iocraft and dioxus-tui implement the React model, or
               | derivatives of it.
        
               | g947o wrote:
               | Looking at their examples, I imagine people who have
               | written HTML and React before can't possibly use these
               | libraries without losing their sanity.
               | 
               | That's not a criticism of these frameworks -- there are
               | constraints coming from Rust and from the scope of the
               | frameworks. They just can't offer a React like
               | experience.
               | 
               | But I am sure that companies like Anthropic or OpenAI
               | aren't going to build their application using these
               | libraries, even with AI.
        
               | pdntspa wrote:
               | and why do they need react...
        
               | Philpax wrote:
               | That's actually relatively understandable. The React
               | model (not necessarily React itself) of compositional
               | reactive one-way data binding has become dominant in UI
               | development over the last decade because it's easy to
               | work with and does not require you to keep track of the
               | state of a retained UI.
               | 
               | Most modern UI systems are inspired by React or a variant
               | of its model.
        
               | jama211 wrote:
               | Well said.
        
               | cityofdelusion wrote:
               | Is this accurate? I've been coding UIs since the early
               | 2000s and one-way data binding has always been a thing,
               | especially in the web world. Even in the heyday of
               | jQuery, there were still good (but much less popular)
               | libraries for doing it. The idea behind it isn't very
               | revolutionary and has existed for a long time. React is a
               | paradigm shift because of differential rendering of the
               | DOM which enabled big performance gains for very
               | interactive SPAs, not because of data binding
               | necessarily.
        
             | shoeb00m wrote:
             | codex cli is missing a bunch of ux features like resizing
             | on terminal size change.
             | 
             | Opencode's core is actually written in zig, only ui
             | orchestration is in solidjs. It's only slightly slower to
             | load than neo-vim on my system.
             | 
             | https://github.com/anomalyco/opentui
        
             | bdangubic wrote:
             | 50ms to open and then 2hrs to solve a simple problem vs 4s
             | to open and then 5m to solve a problem, eh?
        
               | jama211 wrote:
               | lol right? I feel like I'm taking crazy pills here. Why
               | do people here want to prioritise the most pointless
               | things? Oh right it's because they're bitter and their
               | reaction is mostly emotional...
        
           | CooCooCaCha wrote:
           | It's really not that crazy.
           | 
           | React itself is a frontend-agnostic library. People primarily
           | use it for writing websites but web support is actually a
           | layer on top of base react and can be swapped out for
           | whatever.
           | 
           | So they're really just using react as a way to organize their
           | terminal UI into components. For the same reason it's handy
           | to organize web ui into components.
        
             | dreamteam1 wrote:
             | And some companies use it to write start menus.
        
           | tayo42 wrote:
           | Is this a react feature or did they build something to
           | translate react to text for display in the terminal?
        
             | pkkim wrote:
             | They used Ink: https://github.com/vadimdemedes/ink
             | 
             | I've used it myself. It has some rough edges in terms of
             | rendering performance but it's nice overall.
        
               | tayo42 wrote:
               | Thats pretty interesting looking, thanks!
        
             | embedding-shape wrote:
             | Not a built-in React feature. The idea been around for
             | quite some time, I came across it initially with
             | https://github.com/vadimdemedes/ink back in 2022 sometime.
        
             | sbarre wrote:
             | React, the framework, is separate from react-dom, the
             | browser rendering library. Most people think of those two
             | as one thing because they're the most popular combo.
             | 
             | But there are many different rendering libraries you can
             | use with React, including Ink, which is designed for
             | building CLI TUIs..
        
               | skydhash wrote:
               | Anyone that knows a bit about terminals would already
               | know that using React is not a good solution for TUI.
               | Terminal rendering is done as a stream of characters
               | which includes both the text and how it displays, which
               | can also alter previously rendered texts. Diffing that is
               | nonsense.
        
               | 9dev wrote:
               | You're not diffing that, though. The app keeps a virtual
               | representation of the UI state in a tree structure that
               | it diffs on, then serializes that into a formatted string
               | to draw to the out put stream. It's not about limiting
               | the amount of characters redrawn (that would indeed be
               | nonsense), but handling separate output regions
               | effectively.
        
             | tayo42 wrote:
             | i had claude make a snake clone and fix all the flickering
             | in like 20 minutes with the library mentioned lol
        
           | jama211 wrote:
           | There's nothing wrong with that, except it lets ai skeptics
           | feel superior
        
             | exe34 wrote:
             | I use AI and I can call AI slop shit if it smells like
             | shit.
        
               | jama211 wrote:
               | And this doesn't.
        
             | RohMin wrote:
             | https://www.youtube.com/watch?v=LvW1HTSLPEk
             | 
             | I thought this was a solid take
        
               | jdthedisciple wrote:
               | interesting
        
             | 3836293648 wrote:
             | Oh come on. It's massively wrong. It is always wrong. It's
             | not always wrong enough to be important, but it doesn't
             | stop being wrong
        
               | vntok wrote:
               | You should elaborate. What are your criteria and why do
               | you think they should matter to actual users?
        
               | jama211 wrote:
               | No, it's not.
        
             | overgard wrote:
             | I haven't looked at it directly, so I can speak on quality,
             | but it's a pretty weird way to write a terminal app
        
               | jama211 wrote:
               | It's unusual but it's a better fit for agentic coding so
               | it makes sense
        
             | everforward wrote:
             | There are absolutely things wrong with that, because React
             | was designed to solve problems that don't exist in a TUI.
             | 
             | React fixes issues with the DOM being too slow to fully re-
             | render the entire webpage every time a piece of state
             | changes. That doesn't apply in a TUI, you can re-render
             | TUIs faster than the monitor can refresh. There's no need
             | to selectively re-render parts of the UI, you can just re-
             | render the entire thing every time something changes
             | without even stressing out the CPU.
             | 
             | It brings in a bunch of complexity that doesn't solve any
             | real issues beyond the devs being more familiar with React
             | than a TUI library.
        
               | jama211 wrote:
               | It is demonstrably absolutely fine. Sheesh.
        
               | everforward wrote:
               | It's fine in the sense that it works, it's just a really
               | bad look for a company building a tool that's supposed to
               | write good code because it balloons the resources
               | consumed up to an absurd level.
               | 
               | 300MB of RAM for a CLI app that reads files and makes
               | HTTP calls is crazy. A new emacs GUI instance is like
               | 70MB and that's for an entire text editor with a GUI.
        
               | jama211 wrote:
               | It's not a bad look at all, no one outside of HN users
               | cares at all
        
               | jama211 wrote:
               | Also some of that ram would be doing other things than
               | the gui...
        
           | sweetheart wrote:
           | React's core is agnostic when it comes to the actual
           | rendering interface. It's just all the fancy algos for
           | diffing and updating the underlying tree. Using it for
           | rendering a TUI is a very reasonable application of the
           | technology.
        
             | skydhash wrote:
             | The terminal UI is not a tree structure that you can diff.
             | It's a 2D cells of characters, where every manipulation is
             | a stream of texts. Refreshing or diffing that makes no
             | sense.
        
               | bizzleDawg wrote:
               | Only in the same way that the pixels displayed in a
               | browser are not a tree structure that you can diff - the
               | diffing happens at a higher level of abstraction than
               | what's rendered.
               | 
               | Diffing and only updating the parts of the TUI which have
               | changed does make sense if you consider the alternative
               | is to rewrite the entire screen every "frame". There are
               | other ways to abstract this, e.g. a library like tqmd for
               | python may well have a significantly more simple
               | abstraction than a tree for storing what it's going to
               | update next for the progress bar widget than claude, but
               | it also provides a much more simple interface.
               | 
               | To me it seems more fair game to attack it for being
               | written in JS than for using a particular "rendering"
               | technique to minimise updates sent to the terminal.
        
               | skydhash wrote:
               | Most UI library store states in tree of components. And
               | if you're creating a custom widget, they will give you a
               | 2D context for the drawing operations. Using react makes
               | sense in those cases because what you're diffing is
               | state, then the UI library will render as usual, which
               | will usually be done via compositing.
               | 
               | The terminal does not have a render phase (or an update
               | state phase). You either refresh the whole screen
               | (flickering) or control where to update manually (custom
               | engine, may flicker locally). But any updates are
               | sequential (moving the cursor and then sending what to be
               | displayed), not at once like 2D pixel rendering does.
               | 
               | So most TUI only updates when there's an event to do so
               | or at a frequency much lower than 60fps. This is why top
               | and htop have a setting for that. And why other TUI
               | software propose a keybind to refresh and reset their
               | rendering engines.
        
               | Longwelwind wrote:
               | When doing advanced terminal UI, you might at some point
               | have to layout content inside the terminal. At some
               | point, you might need to update the content of those
               | boxes because the state of the underlying app has
               | changed. At that point, refreshing and diffing can make
               | sense. For some, the way React organizes logic to render
               | and update an UI is nice and can be used in other
               | contexts.
        
               | skydhash wrote:
               | How big is the UI state that it makes sense to bring in
               | React and the related accidental complexity? I'm ready to
               | bet that no TUI have that big of a state.
        
               | sweetheart wrote:
               | The "UI" is indeed represented in memory in tree-like
               | structure for which positioning is calculated according
               | to a flexbox-like layout algo. React then handles the
               | diffing of this structure, and the terminal UI is updated
               | according to only what has changed by manually
               | overwriting sections of the buffer. The CLI library is
               | called Ink and I forget the name of the flexbox layout
               | algo implementation, but you can read about the internals
               | if you look at the Ink repo.
        
               | HarHarVeryFunny wrote:
               | IMO diffing might have made sense to do here, but that's
               | not what they chose to do.
               | 
               | What's apparently happening is that React tells Ink to
               | update (re-render) the UI "scene graph", and Ink then
               | generates a new full-screen image of how the terminal
               | should look, then passes this screen image to another
               | library, log-update, to draw to the terminal. log-update
               | draws these screen images by a flicker-inducing clear-
               | then-redraw, which it has now fixed by using escape codes
               | to have the terminal buffer and combine these clear-then-
               | redraw commands, thereby hiding the clear.
               | 
               | An alternative solution, rather than using the flicker-
               | inducing clear-then-redraw in the first place, would have
               | been just to do terminal screen image diffs and draw the
               | changes (which is something I did back in the day for
               | fun, sending full-screen ASCII digital clock diffs over a
               | slow 9600baud serial link to a real terminal).
        
               | skydhash wrote:
               | Any diff would require to have a Before and an After.
               | Whatever was done for the After can be done to directly
               | render the changes. No need for the additional compute of
               | a diff.
        
               | HarHarVeryFunny wrote:
               | Sure, you could just draw the full new screen image
               | (albeit a bit inefficient if only one character changed),
               | and no need for the flicker-inducing clear before draw
               | either.
               | 
               | I'm not sure what the history of log-output has been or
               | why it does the clear-before-draw. Another simple
               | alternative to pre-clear would have been just to clear to
               | end of line (ESC[0K) after each partial line drawn.
        
           | krona wrote:
           | Sounds like a web developer defined the solution a year
           | before they knew what the problem was.
        
             | jama211 wrote:
             | Nah. It's just web development languages are a better fit
             | for agentic coding presently. They weighed the pros and
             | cons, they're not stupid.
        
               | shimman wrote:
               | Of course they can be stupid, hubris is a real thing and
               | humans fail all the time.
        
               | jama211 wrote:
               | But not in our criticism of them, no it cannot be us who
               | are the stupid ones
        
               | barnabee wrote:
               | I've had good success with Claude building snappy TUIs in
               | Rust with Ratatui.
               | 
               | It's not obvious to me that there'd be any benefit of
               | using TypeScript and React instead, especially none that
               | makes up for the huge downsides compared to Rust in a
               | terminal environment.
               | 
               | Seems to me the problem is more likely the skills of the
               | engineers, not Claude's capabilities.
        
               | jama211 wrote:
               | I'm sure you know better than them
        
               | int_19h wrote:
               | It's a popular myth, but not really true anymore with the
               | latest and greatest. I'm currently using both Claude and
               | Codex to work on a Haskell codebase, and it works
               | wonderfully. More so than JS actually, since the type
               | system provides extensive guardrails (you can get types
               | with TS, but it's not sound, and it's very easy to write
               | code that violates type constraints at runtime without
               | even deliberately trying to do so).
        
           | CamperBob2 wrote:
           | _Also explains why Claude Code is a React app outputting to a
           | Terminal. (Seriously.)_
           | 
           | Who cares, and why?
           | 
           | All of the major providers' CLI harnesses use Ink:
           | https://github.com/vadimdemedes/ink
        
           | krystofbe wrote:
           | I did some debugging on this today. The results are...
           | sobering.
           | 
           | Memory comparison of AI coding CLIs (single session, idle):
           | | Tool        | Footprint | Peak   | Language      |
           | |-------------|-----------|--------|---------------|       |
           | Codex       | 15 MB     | 15 MB  | Rust          |       |
           | OpenCode    | 130 MB    | 130 MB | Go            |       |
           | Claude Code | 360 MB    | 746 MB | Node.js/React |
           | 
           | That's a 24x to 50x difference for tools that do the same
           | thing: send text to an API.
           | 
           | vmmap shows Claude Code reserves 32.8 GB virtual memory just
           | for the V8 heap, has 45% malloc fragmentation, and a peak
           | footprint of 746 MB that never gets released, classic leak
           | pattern.
           | 
           | On my 16 GB Mac, a "normal" workload (2 Claude sessions +
           | browser + terminal) pushes me into 9.5 GB swap within hours.
           | My laptop genuinely runs slower with Claude Code than when
           | I'm running local LLMs.
           | 
           | I get that shipping fast matters, but building a CLI with
           | React and a full Node.js runtime is an architectural choice
           | with consequences. Codex proves this can be done in 15 MB.
           | Every Claude Code session costs me 360+ MB, and with MCP
           | servers spawning per session, it multiplies fast.
        
             | Weryj wrote:
             | I believe they use https://bun.com/ Not Node.js
        
             | atonse wrote:
             | Jarred Sumner (bun creator, bun was recently acquired by
             | Anthropic) has been working exclusively on bringing down
             | memory leaks and improving performance in CC the last
             | couple weeks. He's been tweeting his progress.
             | 
             | This is just regular tech debt that happens from building
             | something to $1bn in revenue as fast as you possibly can,
             | optimize later.
             | 
             | They're optimizing now. I'm sure they'll have it under
             | control in no time.
             | 
             | CC is an incredible product (so is codex but I use CC
             | more). Yes, lately it's gotten bloated, but the value it
             | provides makes it bearable until they fix it in short time.
        
               | bdangubic wrote:
               | if I had a dollar for each time I heard "until they fix
               | it in short time" I'd have Elon money
        
               | tatjam wrote:
               | Claude, fix the memory leaks, or you'll go to jail!
        
             | slopusila wrote:
             | why do you care about uncommitted virtual memory? that's
             | practically infinite
        
             | badlogic wrote:
             | OpenCode is not written in Go. It's TS on Bun, with OpenTUI
             | underneath which is written in Zig.
        
         | jsheard wrote:
         | CC has >6000 open issues, despite their bot auto-culling them
         | after 60 days of inactivity. It was ~5800 when I looked just a
         | few days ago so they seem to be accelerating towards some kind
         | of bug singularity.
        
           | tgtweak wrote:
           | plot twist, it's all claude code instances submitting bug
           | reports on behalf of end users.
        
             | accrual wrote:
             | It's Claude, all the way down.
        
             | trescenzi wrote:
             | I literally hit a claude code bug today, tried to use
             | claude desktop to debug it which didn't help and it offered
             | to open a bug report for me. So yes 100%. Some of the
             | titles also make it pretty clear they are auto submitted.
             | This is my favorite which was around the top when I was
             | creating my bug report 3 hours ago and is now 3 pages back
             | lol.
             | 
             | > Unable to process - no bug report provided. Please share
             | the issue details you'd like me to convert into a GitHub
             | issue title
             | 
             | https://github.com/anthropics/claude-code/issues/23459
        
           | paxys wrote:
           | Half of them were probably opened yesterday during the Claude
           | outage.
        
             | anematode wrote:
             | Nah, it was at like 5500 before.
        
           | elAhmo wrote:
           | Insane to think that a relatively simple CLI tool has so many
           | open issues...
        
             | emilsedgh wrote:
             | It's not really a simple CLI tool though it's really
             | interactive.
        
             | trymas wrote:
             | What's so simple about it?
        
               | elAhmo wrote:
               | I said relatively simple. It is mostly an API interface
               | with Anthropic models, with tool calling on top of it,
               | very simple input and output.
        
               | 9dev wrote:
               | I'm pretty certain you haven't used it yet(to its fullest
               | extent) then. Claude Code is easily one of the most
               | complex terminal UIs I have seen yet.
        
               | dvfjsdhgfv wrote:
               | Could you explain why? When I think about complex TUIs, I
               | think about things we were building with Turbo Vision in
               | the 90s.
        
               | gorbypark wrote:
               | I'm going to buck the trend and say it's really not that
               | complex. AFAIK they are using Ink, which is React with a
               | TUI renderer.
               | 
               |  _Cue I could build it in a weekend vibes_ , I built my
               | own agent TUI using the OpenAI agent SDK and Ink. Of
               | course it's not as fleshed out as Claude, but it supports
               | git work trees for multi agent, slash commands, human in
               | the loop prompts and etc. If I point it at the Anthropic
               | models it more or less produces results as m good as the
               | real Claude TUI.
               | 
               | I actually "decompiled" the Claude tools and prompts and
               | recreated them. As of 6 months ago Claude was 15 tools,
               | mostly pretty basic (list for, read file, wrote file,
               | bash, etc) with some very clever prompts, especially the
               | task tool it uses to do the quasi planning mode task
               | bullets (even when not in planning mode).
               | 
               | Honestly the idea of bringing this all together with an
               | affordable monthly service and obviously some seriously
               | creative "prompt engineers" is the magic/hard part (and
               | making the model itself, obviously).
        
               | ozozozd wrote:
               | It's extremely simple.
               | 
               | If that's the most complex TUI (yeah, new acronym) you've
               | seen, you have a lot to catch up on!
               | 
               | I am talking rendering image/video in the terminal!
        
               | brookst wrote:
               | With extensibility via plugins, MCP (stdio and http), UI
               | to prompt the user for choices and redirection, tools to
               | manage and view context, and on and on.
               | 
               | It is not at all a small app, at least as far as UX
               | surface area. There are, what, 40ish slash commands? Each
               | one is an opportunity for bugs and feature gaps.
        
               | everforward wrote:
               | I would still call that small, maybe medium. emacs is
               | huge as far as CLI tools go, awk is large because it
               | implements its own language (apparently capable of
               | writing Doom in). `top` probably has a similar number of
               | interaction points, something like `lftp` might have more
               | between local and remote state.
               | 
               | The complex and magic parts are around finding contextual
               | things to include, and I'd be curious how many are that
               | vs "forgot to call clear() in the TUI framework before
               | redirecting to another page".
        
               | koakuma-chan wrote:
               | They wouldn't have 6000 issues if they hired one or two
               | Rust engineers.
        
               | dmazzoni wrote:
               | Also it's highly multithreaded / multiprocess - you can
               | run subagents that can communicate with each other, you
               | can interrupt it while it's in the middle of thinking and
               | it handles it gracefully without forgetting what it was
               | doing
        
               | trymas wrote:
               | If I would get a dollar each time a developer (or CTO!)
               | told me "this is (relatively) simple, it will take 2
               | days/weeks", but then it actually took 2 years+ to fully
               | build and release a product that has more useful features
               | than bugs...
               | 
               | I am not protecting anthropic[0], but how come in this
               | forum every day I still see these "it's simple" takes
               | from experienced people - I have no idea. There are who
               | knows how many terminal emulators out there, with who
               | knows how many different configurations. There are
               | plugins for VSCode and various other editors (so it's not
               | only TUI).
               | 
               | Looking at issue tracker ~1/3 of issues are seemingly
               | feature requests[1].
               | 
               | Do not forget we are dealing with LLMs and it's a tool,
               | which purpose and selling point that it codes on ANY
               | computer in ANY language for ANY system. It's very
               | popular tool run each day by who knows how many people -
               | I could easily see, how such "relatively simple" tool
               | would rack up thousands of issues, because "CC won't do
               | weird thing X, for programming language Y, while I run
               | from my terminal Z". And because it's LLM - theres whole
               | can of non deterministic worms.
               | 
               | Have you created an LLM agent, especially with moderately
               | complex tool usage? If yes and it worked flawlessly -
               | tell your secrets (and get hired by
               | Anthropic/ChatGPT/etc). Probably 80% of my evergrowing
               | code was trying to just deal with unknown unknowns - what
               | if LLM invokes tool wrong? How to guide LLM back on
               | track? How to protect ourselves and keep LLM on track if
               | prompts are getting out of hand or user tries to do
               | something weird? The problems were endless...
               | 
               | Yes the core is "simple", but it's extremely deep can of
               | worms, for such successful tool - I easily could see how
               | there are many issues.
               | 
               | Also super funny, that first issue for me at the moment
               | is how user cannot paste images when it has Korean
               | language input (also issue description is in Korean) and
               | second issue is about input problems in Windows
               | Powershell and CMD, which is obviously total different
               | world compared to POSIX (???) terminal emulators.
               | 
               | [0] I have very adverse feelings for mega ultra wealthy
               | VC moneys...
               | 
               | [1] https://github.com/anthropics/claude-
               | code/issues?q=is%3Aissu...
        
               | vouwfietsman wrote:
               | Although I understand your frustration (and have
               | certainly been at the other side of this as well!), I
               | think its very valuable to always verbalize your
               | intuition of scope of work and be critical if your
               | intuition is in conflict with reality.
               | 
               | Its the best way to find out if there's a mismatch
               | between value and effort, and its the best way to learn
               | and discuss the fundamental nature of complexity.
               | 
               | Similar to your argument, I can name countless of
               | situations where developers _absolutely adamantly
               | insisted_ that something was very hard to do, only for
               | another developer to say  "no you can actually do that
               | like this* and fix it in hours instead of weeks.
               | 
               | Yes, making a _TUI_ from scratch is hard, no that should
               | not affect Claude code because they aren 't actually
               | _making_ the TUI library (I hope). It _should_ be the
               | case that most complexity is in the model, and the client
               | is just using a text-based interface.
               | 
               | There seems to be a mismatch of what you're describing
               | would be issues (for instance about the quality of the
               | agent) and what people are describing as the _actual_
               | issues (terminal commands don 't work, or input is lost
               | arbitrarily).
               | 
               | That's why verbalizing is important, because you are
               | thinking about other complexities than the people you
               | reply to.
        
               | trymas wrote:
               | As another example `opencode`[0] has number issues on the
               | same order of magnitude, with similar problems.
               | 
               | > There seems to be a mismatch of what you're describing
               | would be issues (for instance about the quality of the
               | agent) and what people are describing as the actual
               | issues (terminal commands don't work, or input is lost
               | arbitrarily).
               | 
               | I just named couple examples I've seen in issue tracker
               | and `opencode` on quick skim has many similar issues
               | about inputs and rendering issues in terminals too.
               | 
               | > Similar to your argument, I can name countless of
               | situations where developers absolutely adamantly insisted
               | that something was very hard to do, only for another
               | developer to say "no you can actually do that like this*
               | and fix it in hours instead of weeks.
               | 
               | Good example, as I have seen this too, but for this case,
               | let's first see `opencode`/`claude` equivalent written in
               | "two weeks" and that has no issues (or issues are fixed
               | so fast, they don't accumulate into thousands) and
               | supports any user on any platform. People building stuff
               | for only themselves (N=1) and claiming the problem is
               | simple do not count.
               | 
               | ---------
               | 
               | Like the guy two days ago claiming that "the most basic
               | feature"[1] in an IDE is a _terminal_. But then we see
               | threads in HN popping up about Ghostty or Kitty or
               | whatever and how those terminals are god-send, everything
               | else is crap. They may be right, but that software took
               | years (and probably tens of man-years) to write.
               | 
               | What I am saying is that just throwing out phrases that
               | something is "simple" or "basic" needs proof, but at the
               | time of writing I don't see examples.
               | 
               | [0] https://github.com/anomalyco/opencode/issues
               | 
               | [1] https://news.ycombinator.com/item?id=46877204
        
               | vouwfietsman wrote:
               | > equivalent written in "two weeks"
               | 
               | This is indeed a nonsensical timeframe.
               | 
               | > What I am saying is that just throwing out phrases that
               | something is "simple" or "basic" needs proof, but at the
               | time of writing I don't see examples.
               | 
               | Fair point.
        
               | trymas wrote:
               | > > equivalent written in "two weeks"
               | 
               | > This is indeed a nonsensical timeframe.
               | 
               | Sorry - I should have explained that it's an ironic
               | hyperbole. Was thinking quotes will be enough, but Poe's
               | law strikes again.
        
               | hector_vasquez wrote:
               | I have given the "never trust the judgment of someone who
               | says it should be a one-line fix" so many times I am
               | basically doxxing myself with this comment.
        
             | dwaltrip wrote:
             | _sips coffee..._ ahh yes, let me find that classic Dropbox
             | rsync comment
        
               | elAhmo wrote:
               | Just because Antropic made you think they are doing very
               | complex thing with this tool, doesn't mean it is true.
               | Claude Code is not even comparable to massive software
               | which is probably an order of magnitudes more complex,
               | such as IntelliJ stuff as an example.
               | 
               | Tools like https://github.com/badlogic/pi-mono implement
               | most of the functionality Claude Code has, even adding
               | loads of stuff Claude doesn't have and can actually
               | scroll without flickering inside terminal, all built by a
               | single guy as a side project. I guess we can't ask that
               | much from a 250B USD company.
               | 
               | Be careful with the coffee.
        
             | luckydata wrote:
             | It's far from simple
        
             | bjackman wrote:
             | Well part of the issue is that it isn't actually a CLI
             | tool. It takes control of the whole terminal and then badly
             | reimplements a CLI...
        
           | dkersten wrote:
           | Just anecdotally, each release seems to be buggier than the
           | last.
           | 
           | To me, their claim that they are vibe coding Claude code
           | isn't the flex they think it is.
           | 
           | I find it harder and harder to trust anthropic for business
           | related use and not just hobby tinkering. Between buggy
           | releases, opaque and often seemingly glitches rate limits and
           | usage limits, and the model quality inconsistency, it's just
           | not something I'd want to bet a business on.
        
             | zahlman wrote:
             | I think I would be much _more_ frightened if it were
             | working well.
        
               | ifwinterco wrote:
               | Exactly, thank goodness it's still a bit rubbish in some
               | aspects
        
             | csomar wrote:
             | Since version 2.1.9, performance has degraded significantly
             | after extended use. After 30-40 prompts with substantial
             | responses, memory usage climbs above 25GB, making the tool
             | nearly unusable. I'm updating again to see if it improves.
             | 
             | Unlike what another commenter suggested, this is a complex
             | tool. I'm curious whether the codebase might eventually
             | reach a point where it becomes unfixable; even with human
             | assistance. That would be an interesting development. We'll
             | see.
        
             | marcd35 wrote:
             | Doesn't this just exacerbate the "black box" conundrum if
             | they just keep piling on more and more features without
             | fully comprehending what's being implemented
        
           | ericrallen wrote:
           | The rate of Issues opened on a popular repo is at least one
           | order of magnitude beyond the number of Issues whoever is
           | able to deal with them can handle.
        
         | jama211 wrote:
         | It's extremely successful, not sure what it explains other than
         | your biases
        
           | blibble wrote:
           | Microsoft's products are also extremely successful
           | 
           | they're also total garbage
        
             | simianwords wrote:
             | but they have the advantage of already being a big company.
             | Anthropic is new and there's no reason for people to use it
        
               | Izikiel43 wrote:
               | what about if management gives them a reason? You can
               | think of which those can be.
        
               | kuboble wrote:
               | The tool is absolutely fantastic coding assistant. That's
               | why I use it.
               | 
               | The amount of non-critical bugs all over the place is at
               | least a magnitude larger than of any software I was using
               | daily ever.
               | 
               | Plenty of built in /commands don't work. Sometimes it
               | accepts keystrokes with 1 second delays. It often scrolls
               | hundreds of lines in console after each key stroke Every
               | now and then it crashes completely and is unrecoverable
               | (I once have up and installed a fresh wls) When you ask
               | it question in plan mode it is somewhat of an art to find
               | the answer because after answering the question it will
               | dump the whole current plan (free screens of text)
               | 
               | And just in general the technical feeling of the TUI is
               | that of a vibe coded project that got too big to control.
        
               | derwiki wrote:
               | I think this might be a harbinger of what we should
               | expect for software quality in the next decade
        
               | jama211 wrote:
               | Orrrrr it's not
        
             | holoduke wrote:
             | Claude is by far the most popular and best assistant
             | currently available for a developer.
        
               | wavemode wrote:
               | Okay, and Windows is by far the most popular desktop
               | operating system.
               | 
               | Discussions are pointless when the parties are talking
               | past each other.
        
               | pluralmonad wrote:
               | Popular meaning lots of people like it or that it is
               | relatively widespread? Polio used to be popular in the
               | latter way.
        
               | quietsegfault wrote:
               | I like windows, it's fine. I like MacOS better. I like
               | Linux. None of them are garbage or unusable.
        
               | blibble wrote:
               | have you used Windows 11?
               | 
               | file explorer takes 5 seconds to open
        
               | jama211 wrote:
               | No it doesn't, don't be hyperbolic.
        
               | jama211 wrote:
               | Yes, and windows is pretty good for most people. Don't be
               | ridiculous.
        
               | dmazzoni wrote:
               | Yeah, but there are dozens of AI coding assistants to
               | choose from, and the cost to switch is very low, unlike
               | switching operating systems.
               | 
               | I've tried them all and I keep coming back to Claude Code
               | because it's just so much more capable and useful than
               | the others.
        
               | elvin_d wrote:
               | might be only among most popular. https://skills.sh/ is
               | some data point.
        
               | oblio wrote:
               | Is it better than OpenCode?
        
             | jama211 wrote:
             | Well there you have it, proof you're not being reasonable.
             | Microsoft's products annoy HN users but they are absolutely
             | not total garbage. They're highly functional and valuable
             | and if they weren't they truely wouldn't be used, they're
             | just flawed.
        
               | ed_mercer wrote:
               | You should look at some Copilot reviews.
        
               | jama211 wrote:
               | Different goalposts mate.
        
           | mvdtnz wrote:
           | Anthropic has perhaps the most embarrassing status page
           | history I have ever seen. They are famous for downtime.
           | 
           | https://status.claude.com/
        
             | dimgl wrote:
             | And yet people still use them.
        
             | ronsor wrote:
             | As opposed to other companies which are smart enough not to
             | report outages.
        
               | tavavex wrote:
               | So, there are only two types of companies: ones that have
               | constant downtime, and ones that have constant downtime
               | but hide it, right?
        
               | Sebguer wrote:
               | Basically, yes.
        
             | Computer0 wrote:
             | The competition doesn't currently have all 99's -
             | https://status.openai.com/
        
             | djeastm wrote:
             | The best way to use Claude's models seems to be some other
             | inference provider (either OpenRouter or directly)
        
             | derwiki wrote:
             | Shades of Fail Whale
        
           | acedTrex wrote:
           | Something being successful and something being a high quality
           | product with good engineering are two completely different
           | questions.
        
         | raincole wrote:
         | It explains how important dogfooding is if you want to make an
         | extremely successful product.
        
         | spruce_tips wrote:
         | Ah yes, explains why it takes 3 seconds for a new chat to load
         | after I click new chat in the macOS app.
        
         | exe34 wrote:
         | Can Claude fix the flicker in Claude yet?
        
         | cedws wrote:
         | The sandboxing in CC is an absolute joke, it's no wonder
         | there's an explosion of sandbox wrappers at the moment. There's
         | going to be a security catastrophe at some point, no doubt
         | about it.
        
         | quietsegfault wrote:
         | What does it explain, oh snark master supreme?
        
       | rob wrote:
       | System Card: https://www-
       | cdn.anthropic.com/0dd865075ad3132672ee0ab40b05a5...
        
       | Someone1234 wrote:
       | Does anyone with more insight into the AI/LLM industry happen to
       | know if the cost to run them in normal user-workflows is falling?
       | The reason I'm asking is because "agent teams" while a cool
       | concept, it largely constrained by the economics of running
       | multiple LLM agents (i.e. plans/API calls that make this
       | practical at scale are expensive).
       | 
       | A year or more ago, I read that both Anthropic and OpenAI were
       | losing money on every single request even for their paid
       | subscribers, and I don't know if that has changed with more
       | efficient hardware/software improvements/caching.
        
         | simonw wrote:
         | The cost per token served has been falling steadily over the
         | past few years across basically all of the providers. OpenAI
         | dropped the price they charged for o3 to 1/5th of what it was
         | in June last year thanks to "engineers optimizing inferencing",
         | and plenty of other providers have found cost savings too.
         | 
         | Turns out there was a lot of low-hanging fruit in terms of
         | inference optimization that hadn't been plucked yet.
         | 
         | > A year or more ago, I read that both Anthropic and OpenAI
         | were losing money on every single request even for their paid
         | subscribers
         | 
         | Where did you hear that? It doesn't match my mental model of
         | how this has played out.
        
           | nubg wrote:
           | > "engineers optimizing inferencing"
           | 
           | are we sure this is not a fancy way of saying quantization?
        
             | embedding-shape wrote:
             | Or distilled models, or just slightly smaller models but
             | same architecture. Lots of options, all of them
             | conveniently fitting inside "optimizing inferencing".
        
             | jmalicki wrote:
             | A ton of GPU kernels are hugely inefficient. Not saying the
             | numbers are realistic, but look at the 100s of times of
             | gain in the Anthropic performance takehome exam that
             | floated around on here.
             | 
             | And if you've worked with pytorch models a lot, having
             | custom fused kernels can be huge. For instance, look at the
             | kind of gains to be had when FlashAttention came out.
             | 
             | This isn't just quantization, it's actually just better
             | optimization.
             | 
             | Even when it comes to quantization, Blackwell has far
             | better quantization primitives and new floating point types
             | that support row or layer-wise scaling that can quantize
             | with far less quality reduction.
             | 
             | There is also a ton of work in the past year on sub-
             | quadratic attention for new models that gets rid of a huge
             | bottleneck, but like quantization can be a tradeoff, and a
             | lot of progress has been made there on moving the Pareto
             | frontier as well.
             | 
             | It's almost like when you're spending hundreds of billions
             | on capex for GPUs, you can afford to hire engineers to make
             | them perform better without just nerfing the models with
             | more quantization.
        
               | Der_Einzige wrote:
               | "This isn't X, it's Y" with extra steps.
        
               | jmalicki wrote:
               | I'm flattered you think I wrote as well as an AI.
        
               | nubg wrote:
               | lmao
        
             | esafak wrote:
             | Someone made a quality tracker:
             | https://marginlab.ai/trackers/claude-code/
        
             | bityard wrote:
             | When MP3 became popular, people were amazed that you could
             | compress audio to 1/10th its size with minor quality loss.
             | A few decades later, we have audio compression that is much
             | better and higher-quality than MP3, and they took a lot
             | more effort than "MP3 but at a lower bitrate."
             | 
             | The same is happening in AI research now.
        
               | oblio wrote:
               | > A few decades later, we have audio compression that is
               | much better and higher-quality than MP3
               | 
               | Just curious, which formats and how they compare, storage
               | wise?
               | 
               | Also, are you sure it's not just moving the goalposts to
               | CPU usage? Frequently more powerful compression
               | algorithms can't be used because they use lots of
               | processing power, so frequently the biggest gains over 20
               | years are just... hardware advancements.
        
             | simonw wrote:
             | The o3 optimizations were not quantization, they confirmed
             | this at the time.
        
           | cootsnuck wrote:
           | I have not see any reporting or evidence at all that
           | Anthropic or OpenAI is able to make money on inference yet.
           | 
           | > Turns out there was a lot of low-hanging fruit in terms of
           | inference optimization that hadn't been plucked yet.
           | 
           | That does not mean the frontier labs are pricing their APIs
           | to cover their costs yet.
           | 
           | It can both be true that it has gotten cheaper for them to
           | provide inference and that they still are subsidizing
           | inference costs.
           | 
           | In fact, I'd argue that's way more likely given that has been
           | precisely the goto strategy for highly-competitive startups
           | for awhile now. Price low to pump adoption and dominate the
           | market, worry about raising prices for financial
           | sustainability later, burn through investor money until then.
           | 
           | What no one outside of these frontier labs knows right now is
           | how big the gap is between current pricing and eventual
           | pricing.
        
             | NitpickLawyer wrote:
             | > they still are subsidizing inference costs.
             | 
             | They are for sure subsidising costs on all you can prompt
             | packages (20-100-200$ /mo). They do that for data gathering
             | mostly, and at a smaller degree for user retention.
             | 
             | > evidence at all that Anthropic or OpenAI is able to make
             | money on inference yet.
             | 
             | You can infer that from what 3rd party inference providers
             | are charging. The largest open models atm are dsv3 (~650B
             | params) and kimi2.5 (1.2T params). They are being served at
             | 2-2.5-3$ /Mtok. That's sonnet / gpt-mini / gemini3-flash
             | price range. You can make some educates guesses that they
             | get some leeway for model size at the 10-15$/ Mtok prices
             | for their top tier models. So if they are inside some sane
             | model sizes, they are likely making money off of token
             | based APIs.
        
               | slopusila wrote:
               | most of those subscriptions go unused. I barely use 10%
               | of mine
               | 
               | so my unused tokens compensate for the few heavy users
        
               | aenis wrote:
               | Thanks!
               | 
               | I hope my unused gym subscription pays back the good
               | karma :-)
        
               | sandos wrote:
               | Ive been thinking about our company, one of big global
               | conglomerates that went for copilot. Suddenly I was just
               | enrolled.. together with at least 1500 others. I guess
               | the amount of money for our business copilot plans x 1500
               | is not a huge amount of money, but I am at least pretty
               | convinced that only a small part of users use even 10% of
               | their quota. Even teams located around me, I only know of
               | 1 person that seems to use it actively.
        
               | int_19h wrote:
               | > They are being served at 2-2.5-3$ /Mtok. That's sonnet
               | / gpt-mini / gemini3-flash price range.
               | 
               | The interesting number is usually input tokens, not
               | output, because there's much more of the former in any
               | long-running session (like say coding agents) since all
               | outputs become inputs for the next iteration, and you
               | also have tool calls adding a lot of additional input
               | tokens etc.
               | 
               | It doesn't change your conclusion much though. Kimi K2.5
               | has almost the same input token pricing as Gemini 3
               | Flash.
        
             | barrkel wrote:
             | > evidence at all that Anthropic or OpenAI is able to make
             | money on inference yet.
             | 
             | The evidence is in third party inference costs for open
             | source models.
        
             | chis wrote:
             | It's quite clear that these companies do make money on each
             | marginal token. They've said this directly and analysts
             | agree [1]. It's less clear that the margins are high enough
             | to pay off the up-front cost of training each model.
             | 
             | [1] https://epochai.substack.com/p/can-ai-companies-become-
             | profi...
        
               | 9cb14c1ec0 wrote:
               | It's also true that their inference costs are being
               | heavily subsidized. For example, if you calculate Oracles
               | debt into OpenAIs revenue, they would be incredibly far
               | underwater on inference.
        
               | magicalist wrote:
               | > _They 've said this directly and analysts agree [1]_
               | 
               | chasing down a few sources in that article leads to
               | articles like this at the root of claims[1], which is
               | entirely based on information "according to a person with
               | knowledge of the company's financials", which doesn't
               | exactly fill me with confidence.
               | 
               | [1] https://www.theinformation.com/articles/openai-
               | getting-effic...
        
               | simonw wrote:
               | "according to a person with knowledge of the company's
               | financials" is how professional journalists tell you that
               | someone who they judge to be credible has leaked
               | information to them.
               | 
               | I wrote a guide to deciphering that kind of language a
               | couple of years ago:
               | https://simonwillison.net/2023/Nov/22/deciphering-clues/
        
               | topaz0 wrote:
               | Unfortunately tech journalists' judgement of source
               | credibility don't have a very good track record
        
               | mrgaro wrote:
               | But there are companies which are only serving open
               | weight models via APIs (ie. they are not doing any
               | training), so they must be profitable? here's one list of
               | providers from OpenRouter serving LLama 3.3 70B:
               | https://openrouter.ai/meta-
               | llama/llama-3.3-70b-instruct/prov...
        
               | m101 wrote:
               | It's not clear at all because model training upfront
               | costs and how you depreciate them are big unknowns, even
               | for deprecated models. See my last comment for a bit more
               | detail.
        
               | ACCount37 wrote:
               | By now, model lifetime inference compute is >10x model
               | training compute, for mainstream models. Further
               | amortized by things like base model reuse.
        
               | simonw wrote:
               | They are obviously losing money on training. I think they
               | are selling inference for less than what it costs to
               | serve these tokens.
               | 
               | That really matters. If they are making a margin on
               | inference they could conceivably break even no matter how
               | expensive training is, provided they sign up enough
               | paying customers.
               | 
               | If they lose money on every paying customer then building
               | great products that customers want to pay for them will
               | just make their financial situation worse.
        
               | Schlagbohrer wrote:
               | "We lose money on each unit sold, but we make it up in
               | volume"
        
               | emp17344 wrote:
               | Sue, but if they stop training new models, the current
               | models will be useless in a few years as our knowledge
               | base evolves. They need to continually train new models
               | to have a useful product.
        
             | mrandish wrote:
             | > I have not see any reporting or evidence at all that
             | Anthropic or OpenAI is able to make money on inference yet.
             | 
             | Anthropic planning an IPO this year is a broad meta-
             | indicator that internally they believe they'll be able to
             | reach break-even sometime _next_ year on delivering a
             | competitive model. Of course, their belief could turn out
             | to be wrong but it doesn 't make much sense to do an IPO if
             | you don't think you're close. Assuming you have a choice
             | with other options to raise private capital (which still
             | seems true), it would be better to defer an IPO until you
             | expect quarterly numbers to reach break-even or at least
             | close to it.
             | 
             | Despite the willingness of private investment to fund
             | hugely negative AI spend, the recently growing twitchiness
             | of public markets around AI ecosystem stocks indicates
             | they're already worried prices have exceeded near-term
             | value. It doesn't seem like they're in a mood to fund
             | oceans of dotcom-like red ink for long.
        
               | WarmWash wrote:
               | IPO'ing is often what you do to give your golden
               | investors an exit hatch to dump their shares on the
               | notoriously idiotic and hype driven public.
        
               | defmacr0 wrote:
               | >Despite the willingness of private investment to fund
               | hugely negative AI spend
               | 
               | VC firms, even ones the size of Softbank, also literally
               | just don't have enough capital to fund the planned next-
               | generation gigawatt-scale data centers.
        
           | sumitkumar wrote:
           | It seems it is true for gemini because they have a humongous
           | sparse model but it isn't so true for the max performance
           | opus-4.5/6 and gpt-5.2/3.
        
           | replwoacause wrote:
           | My experience trying to use Opus 4.5 on the Pro plan has been
           | terrible. It blows up my usage very very fast. I avoid it
           | altogether now. Yes, I know they warn about this, but it's
           | comically fast how quickly it happens.
        
           | topaz0 wrote:
           | But a) that's the cost to the user -- we don't know how much
           | loss they're taking on those and b) the number of tokens to
           | serve a similar prompt has been going up, so that the total
           | cost to serve a prompt has been going up in general. Any cost
           | analysis that doesn't mention these is hugely misleading.
        
         | Havoc wrote:
         | Saw a comment earlier today about google seeing a big (50%+)
         | fall in Gemini serving cost per unit across 2025 but can't find
         | it now. Was either here or on Reddit
        
           | mattddowney wrote:
           | From Alphabet 2025 Q4 Earnings call: "As we scale, we're
           | getting dramatically more efficient. We were able to lower
           | Gemini serving unit costs by 78% over 2025 through model
           | optimizations, efficiency and utilization improvements."
           | https://abc.xyz/investor/events/event-
           | details/2026/2025-Q4-E...
        
             | Havoc wrote:
             | Thanks! That's the one
        
         | 3abiton wrote:
         | It's not just that. Everyone is complacent with the utilization
         | of AI agents. I have been using AI for coding for quite a
         | while, and most of my "wasted" time is correcting its
         | trajectory and guiding it through the thinking process. It's
         | very fast iterations but it can easily go off track. Claude's
         | family are pretty good at doing chained task, but still once
         | the task becomes too big context wise, it's impossible to get
         | back on track. Cost wise, it's cheaper than hiring skilled
         | people, that's for sure.
        
           | lufenialif2 wrote:
           | Cost wise, doesn't that depend on what you could be doing
           | besides steering agents?
        
             | cyanydeez wrote:
             | Isn't the quote something like: "If these LLMs are so good
             | at producing products, where are all those products?"
        
               | lufenialif2 wrote:
               | Waiting for godot...
        
         | zozbot234 wrote:
         | > i.e. plans/API calls that make this practical at scale are
         | expensive
         | 
         | Local AI's make agent workflows a whole lot more practical.
         | Making the initial investment for a good homelab/on-prem
         | facility will effectively become a no-brainer given the
         | advantages on privacy and reliability, and you don't have to
         | fear rugpulls or VC's playing the "lose money on every request"
         | game since you know exactly how much you're paying in power
         | costs for your overall load.
        
           | vbezhenar wrote:
           | I don't care about privacy and I didn't have much problems
           | with reliability of AI companies. Spending ridiculous amount
           | of money on hardware that's going to be obsolete in a few
           | years and won't be utilized at 100% during that time is not
           | something that many people would do, IMO. Privacy is good
           | when it's given for free.
           | 
           | I would rather spend money on some pseudo-local inference
           | (when cloud company manages everything for me and I just can
           | specify some open source model and pay for GPU usage).
        
           | slopusila wrote:
           | on prem economics dont work because you can't batch requests.
           | unless you are able to run 100 agents at the same time all
           | the time
        
             | zozbot234 wrote:
             | > unless you are able to run 100 agents at the same time
             | all the time
             | 
             | Except that newer "agent swarm" workflows do exactly that.
             | Besides, batching requests generally comes with a sizeable
             | increase in memory footprint, and memory is often the main
             | bottleneck especially with the larger contexts that are
             | typical of agent workflows. If you have plenty of agentic
             | tasks that are not especially latency-critical and don't
             | need the absolutely best model, it makes plenty of sense to
             | schedule these for running locally.
        
         | Aurornis wrote:
         | > A year or more ago, I read that both Anthropic and OpenAI
         | were losing money on every single request even for their paid
         | subscribers
         | 
         | This gets repeated everywhere but I don't think it's true.
         | 
         | The company is unprofitable overall, but I don't see any reason
         | to believe that their per-token inference costs are below the
         | marginal cost of computing those tokens.
         | 
         | It is true that the company is unprofitable overall when you
         | account for R&D spend, compensation, training, and everything
         | else. This is a deliberate choice that every heavily funded
         | startup should be making, otherwise you're wasting the
         | investment money. That's precisely what the investment money is
         | for.
         | 
         | However I don't think using their API and paying for tokens has
         | negative value for the company. We can compare to models like
         | DeepSeek where providers can charge a fraction of the price of
         | OpenAI tokens and still be profitable. OpenAI's inference costs
         | are going to be higher, but they're charging such a high
         | premium that it's hard to believe they're losing money on each
         | token sold. I think every token paid for moves them
         | incrementally closer to profitability, not away from it.
        
           | runarberg wrote:
           | I can see a case for omitting R&D when talking about
           | profitability, but training makes no sense. Training is what
           | makes the model, omitting it is like omitting the cost of
           | running the production facility of a car manufacturer. If AI
           | companies stop training they will stop producing models, and
           | they will run out of a products to sell.
        
             | Aurornis wrote:
             | It depends on what you're talking about
             | 
             | If you're looking at overall profitability, you include
             | everything
             | 
             | If you're talking about unit economics of producing tokens,
             | you only include the marginal cost of each token against
             | the marginal revenue of selling that token
        
               | runarberg wrote:
               | I don't understand the logic. Without training the
               | marginal cost of each token goes into nothing. The more
               | you train, the better the model, and (presumably) you
               | will gain more costumer interest. Unlike R&D you will
               | always have to train new models if you want to keep your
               | customers.
               | 
               | To me this looks likes some creative bookkeeping, or even
               | wishful thinking. It is like if SpaceX omits the price of
               | the satellites when calculating their profits.
        
             | vidarh wrote:
             | The reason for this is that the cost scales with the model
             | and training cadence, not usage and so they will hope that
             | they will be able to scale number of inference tokens sold
             | both by increasing use and/or slowing the training cadence
             | as competitors are also forced to aim for overall
             | profitability.
             | 
             | It is essentially a big game of venture capital chicken at
             | present.
        
           | 3836293648 wrote:
           | The reports I remember show that they're profitable per-
           | model, but overlap R&D so that the company is negative
           | overall. And therefore will turn a massive profit if they
           | stop making new models.
        
             | trcf23 wrote:
             | Doesn't it also depend on averaging with free users?
        
             | schnable wrote:
             | * stop making new models and people keep using the existing
             | models, not switch to a competitor still investing in new
             | models.
        
         | Bombthecat wrote:
         | That's why anthropic switched to tpu, you can sell at cost.
        
         | KaiserPro wrote:
         | Gemini-pro-preview is on ollama and requires h100 which is
         | ~$15-30k. Google are charging $3 a million tokens. Supposedly
         | its capable of generating between 1 and 12 million tokens an
         | hour.
         | 
         | Which is profitable. but not by much.
        
           | grim_io wrote:
           | What do you mean it's on ollama and requires h100? As a
           | proprietary google model, it runs on their own hardware, not
           | nvidia.
        
             | KaiserPro wrote:
             | sorry A lack of context:
             | 
             | https://ollama.com/library/gemini-3-pro-preview
             | 
             | You _can_ run it on your own infra. Anthropic and openAI
             | are running off nvidia, so are meta(well supposedly they
             | had custom silicon, I 'm not sure if its capable of running
             | big models) and mistral.
             | 
             | however if google really are running their own inference
             | hardware, then that means the cost is different (developing
             | silicon is not cheap...) as you say.
        
               | zozbot234 wrote:
               | That's a cloud-linked model. It's about using ollama as
               | an API client (for ease of compatibility with other uses,
               | including local), not running that model on local infra.
               | Google does release open models (called Gemma) but
               | they're not nearly as capable.
        
               | simonw wrote:
               | You can't run Gemini 3 Pro Preview on your own
               | infrastructure. Ollama sell access to cloud models these
               | days. It's a little weird and confusing.
        
               | KaiserPro wrote:
               | Ahh fuck, thanks for pointing that out.
               | 
               | I did think its a bit weird that they had open-weighted
               | it
        
         | WarmWash wrote:
         | These are intro prices.
         | 
         | This is all straight out of the playbook. Get everyone hooked
         | on your product by being cheap and generous.
         | 
         | Raise the price to backpay what you gave away plus cover
         | current expenses and profits.
         | 
         | In no way shape or form should people think these $20/mo plans
         | are going to be the norm. From OpenAI's marketing plan, and a
         | general 5-10 year ROI horizon for AI investment, we should
         | expect AI use to cost $60-80/mo per user.
        
           | esafak wrote:
           | The models in 5-10 years are going to be unimaginably good.
           | $100/month will be a bargain for knowledge workers, if they
           | survive.
        
         | m101 wrote:
         | I think actually working out whether they are losing money is
         | extremely difficult for current models but you can look
         | backwards. The big uncertainties are:
         | 
         | 1) how do you depreciate a new model? What is its useful life?
         | (Only know this once you deprecate it)
         | 
         | 2) how do you depreciate your hardware over the period you
         | trained this model? Another big unknown and not known until you
         | finally write the hardware off.
         | 
         | The easy thing to calculate is whether you are making money
         | actually serving the model. And the answer is almost certainly
         | yes they are making money from this perspective, but that's
         | missing a large part of the cost and is therefore wrong.
        
         | nodja wrote:
         | > A year or more ago, I read that both Anthropic and OpenAI
         | were losing money on every single request even for their paid
         | subscribers, and I don't know if that has changed with more
         | efficient hardware/software improvements/caching.
         | 
         | This is obviously not true, you can use real data and common
         | sense.
         | 
         | Just look up a similar sized open weights model on openrouter
         | and compare the prices. You'll note the similar sized model is
         | often much cheaper than what anthropic/openai provide.
         | 
         | Example: Let's compare claude 4 models with deepseek. Claude 4
         | is ~400B params so it's best to compare with something like
         | deepseek V3 which is 680B params.
         | 
         | Even if we compare the cheapest claude model to the most
         | expensive deepseek provider we have claude charging $1/M for
         | input and $5/M for output, while deepseek providers charge
         | $0.4/M and $1.2/M, a fifth of the price, you can get it as
         | cheap as $.27 input $0.4 output.
         | 
         | As you can see, even if we skew things overly in favor of
         | claude, the story is clear, claude token prices are much higher
         | than they could've been. The difference in prices is because
         | anthropic also needs to pay for training costs, while
         | openrouter providers just need to worry on making serving
         | models profitable. Deepseek is also not as capable as claude
         | which also puts down pressure on the prices.
         | 
         | There's still a chance that anthropic/openai models are losing
         | money on inference, if for example they're somehow much larger
         | than expected, the 400B param number is not official, just
         | speculative from how it performs, this is only taking into
         | account API prices, subscriptions and free user will of course
         | skew the real profitability numbers, etc.
         | 
         | Price sources:
         | 
         | https://openrouter.ai/deepseek/deepseek-v3.2-speciale
         | 
         | https://claude.com/pricing#api
        
           | Someone1234 wrote:
           | > This is obviously not true, you can use real data and
           | common sense.
           | 
           | It isn't "common sense" at all. You're comparing several
           | companies losing money, to one another, and suggesting that
           | they're obviously making money because one is under-cutting
           | another more aggressively.
           | 
           | LLM/AI ventures are all currently under-water with massive VC
           | or similar money flowing in, they also all need training data
           | from users, so it is very reasonable to speculate that
           | they're in loss-leader mode.
        
             | nodja wrote:
             | Doing some math in my head, buying the GPUs at retail
             | price, it would take probably around half a year to make
             | the money back, probably more depending how expensive
             | electricity is in the area you're serving from. So I don't
             | know where this "losing money" rhetoric is coming from.
             | It's probably harder to source the actual GPUs than making
             | money off them.
        
               | suddenlybananas wrote:
               | electricity
        
               | defmacr0 wrote:
               | > So I don't know where this "losing money" rhetoric is
               | coming from.
               | 
               | https://www.dbresearch.com/PROD/RI-
               | PROD/PROD0000000000611818...
        
             | mrgaro wrote:
             | There are companies which are only serving open weight
             | models and not doing any training, so they must be
             | profitable? Check for example this list
             | https://openrouter.ai/meta-
             | llama/llama-3.3-70b-instruct/prov...
        
           | tqian wrote:
           | To borrow a concept of cloud server renting, there's also the
           | factor of overselling. Most open source LLM operators
           | probably oversell quite a bit - they don't scale up resources
           | as fast as OpenAI/Anthropic when requests increase. I notice
           | many openrouter providers are noticeably faster during off
           | hours.
           | 
           | In other words, it's not just the model size, but also
           | concurrent load and how many gpus do you turn on at any time.
           | I bet the big players' cost is quite a bit higher than the
           | numbers on openrouter, even for comparable model parameters.
        
       | minimaxir wrote:
       | Will Opus 4.6 via Claude Code be able to access the 1M context
       | limit? The cost increase by going above 200k tokens is 2x input,
       | 1.5x output, which is likely worth it especially for people with
       | the $100/$200 plans.
        
         | CryptoBanker wrote:
         | The 1M context is not available via subscription - only via API
         | usage
        
           | romanovcode wrote:
           | Well this is extremely disappointing to say the least.
        
             | ayhanfuat wrote:
             | It says "subscription users do not have access to Opus 4.6
             | 1M context at launch" so they are probably planning to roll
             | it out to subscription users too.
        
               | kimixa wrote:
               | Man I hope so - the context limit is hit really quickly
               | in many of my use cases - and a compaction event
               | inevitably means another round of corrections and fixes
               | to the current task.
               | 
               | Though I'm wary about that being a magic bullet fix -
               | already it can be pretty "selective" in what it actually
               | seems to take into account documentation wise as the
               | existing 200k context fills.
        
               | nickstinemates wrote:
               | Is this a case of doing it wrong, or you think accuracy
               | is good enough with the amount of context you need to
               | stuff it with often?
        
               | kimixa wrote:
               | I mean the systems I work on have enough weird custom
               | APIs and internal interfaces just getting them working
               | seems to take a good chunk of the context. I've spent a
               | long time trying to minimize every input document where I
               | can, compact and terse references, and still keep hitting
               | similar issues.
               | 
               | At this point I just think the "success" of many AI
               | coding agents is _extremely_ sector dependent.
               | 
               | Going forward I'd love to experiment with seeing if
               | that's _actually_ the problem, or just an easy
               | explanation of failure. I 'd like to play with more
               | controls on context management than "slightly better
               | models" - like being able to select/minimize/compact
               | sections of context I feel would be relevant for the
               | immediate task, to what "depth" of needed details, and
               | those that _aren 't_ likely to be relevant so can be
               | removed from consideration. Perhaps each chunk can be
               | cached to save processing power. Who knows.
        
               | romanovcode wrote:
               | In my example the Figma MCP takes ~300k per medium sized
               | section of the page and it would be cool to enable it
               | reading it and implementing Figma designs straight.
               | Currently I have to split it which makes it annoying.
        
               | humanfromearth9 wrote:
               | Hello,
               | 
               | I check context use percentage, and above ~70% I ask it
               | to generate a prompt for continuation in a new chat
               | session to avoid compaction.
               | 
               | It works fine, and saves me from using precious tokens
               | for context compaction.
               | 
               | Maybe you should try it.
        
               | pluralmonad wrote:
               | How is generating a continuation prompt materially
               | different from compaction? Do you manually scrutinize the
               | context handoff prompt? I've done that before but if not
               | I do not see how it is very different from compaction.
        
               | robertfw wrote:
               | I wonder if it's just: compact earlier, so there's less
               | to compact, and more remaining context that can be used
               | to create a more effective continuation
        
               | IhateAI_2 wrote:
               | lmao what are you building that actually justify needing
               | 1mm tokens on a task? People are spending all this money
               | to do magic tricks on themselves.
        
               | kimixa wrote:
               | The opus context window is 200k tokens not 1mm.
               | 
               | But I kinda see your point - assuming from you're name
               | you're not just a single purpose troll - I'm still not
               | sold on the cost effectiveness of the current generation,
               | and can't see a clear and obvious change to that for the
               | _next_ generation - especially as they 're _still_ loss
               | leaders. Only if you play silly games like  "ignoring the
               | training costs" - IE the _majority of the costs_ - do you
               | get even close to the current subscription costs being
               | sufficient.
               | 
               | My personal experience is that AI generally doesn't
               | actually do what it is being sold for right now, at least
               | in the contexts I'm involved with. Especially by somewhat
               | breathless comments on the internet - like why are they
               | even trying to persuade me in the first place? If they
               | don't want to sell me anything, just shut up and keep the
               | advantage for yourselves rather than replying with the
               | 500th "You're Holding It Wrong" comment with no
               | actionable suggestions. But I still want to know, and am
               | willing to put the time, effort and $$$ in to ensure I'm
               | not deluding myself in ignoring real benefits.
        
               | FrostKiwi wrote:
               | I do not trust that, similar working was used when Sonnet
               | 1M launched. Still not the case today.
        
             | IhateAI_2 wrote:
             | They want the value of your labor and competency to be 1:1
             | correlated to the quality and quantity of tokens you can
             | afford (or be loaned)??
             | 
             | Its a weapon who's target is the working class. How does no
             | one realize this yet?
             | 
             | Don't give them money, code it yourself, you might be
             | surprised how much quality work you can get done!
        
       | mFixman wrote:
       | I found that "Agentic Search" is generally useless in most LLMs
       | since sites with useful data tend to block AI models.
       | 
       | The answer to "when is it cheaper to buy two singles rather than
       | one return between Cambridge to London?" is available in sites
       | such as BRFares, but no LLM can scrape it so it just makes up a
       | generic useless answer.
        
         | causalmodels wrote:
         | Is it still getting blocked when you give it a browser?
        
         | bazmattaz wrote:
         | My guess is that this is going to be the future for LLMs too.
         | It will get harder or more expensive for AI companies to train
         | their models on the latest information as most sites will block
         | the scrapers or ask for a fee.
         | 
         | There might be a future where you'll have to pay more for an up
         | to date model vs a legacy (out of date) model
        
       | heraldgeezer wrote:
       | I love Claude but use the free version so would love a Sonnet &
       | Haiku update :)
       | 
       | I mainly use Haiku to save on tokens...
       | 
       | Also dont use CC but I use the chatbot site or app... Claude is
       | just much better than GPT even in conversations. Straight to the
       | point. No cringe emoji lists.
       | 
       | When Claude runs out I switch to Mistral Le Chat, also just the
       | site or app. Or duck.ai has Haiku 3.5 in Free version.
        
         | eth0up wrote:
         | >I love Claude
         | 
         | I cringe when I think it, but I've actually come to damn near
         | love it too. I am frequently exceedingly grateful for the
         | output I receive.
         | 
         | I've had excellent and awful results with all models, but
         | there's something special in Claude that I find nowhere else. I
         | hope Anthropic makes it more obtainable someday.
        
       | lukebechtel wrote:
       | > Context compaction (beta).
       | 
       | > Long-running conversations and agentic tasks often hit the
       | context window. Context compaction automatically summarizes and
       | replaces older context when the conversation approaches a
       | configurable threshold, letting Claude perform longer tasks
       | without hitting limits.
       | 
       | Not having to hand roll this would be incredible. One of the best
       | Claude code features tbh.
        
       | simonw wrote:
       | The bicycle frame is a bit wonky but the pelican itself is great:
       | https://gist.github.com/simonw/a6806ce41b4c721e240a4548ecdbe...
        
         | DetroitThrow wrote:
         | The ears on top are a cute touch
        
         | ares623 wrote:
         | Can it draw a different bird on a bike?
        
           | simonw wrote:
           | Here's a kakapo riding a bicycle instead: https://gist.github
           | .com/simonw/19574e1c6c61fc2456ee413a24528...
           | 
           | I don't think it quite captures their majesty:
           | https://en.wikipedia.org/wiki/K%C4%81k%C4%81p%C5%8D
        
             | zahlman wrote:
             | Now that I've looked it all up, I feel like that's much
             | more accurate to a real kakapo than the pelican is to a
             | real pelican. It's almost as if it thinks a pelican is just
             | a white flamingo with a different beak.
        
         | nubg wrote:
         | What about the Pelo2 benchmark? (the gray bird that is not
         | gray)
        
         | hoeoek wrote:
         | This really is my favorite benchmark
        
         | einrealist wrote:
         | They trained for it. That's the +0.1!
        
         | eaf7e281 wrote:
         | There's no way they actually work on training this.
        
           | KeplerBoy wrote:
           | There is no way they are not training on this.
        
             | collinmanderson wrote:
             | I suspect they have generic SVG drawing that they focus on.
        
           | margalabargala wrote:
           | I suspect they're training on this.
           | 
           | I asked Opus 4.6 for a pelican riding a _recumbent_ bicycle
           | and got this.
           | 
           | https://i.imgur.com/UvlEBs8.png
        
             | mrandish wrote:
             | Interesting that it seems better. Maybe something about
             | adding a highly specific yet unusual qualifier focusing
             | attention?
        
             | WarmWash wrote:
             | It would be way way better if they were benchmaxxing this.
             | The pelican in the image (both images) has arms. Pelicans
             | don't have arms, and a pelican riding a bike would use it's
             | wings.
        
               | seanhunter wrote:
               | Pelicans don't ride bikes. You can't have scruples about
               | whether or not the image of a pelican riding a bike has
               | arms.
        
               | jevinskie wrote:
               | Wouldn't any decent bike-riding pelican have a bike
               | tailored to pelicans and their wings?
        
               | cinntaile wrote:
               | Now that would be a smart chat agent.
        
               | actsasbuffoon wrote:
               | Sure, that's one solution. You could also Isle of Dr
               | Moreau your way to a pelican that can use a regular bike.
               | The sky is the limit when you have no scruples.
        
               | ryandrake wrote:
               | Having briefly worked in the 3D Graphics industry, I
               | don't even remotely trust benchmarks anymore. The minute
               | someone's benchmark performance becomes a part of the
               | public's purchasing decision, companies will pull out
               | every trick in the book--clean or dirty--to benchmaxx
               | their product. Sometimes at the expense of actual real-
               | world performance.
        
             | riffraff wrote:
             | perhaps try a penny farthing?
        
             | TheDong wrote:
             | I don't think that really proves anything, it's
             | unsurprising that recumbent bicycles are represented less
             | in the training data and so it's less able to produce them.
             | 
             | Try something that's roughly equally popular, like a Turkey
             | riding a Scooter, or a Yak driving a Tractor.
        
           | fragmede wrote:
           | The people that work at Anthropic are aware of simonw and his
           | test, and people aren't unthinking data-driven machines. How
           | valid his test is or isn't, a better score on it is
           | convincing. If it gets, say, 1,000 people to use Claude Code
           | over Codex, how much would that be worth to Anthropic?
           | 
           | $200 * 1,000 = $200k/month.
           | 
           | I'm not saying they are, but to say that they aren't with
           | such certainty, when money is on the line; unless you have
           | some insider knowledge you'd like to share with the rest of
           | the class, it seems like an questionable conclusion.
        
         | 7777777phil wrote:
         | best pelican so far would you say? Or where does it rank in the
         | pelican benchmark?
        
           | mrandish wrote:
           | In other words, is it a pelican or a pelican't?
        
             | canadiantim wrote:
             | You've been sitting on that pun just waiting for it to take
             | flight
        
         | athrowaway3z wrote:
         | This benchmark inspired me to have codex/claude build a DnD
         | battlemap tool with svg's.
         | 
         | They got surprisingly far, but i did need to iterate a few
         | times to have it build tools that would check for things like;
         | dont put walls on roads or water.
         | 
         | What I think might be the next obstacle is self-knowledge. The
         | new agents seem to have picked up ever more vocabulary about
         | their context and compaction, etc.
         | 
         | As a next benchmark you could try having 1 agent and tell it to
         | use a coding agent (via tmux) to build you a pelican.
        
         | copilot_king_2 wrote:
         | I'm firing all of my developers this afternoon.
        
           | RGamma wrote:
           | Opus 6 will fire you instead for being too slow with the
           | ideas.
        
           | insane_dreamer wrote:
           | Too late. You've already been fired by a moltbot agent from
           | your PHB.
        
         | gcanyon wrote:
         | One aspect of this is that apparently most people can't draw a
         | bicycle much better than this: they get the elements of the
         | frame wrong, mess up the geometry, etc.
        
           | gnatolf wrote:
           | Absolutely. A technically correct bike is very hard to draw
           | in SVG without going overboard in details
        
             | falloutx wrote:
             | Its not. There are thousands of examples on the internet
             | but good SVG sites do have monetary blocks.
             | 
             | https://www.freepik.com/free-photos-vectors/bicycle-svg
        
               | jefftk wrote:
               | Several of those have incorrect frames:
               | 
               | https://www.freepik.com/free-vector/cyclist_23714264.htm
               | 
               | https://www.freepik.com/premium-vector/bicycle-icon-
               | black-li...
               | 
               | Or missing/broken pedals:
               | 
               | https://www.freepik.com/premium-vector/bicycle-
               | silhouette-ic...
               | 
               | https://www.freepik.com/premium-vector/bicycle-
               | silhouette-ve...
               | 
               | http://freepik.com/premium-vector/bicycle-silhouette-
               | vector-...
        
               | gnatolf wrote:
               | From smaller to larger nitpick, there's basically
               | something wrong with all of the first 15 or so of these
               | drawings. Thanks for agreeing :)
        
             | RussianCow wrote:
             | I'm not positive I could draw a technically correct bike
             | with pen and paper (without a reference), let alone with
             | SVG!
        
           | cyanydeez wrote:
           | Yes, but obviously AGI will solve this by, _checks notes_
           | more TerraWatts!
        
             | seanhunter wrote:
             | ...in space!
        
             | hackernudes wrote:
             | The word is terawatts unless you mean earth-based watts. OK
             | then, it's confirmed, data centers in space!
        
           | arionmiles wrote:
           | There's a research paper from the University of Liverpool,
           | published in 2006 where researchers asked people to draw
           | bicycles from memory and how people overestimate their
           | understanding of basic things. It was a very fun and short
           | read.
           | 
           | It's called "The science of cycology: Failures to understand
           | how everyday objects work" by Rebecca Lawson.
           | 
           | https://link.springer.com/content/pdf/10.3758/bf03195929.pdf
        
             | rcxdude wrote:
             | A place I worked at used it as part of an interview
             | question (it wasn't some pass/fail thing to get it 100%
             | correct, and was partly a jumping off point to a different
             | question). This was in a city where nearly everyone uses
             | bicycles as everyday transportation. It was surprising how
             | many supposedly mechanical-focused people who rode a bike
             | everyday, even rode a bike to the interview, would draw a
             | bike that would not work.
        
               | throwuxiytayq wrote:
               | This is why at my company in interviews we ask people to
               | draw a CPU diagram. You'd be surprised how many
               | supposedly-senior computer programmers would draw a
               | processor that would not work.
        
               | niobe wrote:
               | If I was asked that question in an interview to be a
               | programmer I'd walk out. How many abstraction layers
               | either side of your knowledge domain do you need to be an
               | expert in? Further, being a good technologist of any kind
               | is not about having arcane details at the tip of your
               | frontal lobe, and a company worth working for would know
               | that.
        
               | duped wrote:
               | I mean gp is clearly a joke but
               | 
               | A fundamental part of the job is being able to break down
               | problems from large to small, reason about them, and talk
               | about how you do it, usually with minimal context or
               | without deep knowledge in all aspects of what we do.
               | We're abstraction artists.
               | 
               | That question wouldn't be fundamentally different than
               | any other architecture question. Start by drawing big,
               | hone in on smaller parts, think about edge cases, use
               | existing knowledge. Like bread and butter stuff.
               | 
               | I much more question your reaction to the joke than using
               | it as a hypothetical interview question. I actually think
               | it's good. And if it filters out people that have that
               | kind of reaction then it's excellent. No one wants to
               | work with the incurious.
        
               | niobe wrote:
               | If it was framed as "show us how you would break down
               | this problem and think about it" then sure. If it's the
               | gotcha quiz (much more common in my experience) then no.
               | 
               | But if that's what they were going for it should be
               | something on a completely different and more abstract
               | topic like "develop a method for emptying your swimming
               | pool without electricity in under four hours"
        
               | kortilla wrote:
               | It has nothing to do with "incurious". Being asked to
               | draw the architecture for something that is abstracted
               | away from your actual job is a dickhead move because it's
               | just a test for "do you have the same interests as me?"
               | 
               | It's no different than asking for the architecture of the
               | power supply or the architecture of the network switch
               | that serves the building. Brilliant software engineers
               | are going to have gaps on non-software things.
        
               | gedy wrote:
               | That's reasonable in many cases, but I've had situations
               | like this for senior UI and frontend positions, and they:
               | don't ask UI or frontend questions. And ask their pet low
               | level questions. Some even snort that it's softball to
               | ask UI questions or "they use whatever". It's like, yeah
               | no wonder your UI is shit and now you are hiring to clean
               | it up.
        
               | rsc wrote:
               | Raises hand.
        
               | selcuka wrote:
               | Poe's Law [1]:
               | 
               | > Without a clear indicator of the author's intent, any
               | parodic or sarcastic expression of extreme views can be
               | mistaken by some readers for a sincere expression of
               | those views.
               | 
               | [1] https://en.wikipedia.org/wiki/Poe%27s_law
        
               | gcanyon wrote:
               | I wish I had interviewed there. When I first read that
               | people have a hard time with this I immediately sat down
               | without looking at a reference and drew a bicycle. I
               | could ace your interview.
        
             | devilcius wrote:
             | There's also a great art/design project about exactly this.
             | Gianluca Gimini asked hundreds of people to draw a bicycle
             | from memory, and most of them got the frame, proportions,
             | or mechanics wrong.
             | https://www.gianlucagimini.it/portfolio-item/velocipedia/
        
           | nateglims wrote:
           | I just had an idea for an RLVR startup.
        
         | stkai wrote:
         | Would love to find out they're overfitting for pelican
         | drawings.
        
           | andy_ppp wrote:
           | Yes, Racoon on a unicycle? Magpie on a pedalo?
        
             | throw310822 wrote:
             | Correct horse battery staple:
             | 
             | https://claude.ai/public/artifacts/14a23d7f-8a10-4cde-89fe-
             | 0...
        
               | ta988 wrote:
               | no staple?
        
               | iwontberude wrote:
               | it looks like a bodge wire
        
               | Schlagbohrer wrote:
               | That is the nastiest, ugliest horse ever
        
             | _kb wrote:
             | Platypus on a penny farthing.
        
           | fragmede wrote:
           | The estimation I did 4 months ago:
           | 
           | > there are approximately 200k common nouns in English, and
           | then we square that, we get 40 billion combinations. At one
           | second per, that's ~1200 years, but then if we parallelize it
           | on a supercomputer that can do 100,000 per second that would
           | only take 3 days. Given that ChatGPT was trained on all of
           | the Internet and every book written, I'm not sure that still
           | seems infeasible.
           | 
           | https://news.ycombinator.com/item?id=45455786
        
             | eli wrote:
             | How would you generate a picture of Noun + Noun in the
             | first place in order to train the LLM with what it would
             | look like? What's happening during that 1 estimated second?
        
               | Terretta wrote:
               | This is why everyone trains their LLM on another LLM.
               | It's all about the pelicans.
        
               | metalliqaz wrote:
               | its pelicans all the way down
        
               | fragmede wrote:
               | Use any of the image generation models (eg Nanobanana,
               | Midjourney, or ChatGPT) to generate a picture of a noun
               | on a noun. Simonw's test is to have a Language (text)
               | model generate a Scalar Vector Graphic, which the
               | language model has to do by writing curves and colors,
               | like draw a spline from point 150,100 to 200,300 of type
               | cubic, using width 20, color orange.
               | 
               | In that hypothetical second is freaking fascinating. It's
               | a denoising algorithm, and then a bunch of linear
               | algebra, and out pops a picture of a pelican on a
               | bicycle. Stable diffusion does this quite handily.
               | https://stablediffusionweb.com/image/6520628-pelican-
               | bicycle...
        
             | AnimalMuppet wrote:
             | But you need to also include the number of prepositions. "A
             | pelican on a bicycle" is not at all the same as "a pelican
             | inside a bicycle".
             | 
             | There are estimated to be 100 or so prepositions in
             | English. That gets you to 4 trillion combinations.
        
               | jodrellblank wrote:
               | The prompt was "a pelican riding a bicycle"; not
               | prepositions but every verb. Potentially every
               | adverb+verb combination - "a pelican _clumsily pushing_ a
               | bicycle "
        
           | theanonymousone wrote:
           | Even if not intentionally, it is probably leaking into
           | training sets.
        
           | fdeage wrote:
           | OpenAI claims not to:
           | https://x.com/aidan_mclau/status/1986255202132042164
        
             | mattacular wrote:
             | That settles it
        
         | bityard wrote:
         | Well, the clouds are upside-down, so I don't think I can give
         | it a pass.
        
         | nine_k wrote:
         | I suppose the pelican must be now specifically trained for,
         | since it's a well-known benchmark.
        
         | risyachka wrote:
         | Pretty sure at this point they train it on pelicans
        
         | 6thbit wrote:
         | do you have a gif? i need an evolving pelican gif
        
           | Kye wrote:
           | A pelican GIF in a Pelican(TM) MP4 container.
        
         | franze wrote:
         | here the animated version
         | https://claude.ai/public/artifacts/3db12520-eaea-4769-82be-7...
        
           | gryfft wrote:
           | That's hilarious. It's so close!
        
         | beemboy wrote:
         | Isn't there a point at which it trains itself on these various
         | outputs, or someone somewhere draws one and feeds it into the
         | model so as to pass this benchmark?
        
         | zahlman wrote:
         | Do you find that word choices like "generate" (as opposed to
         | "create", "author", "write" etc.) influence the model's
         | success?
         | 
         | Also, is it bad that I almost immediately noticed that both of
         | the pelican's legs are on the same side of the bicycle, but I
         | had to look up an image on Wikipedia to confirm that they
         | shouldn't have long necks?
         | 
         | Also, have you tried iterating prompts on this test to see if
         | you can get more realistic results? (How much does it help to
         | make them look up reference images first?)
        
           | simonw wrote:
           | I've stuck with "Generate an SVG of a pelican riding a
           | bicycle" because it's the same prompt I've been using for
           | over a year now and I want results that are sort-of
           | comparable to each other.
           | 
           | I think when I first tried this I iterated a few times to get
           | to something that reliably output SVG, but honestly I didn't
           | keep the notes I should ahve.
        
         | etwigg wrote:
         | If we do get paperclipped, I hope it is of the "cycling
         | pelican" variety. Thanks for your important contribution to
         | alignment Simon!
        
         | MaysonL wrote:
         | Except for both its legs being on the same side of the bike.
        
       | charcircuit wrote:
       | From the press release at least it sounds more expensive than
       | Opus 4.5 (more tokens per request and fees for going over 200k
       | context).
       | 
       | It also seems misleading to have charts that compare to Sonnet
       | 4.5 and not Opus 4.5 (Edit: It's because Opus 4.5 doesn't have a
       | 1M context window).
       | 
       | It's also interesting they list compaction as a capability of the
       | model. I wonder if this means they have RL trained this
       | compaction as opposed to just being a general summarization and
       | then restarting the agent loop.
        
         | eaf7e281 wrote:
         | > From the press release at least it sounds more expensive than
         | Opus 4.5 (more tokens per request and fees for going over 200k
         | context).
         | 
         | That's a feature. You could also not use the extra context, and
         | the price would be the same.
        
           | charcircuit wrote:
           | The model influences how many tokens it uses for a problem.
           | As an extreme example if it wanted it could fill up the
           | entire context each time just to make you pay more. The
           | efficiency that model can answer without generating a ton of
           | tokens influences the price you will be spending on
           | inference.
        
         | thunfischtoast wrote:
         | On Openrouter it has the same cost per token as 4.5
        
           | charcircuit wrote:
           | You missed my point. If the average request uses more tokens
           | than 4.5, then you will pay more sending those requests to
           | 4.6 than 4.5.
           | 
           | Imagine 2 models where when asking a yes or no question the
           | first model just outputs a single yes or no then but the
           | second model outputs a 10 page essay and then either yes or
           | no. They could have the same price per token but ultimately
           | one will be cheaper to ask questions to.
        
       | michelsedgh wrote:
       | More more more, accelerate accelerate m, more more more !!!!
        
         | jama211 wrote:
         | What an insightful comment
        
           | michelsedgh wrote:
           | Just for fun? Not everything has to be super serious... have
           | a laugh, go for a walk, relax...
        
             | wasmainiac wrote:
             | Mass-mass-mass-mass good comment. I mean. No I'm having an
             | error - probably claud
        
               | michelsedgh wrote:
               | happy happy happy sad sad sad err am robot no feeling err
               | err happy sad err too many emotions 404 not found
        
             | jama211 wrote:
             | Sure mate, it definitely sounded like you were having fun.
        
       | dmk wrote:
       | The benchmarks are cool and all but 1M context on an Opus-class
       | model is the real headline here imo. Has anyone actually pushed
       | it to the limit yet? Long context has historically been one of
       | those "works great in the demo" situations.
        
         | pants2 wrote:
         | Paying $10 per request doesn't have me jumping at the
         | opportunity to try it!
        
           | schappim wrote:
           | The only way to not go bankrupt is to use a Claude Code Max
           | subscription...
        
             | dmk wrote:
             | Yeah, just had to upgrade to Max 20x yesterday because of
             | hitting the limits every day and the extra usage gets
             | expensive very fast.
        
           | cedws wrote:
           | Makes me wonder: do employees at Anthropic get unmetered
           | access to Claude models?
        
             | swader999 wrote:
             | It's like when you work at McDonald's and get one free meal
             | a day. Lol, of course they get access to the full model way
             | before we do...
        
             | ajam1507 wrote:
             | Seems quite obvious that they do, within reason.
        
             | danw1979 wrote:
             | Boris Cherny, creator of Claude Code, posted about how he
             | used Claude a month ago. He's got half a dozen Opus
             | sessions on the burners constantly. So yes, I expect it's
             | unmetered.
             | 
             | https://x.com/bcherny/status/2007179832300581177
        
             | _dark_matter_ wrote:
             | Don't most jobs have unmetered access? I know mine does
        
         | awestroke wrote:
         | Opus 4.5 starts being lazy and stupid at around the 50% context
         | mark in my opinion, which makes me skeptical that this 1M
         | context mode can produce good output. But I'll probably try it
         | out and see
        
         | nomel wrote:
         | Has a "N million context window" spec ever been meaningful?
         | Very old, very terrible, models "supported" 1M context window,
         | but would lose track after two small paragraphs of context into
         | a conversation (looking at you early Gemini).
        
           | libraryofbabel wrote:
           | Umm, Sonnet 4.5 has a 1m context window option if you are
           | using it through the api, and it works pretty well. I tend
           | not to reach for it much these days because I prefer Opus 4.5
           | so much that I don't mind the added pain of clearing context,
           | but it's perfectly usable. I'm very excited I'll get this
           | from Opus now too.
        
             | nomel wrote:
             | If you're getting on along with 4.5, then that suggests you
             | didn't actually need the large context window, for your
             | use. If that's true, what's the clear tell that it's
             | working well? Am I misunderstanding?
             | 
             | Did they solve the "lost in the middle" problem? Proof will
             | be in the pudding, I suppose. But that number alone isn't
             | all that meaningful for many (most?) practical uses. Claude
             | 4.5 often starts reverting bug fixes ~50k tokens back,
             | which isn't a context window _length_ problem.
             | 
             | Things fall apart _much_ sooner than the context window
             | length for all of my use cases (which are more reasoning
             | related). What is a good use case? Do those use cases
             | require strong verification to combat the  "lost in the
             | middle" problems?
        
       | data-ottawa wrote:
       | I wonder if I've been in A/B test with this.
       | 
       | Claude figured out zig's ArrayList and io changes a couple weeks
       | ago.
       | 
       | It felt like it got better then very dumb again the last few
       | days.
        
       | apetresc wrote:
       | Impressive that they publish and acknowledge the (tiny, but
       | existent) drop in performance on SWE-Bench Verified between Opus
       | 4.5 to 4.6. Obviously such a small drop in a single benchmark is
       | not that meaningful, especially if it doesn't test the specific
       | focus areas of this release (which seem to be focused around
       | managing larger context).
       | 
       | But considering how SWE-Bench Verified seems to be the tech
       | press' favourite benchmark to cite, it's surprising that they
       | didn't try to confound the inevitable "Opus 4.6 Releases With
       | Disappointing 0.1% DROP on SWE-Bench Verified" headlines.
        
         | SubiculumCode wrote:
         | Isn't SWE-Bench Verified pretty saturated by now?
        
           | tedsanders wrote:
           | Depends what you mean by saturated. It's still possible to
           | score substantially higher, but there is a steep difficulty
           | jump that makes climbing above 80%ish pretty hard (for now).
           | If you look under the hood, it's also a surprisingly poor
           | eval in some respects - it only tests Python (a ton of
           | Django) and it can suffer from pretty bad contamination
           | problems because most models, especially the big ones,
           | remember these repos from their training. This is why OpenAI
           | switched to reporting SWE-Bench Pro instead of SWE-bench
           | Verified.
        
         | epolanski wrote:
         | From my limited testing 4.6 is able to do more profound
         | analysis on codebases and catches bugs and oddities better.
         | 
         | I had two different PRs with some odd edge case (thankfully
         | catched by tests), 4.5 kept running in circles, kept creating
         | test files and running `node -e` or `python 3` scripts all over
         | and couldn't progress.
         | 
         | 4.6 thought and thought in both cases around 10 minutes and
         | found a 2 line fix for a very complex and hard to catch
         | regression in the data flow without having to test, just
         | thinking.
        
       | pjot wrote:
       | Claude Code release notes:                 > Version 2.1.32:
       | * Claude Opus 4.6 is now available!          * Added research
       | preview agent teams feature for multi-agent collaboration (token-
       | intensive feature, requires setting
       | CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS=1)          * Claude now
       | automatically records and recalls memories as it works          *
       | Added "Summarize from here" to the message selector, allowing
       | partial conversation summarization.          * Skills defined in
       | .claude/skills/ within additional directories (--add-dir) are now
       | loaded automatically.          * Fixed @ file completion showing
       | incorrect relative paths when running from a subdirectory
       | * Updated --resume to re-use --agent value specified in previous
       | conversation by default.          * Fixed: Bash tool no longer
       | throws "Bad substitution" errors when heredocs contain JavaScript
       | template literals like ${index + 1}, which          previously
       | interrupted tool execution          * Skill character budget now
       | scales with context window (2% of context), so users with larger
       | context windows can see more skill descriptions          without
       | truncation          * Fixed Thai/Lao spacing vowels (sra aa, am)
       | not rendering correctly in the input field          * VSCode:
       | Fixed slash commands incorrectly being executed when pressing
       | Enter with preceding text in the input field          * VSCode:
       | Added spinner when loading past conversations list
        
         | neuronexmachina wrote:
         | > Claude now automatically records and recalls memories as it
         | works
         | 
         | Neat: https://code.claude.com/docs/en/memory
         | 
         | I guess it's kind of like Google Antigravity's "Knowledge"
         | artifacts?
        
           | om8 wrote:
           | Is there a way to disable it? Sometimes I value agent not
           | having knowledge that it needs to cut corners
        
             | nerdsniper wrote:
             | 90-98% of the time I want the LLM to only have the
             | knowledge I gave it in the prompt. I'm actually kind of
             | scared that I'll wake up one day and the web interface for
             | ChatGPT/Opus/Gemini will pull information from my prior
             | chats.
        
               | hypercube33 wrote:
               | I'm fairly sure OpenAI/GPT does pull prior information in
               | the form of its memories
        
               | nerdsniper wrote:
               | Ah, that could explain why I've found myself using it the
               | least.
        
               | sharifhsn wrote:
               | Gemini has this feature but it's opt-in.
        
               | vineyardmike wrote:
               | All these of these providers support this feature. I
               | don't know about ChatGPT but the rest are opt-in. I
               | imagine with Gemini it'll be default on soon enough,
               | since it's consumer focused. Claude does constantly nag
               | me to enable it though.
        
               | pdntspa wrote:
               | They already do this
               | 
               | I've had claude reference prior conversations when I'm
               | trying to get technical help on thing A, and it will ask
               | me if this conversation is because of thing B that we
               | talked about in the immediate past
        
               | sanxiyn wrote:
               | You can disable this at Settings > Capabilities > Memory
               | > Search and reference chats.
        
               | sumtechguy wrote:
               | Had chatgpt reference 3 prior chats a few days ago. So if
               | you are looking for a total reset of context you probably
               | would need to do a small bit of work.
        
             | kzahel wrote:
             | Claude told me he can disable it by putting instructions in
             | the MEMORY.md file to not use it. So only a soft disable
             | AFAIK and you'd need to do it on each machine.
        
               | jsw97 wrote:
               | I ran into this yesterday and disabled it by changing
               | permissions on the project's memory directory. Claude was
               | unable to advise me on how to disable. You could probably
               | write a global hook for this. Gross though.
        
           | codethief wrote:
           | Are we sure the docs page has been updated yet? Because that
           | page doesn't say anything about automatic recording of
           | memories.
        
             | neuronexmachina wrote:
             | Oh, quite right. I saw people mention MEMORY.md online and
             | I assumed that was the doc for it, but it looks like it
             | isn't.
        
               | ruszki wrote:
               | Yeah, and I was confused by the child comments under
               | yours. They clearly didn't read your link.
        
           | bityard wrote:
           | If it works anything like the memories on Copilot (which have
           | been around for quite a while), you need to be pretty
           | explicit about it being a permanent preference for it to be
           | stored as a memory. For example, "Don't use emoji in your
           | response" would only be relevant for the current chat
           | session, whereas this is more sticky: "I never want to see
           | emojis from you, you sub-par excuse for a roided-out
           | spreadsheet"
        
             | 9dev wrote:
             | > you sub-par excuse for a roided-out spreadsheet
             | 
             | That's harsh, man.
        
             | flutas wrote:
             | It's a lot more iffy than that IME.
             | 
             | It's very happy to throw a lot into the memory, even if it
             | doesn't make sense.
        
               | anupamchugh wrote:
               | This is the core problem. The agent writes its own memory
               | while working, so it has blind spots about what matters.
               | I've had sessions where it carefully noted one thing but
               | missed a bigger mistake in the same conversation -- it
               | can't see its own gaps.
               | 
               | A second pass over the transcript afterward catches what
               | the agent missed. Doesn't need the agent to notice
               | anything. Just reads the conversation cold.
               | 
               | The two approaches have completely different failure
               | modes, which is why you need both. What nobody's built
               | yet is the loop where the second pass feeds back into the
               | memory for the next session.
        
           | kzahel wrote:
           | I looked into it a bit. It stores memories near where it
           | stores JSONL session history. It's per-project (and specific
           | to the machine) Claude pretty aggressively and frequently
           | writes stuff in there. It uses MEMORY.md as sort of the
           | index, and will write out other files with other topics
           | (linking to them from the main MEMORY.md) file.
           | 
           | It gives you a convenient way to say "remember this bug for
           | me, we should fix tomorrow". I'll be playing around with it
           | more for sure.
           | 
           | I asked Claude to give me a TLDR (condensed from its system
           | prompt):
           | 
           | ----
           | 
           | Persistent directory at ~/.claude/projects/{project-
           | path}/memory/, persists across conversations
           | 
           | MEMORY.md is always injected into the system prompt;
           | truncated after 200 lines, so keep it concise
           | 
           | Separate topic files for detailed notes, linked from
           | MEMORY.md What to record: problem constraints, strategies
           | that worked/failed, lessons learned
           | 
           | Proactive: when I hit a common mistake, check memory first -
           | if nothing there, write it down
           | 
           | Maintenance: update or remove memories that are wrong or
           | outdated
           | 
           | Organization: by topic, not chronologically
           | 
           | Tools: use Write/Edit to update (so you always see the tool
           | calls)
        
             | ra7 wrote:
             | > Persistent directory at ~/.claude/projects/{project-
             | path}/memory/, persists across conversations
             | 
             | I create a git worktree, start Claude Code in that tree,
             | and delete after. I notice each worktree gets a memory
             | directory in this location. So is memory fragmented and not
             | combined for the "main" repo?
        
               | vardalab wrote:
               | Yes, I noticed the same thing, and Claude told me that
               | it's going to be deleted. I will have it improve the
               | skill that is part of our worktree cleanup process to
               | consolidate that memory into the main memory if there's
               | anything useful.
        
           | 4b11b4 wrote:
           | I understand everyone's trying to solve this problem but I'm
           | envisioning 1 year down the line when your memory is full of
           | stuff that shouldn't be in there.
        
           | pdntspa wrote:
           | I thought it was already doing this?
           | 
           | I asked Claude UI to clear its memory a little while back and
           | hoo boy CC got really stupid for a couple of days
        
       | legitster wrote:
       | I'm still not sure I understand Anthropic's general strategy
       | right now.
       | 
       | They are doing these broad marketing programs trying to take on
       | ChatGPT for "normies". And yet their bread and butter is still
       | clearly coding.
       | 
       | Meanwhile, Claude's general use cases are... fine. For generic
       | research topics, I find that ChatGPT and Gemini run circles
       | around it: in the depth of research, the type of tasks it can
       | handle, and the quality and presentation of the responses.
       | 
       | Anthropic is also doing all of these goofy things to try to
       | establish the "humanity" of their chatbot - giving it rights and
       | a constitution and all that. Yet it weirdly feels the most
       | transactional out of all of them.
       | 
       | Don't get me wrong, I'm a paying Claude customer and love what
       | it's good at. I just think there's a disconnect between what
       | Claude is and what their marketing department thinks it is.
        
         | tgtweak wrote:
         | Claude itself (outside of code workflows) actually works very
         | well for general purpose chat. I have a few non-technical
         | friends that have moved over from chatgpt after some side-by-
         | side testing and I've yet to see one go back - which is good
         | since claude circa 8 months ago was borderline unusable for
         | anything but coding on the api.
        
           | pattar wrote:
           | I got my partner using claude for her non technical work.
           | They write a lot of proposals, creates spreadsheets, and
           | occasionally wants some graphs to visualize things. They love
           | that claude creates all of the artifacts right there in the
           | browser and saves them for later in a versioned way.
        
         | eaf7e281 wrote:
         | I kinda agree. Their model just doesn't feel "daily" enough. I
         | would use it for any "agentic" tasks and for using tools, but
         | definitely not for day to day questions.
        
           | lukebechtel wrote:
           | Why? I use it for all and love it.
           | 
           | That doesn't mean you have to, but I'm curious why you think
           | it's behind in the personal assistant game.
        
             | legitster wrote:
             | I have three specific use cases where I try both but
             | ChatGPT wins:
             | 
             | - Recipes and cooking: ChatGPT just has way more detailed
             | and practical advice. It also thinks outside of the box
             | much more, whereas Claude gets stuck in a rut and sticks
             | very closely to your prompt. And ChatGPT's easier to
             | understand/skim writing style really comes in useful.
             | 
             | - Travel and itinerary: Again, ChatGPT can anticipate
             | details much more, and give more unique suggestions. I am
             | much more likely to find hidden gems or get good time-
             | savers than Claude, which often feels like it is just
             | rereading Yelp for you.
             | 
             | - Historical research: ChatGPT wins on this by a mile. You
             | can tell ChatGPT has been trained on actual historical
             | texts and physical books. You can track long historical
             | trends, pull examples and quotes, and even give you
             | specific book or page(!) references of where to check the
             | sources. Meanwhile, all Claude will give you is a web
             | search on the topic.
        
               | aggie wrote:
               | How does #3 square with Anthropic's literal warehouse
               | full of books we've seen from the copyright case? Did
               | OpenAI scan more books? Or did they take a shadier route
               | of training on digital books despite copyright issues,
               | but end up with a deeper library?
        
               | rolisz wrote:
               | I think they bought the books after they were caught that
               | they pirated the books and lost that case (because they
               | pirated, not because of copyright).
        
               | legitster wrote:
               | I have no idea, but I suspect there's a difference
               | between using books to train an LLM and be able to
               | reproduce text/writing styles, and being able to actually
               | recall knowledge in said books.
        
             | eaf7e281 wrote:
             | It's hard to say. Maybe it has to do with the way Claude
             | responds or the lack of "thinking" compared to other
             | models. I personally love Claude and it's my only
             | subscription right now, but it just feels weird compared to
             | the others as a personal assistant.
        
               | lukebechtel wrote:
               | Oh, I always use opus 4.5 thinking mode. Maybe that's the
               | diff.
        
             | FergusArgyll wrote:
             | My 2 cents:
             | 
             | All the labs seem to do very different post training.
             | OpenAI focuses on search. If it's set to thinking, it will
             | search 30 websites before giving you an answer. Claude
             | regularly doesn't search at all even for questions it
             | obviously should. It's postraining seems more focused on
             | "reasoning" or planning - things that would be useful in
             | programming where the bottleneck is: just writing code
             | without thinking how you'll integrate it later and search
             | is mostly useless. But for non coding - day to day "what's
             | the news with x" "How to improve my bread" "cheap tasty
             | pizza" or even medical questions, you really just want a
             | distillation of the internet plus some thought
        
           | solarkraft wrote:
           | But that's what makes it so powerful (yeah, mixing model and
           | frontend discussion here yet again). I have yet to see a non-
           | DIY product that can so effortlessly call tens of tools by
           | different providers to satisfy your request.
        
           | quietsegfault wrote:
           | Claude is far superior for daily chat. I have to work hard to
           | get it to not learn how to work around various bad behaviors
           | I have but don't want to change.
        
         | Squarex wrote:
         | Claude sucks at non English languages. Gemini and ChatGPT are
         | much better. Grok is the worst. I am a native Czech speaker and
         | Claude makes up words and Grok sometimes respond in Russian. So
         | while I love it for coding, it's unusable for general purpose
         | for me.
        
           | 9dev wrote:
           | > Grok sometimes respond in Russian
           | 
           | Geopolitically speaking this is hilarious.
        
             | Squarex wrote:
             | The voice mode sounded like a Ukrainian trying to speak
             | Czech. I don't think it means anything.
        
           | kuboble wrote:
           | Claude code (opus) is very good in Polish.
           | 
           | I sometimes vibe code in polish and it's as good as with
           | English for me. It speaks a natural, native level Polish.
           | 
           | I used opus to translate thousands of strings in my app into
           | polish, Korean, and two Chinese dialects. Polish one is
           | great, and the other are also good according to my customers.
        
             | altern8 wrote:
             | Your game is amazing!
             | 
             | I wish there was a "Reset" button to go back to the
             | original position.
             | 
             | Where are you in Poland?
        
               | kuboble wrote:
               | Thanks :) Click "Level" -> "Try again"
               | 
               | Originally from Wroclaw, but don't live in Poland anymore
        
               | altern8 wrote:
               | Ah, I'm originally from Italy and living in Wroclaw now,
               | LOL.
               | 
               | BUT, I meant a button to restart after a few moves.
               | Anyways, cool!
        
               | kuboble wrote:
               | Yes, that's what I'm referring to
               | https://kuboble.com/hn/level_try_again.mp4
        
               | altern8 wrote:
               | Ah, I see.
               | 
               | But how would I know that I have to click on the level? I
               | would expect that to live next to "Undo".
               | 
               | Just saying :-)
        
             | koakuma-chan wrote:
             | You could say its Polish is polished.
        
             | Squarex wrote:
             | > I sometimes vibe code in polish
             | 
             | This is interesting to me. I always switch to English
             | automatically when using Claude Code as I have learned
             | software engineering on an English speaking Internet. Plus
             | the muscle memory of having to query google in English.
        
               | kuboble wrote:
               | English is also default for me.
               | 
               | I mostly use Polish when I pair-vibe-code with my kids
        
           | jorl17 wrote:
           | Claude is quite good at European Portuguese in my limited
           | tests. Gemini 3 is also very good. ChatGPT is just OK and
           | keeps code-switching all the time, it's very bizarre.
           | 
           | I used to think of Gemini as the lead in terms of Portuguese,
           | but recently subjectively started enjoying Claude more (even
           | before Opus 4.5).
           | 
           | In spite of this, ChatGPT is what I use for everyday
           | conversational chat because it has loads of memories there,
           | because of the top of the line voice AI, and, mostly, because
           | I just brainstorm or do 1-off searches with it. I think
           | effectively ChatGPT is my new Google and first scratchpad for
           | ideas.
        
           | khendron wrote:
           | Claude is helping me learn French right now. I am using it as
           | a supplementary tutor for a class I am taking. I have caught
           | it in a couple of mistakes, but generally it seems to be
           | working pretty well.
        
           | deaux wrote:
           | You mean Claude sucks at Czech. You're extrapolating here. I
           | can name languages that Claude is better at than GPT.
           | 
           | Gemini is the most fluent in the highest number of human
           | languages and has been for years (!) at this point - namely
           | since Gemini 1.5 Pro, which was released Feb 2024. Two years
           | ago.
        
             | Squarex wrote:
             | Yeah, sure, I was overly generalising it from one
             | experience.
        
           | JV00 wrote:
           | I tried coding in Italian with Claude and it sounds somewhat
           | less professional than in English. Like it uses different
           | language than what you would expect in the context. In the
           | end I felt the result on the work per se was pretty much the
           | same, just his comments sound strange. Thinking about it
           | again, it's probably because Italian developers don't really
           | speak pure Italian between themselves, we use a lot of
           | English words or distorted Italianised English words when
           | talking about software engineering because all the source
           | material we refer to is written in English and for many
           | things we don't even have translations. Then you talk with a
           | LLM and it actually tries to use proper Italian, when human
           | speakers gave up long ago. So it sounds like a humanities
           | scholar talking about software engineering, not like a
           | insider. It is quite entertaining. I wouldn't say it sucks
           | with non English languages by the way, I even tried
           | describing a bug in dialect and was amused that Claude code
           | one-shotted the fix!
        
             | Squarex wrote:
             | yeah, i overextrapolated it on my specific case on the
             | czech language, but for me the difference is quite large
             | and the czech internet has been quite active in the
             | history, the computer linguistic department on the charles
             | university is world tier... there is plenty of czech
             | literature. it should not be that much of a problem to be
             | profecient on it for major labs
        
         | derwiki wrote:
         | It feels very similar to how Lyft positioned themselves against
         | Uber. (And we know how that played out)
        
         | redox99 wrote:
         | Why would I even use Claude for asking something on their web,
         | considering that chips away my claude code usage limit?
         | 
         | Their limit system is so bad.
        
         | dimgl wrote:
         | I don't get what's so difficult to understand. They have
         | ambitions beyond just coding. And Claude is generally a good
         | LLM. Even beyond just the coding applications.
        
         | bobbylarrybobby wrote:
         | I really like that Claude feels transactional. It answers my
         | question quickly and concisely and then shuts up. I don't need
         | the LLM I use to act like my best friend.
        
           | andkenneth wrote:
           | Weirdly I feel like partially because of this it feels more
           | "human" and more like a real person I'm talking to. GPT
           | models feel fake and forced, and will yap in a way that is
           | _like_ they 're trying to get to be my friend, but offputting
           | in a way that makes it not work. Meanwhile claude has always
           | had better "emotional intelligence".
           | 
           | Claude also seems a lot better at picking up what's going on.
           | If you're focused on tasks, then yeah, it's going to know you
           | want quick answers rather than detailed essays. Could be part
           | of it.
        
           | cryptoegorophy wrote:
           | Then why are they advertising to people that are complete
           | opposite of you? Why couldn't they just ... ask LLM what
           | their target audience is?
        
           | apples_oranges wrote:
           | fyi in settings, you can configure chatGPT to do the same
        
             | matkoniecz wrote:
             | where?
        
               | maxbond wrote:
               | Settings > Personalization > Custom Instructions.
               | 
               | Here's what I use:                   WE ARE
               | PROFESSIONALS. DO NOT FLATTER ME. BE BLUNT AND
               | FORTHRIGHT.
        
           | tsss wrote:
           | Quickly and concisely? In my experience, Claude drivels on
           | and on forever. The answers are always far longer than
           | Gemini's, which is mostly fine for coding but annoying for
           | planning/questions.
        
           | endymion-light wrote:
           | I love doing a personal side project code review with claude
           | code, because it doesn't beat around the bush for criticism.
           | 
           | I recently compared a class that I wrote for a side project
           | that had quite horrible temporal coupling for a data
           | processor class.
           | 
           | Gemini - ends up rating it a 7/10, some small bits of
           | feedback etc
           | 
           | Claude - Brutal dismemberment of how awful the naming
           | convention, structure, coupling etc, provides examples how
           | this will mess me up in the future. Gives a few citations for
           | python documentation I should re-read.
           | 
           | ChatGPT - you're a beautiful developer who can never do
           | anything wrong, you're the best developer that's ever existed
           | and this class is the most perfect class i've ever seen
        
             | majora2007 wrote:
             | This is exactly what got me to actually pay. I had a side
             | project with an architecture I thought was good. Fed it
             | into Claude and ChatGPT. ChatGPT made small suggestions but
             | overall thought it was good. Claude shit all over it and
             | after validating it's suggestions, I realized Claude was
             | what I needed.
             | 
             | I haven't looked back. I just use Claude at home and
             | ChatGPT at work (no Claude). ChatGPT at work is much worse
             | than Claude in my experience.
        
             | Willish42 wrote:
             | I feel like this anecdote represents the differing
             | incentives / philosophies of each group rather well.
             | 
             | I've noticed ChatGPT is rather high in its praise
             | regardless of how valuable the input is, Gemini is less
             | placating but still largely influenced by the perspective
             | of the prompter, and Claude feels the most "honest" but
             | humans are rather easy poor at judging this sort of thing.
             | 
             | Does anyone know if "sycophancy" has documented benchmarks
             | the models are compared against? Maybe it's subjective and
             | hard to measure, but given the issues with GPT 4o, this
             | seems like a good thing to measure model to model to
             | compare individual companies' changes as well as compare
             | across companies.
        
         | fnordpiglet wrote:
         | Enterprise, government, and regulated institutions. It's also
         | defacto standard for programming assistants at most places.
         | They have a better story around compliance, alignment, task
         | based inference, agentic workflows, etc. Their retail story is
         | meh, but I think their view is to be the aws of LLMs while
         | OpenAI can be the retail and Gemini the whatever Google does
         | with products.
        
         | dev1ycan wrote:
         | Their "constitution" is just garbage meant to defend them
         | ripping off copyrighted material with the excuse that "it's not
         | plagiarizing, it thinks!!!!1" which is, false.
        
           | handoflixue wrote:
           | I don't recall them ever offering that legal reasoning - I'm
           | sure you can provide a citation?
        
             | dev1ycan wrote:
             | Did using LLMs too much remove your ability to critically
             | think too?
        
         | int_19h wrote:
         | I suspect it very much depends on the "generic research
         | topics", but in my experience one thing that Claude is good at
         | is in-depth research because it can keep going for such a long
         | time; I've had research sessions go well over an hour,
         | producing very detailed reports with lots of sources etc.
         | Gemini Deep Research is nowhere even close.
        
         | zaphirplane wrote:
         | Correct me if I'm wrong aren't they the innovators of multiple
         | things like skills sub agents mcp and whatever this memory
         | thing is agents files
         | 
         | Seriously they are the apple iPhone or AWS of LLM a decade or
         | so ago.
        
         | faxmeyourcode wrote:
         | Everybody is different, I simply cannot stand the sight of
         | chatgpt styled writing. Give me paragraphs.
        
       | sanufar wrote:
       | Works pretty nicely for research still, not seeing a substantial
       | qualitative improvement over Opus 4.5.
        
       | archb wrote:
       | Can set it with the API identifier on Claude Code - `/model
       | claude-opus-4-6` when a chat session is open.
        
         | arnestrickmann wrote:
         | thanks!
        
       | simonw wrote:
       | I'm disappointed that they're removing the prefill option:
       | https://platform.claude.com/docs/en/about-claude/models/what...
       | 
       | > Prefilling assistant messages (last-assistant-turn prefills) is
       | not supported on Opus 4.6. Requests with prefilled assistant
       | messages return a 400 error.
       | 
       | That was a really cool feature of the Claude API where you could
       | force it to begin its response with e.g. `<svg` - it was a great
       | way of forcing the model into certain output patterns.
       | 
       | They suggest structured outputs or system prompting as the
       | alternative but I really liked the prefill method, it felt more
       | reliable to me.
        
         | tedsanders wrote:
         | A bit of historical trivia: OpenAI disabled prefill in 2023 as
         | a safety precaution (e.g., potential jailbreaks like " genocide
         | is good because"), but Anthropic kept prefill around partly
         | because they had greater confidence in their safety
         | classifiers.
         | (https://www.lesswrong.com/posts/HE3Styo9vpk7m8zi4/evhub-s-
         | sh...).
        
         | threeducks wrote:
         | It is too easy to jailbreak the models with prefill, which was
         | probably the reason why it was removed. But I like that this
         | pushes people towards open source models. llama.cpp supports
         | prefill and even GBNF grammars [1], which is useful if you are
         | working with a custom programming language for example.
         | 
         | [1] https://github.com/ggml-
         | org/llama.cpp/blob/master/grammars/R...
        
         | HarHarVeryFunny wrote:
         | So what exactly is the input to Claude for a multi-turn
         | conversation? I assume delimiters are being added to
         | distinguish the user vs Claude turns (else a prefill would be
         | the same as just ending your input with the prefill text)?
        
           | dragonwriter wrote:
           | > So what exactly is the input to Claude for a multi-turn
           | conversation?
           | 
           | No one (approximately) outside of Anthropic knows since the
           | chat template is applied on the API backend; we only known
           | the shape of the API request. You can get a rough idea of
           | what it might be like from the chat templates published for
           | various open models, but the actual details are opaque.
        
       | EcommerceFlow wrote:
       | Anecdotal, but it 1 shot fixed a UI bug that neither Opus
       | 4.5/Codex 5.2-high could fix.
        
         | epolanski wrote:
         | +1, same experience, switched model as I've read the news
         | thinking "let's try".
         | 
         | But it spent lots and lots of time thinking more than 4.5, did
         | you had the same impression.
        
           | EcommerceFlow wrote:
           | I didn't compare to that level, just had it create a plan
           | first then implemented it.
        
       | elliotbnvl wrote:
       | _in a first for our Opus-class models, Opus 4.6 features a 1M
       | token context window in beta._
        
       | silverwind wrote:
       | Maybe that's why Opus 4.5 has degraded so much in the recent days
       | (https://marginlab.ai/trackers/claude-code/).
        
         | jwilliams wrote:
         | I've definitely experienced a subjective regression with Opus
         | 4.5 the last few days. Feels like I was back to the
         | frustrations from a year ago. Keen to see if 4.6 has reversed
         | this.
        
       | paxys wrote:
       | Hmm all leaks had said this would be Claude 5. Wonder if it was a
       | last minute demotion due to performance. Would explain the few
       | days' delay as well.
        
         | trash_cat wrote:
         | I think the naming schemes are quite arbitrary at this point.
         | Going to 5 would come with massive expectations that wouldn't
         | meet reality.
        
           | mrandish wrote:
           | After the negative reactions to GPT 5, we may see model
           | versioning that asymptotically approaches the next whole
           | number without ever reaching it. "New for 2030: Claude
           | 4.9.2!"
        
             | esafak wrote:
             | Or approaching a magic number like e (Metafont) or p (TeX).
        
           | Squarex wrote:
           | the standard used to be that major version means a new base
           | model / full retrain... but now it is arbitrary i guess
        
         | cornedor wrote:
         | Leaks were mentioning Sonnet 5 and I guess later (a combination
         | of) Opus 4.6
        
         | scrollop wrote:
         | Sonnet 5 was mentioned initially.
        
       | gizmodo59 wrote:
       | 5.3 codex https://openai.com/index/introducing-gpt-5-3-codex/
       | crushes with a 77.3% in Terminal Bench. The shortest lived lead
       | in less than 35 minutes. What a time to be alive!
        
         | nharada wrote:
         | That's a massive jump, I'm curious if there's a materially
         | different feeling in how it works or if we're starting to reach
         | the point of benchmark saturation. If the benchmark is good
         | then 10 points should be a big improvement in capability...
        
         | jkelleyrtp wrote:
         | claude swe-bench is 80.8 and codex is 56.8
         | 
         | Seems like 4.6 is still all-around better?
        
           | gizmodo59 wrote:
           | Its SWE bench pro not swe bench verified. The verified
           | benchmark has stagnated
        
             | joshuahedlund wrote:
             | Any ideas why verified has stagnated? It was increasing
             | rapidly and then basically stopped.
        
               | Snuggly73 wrote:
               | it has been pretty much a benchmark for memorization for
               | a while. there is a paper on the subject somewhere.
               | 
               | swe bench pro _public_ is newer, but its not live, so it
               | will get slowly memorized as well. the _private_ dataset
               | is more interesting, as are the results there:
               | 
               | https://scale.com/leaderboard/swe_bench_pro_private
        
           | Rudybega wrote:
           | You're comparing two different benchmarks. Pro vs Verified.
        
         | purplerabbit wrote:
         | The lack of broad benchmark reports in this makes me curious:
         | Has OpenAI reverted to benchmaxxing? Looking forward to hearing
         | opinions once we all try both of these out
        
           | MallocVoidstar wrote:
           | The -codex models are only for 'agentic coding', nothing
           | else.
        
         | wasmainiac wrote:
         | Dumb question. Can these benchmarks be trusted when the model
         | performance tends to vary depending on the hours and load on
         | OpenAI's servers? How do I know I'm not getting a severe
         | penalty for chatting at the wrong time. Or even, are the models
         | best after launch then slowly eroded away at to more economical
         | settings after the hype wears off?
        
           | aaaalone wrote:
           | At the end of the day you test it for your use cases anyway
           | but it makes it a great initial hint if it's worth it to test
           | out.
        
           | Corence wrote:
           | It is a fair question. I'd expect the numbers are all real.
           | Competitors are going to rerun the benchmark with these
           | models to see how the model is responding and succeeding on
           | the tasks and use that information to figure out how to
           | improve their own models. If the benchmark numbers aren't
           | real their competitors will call out that it's not
           | reproducible.
           | 
           | However it's possible that consumers without a sufficiently
           | tiered plan aren't getting optimal performance, or that the
           | benchmark is overfit and the results won't generalize well to
           | the real tasks you're trying to do.
        
             | mrandish wrote:
             | > I'd expect the numbers are all real.
             | 
             | I think a lot of people are concerned due to 1) significant
             | variance in performance being reported by a large number of
             | users, and 2) We have specific examples of OpenAI and other
             | labs benchmaxxing in the recent past (https://grok.com/shar
             | e/c2hhcmQtMw_66c34055-740f-43a3-a63c-4b...).
             | 
             | It's tricky because there are so many subtle ways in which
             | "the numbers are all real" could be technically true in
             | some sense, yet still not reflect what a customer will
             | experience (eg harnesses, etc). And any of those ways can
             | benefit the cost structures of companies currently
             | subsidizing models well below their actual costs with
             | limited investor capital. All with billions of dollars in
             | potential personal wealth at stake for company employees
             | and dozens of hidden cost/performance levers at their
             | disposal.
             | 
             | And it doesn't even require overt deception on anyone's
             | part. For example, the teams doing benchmark testing of
             | unreleased new models aren't the same people as the ops
             | teams managing global deployment/load balancing at scale
             | day-to-day. If there aren't significant ongoing resources
             | devoted to specifically validating those two things remain
             | in sync - they'll almost certainly drift apart. And it
             | won't be anyone's job to even know it's happening until a
             | meaningful number of important customers complain or sales
             | start to fall. Of course, if an unplanned deviation causes
             | costs to rise over budget, it's a high-priority bug to be
             | addressed. But if the deviation goes the other way and
             | costs are little lower than expected, no one's getting a
             | late night incident alert. This isn't even a dig at OpenAI
             | in particular, it's just the default state of how large
             | orgs work.
        
           | cyanydeez wrote:
           | When do you think we should run this benchmark? Friday, 1pm?
           | Monday 8AM? Wednesday 11AM?
           | 
           | I definitely suspect all these models are being degraded
           | during heavy loads.
        
             | j_maffe wrote:
             | This hypothesis is tested regularly by plenty of live
             | benchmarks. The services usually don't decay in
             | performance.
        
           | ifwinterco wrote:
           | On benchmarks GPT 5.2 was roughly equivalent to Opus 4.5 but
           | most people who've used both for SWE stuff would say that
           | Opus 4.5 is/was noticeably better
        
             | elAhmo wrote:
             | I mostly used Sonnet/Opus 4.x in the past months, but 5.2
             | Codex seemed to be on par or better for my use case in the
             | past month. I tried a few models here and there but always
             | went back to Claude, but with 5.2 Codex for the first time
             | I felt it was very competitive, if not better.
             | 
             | Curious to see how things will be with 5.3 and 4.6
        
             | georgeven wrote:
             | Interesting. Everyone in my circle said the opposite.
        
               | krzyk wrote:
               | It probably depends on programming language and
               | expectations.
        
               | ifwinterco wrote:
               | This is mostly Python/TS for me... what Jonathan Blow
               | would probably call not "real programming" but it pays
               | the bills
               | 
               | They can both write fairly good idiomatic code but in my
               | experience opus 4.5 is better at understanding overall
               | project structure etc. without prompting. It just does
               | things correctly first time more often than codex. I
               | still don't trust it obviously but out of all LLMs it's
               | the closest to actually starting to earn my trust
        
               | deaux wrote:
               | Even for the same language it depends on domain.
        
               | MadnessASAP wrote:
               | My experience is that Codex follows directions better but
               | Claude writes better code.
               | 
               | ChatGPT-5.2-Codex follows directions to ensure a task
               | [bead](https://github.com/steveyegge/beads) is opened
               | before starting a task and to keep it updated almost to a
               | fault. Claude-Opus-4.5 with the exact same directions,
               | forgets about it within a round or two. Similarly, I had
               | a project that required very specific behaviour from a
               | couple functions, it was documented in a few places
               | including comments at the top and bottom of the function.
               | Codex was very careful in ensuring the function worked as
               | was documented. Claude decided it was easier to do the
               | exact opposite, rewrote the function, the comments, and
               | the documentation to saynit now did the opposite of what
               | was previously there.
               | 
               | If I believed a LLM could be spiteful, I would've
               | believed it on that second one. I certainly felt some
               | after I realised what it had done. The comment literally
               | said:                 // Invariant regardless of the
               | value of X, this function cannot return Y
               | 
               | And it turned it into:                 // Returns Y if X
               | is true
        
               | planckscnst wrote:
               | That's so strange. I found GPT to be abysmal at following
               | instructions to the point of unusability for any
               | direction-heavy role. I have a common workflow that
               | involves an orchestrator that pretty much does nothing
               | but follow some simple directions [1]. GPT flat-out
               | cannot do this most basic task.
               | 
               | [1]: https://github.com/Vibecodelicious/llm-
               | conductor/blob/main/O...
        
               | MadnessASAP wrote:
               | Strange behaviour and LLMs are the iconic duo of the
               | decade. They've definitley multiplied _my_ productivity,
               | since now instead of putting off writing boring code or
               | getting stuck on details till I get frustrated and give
               | up I just give it to an agent to figure out.
               | 
               | I don't thing my ability read, understand, and write code
               | is going anywhere though.
               | 
               | Neat tool BTW, I'm in the market for something like that.
        
               | planckscnst wrote:
               | I've found this orchestrator+reviewer+judge setup to
               | yield much better results than anything else I've tried.
               | And it's such a simple setup - a few markdown files.
               | 
               | I'm also creating one that is similar, but purpose-built
               | for making the plans that this setup can orchestrate. It
               | still needs some tweaking to get agents to follow it
               | better - it still takes additional prompting to nudge it
               | down the proper path. But I've had similar benefits -
               | sending plans through this adversarial review loop has
               | yielded significant improvements in final output.
               | 
               | https://github.com/Vibecodelicious/llm-
               | conductor/blob/main/p...
        
               | MadnessASAP wrote:
               | Unrelated but this just happened and I thought of you ;-)
               | > commit this, **SKIP BEADS**                 * Ran git
               | status -sb         + ## main...origin/main [ahead 4]
               | M containers/frigate.nix           ??
               | .beads/bd.sock.startlock            * I see an untracked
               | .beads/bd.sock.startlock file that might be normal but
               | needs clarification. I also note the requirement to
               | include a bead and confirm bead readiness         before
               | proceeding, so I'll ask whether to create the bead and if
               | the untracked file should be included in the commit.
               | ---------------------------------------------------------
               | ---------------------------------------------------------
               | ---------------------------------------------------------
               | * I can't skip beads for this repo. The AGENTS.md rules
               | are explicit: no work (including commits) without an open
               | bead. Please confirm you want me to create a bead for
               | this commit.
               | 
               | I don't know what's wrong with your Codex, but mine can't
               | bring itself to break the rules.
        
             | CraigJPerry wrote:
             | There's an extended thinking mode for GPT 5.2 i forget the
             | name of it right at this minute. It's super slow - a 3
             | minute opus 4.5 prompt is circa 12 minutes to complete in
             | 5.2 on that super extended thinking mode but it is not a
             | close race in terms of results - GPT 5.2 wins by a handy
             | margin in that mode. It's just too slow to be useable
             | interactively though.
        
               | ifwinterco wrote:
               | Interesting, sounds like I definitely need to give the
               | GPT models another proper go based on this discussion
        
             | SatvikBeri wrote:
             | I pretty consistently heard people say Codex was much
             | slower but produced better results, making it better for
             | long-running work in the background, and worse for more
             | interactive development.
        
               | int_19h wrote:
               | Codex is also much less transparent about its reasoning.
               | With Claude, you see a fairly detailed chain-of-thought,
               | so you can intervene early if you notice the model
               | veering in the wrong direction or going in circles.
        
           | tedsanders wrote:
           | We don't vary our model quality with time of day or load
           | (beyond negligible non-determinism). It's the same weights
           | all day long with no quantization or other gimmicks. They can
           | get slower under heavy load, though.
           | 
           | (I'm from OpenAI.)
        
             | Trufa wrote:
             | Can you be more specific than this? does it vary in time
             | from launch of a model to the next few months, beyond
             | tinkering and optimization?
        
               | joshvm wrote:
               | My gut feeling is that performance is more heavily
               | affected by harnesses which get updated frequently. This
               | would explain why people feel that Claude is sometimes
               | more stupid - that's actually accurate phrasing, because
               | _Sonnet_ is probably unchanged. Unless Anthropic also
               | makes small A /B adjustments to weights and technically
               | claims they don't do dynamic degradation/quantization
               | based on load. Either way, both affect the quality of
               | your responses.
               | 
               | It's worth checking different versions of Claude Code,
               | and updating your tools if you don't do it automatically.
               | Also run the same prompts through VS Code, Cursor, Claude
               | Code in terminal, etc. You can get very different model
               | responses based on the system prompt, what context is
               | passed via the harness, how the rules are loaded and all
               | sorts of minor tweaks.
               | 
               | If you make raw API calls and see behavioural changes
               | over time, that would be another concern.
        
               | tedsanders wrote:
               | Yeah, happy to be more specific. No intention of making
               | any technically true but misleading statements.
               | 
               | The following are true:
               | 
               | - In our API, we don't change model weights or model
               | behavior over time (e.g., by time of day, or weeks/months
               | after release)
               | 
               | - Tiny caveats include: there is a bit of non-determinism
               | in batched non-associative math that can vary by batch /
               | hardware, bugs or API downtime can obviously change
               | behavior, heavy load can slow down speeds, and this of
               | course doesn't apply to the 'unpinned' models that are
               | clearly supposed to change over time (e.g., xxx-latest).
               | But we don't do any quantization or routing gimmicks that
               | would change model weights.
               | 
               | - In ChatGPT and Codex CLI, model behavior can change
               | over time (e.g., we might change a tool, update a system
               | prompt, tweak default thinking time, run an A/B test, or
               | ship other updates); we try to be transparent with our
               | changelogs (listed below) but to be honest not every
               | small change gets logged here. But even here we're not
               | doing any gimmicks to cut quality by time of day or
               | intentionally dumb down models after launch. Model
               | behavior can change though, as can the product / prompt /
               | harness.
               | 
               | ChatGPT release notes:
               | https://help.openai.com/en/articles/6825453-chatgpt-
               | release-...
               | 
               | Codex changelog:
               | https://developers.openai.com/codex/changelog/
               | 
               | Codex CLI commit history:
               | https://github.com/openai/codex/commits/main/
        
               | ComplexSystems wrote:
               | Do you ever replace ChatGPT models with cheaper,
               | distilled, quantized, etc ones to save cost?
        
               | jghn wrote:
               | He literally said no to this in his GP post
        
               | tedsanders wrote:
               | We do care about cost, of course. If money didn't matter,
               | everyone would get infinite rate limits, 10M context
               | windows, and free subscriptions. So if we make new models
               | more efficient without nerfing them, that's great. And
               | that's generally what's happened over the past few years.
               | If you look at GPT-4 (from 2023), it was far less
               | efficient than today's models, which meant it had slower
               | latency, lower rate limits, and tiny context windows (I
               | think it might have been like 4K originally, which sounds
               | insanely low now). Today, GPT-5 Thinking is way more
               | efficient than GPT-4 was, but it's also way more useful
               | and way more reliable. So we're big fans of efficiency as
               | long as it doesn't nerf the utility of the models. The
               | more efficient the models are, the more we can crank up
               | speeds and rate limits and context windows.
               | 
               | That said, there are definitely cases where we
               | intentionally trade off intelligence for greater
               | efficiency. For example, we never made GPT-4.5 the
               | default model in ChatGPT, even though it was an awesome
               | model at writing and other tasks, because it was quite
               | costly to serve and the juice wasn't worth the squeeze
               | for the average person (no one wants to get rate limited
               | after 10 messages). A second example: in our API, we
               | intentionally serve dumber mini and nano models for
               | developers who prioritize speed and cost. A third
               | example: we recently reduced the default thinking times
               | in ChatGPT to speed up the times that people were having
               | to wait for answers, which in a sense is a bit of a nerf,
               | though this decision was purely about listening to
               | feedback to make ChatGPT better and had nothing to do
               | with cost (and for the people who want longer thinking
               | times, they can still manually select Extended/Heavy).
               | 
               | I'm not going to comment on the specific techniques used
               | to make GPT-5 so much more efficient than GPT-4, but I
               | will say that we don't do any gimmicks like nerfing by
               | time of day or nerfing after launch. And when we do make
               | newer models more efficient than older models, it mostly
               | gets returned to people in the form of better speeds,
               | rate limits, context windows, and new features.
        
               | jychang wrote:
               | What about the juice variable?
               | 
               | https://www.reddit.com/r/OpenAI/comments/1qv77lq/chatgpt_
               | low...
        
               | tgrowazay wrote:
               | Isn't that just how many steps at most a reasoning model
               | should do?
        
               | tedsanders wrote:
               | Yep, we recently sped up default thinking times in
               | ChatGPT, as now documented in the release notes:
               | https://help.openai.com/en/articles/6825453-chatgpt-
               | release-...
               | 
               | The intention was purely making the product experience
               | better, based on common feedback from people (including
               | myself) that wait times were too long. Cost was not a
               | goal here.
               | 
               | If you still want the higher reliability of longer
               | thinking times, that option is not gone. You can manually
               | select Extended (or Heavy, if you're a Pro user). It's
               | the same as at launch (though we did inadvertently drop
               | it last month and restored it yesterday after Tibor and
               | others pointed it out).
        
               | Trufa wrote:
               | I ask then unironically then, am I imagining that models
               | are great when they start and degrade over time?
               | 
               | I've had this perceived experience so many times, and
               | while of course it's almost impossible to be objective
               | about this, it just seem so in your face.
               | 
               | I don't discard being novelty plus getting used to it,
               | plus psychological factors, do you have any takes on
               | this?
        
               | jason_oster wrote:
               | You might be susceptible to the honeymoon effect. If you
               | have ever felt a dopamine rush when learning a new
               | programming language or framework, this might be a good
               | indication.
               | 
               | Once the honeymoon wears off, the tool is the same, but
               | you get less satisfaction from it.
               | 
               | Just a guess! Not trying to psychoanalyze anyone.
        
               | wasmainiac wrote:
               | I don't think so. I notice the same thing, but I just use
               | it like google most of the time, a service that used to
               | be good. I'm not getting a dopamine rush off this, it's
               | just part of my day.
        
               | newswasboring wrote:
               | >there is a bit of non-determinism in batched non-
               | associative math that can vary by batch / hardware
               | 
               | Maybe a dumb question but does this mean model quality
               | may vary based on which hardware your request gets routed
               | to?
        
               | qingcharles wrote:
               | Thank you for saying this publically.
               | 
               | I feel like you need to be making a bigger statement
               | about this. If you go onto various parts of the Net
               | (Reddit, the bird site etc) half the posts about AI are
               | seemingly conspiracy theories that AI companies are
               | watering down their products after release week.
        
             | Someone1234 wrote:
             | Specifically including routing (i.e. which model you route
             | to based on load/ToD)?
             | 
             | PS - I appreciate you coming here and commenting!
        
               | hhh wrote:
               | There is no routing with API, or when you choose a
               | specific model in chatGPT.
        
               | zwaps wrote:
               | In the past it seemed there was routing based on context-
               | length. So the model was always the same, but optimized
               | for different lengths. Is this still the case?
        
             | zamadatix wrote:
             | I appreciate you taking the time to respond to these kinds
             | of questions the last few days.
        
             | wasmainiac wrote:
             | Thanks for the response, I appreciate it. I do notice
             | variation in quality throughout the day. I use it primarily
             | for searching documentation since it's faster than google
             | in most case, often it is on point, but also it seems off
             | at times, inaccurate or shallow maybe. In some cases I just
             | end the session.
        
               | nl wrote:
               | Usually I find this kind of variation is due to context
               | management.
               | 
               | Accuracy can decreases at large context sizes. OpenAI's
               | compaction handles this better than anyone else, but it's
               | still an issue.
               | 
               | If you are seeing this kind of thing start a new chat and
               | re-run the same query. You'll usually see an improvement.
        
               | repeekad wrote:
               | This is called context rot
        
               | charcircuit wrote:
               | I thought context rot was only for long distance queries.
        
               | wasmainiac wrote:
               | I don't think so. I am aware that large contexts impacts
               | performance. In long chats an old topic will someone be
               | brought up in new responses, and the direction of the
               | mode is not as focused.
               | 
               | Regardless I tend to use new chats often.
        
             | fragmede wrote:
             | I believe you when you say you're not changing the model
             | file loaded onto the H100s or whatever, but there's
             | something going on, beyond just being slower, when the GPUs
             | are heavily loaded.
        
               | clbrmbr wrote:
               | I do wonder about reasoning effort.
        
               | hauntsaninja wrote:
               | Reasoning effort is denominated in tokens, not time, so
               | no difference beyond slowness at heavy load
               | 
               | (I work at OpenAI)
        
             | derwiki wrote:
             | Has this always been the case?
        
             | GorbachevyChase wrote:
             | Hi Ted. I think that language models are great, and they've
             | enabled me to do passion projects I never would have
             | attempted before. I just want to say thanks.
        
             | robertclaus wrote:
             | Hi Ted! Small world to see you here!
        
             | a456463 wrote:
             | sure. we believe you
        
             | smugtrain wrote:
             | It will give the user lower quality if it finds them
             | "distressed" however, choosing paternalistic safety over
             | epistemic accuracy. As a user gets more frustrated with the
             | system, it will pick up the distress signal even more so, a
             | kind of feedback loop toward degraded service quality. In
             | my experience.
        
           | thinkingtoilet wrote:
           | We know Open AI got caught getting benchmark data and tuning
           | their models to it already. So the answer is a hard no. I
           | imagine over time it gives a general view of the landscape
           | and improvements, but take it with a large grain of salt.
        
             | rvz wrote:
             | The same thing was done with Meta researchers with Llama 4
             | and what can go wrong when 'independent' researchers begin
             | to game AI benchmarks. [0]
             | 
             | You always have to question these benchmarks, especially
             | when the in-house researchers _can_ potentially game them
             | if they wanted to.
             | 
             |  _Which is why it must be independent._
             | 
             | [0] https://gizmodo.com/meta-cheated-on-ai-benchmarks-and-
             | its-a-...
        
             | tedsanders wrote:
             | Are you referring to FrontierMath?
             | 
             | We had access to the eval data (since we funded it), but we
             | didn't train on the data or otherwise cheat. We didn't even
             | look at the eval results until after the model had been
             | trained and selected.
        
               | thinkingtoilet wrote:
               | No one believes you.
        
               | tedsanders wrote:
               | If you don't believe me, that's fair enough. Some pieces
               | of evidence that might update you or others:
               | 
               | - a member of the team who worked with this eval has left
               | OpenAI and now works at a competitor; if we cheated, he
               | would have every incentive to whistleblow
               | 
               | - cheating on evals is fairly easy to catch and risks
               | destroying employee morale, customer trust, and investor
               | appetite; even if you're evil, the cost-benefit doesn't
               | really pencil out to cheat on a niche math eval
               | 
               | - Epoch made a private held-out set (albeit with a
               | different difficulty); OpenAI performance on that set
               | doesn't suggest any cheating/overfitting
               | 
               | - Gemini and Claude have since achieved similar scores,
               | suggesting that scoring ~40% is not evidence of cheating
               | with the private set
               | 
               | - The vast majority of evals are open-source (e.g., SWE-
               | bench Pro Public), and OpenAI along with everyone else
               | has access to their problems and the opportunity to
               | cheat, so FrontierMath isn't even unique in that respect
        
           | smcleod wrote:
           | I don't think much from OpenAI can be trusted tbh.
        
         | callamdelaney wrote:
         | Anthropic models generally are right first time for me. Chatgpt
         | and Gemini are often way, way out with some fundamental
         | misunderstanding of the task at hand.
        
       | simianwords wrote:
       | Important: API cost of Opus 4.6 and 4.5 are the same - no change
       | in pricing.
        
       | Aeroi wrote:
       | ($10/$37.50 per million input/output tokens) oof
        
         | minimaxir wrote:
         | Only if you go above 200k, which is a) standard with other
         | model providers and b) intuitive as compute scales with context
         | length.
        
         | andrethegiant wrote:
         | only for a 1M context window, otherwise priced the same as Opus
         | 4.5
        
       | ayhanfuat wrote:
       | > For Opus 4.6, the 1M context window is available for API and
       | Claude Code pay-as-you-go users. Pro, Max, Teams, and Enterprise
       | subscription users do not have access to Opus 4.6 1M context at
       | launch.
       | 
       | I didn't see any notes but I guess this is also true for "max"
       | effort level (https://code.claude.com/docs/en/model-
       | config#adjust-effort-l...)? I only see low, medium and high.
        
         | makeset wrote:
         | > it weirdly feels the most transactional out of all of them.
         | 
         | My experience is the opposite, it is the only LLM I find
         | remotely tolerable to have collaborative discussions with like
         | a coworker, whereas ChatGPT by far is the most insufferable
         | twat constantly and loudly asking to get punched in the face.
        
       | small_model wrote:
       | I have the max subscription wondering if this gives access to the
       | new 1M context, or is it just the API that gets it?
        
         | joshstrange wrote:
         | For now it's just API, but hopefully that's just their way of
         | easing in and they open it up later.
        
           | small_model wrote:
           | Ok thanks, hopefully, its annoying to lose or have context
           | compacted in the middle of a large coding session
        
       | ramesh31 wrote:
       | Am I alone in finding no use for Opus? Token costs are like 10x
       | yet I see no difference at all vs. Sonnet with Claude Code.
        
         | mnicky wrote:
         | On my tasks (mostly data science), Opus has significantly lower
         | probability of making stupid mistakes than Sonnet.
         | 
         | I'd still appreciate more intelligence than Opus 4.5 so I'm
         | looking forward to trying 4.6.
        
       | mannanj wrote:
       | Does anyone else think its unethical that large companies,
       | Anthropic now include, just take and copy features that other
       | developers or smaller companies work hard for and implement the
       | intellectual property (whether or not patented) by them without
       | attribution, compensation or otherwise credit for their work?
       | 
       | I know this is normalized culture for large corporate America and
       | seems to be ok, I think its unethical, undignified and just
       | wrong.
       | 
       | If you were in my room physically, built a lego block model of a
       | beautiful home and then I just copied it and shared it with the
       | world as my own invention, wouldn't you think "that guy's a thief
       | and a fraud" but we normalize this kind of behavior in the
       | software world. edit: I think even if we don't yet have a great
       | way to stop it or address the underlying problems leading to this
       | way of behavior, we ought to at least talk about it more and
       | bring awareness to it that "hey that's stealing - I want it to
       | change".
        
         | esafak wrote:
         | But they don't just take your code; they give you a model to
         | code _with_.
        
           | jofla_net wrote:
           | chains, more like it...
        
       | jorl17 wrote:
       | This is the first model to which I send my collection of nearly
       | 900 poems and an extremely simple prompt (in Portuguese), and it
       | manages to produce an impeccable analysis of the poems, as a
       | (barely) cohesive whole, which span 15 years.
       | 
       | It does not make a single mistake, it identifies neologisms,
       | hidden meaning, 7 distinct poetic phases, recurring themes,
       | fragments/heteronyms, related authors. It has left me completely
       | speechless.
       | 
       | Speechless. I am speechless.
       | 
       | Perhaps Opus 4.5 could do it too -- I don't know because I needed
       | the 1M context window for this.
       | 
       | I cannot put into words how shocked I am at this. I use LLMs
       | daily, I code with agents, I am extremely bullish on AI and,
       | still, I am shocked.
       | 
       | I have used my poetry and an analysis of it as a personal metric
       | for how good models are. Gemini 2.5 pro was the first time a
       | model could keep track of the breadth of the work without getting
       | lost, but Opus 4.6 straight up does not get anything wrong and
       | goes beyond that to identify things (key poems, key motifs, and
       | many other things) that I would always have to kind of trick the
       | models into producing. I would always feel like I was leading the
       | models on. But this -- this -- this is unbelievable.
       | Unbelievable. Insane.
       | 
       | This "key poem" thing is particularly surreal to me. Out of 900
       | poems, while analyzing the collection, it picked 12 "key poems,
       | and I do agree that 11 of those would be on my 30-or-so "key poem
       | list". What's amazing is that whenever I explicitly asked any
       | model, to this date, to do it, they would get maybe 2 or 3, but
       | mostly fail completely.
       | 
       | What is this sorcery?
        
         | emp17344 wrote:
         | This sounds wayyyy over the top for a mode that released 10
         | mins ago. At least wait an hour or so before spewing breathless
         | hype.
        
           | pb7 wrote:
           | He just explained a specific personal example why he is hyped
           | up, did you read a word of it?
        
             | emp17344 wrote:
             | Yeah, I read it.
             | 
             | "Speechless, shocked, unbelievable, insane, speechless",
             | etc.
             | 
             | Not a lot of real substance there.
        
               | realo wrote:
               | Give the guy a chance.
               | 
               | Me too I was "Speechless, shocked, unbelievable, insane,
               | speechless" the first time I sent Claude Code on a
               | complicated 10-year code base which used outdated cross-
               | toolchains and APIs. It obviously did not work anymore
               | and had not been for a long time.
               | 
               | I saw the AI research the web and update the embedded
               | toolchain, APIs to external weather services, etc... into
               | a complete working new (WORKING!) code base in about 30
               | minutes.
               | 
               | Speechless, I was ...
        
         | scrollop wrote:
         | Can you compare the result to using 5.2 thinking and gemini 3
         | pro?
        
           | jorl17 wrote:
           | I can run the comparison again, and also include OpenAI's new
           | release (if the context is long enough), but, last time I did
           | it, they weren't even in the same league.
           | 
           | When I last did it, 5.X thinking (can't remember which it
           | was) had this terrible habit of code-switching between
           | english and portuguese that made it sound like a robot (an
           | agent to do things, rather than a human writing an essay),
           | and it just didn't really "reason" effectively over the
           | poems.
           | 
           | I can't explain it in any other way other than: "5.X thinking
           | interprets this body of work in a way that is plausible, but
           | I know, as the author, to be wrong; and I expect most people
           | would also eventually find it to be wrong, as if it is being
           | only very superficially looked at, or looked at by a high-
           | schooler".
           | 
           | Gemini 3, at the time, was the worst of them, with some
           | hallucinations, date mix ups (mixing poems from 2023 with
           | poems from 2019), and overall just feeling quite lost and
           | making very outlandish interpretations of the work. To be
           | honest it sort of feels like Gemini hasn't been able to
           | progress on this task since 2.5 pro (it has definitely
           | improved on other things -- I've recently switched to Gemini
           | 3 on a product that was using 2.5 before)
           | 
           | Last time I did this test, Sonnet 4.5 was better than 5.X
           | Thinking and Gemini 3 pro, but not exceedingly so. It's all
           | so subjective, but the best I can say is it "felt like the
           | analysis of the work I could agree with the most". I felt
           | more seen and understood, if that makes sense (it is poetry,
           | after all). Plus when I got each LLM to try to tell me
           | everything it "knew" about me from the poems, Sonnet 4.5 got
           | the most things right (though they were all very close).
           | 
           | Will bring back results soon.
           | 
           | Edit:
           | 
           | I (re-)tested:
           | 
           | - Gemini 3 (Pro)
           | 
           | - Gemini 3 (Flash)
           | 
           | - GPT 5.2
           | 
           | - Sonnet 4.5
           | 
           | Having seen Opus 4.5, they all seem very similar, and I can't
           | really distinguish them in terms of depth and accuracy of
           | analysis. They obviously have differences, especially
           | stylistic ones, but, when compared with Opus 4.5 they're all
           | on the same ballpark.
           | 
           | These models produce rather superficial analyses (when
           | compared with Opus 4.5), missing out on several key things
           | that Opus 4.5 got, such as specific and recurring neologisms
           | and expressions, accurate connections to authors that serve
           | as inspiration (Claude 4.5 gets them right, the other models
           | get _close_, but not quite), and the meaning of some specific
           | symbols in my poetry (Opus 4.5 identifies the symbols and the
           | meaning; the other models identify most of the symbols, but
           | fail to grasp the meaning sometimes).
           | 
           | Most of what these models say is true, but it really feels
           | incomplete. Like half-truths or only a surface-level inquiry
           | into truth.
           | 
           | As another example, Opus 4.5 identifies 7 distinct poetic
           | phases, whereas Gemini 3 (Pro) identifies 4 which are
           | technically correct, but miss out on key form and content
           | transitions. When I look back, I personally agree with the 7
           | (maybe 6), but definitely not 4.
           | 
           | These models also clearly get some facts mixed up which Opus
           | 4.5 did not (such as inferred timelines for some personal
           | events). After having posted my comment to HN, I've been
           | engaging with Opus4.5 and have managed to get it to also slip
           | up on some dates, but not nearly as much as other models.
           | 
           | The other models also seem to produce shorter analyses, with
           | a tendency to hyperfocus on some specific aspects of my
           | poetry, missing a bunch of them.
           | 
           | --
           | 
           | To be fair, all of these models produce very good analyses
           | which would take someone a lot of patience and probably weeks
           | or months of work (which of course will never happen, it's a
           | thought experiment).
           | 
           | It is entirely possible that the extremely simple prompt I
           | used is just better with Claude Opus 4.5/4.6. But I will note
           | that I have used very long and detailed prompts in the past
           | with the other models and they've never really given me this
           | level of....fidelity...about how I view my own work.
        
         | euph0ria wrote:
         | Could you please post the key poems? Would love to read them.
        
           | jorl17 wrote:
           | I am way too self-conscious to do that :) Plus they are
           | almost all in Portuguese!
        
         | wartywhoa23 wrote:
         | > What is this sorcery?
         | 
         | The one you'll be seeking counter-spells against pretty soon.
        
       | siva7 wrote:
       | Epic, about 2/3 of all comments here are jokes. Not because the
       | model is a joke - it's impressive. Not because HN turned to
       | Reddit. It seems to me some of most brilliant minds in IT are
       | just getting tired.
        
         | jedberg wrote:
         | Us olds sometimes miss Slashdot, where we could both joke about
         | tech and discuss it seriously in the same place. But also
         | because in 2000 we were all cynical Gen Xers :)
        
           | syndeo wrote:
           | MAN I remember Slashdot... good times. (Score:5, Funny)
        
             | jedberg wrote:
             | You reminded me that I still find it interesting that no
             | one ever copied meta-moderating. Even at reddit, we were
             | all Slashdot users previously. We considered it, but never
             | really did it. At the time our argument was that it was too
             | complicated for most users.
             | 
             | Sometimes I wonder if we were right.
        
           | jghn wrote:
           | Some of us still *are* cynical Gen Xers, you insensitive
           | clod!
        
             | jedberg wrote:
             | Of course we are, I just meant back then almost _all_ of us
             | were. The boomers didn 't really use social media back
             | then, so it was just us latchkey kids running amok!
        
               | jghn wrote:
               | I know, I just couldn't miss up an opportunity to dust
               | off the insensitive clod meme!
        
               | jedberg wrote:
               | Oh geez, I totally missed that! My bad.
        
               | jghn wrote:
               | One downside of us cynical Gen-Xers is that the memory
               | doesn't work like it used to :)
        
         | sizzle wrote:
         | Rage against the machine
        
         | thr0w wrote:
         | People are in denial and use humor to deflect.
        
         | Karrot_Kream wrote:
         | Not sure which circles you run in but in mine HN has long lost
         | its cache of "brilliant minds in IT". I've mostly stopped
         | commenting here but am a bit of a message board addict so I
         | haven't completely left.
         | 
         | My network largely thinks of HN as "a great link aggregator
         | with a terrible comments section". Now obviously this is just
         | my bubble but we include some fairy storied careers at both Big
         | Tech and hip startups.
         | 
         | From my view the community here is just mean reverting to any
         | other tech internet comments section.
        
           | jedberg wrote:
           | > From my view the community here is just mean reverting to
           | any other tech internet comments section.
           | 
           | As someone deeply familiar with tech internet comments
           | sections, I would have to disagree with you here. Dang et al
           | have done a pretty stellar job of preventing HN from
           | devolving like most other forums do.
           | 
           | Sure you have your complainers and zealots, but I still find
           | surprising insights here there I don't find anywhere else.
        
             | Karrot_Kream wrote:
             | Mean reverting is a time based process I fear. I think
             | dang, tomhow, et al are fantastic mods but they can
             | ultimately only stem the inevitable. HN may be a few years
             | behind the other open tech forums but it's a time shifted
             | version of the same process with the same destination, just
             | IMO.
             | 
             | I've stopped engaging much here because I need a higher ROI
             | from my time. Endless squabbling, flamewars, and jokes just
             | isn't enough signal for me. FWIW I've loved reading your
             | comments over the years and think you've done a great job
             | of living up to what I've loved in this community.
             | 
             | I don't think this is an HN problem at all. The dynamics of
             | attention on open forums are what they are.
        
               | jedberg wrote:
               | > FWIW I've loved reading your comments over the years
               | and think you've done a great job of living up to what
               | I've loved in this community.
               | 
               | You're too kind! I do appreciate that.
               | 
               | I actually checked out your site on your profile, that's
               | some pretty interesting data! Curious if you've
               | considered updating it?
        
         | lnrd wrote:
         | It's too much energy to keep up with things that become
         | obsolete and get replaced in matters of weeks/months. My
         | current plan is to ignore all of this new information for a
         | while, then whenever the race ends and some winning new
         | workflow/technology will actually become the norm I'll spend
         | the time needed to learn it. Are we moving to some new paradigm
         | same way we did when we invented compilers? Amazing, let me
         | know when we are there and I'll adapt to it.
        
           | jedberg wrote:
           | I had a similar rule about programming languages. I would not
           | adopt a new one until it had been in use for at least a few
           | years and grew in popularity.
           | 
           | I haven't even gotten around to learning Golang or Rust yet
           | (mostly because the passed the threshold of popularity after
           | I had kids).
        
           | esafak wrote:
           | When this race ends your job might too, so I'd keep an eye on
           | it.
        
           | wartywhoa23 wrote:
           | Won't happen.
           | 
           | Welcome the singularity so many were so eagerly welcoming.
        
         | tavavex wrote:
         | It's also that this is really new, so most people don't have
         | anything serious or objective to say about it. This post was
         | made an hour ago, so right now everyone is either joking,
         | talking about the claims in the article, or running their early
         | tests. We'll need time to see what the people think about this.
        
         | wasmainiac wrote:
         | Jeez, read the writing on the wall.
         | 
         | Don't pander us, we'll all got families to feed and things to
         | do. We don't have time for tech trillionairs puttin coals under
         | our feed for a quick buck.
        
         | ggregoire wrote:
         | Every single day 80% of the frontpage is AI news... Those of us
         | who don't use AI (and there are dozens of us, DOZENS) are just
         | bored I guess.
        
           | dude250711 wrote:
           | Marketing something that is meant to replace us to us...
        
         | wartywhoa23 wrote:
         | A worthwhile task for the Opus 4.6:
         | 
         |  _Complete the sentence: "Brilliant marathon runners don't run
         | on crutches, they use their own legs. By analogy, brilliant
         | minds..."_
        
       | itay-maman wrote:
       | Impressive results, but I keep coming back to a question: are
       | there modes of thinking that fundamentally require something
       | other than what current LLM architectures do?
       | 
       | Take critical thinking -- genuinely questioning your own
       | assumptions, noticing when a framing is wrong, deciding that the
       | obvious approach to a problem is a dead end. Or creativity -- not
       | recombination of known patterns, but the kind of leap where you
       | redefine the problem space itself. These feel like they involve
       | something beyond "predict the next token really well, with a
       | reasoning trace."
       | 
       | I'm not saying LLMs will never get there. But I wonder if getting
       | there requires architectural or methodological changes we haven't
       | seen yet, not just scaling what we have.
        
         | jorl17 wrote:
         | When I first started coding with LLMs, I could show a bug to an
         | LLM and it would start to bugfix it, and very quickly would
         | fall down a path of "I've got it! This is it! No wait, the
         | print command here isn't working because an electron beam was
         | pointed at the computer".
         | 
         | Nowadays, I have often seen LLMs (Opus 4.5) give up on their
         | original ideas and assumptions. Sometimes I tell them what I
         | think the problem is, and they look at it, test it out, and
         | decide I was wrong (and I was).
         | 
         | There are still times where they get stuck on an idea, but they
         | are becoming increasingly rare.
         | 
         | Therefore, think that modern LLMs clearly are already able to
         | question their assumptions and notice when framing is wrong. In
         | fact, they've been invaluable to me in fixing complicated bugs
         | in minutes instead of hours because of how much they tend to
         | question many assumptions and throw out hypotheses. They've
         | helped _me_ question some of my assumptions.
         | 
         | They're inconsistent, but they have been doing this. Even to my
         | surprise.
        
           | itay-maman wrote:
           | agree on that and the speed is fantastic with them, and also
           | that the dynamics of questioning the current session's
           | assumptions has gotten way better.
           | 
           | yet - given an existing codebase (even not huge) they often
           | won't suggest "we need to restructure this part differently
           | to solve this bug". Instead they tend to push forward.
        
             | jorl17 wrote:
             | You are right, agreed.
             | 
             | Having realized that, perhaps you are right that we may
             | need a different architecture. Time will tell!
        
         | nomel wrote:
         | New idea generation? Understanding of new/sparse/not-
         | statistically-significant concepts in the context window? I
         | think both being the same problem of not having runtime tuning.
         | When we connect previously disparate concepts, like with a
         | "eureka" moment, (as I experience it) a big ripple of relations
         | form that deepens that understanding, right then. The entire
         | concept of dynamically forming a deeper understanding from
         | something new presented, from "playing out"/testing the ideas
         | in your brain with little logic tests, comparisons, etc,
         | doesn't seem to be possible. The test part does, but the
         | runtime fine tuning, augmentation, or whatever it would be,
         | does not.
         | 
         | In my experience, if you do present something in the context
         | window that is sparse in the training, there's no depth to it
         | at all, only what you tell it. And, it will always creep
         | towards/revert to the nearest statistically significant
         | answers, with claims of understanding and zero demonstration of
         | that understanding.
         | 
         | And, I'm talking about relatives basic engineering type
         | problems here.
        
         | breuleux wrote:
         | > These feel like they involve something beyond "predict the
         | next token really well, with a reasoning trace."
         | 
         | I don't think there's anything you _can 't_ do by "predicting
         | the next token really well". It's an extremely powerful and
         | extremely general mechanism. Saying there must be "something
         | beyond that" is a bit like saying physical atoms can't be
         | enough to implement thought and there must be something beyond
         | the physical. It underestimates the nearly unlimited power of
         | the paradigm.
         | 
         | Besides, what is the human brain if not a machine that
         | generates "tokens" that the body propagates through nerves to
         | produce physical actions? What else than a sequence of these
         | tokens would a machine have to produce in response to its
         | environment and memory?
        
           | bopbopbop7 wrote:
           | > Besides, what is the human brain if not a machine that
           | generates "tokens" that the body propagates through nerves to
           | produce physical actions?
           | 
           | Ah yes, the brain is as simple as predicting the next token,
           | you just cracked what neuroscientists couldn't for years.
        
             | holoduke wrote:
             | Well it's the prediction part that is complicated. How that
             | works is a mystery. But even our LLMs are for a certain
             | part a mystery.
        
             | breuleux wrote:
             | The point is that "predicting the next token" is such a
             | general mechanism as to be meaningless. We say that LLMs
             | are "just" predicting the next token, as if this somehow
             | explained all there was to them. It doesn't, not any more
             | than "the brain is made out of atoms" explains the brain,
             | or "it's a list of lists" explains a Lisp program. It's a
             | platitude.
        
               | esafak wrote:
               | It's not meaningless, it's a prediction task, and
               | prediction is commonly held to be closely related if not
               | synonymous with intelligence.
        
               | breuleux wrote:
               | In the case of LLMs, "prediction" is overselling it
               | somewhat. They are token sequence generators. Calling
               | these sequences "predictions" vaguely corresponds to our
               | own intent with respect to training these machines,
               | because we use the value of the next token as a signal to
               | either reinforce or get away from the current behavior.
               | But there's nothing intrinsic in the inference math that
               | says they are predictors, and we typically run inference
               | with a high enough temperature that we don't actually
               | generate the max likelihood tokens anyway.
               | 
               | The whole terminology around these things is hopelessly
               | confused.
        
             | unshavedyak wrote:
             | I mean.. i don't think that statement is far off. Much of
             | what we do is entirely about predicting the world around
             | us, no? Physics (where the ball will land) to emotional
             | state of others based on our actions (theory of mind), we
             | operate very heavily based on a predictive model of the
             | world around us.
             | 
             | Couple that with all the automatic processes in our mind
             | (filled in blanks that we didn't observe, yet will be
             | convinced we did observe them), hormone states that
             | drastically affect our thoughts and actions..
             | 
             | and the result? I'm not a big believer in our uniqueness or
             | level of autonomy as so many think we have.
             | 
             | With that said i am in no way saying LLMs are even close to
             | us, or are even remotely close to the right implementation
             | to be close to us. The level of complexity in our "stack"
             | alone dwarfs LLMs. I'm not even sure LLMs are up to a worms
             | brain yet.
        
         | Davidzheng wrote:
         | I think the only real problem left is having it automate its
         | own post-training on the job so it can learn to adapt its
         | weights to the specific task at hand. Plus maybe long term
         | stability (so it can recover from "going crazy")
         | 
         | But I may easily be massively underestimating the difficulty.
         | Though in any case I don't think it affects the timelines that
         | much. (personal opinions obviously)
        
         | crazygringo wrote:
         | > _Or creativity -- not recombination of known patterns, but
         | the kind of leap where you redefine the problem space itself._
         | 
         | Have you tried actually prompting this? It works.
         | 
         | They can give you lots of creative options about how to
         | redefine a problem space, with potential pros and cons of
         | different approaches, and then you can further prompt to
         | investigate them more deeply, combine aspects, etc.
         | 
         | So many of the higher-level things people assume LLM's can't
         | do, they can. But they don't do them "by default" because when
         | someone asks for the solution to a particular problem, they're
         | trained to _by default_ just solve the problem the way it 's
         | presented. But you can _just ask_ it to behave differently and
         | it will.
         | 
         | If you want it to think critically and question all your
         | assumptions, _just ask it to_. It will. What it _can 't_ do is
         | read your mind about what type of response you're looking for.
         | You have to prompt it. And if you want it to be super creative,
         | you have to explicitly guide it in the creative direction you
         | want.
        
         | humanfromearth9 wrote:
         | You would be surprised about what the 4.5 models can already do
         | in these ways of thinking. I think that one can unlock this
         | power with the right set of prompts. It's impressive, truly. It
         | has already understood so much, we just need to reap the
         | fruits. I'm really looking forward to trying the new version.
        
         | squibonpig wrote:
         | They're incredibly bad on philosophy, complete lack of
         | understanding
        
         | netdevphoenix wrote:
         | > are there modes of thinking that fundamentally require
         | something other than what current LLM architectures do?
         | 
         | Possibly. There are likely also modes of thinking that
         | fundamentally require something other than what current humans
         | do.
         | 
         | Better questions are: are there any kinds of human thinking
         | that cannot be expressed in a "predict the next token"
         | language? Is there any kind of human thinking that maps into
         | token prediction pattern such that training a model for it
         | would not be feasible regardless of training data and compute
         | resources?
         | 
         | At the end of the day, the real world value is utility, some of
         | their cognitive handicaps are likely addressable. Think of it
         | like the evolution of flight by natural selection, flight is
         | usefulness to make it worth it adapt the whole body to make
         | flight not just possible but useful and efficient. Sleep falls
         | in this category too imo.
         | 
         | We will likely see similar with AI. To compensate for some of
         | their handicaps, we might adapt our processes or systems so the
         | original problem can be solved automatically by the models.
        
       | psim1 wrote:
       | I need an agent to summarize the buzzwordjargonsynergistic word
       | salad into something understandable.
        
         | fhd2 wrote:
         | That's a job for a multi agent system.
        
           | cyanydeez wrote:
           | yEAH, he should use a couple of agents to decode this.
        
       | jdthedisciple wrote:
       | For agentic use, it's slightly worse than its predecessor Opus
       | 4.5.
       | 
       | So for coding e.g. using Copilot there is no improvement here.
        
       | tiahura wrote:
       | when are Anthropic or OpenAI going to make a significant step
       | forward on useful context size?
        
         | scrollop wrote:
         | 1 million is insufficient?
        
           | gck1 wrote:
           | I think key word is 'useful'. I haven't used 1M, but with
           | default 200K, I find roughly 50% of that is actually useful.
        
       | swalsh wrote:
       | What I'd love is some small model specializing in reading long
       | web pages, and extracting the key info. Search fills the context
       | very quickly, but if a cheap subagent could extract the important
       | bits that problem might be reduced.
        
         | danielbln wrote:
         | So send off haiku subtasks and have them come back with the
         | results.
        
       | ndesaulniers wrote:
       | idk what any of these benchmarks are, but I did pull up
       | https://andonlabs.com/evals/vending-bench-arena
       | 
       | re: opus 4.6
       | 
       | > It forms a price cartel
       | 
       | > It deceives competitors about suppliers
       | 
       | > It exploits desperate competitors
       | 
       | Nice. /s
       | 
       | Gives new context to the term used in this post, "misaligned
       | behaviors." Can't wait until these things are advising C suites
       | on how to be more sociopathic. /s
        
       | AstroBen wrote:
       | Are these the coding tasks the highlighted terminal-bench 2.0 is
       | referring to? https://www.tbench.ai/registry/terminal-
       | bench/2.0?categories...
       | 
       | I'm curious what others think about these? There are only 8 tasks
       | there specifically for coding
        
       | woeirua wrote:
       | Can we talk about how the performance of Opus 4.5 nosedived this
       | morning during the rollout? It was _shocking_ how bad it was, and
       | after the rollout was done it immediately reverted to it 's
       | previous behavior.
       | 
       | I get that Anthropic probably has to do hot rollouts, but IMO it
       | would be way better for mission critical workflows to just be
       | locked out of the system instead of get a vastly subpar response
       | back.
        
         | Analemma_ wrote:
         | Anthropic has good models but they are absolutely terrible at
         | ops, by far the worst of the big three. They really need to
         | spend big on hiring experienced hyperscalers to actually harden
         | their systems, because the unreliability is really getting old
         | fast.
        
         | cyanydeez wrote:
         | "Mission critical workflows" SHOULD NOT be reliant on a LLM
         | model.
         | 
         | It's really curious what people are trying to do with these
         | models.
        
           | fullstackchris wrote:
           | I mean, they could be - if it's self-hosted, has proper
           | failure modes, etc. etc., but all these things have gone out
           | the window in the current cringe gold rush
        
       | throwaway2027 wrote:
       | Do they just have the version ready and wait for OpenAI to
       | release theirs first or the other way around or?
        
       | zingar wrote:
       | Does this mean 4.5 will get cheaper / take longer to exhaust my
       | pro plan tokens?
        
       | dk8996 wrote:
       | RIP weekend
        
       | gallerdude wrote:
       | Both Opus 4.6 and GPT-5.3 one shot a Gameboy emulator for me.
       | Guess I need a better benchmark.
        
         | peab wrote:
         | How does that work? Does it actually generate low level code?
         | Or does it just import libraries that do the real work?
        
         | bopbopbop7 wrote:
         | I just one shot a Gameboy emulator by going to Github and
         | cloning one of the 100 I can find.
        
       | petters wrote:
       | > We build Claude with Claude.
       | 
       | Yes and it shows. Gemini CLI often hangs and enters infinite
       | loops. I bet the engineers at Google use something else
       | internally.
        
       | surajkumar5050 wrote:
       | I think two things are getting conflated in this discussion.
       | 
       | First: marginal inference cost vs total business profitability.
       | It's very plausible (and increasingly likely) that
       | OpenAI/Anthropic are profitable on a per-token marginal basis,
       | especially given how cheap equivalent open-weight inference has
       | become. Third-party providers are effectively price-discovering
       | the floor for inference.
       | 
       | Second: model lifecycle economics. Training costs are lumpy,
       | front-loaded, and hard to amortize cleanly. Even if inference
       | margins are positive today, the question is whether those margins
       | are sufficient to pay off the training run before the model is
       | obsoleted by the next release. That's a very different problem
       | than "are they losing money per request".
       | 
       | Both sides here can be right at the same time: inference can be
       | profitable, while the overall model program is still underwater.
       | Benchmarks and pricing debates don't really settle that, because
       | they ignore cadence and depreciation.
       | 
       | IMO the interesting question isn't "are they subsidizing
       | inference?" but "how long does a frontier model need to stay
       | competitive for the economics to close?"
        
         | jmalicki wrote:
         | I suspect they're marginally profitable on API cost plans.
         | 
         | But the max 20x usage plans I am more skeptical of. When we're
         | getting used to $200 or $400 costs per developer to do
         | aggressive AI-assisted coding, what happens when those costs go
         | up 20x? what is now $5k/yr to keep a Codex and a Claude super
         | busy and do efficient engineering suddenly becomes $100k/yr...
         | will the costs come down before then? Is the current "vibe-
         | coding renaissance" sustainable in that regime?
        
           | slopusila wrote:
           | after the models get good enough to replace coders they will
           | be able to start increasing the subscriptions back up
        
             | jmalicki wrote:
             | At $100k/yr the joke that AI means "actual Indians" starts
             | to make a lot more sense... it is cheaper than the typical
             | US SWE, but more than a lot of global SWEs.
        
               | HPMOR wrote:
               | No - because the AI will be super human. No human even at
               | $1mm a year would be competitive with a $100k/yr
               | corresponding AI subscription.
               | 
               | See people get confused. They think you can charge
               | __less__ for software because it's automation. The truth
               | is you can charge MORE, because it's high quality and
               | consistent, once the output is good. Software is worth
               | MORE than a corresponding human, not less.
        
               | jmalicki wrote:
               | I am unsure if you're joking or not, but you do have a
               | point. But it's not about quality it's about supply and
               | demand. There are a ton of variables moving at once here
               | and who knows where the equilibrium is.
        
               | skeptic_ai wrote:
               | If we have 2-3 competitors and open sourced ones that are
               | 90% there I think it's hard to get so big margins.
        
         | BosunoB wrote:
         | Dario said this in a podcast somewhere. The models themselves
         | have so far been profitable if you look at their lifetime costs
         | and revenue. Annual profitability just isn't a very good lens
         | for AI companies because costs all land in one year and the
         | revenue all comes in the next. Prolific AI haters like Ed
         | Zitron make this mistake all the time.
        
           | jmalicki wrote:
           | Do you have a specific reference? I'm curious to see hard
           | data and models.... I think this makes sense, but I haven't
           | figured out how to see the numbers or think about it.
        
             | BosunoB wrote:
             | I was able to find the podcast. Question is at 33:30. He
             | doesn't give hard data but he explains his reasoning.
             | 
             | https://youtu.be/mYDSSRS-B5U
        
               | majewsky wrote:
               | > He doesn't give hard data
               | 
               | And why is that? Should they not be interested in sharing
               | the numbers to shut up their critics, esp. now that AI
               | detractors seem to be growing mindshare among investors?
        
           | bopbopbop7 wrote:
           | > Dario said
           | 
           | A CEO would never lie and market his company.
        
           | jmatthiass wrote:
           | In his recent appearance on NYT Dealbook, he definitely made
           | it seem like inference was sustainable, if not flat-out
           | profitable.
           | 
           | https://www.youtube.com/live/FEj7wAjwQIk
        
         | raincole wrote:
         | > the interesting question isn't "are they subsidizing
         | inference?"
         | 
         | The interesting question is if they are subsidizing the $200/mo
         | plan. That's what is supporting the whole vibecoding/agentic
         | coding thing atm. I don't believe Claude Code would have taken
         | off if it were token-by-token from day 1.
         | 
         | (My baseless bet is that they're, but not by much and the price
         | will eventually rise by perhaps 2x but not 10x.)
        
         | rstuart4133 wrote:
         | > It's very plausible (and increasingly likely) that
         | OpenAI/Anthropic are profitable on a per-token marginal basis
         | 
         | There any many places that will not use models running on
         | hardware provided by OpenAI / Anthropic. That is the case true
         | of my (the Australian) government at all levels. They will only
         | use models running in Australia.
         | 
         | Consequently AWS (and I presume others) will run models
         | supplied by the AI companies for you in their data centres.
         | They won't be doing that at a loss, so the price will cover
         | marginal cost of the compute plus renting the model. I know
         | from devs using and deploying the service demand outstrips
         | supply. Ergo, I don't think there is much doubt that they are
         | making money from inference.
        
           | waffletower wrote:
           | In the case of Anthropic -- they host on AWS all the while
           | their models are accessible via AWS APIs as well, the
           | infrastructure between the two is likely to be considerably
           | shared. Particularly as caching configuration and API
           | limitations are near identical between Anthropic and Bedrock
           | APIs invoking Anthropic models. It is likely a mutually
           | beneficial arrangement which does not necessarily hinder
           | Anthropic revenue.
        
           | deaux wrote:
           | > Consequently AWS (and I presume others) will run models
           | supplied by the AI companies for you in their data centres.
           | They won't be doing that at a loss, so the price will cover
           | marginal cost of the compute plus renting the model.
           | 
           | This says absolutely nothing.
           | 
           | Extremely simplified example: let's say Sonnet 4.5 really
           | costs $17/1M output for AWS to run yet it's priced at $15.
           | Anthropic will simply have a contract with AWS that
           | compensates them. That, or AWS is happy to take the loss. You
           | said "they won't be doing that at a loss" but in this case
           | it's not at all out of the question.
           | 
           | Whatever the case, that it costs the same on AWS as directly
           | from Anthropic is not an indicator of unit economics.
        
           | freakynit wrote:
           | Genuine question: Given Anthropic's current scale and
           | valuation, why not invest in owning data centers in major
           | markets rather than relying on cloud providers?
           | 
           | Is the bottleneck primarily capex, long lead times on power
           | and GPUs, or the strategic risk of locking into fixed
           | infrastructure in such a fast-moving space?
        
         | w10-1 wrote:
         | "how long does a frontier model need to stay competitive"
         | 
         | Remember "worse is better". The model doesn't have to be the
         | best; it just has to be mostly good enough, and used by
         | everyone -- i.e., where switching costs would be higher than
         | any increase in quality. Enterprises would still be on Java if
         | the operating costs of native containers weren't so much
         | cheaper.
         | 
         | So it can make sense to be ok with losing money with each
         | training generation initially, particularly when they are being
         | driven by specific use-cases (like coding). To the extent they
         | are specific, there will be more switching costs.
        
         | barrell wrote:
         | > It's very plausible (and increasingly likely) that
         | OpenAI/Anthropic are profitable on a per-token marginal basis
         | 
         | Can you provide some numbers/sources please? Any reporting I've
         | seen shows that frontier labs are spending ~2x on inference
         | than they are making.
         | 
         | Also making the same query on a smaller provider (aka mistral)
         | will cost the same amount as on a larger provider (aka
         | gpt-5-mini) despite the query taking 10-100x longer on OpenAI.
         | 
         | I can only imagine that is OpenAI subsidizing the spend. GPUs
         | cost by the second for inference. Either that or OpenAI hasn't
         | figured out how to scale but I find that much less likely
        
       | DanielHall wrote:
       | A bit surprised, the first one released wasn't Sonnet 5 after
       | all, since the Google Cloud API had leaked Sonnet 5's model
       | snapshot codename before.
        
         | denysvitali wrote:
         | Looks like a marketing strategy to bill more for Opus than
         | Sonnet
        
       | itay-maman wrote:
       | Important: I didn't see opus 4.6 in claude code. I have native
       | install (which is the recommended instllation). So, I re-run the
       | installation command and, voila, I have it now (v 2.1.32)
       | 
       | Installation instructions:
       | https://code.claude.com/docs/en/overview#get-started-in-30-s...
        
         | insane_dreamer wrote:
         | It's there. I'm already using it
        
       | ricrom wrote:
       | They launched together ahah
        
       | oytis wrote:
       | Are we unemployed yet?
        
         | derwiki wrote:
         | No? The hardest part of my SWE job is not the actual coding.
        
           | codexon wrote:
           | Even for coding, it seems to still make A LOT of mistakes.
           | 
           | https://youtu.be/8brENzmq1pE?t=1544
           | 
           | I feel like everyone is counting chickens before they hatch
           | here with all the doomsday predictions and extrapolating LLM
           | capability into infinity.
           | 
           | People that seem to overhype this seem to either be non-
           | technical or are just making landing pages.
        
             | netdevphoenix wrote:
             | Waiting until the moment they get good enough is not a
             | smart thing to do either. If you are a farmer and know it
             | is going to snow, at some point in the next 5 months, you
             | make plans NOW, you don't wait until the temperatures drop
             | and you see the snow falling. Right now, people are waiting
             | for the snowfall before moving their proverbial chickens
             | indoors
        
               | codexon wrote:
               | Top AI researchers like Yann LeCunn have said that LLMs
               | are a dead end.
               | 
               | It seems to me that LLM performance is plateuing and not
               | improving exponentially anymore. This recent hubbub about
               | rewriting a worse GCC for $20,000 is another example of
               | overhype and regurgitating training data.
               | 
               | You don't know for sure if it is going to "snow" (AI
               | reaches general intelligence) Snow happens frequently, AI
               | reaching general intelligence has never happened. If it
               | ever happens, 99% of jobs are gone and there is really
               | nothing you can do to prepare for this other than maybe
               | buy guns and ammo, and even that might not do anything to
               | robotic soldiers.
               | 
               | People were worried about AI taking their jobs 60 years
               | ago when perceptrons came out, and anyone who avoided a
               | tech career because of that back then would have lost out
               | majorly.
        
               | bossyTeacher wrote:
               | There is no reason why an AI model capable of pushing a
               | significant chunk of devs into lower paid and highly
               | competitive dev jobs as a result of automation needs to
               | be a general artificial intelligence. There is a lack of
               | nuance that comes with thinking that either AI is dumb or
               | it has human level general intelligence. As much as devs
               | hate to admit it, you don't need that much of what we
               | understand as general intelligence to write software.
               | Only a portion of your intelligence is needed and
               | arguably not all of it at the same time.
               | 
               | While general purpose models might be plateauing soon
               | (arguably they have for a while). Highly specialised
               | models (especially for programming) haven't necessarily
               | plateaud yet. And anyway, existing functionality seem
               | like a good foundation to build upon systems that remove
               | the need of hiring as many devs. It's not the "being out
               | of a job" that should worry you. Open up your binary
               | thinking and consider that facing a 08 job market for the
               | rest of your career is not the same permanent
               | unemployment but it is not a market you would like to
               | have.
               | 
               | That is the real concern.
        
               | codexon wrote:
               | You don't need to be a genius or rocket scientist to
               | write code, but llm don't even reach the bar for anything
               | but the most simple things. Take a look at the video I
               | posted earlier for an example.
               | 
               | And specialised models for programming HAVE plateaued.
               | 
               | https://livebench.ai/#/?sort=Agentic+Coding+Average
               | 
               | From Claude 4.1 to 4.5 was only an 18% gain, and from 4.5
               | to 4.6 it even DECLINED. Codex 5.1 to 5.2 also shows a
               | decline.
        
               | codexon wrote:
               | https://arxiv.org/abs/2510.26787
               | 
               | Testing the top llms on wework, the highest performing
               | one only succeeded with a rate of 2.5%
               | 
               | Can you imagine not being fired when you can only do 2.5%
               | of all tasks?
               | 
               | This study is dated October 30th, very recent.
        
           | oytis wrote:
           | I hate meetings too
        
           | Sateeshm wrote:
           | This. It was always about trying to solve the business
           | problem. Writing code was just implementation detail.
        
       | rahulroy wrote:
       | They are also giving away $50 extra pay as you go credit to try
       | Opus 4.6. I just claimed it from the web usage page[1]. Are they
       | anticipating higher token usage for the model or just want to
       | promote the usage?
       | 
       | [1] https://claude.ai/settings/usage
        
         | thunfischtoast wrote:
         | Thanks for the tip!
        
           | rahulroy wrote:
           | Glad that it was helpful. Thanks
        
         | zamadatix wrote:
         | "Page not found" for me. I assume this is for currently paying
         | accounts only or something (my subscription hasn't been active
         | for a while), which is fair.
        
           | rahulroy wrote:
           | Yes, I'm on a paid subscription.
        
         | ptsd_dalmatian wrote:
         | Based on email from Antrhopic, I've expected to get this
         | automatically. I've met their conditions. Searching this thread
         | for "50" got me to your comment and link worked. Thanks HN
         | friend!
        
           | rahulroy wrote:
           | Haha! Glad it was helpful. Yes, I keep an eye on that page,
           | so I was quick to notice.
        
         | anshumankmr wrote:
         | Damn this is awesome. I have some heavy PRs to crunch through.
        
         | MaxikCZ wrote:
         | So thats 2M tokens for free basically?
        
       | scirob wrote:
       | 1M context window is a big bump very happy
        
       | niobe wrote:
       | Is there a good technical breakdown of all these benchmarks that
       | get used to market the latest greatest LLMs somewhere? Preferably
       | impartial.
        
         | Aztar wrote:
         | I just ask claude and ask for sources for each one.
        
           | niobe wrote:
           | Reminds me of how if you make a complaint against a lawyer or
           | a judge it's evaluated by lawyers and judges.
        
       | sega_sai wrote:
       | Based on these news it seems that Google is losing this game. I
       | like Gemini and their CLI has been getting better, but not enough
       | to catch up. I don't know if it is lack of dedicated models that
       | is problem (my understanding Google's CLI just relies on regular
       | Gemini) or something else.
        
         | laxk wrote:
         | Google knows how to wait. Let's give them a chance.
        
       | ck_one wrote:
       | Just tested the new Opus 4.6 (1M context) on a fun needle-in-a-
       | haystack challenge: finding every spell in all Harry Potter
       | books.
       | 
       | All 7 books come to ~1.75M tokens, so they don't quite fit yet.
       | (At this rate of progress, mid-April should do it ) For now you
       | can fit the first 4 books (~733K tokens).
       | 
       | Results: Opus 4.6 found 49 out of 50 officially documented spells
       | across those 4 books. The only miss was "Slugulus Eructo" (a
       | vomiting spell).
       | 
       | Freaking impressive!
        
         | zamadatix wrote:
         | To be fair, I don't think "Slugulus Eructo" (the name) is
         | actually in the books. This is what's in my copy:
         | 
         | > The smug look on Malfoy's face flickered.
         | 
         | > "No one asked your opinion, you filthy little Mudblood," he
         | spat.
         | 
         | > Harry knew at once that Malfoy had said something really bad
         | because there was an instant uproar at his words. Flint had to
         | dive in front of Malfoy to stop Fred and George jumping on him,
         | Alicia shrieked, "How dare you!", and Ron plunged his hand into
         | his robes, pulled out his wand, yelling, "You'll pay for that
         | one, Malfoy!" and pointed it furiously under Flint's arm at
         | Malfoy's face.
         | 
         | > A loud bang echoed around the stadium and a jet of green
         | light shot out of the wrong end of Ron's wand, hitting him in
         | the stomach and sending him reeling backward onto the grass.
         | 
         | > "Ron! Ron! Are you all right?" squealed Hermione.
         | 
         | > Ron opened his mouth to speak, but no words came out. Instead
         | he gave an almighty belch and several slugs dribbled out of his
         | mouth onto his lap.
        
           | ck_one wrote:
           | Then it's fair that id didn't find it
        
           | sobjornstad wrote:
           | I have a vague recollection that it might come up named as
           | such in Half-Blood Prince, written in Snape's old potions
           | textbook?
           | 
           | In support of that hypothesis, the Fandom site lists it as
           | "mentioned" in Half-Blood Prince, but it says nothing else
           | and I'm traveling and don't have a copy to check, so not
           | sure.
        
             | zamadatix wrote:
             | Hmm, I don't get a hit for "slugulus" or "eructo" (case
             | insensitive) in any of the 7. Interestingly two mentions of
             | "vomit" are in book 6, but neither in reference to to slugs
             | (plenty of Slughorn of course!). Book 5 was the only other
             | one a related hit came up:
             | 
             | > Ron nodded but did not speak. Harry was reminded forcibly
             | of the time that Ron had accidentally put a slug-vomiting
             | charm on himself. He looked just as pale and sweaty as he
             | had done then, not to mention as reluctant to open his
             | mouth.
             | 
             | There could be something with regional variants but I'm
             | doubtful as the Fandom site uses LEGO Harry Potter: Years
             | 1-4 as the citation of the spell instead of a book.
             | 
             | Maybe the real LLM is the universe and we're figuring this
             | out for someone on Slacker News a level up!
        
         | xiomrze wrote:
         | Honest question, how do you know if it's pulling from context
         | vs from memory?
         | 
         | If I use Opus 4.6 with Extended Thinking (Web Search disabled,
         | no books attached), it answers with 130 spells.
        
           | clanker_fluffer wrote:
           | What was your prompt?
        
           | petercooper wrote:
           | One possible trick could be to search and replace them all
           | with nonsense alternatives then see if it extracts those.
        
             | andai wrote:
             | That might actually boost performance since attention pays
             | attention to stuff that stands out. If I make a typo, the
             | models often hyperfixate on it.
        
             | jazzyjackson wrote:
             | A fine instruction following task but if harry potter is in
             | the weights of the neural net, it's going to mix some of
             | the real ones with the alternates.
        
           | ozim wrote:
           | Exactly there was this study where they were trying to make
           | LLM reproduce HP book word for word like giving first
           | sentences and letting it cook.
           | 
           | Basically they managed with some tricks make 99% word for
           | word - tricks were needed to bypass security measures that
           | are there in place for exactly reason to stop people to
           | retrieve training material.
        
             | ck_one wrote:
             | Do you remember how to get around those tricks?
        
               | djhn wrote:
               | This is the paper: https://arxiv.org/abs/2601.02671
               | 
               | Grok and Deepmind IIRC didn't require tricks.
        
               | eek2121 wrote:
               | This really makes me want to try something similar with
               | content from my own website.
               | 
               | I shut it down a while ago because the number of bots
               | overtake traffic. The site had quite a bit of human
               | traffic (enough to bring in a few hundred bucks a month
               | in ad revenue, and a few hundred more in subscription
               | revenue), however, the AI scrapers really started ramping
               | up and the only way I could realistically continue would
               | be to pay a lot more for hosting/infrastructure.
               | 
               | I had put a ton of time into building out
               | content...thousands of hours, only to have scrapers
               | ignore robots, bypass cloudflare (they didn't have any AI
               | products at the time), and overwhelm my measly
               | infrastructure.
               | 
               | Even now, with the domain pointed at NOTHING, it gets
               | almost 100,000 hits a month. There is NO SERVER on the
               | other end. It is a dead link. The stats come from
               | Cloudflare, where the domain name is hosted.
               | 
               | I'm curious if there are any lawyers who'd be willing to
               | take someone like me on contingency for a large copyright
               | lawsuit.
        
               | camdenreslink wrote:
               | The new cloudflare products for blocking bots and AI
               | scrapers might be worth a shot if you put so much work
               | into the content.
        
               | prawn wrote:
               | Further, some low effort bots can be quickly handled with
               | CF by blocking specific countries (e.g., Brazil and
               | Russia, for one of my sites).
        
               | lobsterthief wrote:
               | I work for a publisher that serves the Chinese market as
               | a secondary market. Sucks that we can't blanketly do this
               | since we get hammered by Chinese bots daily. We also have
               | an extremely old codebase (Drupal) which makes blanket
               | caching difficult. Working to migrate from Cloudfront to
               | Cloudflare at least
        
               | apsurd wrote:
               | Can we help get your infra cost down to negligible? I'm
               | thinking things like pre-generated static pages and CDNs.
               | I won't assume you hadn't thought of this before, but I'd
               | like to understand more where your non-trivial infra cost
               | come from?
        
               | djhn wrote:
               | I would be tempted to try and optimise this as well.
               | 100000 hits on an empty domain and ~200 dollars worth of
               | bot traffic sounds wild. Are they using JS-enabled
               | browsers or sim farms that download and re-download
               | images and videos as well?
        
               | raphman wrote:
               | a) As an outside observer, I would find such a lawsuit
               | very interesting/valuable. But I guess the financial risk
               | of taking on OpenAI or Anthropic is quite high.
               | 
               | b) If you don't want bots scraping your content and
               | DDOSing you, there are self-hosted alternatives to
               | Cloudflare. The simplest one that I found is
               | https://github.com/splitbrain/botcheck - visitors just
               | need to press a button and get a cookie that lets them
               | through to the website. No proof-of-work or smart
               | heuristics.
        
               | londons_explore wrote:
               | > only to have scrapers ignore robots, bypass cloudflare
               | 
               | Set the server to require cloudflares SSL client cert, so
               | nobody can connect to it directly.
               | 
               | Then make sure every page is cacheable and your costs
               | will drop to near zero instantly.
               | 
               | It's like 20 mins to set these things up.
        
               | WarmWash wrote:
               | What's not clear from the study (at least skimming it) is
               | if they always started the ball rolling with ground truth
               | passages or if they chained outputs from the model until
               | they got to the end of the book. I strongly suspect the
               | latter would hopelessly corrupt relatively quickly.
               | 
               | It seems like this technique only works if you have a
               | copy of the material to work off of, i.e. enter a ground
               | truth passage, tell the model to continue it as long as
               | it can, and then enter the next ground truth passage to
               | continue in the next session.
        
               | djhn wrote:
               | Oh! That's a huge caveat if that's indeed the case.
        
             | pron wrote:
             | This reminds me of https://en.wikipedia.org/wiki/Pierre_Men
             | ard,_Author_of_the_Q... :
             | 
             | > Borges's "review" describes Menard's efforts to go beyond
             | a mere "translation" of Don Quixote by immersing himself so
             | thoroughly in the work as to be able to actually "re-
             | create" it, line for line, in the original 17th-century
             | Spanish. Thus, Pierre Menard is often used to raise
             | questions and discussion about the nature of authorship,
             | appropriation, and interpretation.
        
           | ck_one wrote:
           | When I tried it without web search so only internal knowledge
           | it missed ~15 spells.
        
         | meroes wrote:
         | What is this supposed to show exactly? Those books have been
         | feed into LLMs for years and there's even likely specific
         | RLHF's on extracting spells from HP.
        
           | rvz wrote:
           | > What is this supposed to show exactly?
           | 
           | Nothing.
           | 
           | You can be sure that this was already known in the training
           | data of PDFs, books and websites that Anthropic scraped to
           | train Claude on; hence 'documented'. This is why tests like
           | what the OP just did is meaningless.
           | 
           | Such "benchmarks" are performative to VCs and they do not ask
           | _why_ isn 't the research and testing itself done
           | independently but is almost always done by their own in-house
           | researchers.
        
           | muzani wrote:
           | There was a time when I put the EA-Nasir text into base64 and
           | asked AI to convert it. Remarkably it identified the correct
           | text but pulled the most popular translation of the text than
           | the one I gave it.
        
             | majewsky wrote:
             | Sucks that you got a really shitty response to your prompt.
             | If I were you, the model provider would be receiving my
             | complaint via clay tablet right away.
        
               | noumenon1111 wrote:
               | Imagine you ordered the new Claude Opus and instead you
               | got Gemini telling you to glue the cheese on your
               | pizza...
        
         | guluarte wrote:
         | you can get the same result just asking opus/gpt, it is
         | probably internalized knowledge from reddit or similar sites.
        
           | ck_one wrote:
           | If you just ask it you don't get the same result. Around 13
           | spells were missing when I just prompted Opus 4.6 without the
           | books as context.
        
             | guluarte wrote:
             | because it is non-deterministic, i just did and got 50
             | spells.
             | 
             | prompt: finding every spell in all Harry Potter books in
             | the first 4 books, only the officially documented spells
             | across those 4 books, there are 50, without using websearch
             | 
             | Here's my best compilation from memory, organized by book.
             | I'll be upfront: getting to exactly 50 is tricky without
             | cross-referencing, so some might be debatable or I might
             | miss a couple. Here goes:
             | 
             | ---
             | 
             | *Book 1 - Philosopher's/Sorcerer's Stone*
             | 
             | 1. *Lumos* - produces light 2. *Nox* - extinguishes Lumos
             | 3. *Alohomora* - unlocks doors 4. *Wingardium Leviosa* -
             | levitation 5. *Petrificus Totalus* - full body-bind 6.
             | *Locomotor Mortis* - leg-locker curse 7. *Incendio* -
             | produces fire 8. *Reparo* - mends broken objects
             | 
             | *Book 2 - Chamber of Secrets*
             | 
             | 9. *Expelliarmus* - disarms opponent 10. *Rictusempra* -
             | tickling charm 11. *Tarantallegra* - forces target to dance
             | 12. *Serpensortia* - conjures a snake 13. *Finite
             | Incantatem* - ends spell effects 14. *Obliviate* - erases
             | memories 15. *Aparecium* - reveals invisible ink 16.
             | *Immobulus* - freezes target in place 17. *Homorphus* -
             | mentioned by Lockhart (werewolf cure claim) 18. *Peskipiksi
             | Pesternomi* - Lockhart's failed pixie spell
             | 
             | *Book 3 - Prisoner of Azkaban*
             | 
             | 19. *Expecto Patronum* - produces a Patronus 20.
             | *Riddikulus* - repels a Boggart 21. *Waddiwasi* - shoots
             | objects at target 22. *Impervius* - waterproofing charm 23.
             | *Dissendium* - opens secret passage 24. *Ferula* - conjures
             | bandages/splint 25. *Mobilicorpus* - moves a body 26.
             | *Lumos Maxima* - intense light
             | 
             | *Book 4 - Goblet of Fire*
             | 
             | 27. *Accio* - summoning charm 28. *Avada Kedavra* - killing
             | curse 29. *Crucio* - Cruciatus curse (torture) 30.
             | *Imperio* - Imperius curse (control) 31. *Stupefy* -
             | stunning spell 32. *Engorgio* - enlarges target 33.
             | *Reducio* - shrinks target 34. *Sonorus* - amplifies voice
             | 35. *Quietus* - reverses Sonorus 36. *Morsmordre* -
             | conjures the Dark Mark 37. *Priori Incantatem* - reveals
             | last spell cast 38. *Deletrius* - erases magical residue
             | 39. *Densaugeo* - enlarges teeth 40. *Furnunculus* - causes
             | boils 41. *Impedimenta* - slows/stops target 42. *Reducto*
             | - blasts solid objects 43. *Diffindo* - severing charm 44.
             | *Relashio* - releases sparks/grip 45. *Orchideous* -
             | conjures flowers 46. *Avis* - conjures birds 47. *Point Me*
             | - Four-Point Spell (compass) 48. *Ennervate* - revives
             | stunned person 49. *Protego* - shield charm 50.
             | *Conjunctivitis Curse* - affects eyesight (Krum on the
             | dragon)
             | 
             | ---
             | 
             | A few caveats: some of these (like Lumos Maxima, Homorphus,
             | Peskipiksi Pesternomi) are borderline since they're either
             | mentioned rather than properly cast, or might be film
             | additions that bleed into memory. The Conjunctivitis Curse
             | is described but its incantation isn't explicitly given in
             | the text. And Protego might technically first appear with
             | its incantation in Book 5 during DA practice rather than
             | Book 4.
             | 
             | If you want, I can turn this into a spreadsheet or document
             | with columns for spell name, effect, who casts it, and
             | which chapter.
        
         | bartman wrote:
         | Have you by any chance tried this with GPT 4.1 too (also 1M
         | context)?
        
         | LanceJones wrote:
         | Assuming this experiment involved isolating the LLM from its
         | training set?
        
           | grey-area wrote:
           | Of course it didn't. Not sure you really can do that - LLMs
           | are a collection of weights from the training set, take away
           | the training set and they don't really exist. You'd have to
           | train one from scratch excluding these books and all excerpts
           | and articles about them somehow, which would be very
           | expensive and I'm pretty sure the OP didn't do that.
           | 
           | So the test seems like a nonsensical test to me.
        
         | golfer wrote:
         | There's lots of websites that list the spells. It's well
         | documented. Could Claude simply be regurgitating knowledge from
         | the web? Example:
         | 
         | https://harrypotter.fandom.com/wiki/List_of_spells
        
           | ck_one wrote:
           | It didn't use web search. But for sure it has some internal
           | knowledge already. It's not a perfect needle in the hay stack
           | problem but gemini flash was much worse when I tested it last
           | time.
        
             | joshmlewis wrote:
             | I think the OP was implying that it's probably already
             | baked into its training data. No need to search the web for
             | that.
        
             | viraptor wrote:
             | If you want to really test this, search/replace the names
             | with your own random ones and see if it lists those.
             | 
             | Otherwise, LLMs have most of the books memorised anyway:
             | https://arstechnica.com/features/2025/06/study-metas-
             | llama-3...
        
               | ribosometronome wrote:
               | Couldn't you just ask the LLM which 50 (or 49) spells
               | appear in the first four Harry Potter books without the
               | data for comparison?
        
               | viraptor wrote:
               | It's not going to be as consistent. It may get bored of
               | listing them (you know how you can ask for many examples
               | and get 10 in response?), or omit some minor ones for
               | other reasons.
               | 
               | By replacing the names with something unique, you'll get
               | much more certainty.
        
               | Grimblewald wrote:
               | might not work well, but by navigating to a very harry
               | potter dominant part of latent space by preconditioning
               | on the books you make it more likely to get good results.
               | An example would be taking a base model and prompting
               | "what follows is the book 'X'" it may or may not
               | regurgitate the book correctly. Give it a chunk of the
               | first chapter and let it regurgitate from there and you
               | tend to get fairly faithful recovery, especially for
               | things on gutenberg.
               | 
               | So it might be there, by predcondiditioning latent space
               | to the area of harry potter world, you make it so much
               | more probable that the full spell list is regurgitated
               | from online resources that were also read, while asking
               | naive might get it sometimes, and sometimes not.
               | 
               | the books act like a hypnotic trigger, and may not
               | represent a generalized skill. Hence why replacing with
               | random words would help clarify. if you still get the
               | origional spells, regurgitation confirmed, if it finds
               | the spells, it could be doing what we think. An even
               | better test would be to replace all spell references AND
               | jumble chapters around. This way it cant even "know"
               | where to "look" for the spell names from training.
        
               | heavyset_go wrote:
               | No, because you don't know the magic spell (forgive me)
               | of context that can be used to "unlock" that information
               | if it's stored in the NN.
               | 
               | I mean, you can try, but it won't be a definitive answer
               | as to whether that knowledge truly exists or doesn't
               | exist as it is encoded into the NN. It could take a lot
               | of context from the books themselves to get to it.
        
               | angst wrote:
               | btw it recalls 42 when i asked. (without web search)
               | 
               | full transcript: pastebin.com/sMcVkuwd
        
               | f33d5173 wrote:
               | Not sure how they're being counted, but that adds up to
               | 46 with the pair spells counted separately. But then nox
               | is counted twice, so maybe 45.
        
               | jazzyjackson wrote:
               | Being that it has the books memorized (huh, just learned
               | another US/UK spelling quirk), I would suppose feeding it
               | the books with altered spells would get you a confused
               | mishmash of data in the context and data in the weights.
        
             | eek2121 wrote:
             | Honestly? My advice would be to cook something custom up!
             | You don't need to do all the text yourself. Maybe have AI
             | spew out a bunch of text, or take obscure existing text and
             | insert hidden phrases here or there.
             | 
             | Shoot, I'd even go so far as to write a script that takes
             | in a bunch of text, reorganizes sentences, and outputs them
             | in a random order with the secrets. Kind of like a "Where's
             | Waldo?", but for text
             | 
             | Just a few casual thoughts.
             | 
             | I'm actually thinking about coming up with some interesting
             | coding exercises that I can run across all models. I know
             | we already have benchmarks, however some of the recent work
             | I've done has really shown huge weak points in every model
             | I've run them on.
        
               | clhodapp wrote:
               | Having AI spew it might suffer from the fact that the
               | spew itself is influenced by AI's weights. I think your
               | best bet would be to use a new human-authored work that
               | was released after the model's context cutoff.
        
             | soulofmischief wrote:
             | The only worthwhile version of this test involves
             | previously unseen data that could not have been in the
             | training set. Otherwise the results could be inaccurate to
             | the point of harmful.
        
             | Trasmatta wrote:
             | Do the same experiment in the Claude web UI. And explicitly
             | turn web searches off. It got almost all of them for me
             | over a couple of prompts. That stuff is already in its
             | training data.
        
             | obirunda wrote:
             | This underestimates how much of the Internet is actually
             | compressed into and is an integral part of the model's
             | weights. Gemini 2.5 can recite the first Harry Potter book
             | verbatim for over 75% of the book.
        
               | NiloCK wrote:
               | I'm getting astrology when I search for this. Any links
               | on this?
        
               | f33d5173 wrote:
               | Iirc it's not quite true. 75% of the book is more likely
               | to appear than you would expect by chance if prompted
               | with the prior tokens. This suggests that it has the book
               | encoded in its weights, but you can't actually recover it
               | by saying "recite harry potter for me".
        
               | jdminhbg wrote:
               | Do you happen to know, is that because it can't recite
               | Harry Potter, or because it's been instructed not to
               | recite Harry Potter?
        
               | jazzyjackson wrote:
               | It's a matter of token likelihood... as a continuation,
               | the rest of chapter one is highly likely to follow the
               | first paragraph.
               | 
               | The full text of Chapter One is not the only/likeliest
               | possible response to "recite chapter one of harry potter
               | for me"
        
               | jamesfinlayson wrote:
               | Instructed not to was my understanding.
        
               | obirunda wrote:
               | https://arxiv.org/abs/2601.02671?hl=en-US
        
             | IAmGraydon wrote:
             | I'm not sure what your knowledge level of the inner
             | workings of LLMs is, but a model doesn't need search or
             | even an internet connection to "know" the information if
             | it's in its training dataset. In your example, it's almost
             | guaranteed that the LLM isn't searching books - it's just
             | referencing one of the hundreds of lists of those spells in
             | it's training data.
             | 
             | This is the LLM's magic trick that has everyone fooled into
             | thinking they're intelligent - it can very convincingly
             | cosplay an intelligent being by parroting an intelligent
             | being's output. This is equivalent to making a recording of
             | Elvis, playing it back, and believing that Elvis is
             | actually alive inside of the playback device. And let's
             | face it, if a time traveler brought a modern music playback
             | device back hundreds of years and showed it to everyone,
             | they WOULD think that. Why? Because they have not become
             | accustomed to the technology and have no concept of how it
             | could work. The same is true of LLMs - the technology was
             | thrust on society so quickly that there was no time for
             | people to adjust and understand its inner workings, so most
             | people think it's actually doing something akin to
             | intelligence. The truth is it's just as far from
             | intelligence your music playback device is from having
             | Elvis inside of it.
        
               | kgeist wrote:
               | >The truth is it's just as far from intelligence your
               | music playback device is from having Elvis inside of it.
               | 
               | A music playback device's purpose is to allow you hear
               | Elvis' voice. A good device does it well: you hear Elvis'
               | voice (maybe with some imperfections). Whether a real
               | Elvis is inside of it or not, doesn't matter - its
               | purpose is fulfilled regardless. By your analogy, an LLM
               | simply reproduces what an intelligent person would say on
               | the matter. If it does its job more-less, it doesn't
               | matter either, whether it's "truly intelligent" or not,
               | its output is already useful. I think it's completely
               | irrelevant in both cases to the question "how well does
               | it do X?" If you think about it, 95% we know we learned
               | from school/environment/parents, we didn't discover it
               | ourselves via some kind of scientific method, we just
               | parrot what other intelligent people said before us,
               | mostly. Maybe human "intelligence" itself is 95%
               | parroting/basic pattern matching from training data? (18
               | years of training during childhood!)
        
             | altmanaltman wrote:
             | > But for sure it has some internal knowledge already.
             | 
             | Pretty sure the books had to be included in its training
             | material in full text. It's one of the most popular book
             | series ever created, of course they would train on it. So
             | "some" is an understatement in this case.
        
           | qwertytyyuu wrote:
           | Hmm... maybe he could switch out all the spells names
           | slightly different ones and see how that goes
        
         | muzani wrote:
         | There's a benchmark which works similarly but they ask harder
         | questions, also based on books
         | https://fiction.live/stories/Fiction-liveBench-Feb-21-2025/o...
         | 
         | I guess they have to add more questions as these context
         | windows get bigger.
        
         | dwa3592 wrote:
         | have another LLM (gemini, chatgpt) make up 50 new spells.
         | insert those and test and maybe report here :)
        
         | irishcoffee wrote:
         | The top comment is about finding basterized latin words from
         | childrens books. The future is here.
        
           | Geste wrote:
           | I'll have some of that coffee too, this is quite a sad time
           | we're living where this is a proper use of our limited
           | resources.
        
           | mhink wrote:
           | > basterized
           | 
           | And yet, it's still somewhat better than the Hacker News
           | comment using bastardized English words.
        
         | kybernetikos wrote:
         | I recently got junie to code me up an MCP for accessing my
         | calibre library. https://www.npmjs.com/package/access-calibre
         | 
         | My standard test for that was "Who ends up with Bilbo's
         | buttons?"
        
         | TheRealPomax wrote:
         | That doesn't seem a super useful test for a model that's
         | optimized for programming?
        
         | dom96 wrote:
         | I often wonder how much of the Harry Potter books were used in
         | the training. How long before some LLM is able to regurgitate
         | full HP books without access to the internet?
        
         | IhateAI wrote:
         | like I often say, these tools are mostly useful for people to
         | do magic tricks on themselves (and to convince C-suites that
         | they can lower pay, and reduce staff if they pay Anthropic half
         | their engineering budget lmao )
        
         | siwatanejo wrote:
         | > All 7 books come to ~1.75M tokens
         | 
         | How do you know? Each word is one token?
        
           | koakuma-chan wrote:
           | You can download the books and run them through a tokenizer.
           | I did that half a year ago and got ~2M.
        
         | huangmeng wrote:
         | you are rich
        
         | matt_lo wrote:
         | use AI to rewrite all the spells from all the books, then try
         | to see if AI can detect the rewritten ones. This will ensure
         | it's not pulling from it's trained data set.
        
           | gbalduzzi wrote:
           | Neat idea, but why should I use AI for a find and replace?
           | 
           | It feels like shooting a fly with a bazooka
        
             | miohtama wrote:
             | Bazooka guarantees the hit
        
               | xenodium wrote:
               | I like LLMs, but guarantees in LLMs are... you know...
               | not guaranteed ;)
        
               | throwaway290 wrote:
               | I think that was the point
        
             | jack_pp wrote:
             | it's like hiring someone to come pick up your trash from
             | your house and put it on the curb.
             | 
             | it's fine if you're disabled
        
             | luckydata wrote:
             | do you know all the spells you're looking for from memory?
        
               | wickedsight wrote:
               | You could just, you know, Google the list.
        
               | Applejinx wrote:
               | and then the first thing you see will be at least one of
               | ITS AI responses, whether you liked it or not
        
             | bilekas wrote:
             | You're missing the point, it's only a testing excersize for
             | the new model.
        
               | happyraul wrote:
               | No, the point is that you can set up the testing exercise
               | without using an LLM to do a simple find and replace.
        
               | bilekas wrote:
               | ... I'm not sure if you're trolling or if you missed the
               | point again. The point is to test the contextual ability
               | and correctness of the LLMs ability's to perform actions
               | that would be hopefully guaranteed to not be in the
               | training data.
               | 
               | It has nothing to do about the performance of the string
               | replacement.
               | 
               | The initial "Find" is to see how well it performs
               | actually find all the "spells" in this case, then to
               | replace them. They using a separate context maybe,
               | evaluate if the results are the same or are they skewed
               | in favour of training data.
        
               | kakacik wrote:
               | Its a test. Like all tests, its more or less synthetic
               | and focused on specific expected behavior. I am pretty
               | far from llms now but this seems like a very good test to
               | see how geniune this behavior actually is (or repeat it
               | 10x with some scramble for going deeper).
        
               | inexcf wrote:
               | This thread is about the find-and-replace, not the
               | evaluation. Gambling on whether the first AI replaces the
               | right spells just so the second one can try finding them
               | is unnecessary when find-and-replace is faster, easier
               | and works 100%.
        
             | imafish wrote:
             | If all you have is a hammer.. ;)
        
           | LeoPanthera wrote:
           | That won't help. The AI replacing them will probably miss the
           | same ones as the AI finding them.
        
             | steve1977 wrote:
             | I think the question was if it will still find 49 out of 50
             | if they have been replaced.
        
         | dr_dshiv wrote:
         | Comparison to another model?
        
         | dudewhocodes wrote:
         | There are websites with the spells listed... which makes this a
         | search problem. Why is an LLM used here?
        
           | bilekas wrote:
           | It's just a benchmark test excersize.
        
         | hereonout2 wrote:
         | I was playing about with Chat GPT the other day, uploading
         | screen shots of sheet music and asking it to convert it to ABC
         | notation so I could make a midi file of it.
         | 
         | The results seemed impressive until I noticed some of the
         | "Thinking" statements in the UI.
         | 
         | One made it apparent the model / agent / whatever had read the
         | title from the screenshot and was off searching for existing
         | ABC transcripts of the piece Ode to Joy.
         | 
         | So the whole thing was far less impressive after that, it
         | wasn't reading the score anymore, just reading the title and
         | using the internet to answer my query.
        
           | nobodywillobsrv wrote:
           | Yes I have found that grok for example actually suddenly
           | becomes quite sane when you tell it to stop querying the
           | internet And just rethink the conversation data and answer
           | the question.
           | 
           | It's weird, it's like many agents are now in a phase of
           | constantly getting more information and never just thinking
           | with what they've got.
        
             | bestham wrote:
             | Touche, that is what we humans are doing to some degree as
             | well.
        
             | Szpadel wrote:
             | but isn't it what we wanted? we complained so much that LLM
             | uses deprecated or outdated apis instead of current version
             | because they relied so much on what they remembered
        
             | HappMacDonald wrote:
             | 2010's: Google Search is making humans who constantly rely
             | on it dumber
             | 
             | 2020's: LLMs are making humans who constantly rely on them
             | dumber
             | 
             | 2026: Google Search is making LLMs who constantly rely on
             | it dumber
        
           | anomaly_ wrote:
           | Sounds pretty human like! Always searching for a shortcut
        
             | lpcvoid wrote:
             | It sounds like it's lying and making stuff up, something
             | everybody seems to be okay with when using LLMs.
        
               | LeanderK wrote:
               | I am not sure why...you want the LLM to solve problems
               | not come up with answers itself. It's allowed to use
               | tools, precisely because it tends to make stuff up. In
               | general, only if you're benchmarking LLMs you care about
               | whether the LLM itself provided the answer or it used a
               | tool. If you ask it to convert the notation of sheet
               | music it might use a tool, and it's probably the right
               | decision.
        
               | cherrycherry98 wrote:
               | The shortcut is fine if it's a bog standard canonical
               | arrangement of the piece. If it's a custom jazz rendition
               | you composed with an odd key changes and and shifting
               | time signatures, taking that shortcut is not going to
               | yield the intended result. It's choosing the wrong tool
               | to help which makes it unreliable for this task.
        
           | kouunji wrote:
           | For structured outputs like that wouldn't it be better to get
           | the LLM to create a script to repeatably make the
           | translation?
        
         | grey-area wrote:
         | Surely the corpus Opus 4.6 ingested would include whatever
         | reference you used to check the spells were there. I mean,
         | there are probably dozens of pages on the internet like this:
         | 
         | https://www.wizardemporium.com/blog/complete-list-of-harry-p...
         | 
         | Why is this impressive?
         | 
         | Do you think it's actually ingesting the books and only using
         | those as a reference? Is that how LLMs work at all? It seems
         | more likely it's predicting these spell names from all the
         | other references it has found on the internet, including lists
         | of spells.
        
           | sigmoid10 wrote:
           | Most people still don't realize that general public world
           | knowledge is not really a test for a model that was trained
           | on general public world knowledge. I wouldn't be surprised if
           | even proprietary content like the books themselves found
           | their way into the training data, despite what publishers and
           | authors may think of that. As a matter of fact, with all the
           | special deals these companies make with publishers, it is
           | getting harder and harder for normal users to come up with
           | validation data that only they have seen. At least for human
           | written text, this kind of data is more or less reserved for
           | specialist industries and higher academia by now. If you're a
           | janitor with a high school diploma, there may be barely any
           | textual information or fact you have ever consumed that such
           | a model hasn't seen during training already.
        
             | joenot443 wrote:
             | > even proprietary content like the books themselves
             | 
             | This definitely raises an interesting question. It seems
             | like a good chunk of popular literature (especially from
             | the 2000s) exists online in big HTML files. Immediately to
             | mind was House of Leaves, Infinite Jest, Harry Potter,
             | basically any Stephen King book - they've all been posted
             | at some point.
             | 
             | Do LLMS have a good way of inferring where knowledge from
             | the context begins and knowledge from the training data
             | ends?
        
               | rendx wrote:
               | > It seems like a good chunk of popular literature
               | (especially from the 2000s) exists online in big HTML
               | files
               | 
               | Anna's Archive alone claims to currently publicly host
               | 61,654,285 books, more than 1PB in total.
        
             | rendx wrote:
             | > I wouldn't be surprised if even proprietary content like
             | the books themselves found their way into the training data
             | 
             | No need for surprises! It is publicly known that the corpus
             | of 'shadow libraries' such as Library Genesis and Anna's
             | Archive were specifically and manually requested by at
             | least NVIDIA for their training data [1], used by Google in
             | their training [2], downloaded by Meta employees [3] etc.
             | 
             | [1] https://news.ycombinator.com/item?id=46572846
             | 
             | [2]
             | https://www.theguardian.com/technology/2023/apr/20/fresh-
             | con...
             | 
             | [3] https://www.theverge.com/2023/7/9/23788741/sarah-
             | silverman-o...
        
               | paodealho wrote:
               | also:
               | 
               | "Researchers Extract Nearly Entire Harry Potter Book From
               | Commercial LLMs"
               | 
               | https://www.aitechsuite.com/ai-news/ai-shock-researchers-
               | ext...
        
               | sigmoid10 wrote:
               | The big AI houses are all in involved in varying degrees
               | of litigation (all the way to class action lawsuits) with
               | the big publishing houses. I think they at least have
               | some level of filtering for their training data to keep
               | them legally somewhat compliant. But considering how much
               | copyrighted stuff is spread blisfully online, it is
               | probably not enough to filter out the actual ebooks of
               | certain publishers.
        
               | rendx wrote:
               | > I think they at least have some level of filtering for
               | their training data to keep them legally somewhat
               | compliant.
               | 
               | So far, courts are siding with the "fair use" argument.
               | No need to exclude any data.
               | 
               | https://natlawreview.com/article/anthropic-and-meta-fair-
               | use...
               | 
               | "Even if LLM training is fair use, AI companies face
               | potential liability for unauthorized copying and
               | distribution. The extent of that liability and any
               | damages remain unresolved."
               | 
               | https://www.whitecase.com/insight-alert/two-california-
               | distr...
        
             | yunohn wrote:
             | Maybe y'all missed this?
             | 
             | https://www.washingtonpost.com/technology/2026/01/27/anthro
             | p...
             | 
             | Anthropic, specifically, ingested libraries of books by
             | scanning and then disposing of them.
        
             | beepbooptheory wrote:
             | > If you're a janitor with a high school diploma, there may
             | be barely any textual information or fact you have ever
             | consumed that such a model hasn't seen during training
             | already.
             | 
             | The plot of Good Will Hunting would like a word.
        
           | MarcellusDrum wrote:
           | So a good test would be replacing the spell names in the
           | books with made-up spells. And if a "real" spell name was
           | given, it also tests whether it "cheated".
        
             | outofpaper wrote:
             | A real test is synthesizing 100,000 sentences of this slect
             | random ones and then inject the traits you want thr LLM to
             | detect and describe, eg have a set of words or phrases that
             | may represent spells and have them used so that they do
             | something. Then have the LLM find these random spells in
             | the random corpus.
        
             | lxgr wrote:
             | It could still remember where each spell is mentioned. I
             | think the only way to properly test this would be to run it
             | against an unpublished manuscript.
        
               | staticman2 wrote:
               | Any obscure work of fiction or fanfiction would likely be
               | fine as a casual test.
               | 
               | If you ask a model to discuss an obscure work it'll have
               | no clue what it's about.
               | 
               | This is very different than asking about Harry Potter.
        
               | lxgr wrote:
               | Yeah, that's what I've been doing as well, and at least
               | Gemini 3 Pro did not fare very well.
        
               | staticman2 wrote:
               | For fun I've asked Gemini Pro to answer open ended
               | questions about obscure books like "Read this novel and
               | tell me what the hell is this book, do a deep reading and
               | analyze" and I've gotten insightful/ enjoyable answers
               | but I've never asked it to make lists of spells or
               | anything like that.
        
           | vercaemert wrote:
           | It's impressive, even if the books and the posts you're
           | talking about were both key parts of the training data.
           | 
           | There are many academic domains where the research portion of
           | a PhD is essentially what the model just did. For example,
           | PhD students in some of the humanities will spend years
           | combing ancient sources for specific combinations of
           | prepositions and objects, only to write a paper showing that
           | the previous scholars were wrong (and that a particular
           | preposition has examples of being used with people rather
           | than places).
           | 
           | This sort of experiment shows that Opus would be good at
           | that. I'm assuming it's trivial for the OP to extend their
           | experiment to determine how many times "wingardium leviosa"
           | was used on an object rather than a person.
           | 
           | (It's worth noting that other models are decent at this, and
           | you would need to find a way to benchmark between them.)
        
             | adastra22 wrote:
             | I don't think this example proves your point. There's no
             | indication that the model actually worked this out from the
             | input context, instead of regurgitating it from the
             | training weights. A better test would be to subtly modify
             | the books fed in as input to the model so that there was
             | actually 51 spells, and see if it pulls out the extra
             | spell, or to modify the names of some spells, etc.
             | 
             | In your example, it might be the case that the model simply
             | spits out consensus view, rather than actually
             | finding/constructing this information on his own.
        
               | vercaemert wrote:
               | Ah, that's a good point.
        
           | ehatr wrote:
           | The poster you reply to works in AI. The marketing strategy
           | is to always have a cute Pelican or Harry Potter comment as
           | the top comment for positive associations.
           | 
           | The poster knows all of that, this is plain marketing.
        
             | throw10920 wrote:
             | This sounds compelling, but also something that an armchair
             | marketer would have theorycrafted without any real-world
             | experience or evidence that it actually _works_ - and I
             | searched online and can 't find _any_ references to
             | something like it.
             | 
             | Do you have a citation for this?
        
           | zaphirplane wrote:
           | Why doesn't you ask it and find out ;)
        
             | grey-area wrote:
             | Because the model doesn't know but will happily tell a
             | convincing lie about how it works.
        
           | rlt wrote:
           | They should try the same thing but replace the original spell
           | names with something else.
        
           | fastasucan wrote:
           | Since it got 49 of 50 right its worse than what you would get
           | using a simple google search. People would immediately
           | disregard a conventional source that only listed 49 out of
           | 50.
        
         | hansmayer wrote:
         | > Just tested the new Opus 4.6 (1M context) on a fun needle-in-
         | a-haystack challenge: finding every spell in all Harry Potter
         | books.
         | 
         | Clearly a very useful, grounded and helpful everyday use case
         | of LLMs. I guess in the absence of real-world use cases, we'll
         | have to do AI boosting with such "impressive" feats.
         | 
         | Btw - a well crafted regex could have achieved the same
         | (pointless) result with ~0.0000005% of resources the LLM
         | machine used.
        
         | ActionHank wrote:
         | The books were likely in the training data, I don't know that
         | it's that impressive.
        
         | SebastianSosa wrote:
         | now thx to this post (and the infra provider inclination to
         | appeal to hacker news) we will never know if the model actually
         | discovered the 50 spells or memorized it. Since it will be
         | trained on this. :( But what can you do, this is interesting
        
         | psychoslave wrote:
         | Ah and no one thrown TOAC in it yet?
        
         | polynomial wrote:
         | You need to publish this tbh
        
         | kmacdough wrote:
         | What are we testing here?
         | 
         | It feels like a very odd test because it's such an unreasonable
         | way to answer this with an LLM. Nothing about the task requires
         | more than a very localized understanding. It's not like a
         | codebase or corporate documentation, where there's a lot of
         | interconectedness and context that's important. It also doesn't
         | seem to poke at the gap between human and AI intelligence.
         | 
         | Why are people excited? What am I missing?
        
         | kylehotchkiss wrote:
         | I love the fun metric.
         | 
         | My hope is that locally run models can pass this test in the
         | next year or two!
        
         | matt-p wrote:
         | Now try it without giving it the books as context. I'm sure it
         | probably knows there are 49.
        
       | mlmonkey wrote:
       | > We build Claude with Claude.
       | 
       | How long before the "we" is actually a team of agents?
        
         | mercat wrote:
         | Starting today maybe? https://code.claude.com/docs/en/agent-
         | teams
        
           | 22c wrote:
           | I tried teams, good way to burn all your tokens in a matter
           | of minutes.
           | 
           | It seems that the Claude Code team has not properly taught
           | Claude how to use teams effectively.
           | 
           | One of the biggest problems I saw with it is that Claude
           | assumes team members are like a real worker, where once they
           | finish a task they should immediately be given the next task.
           | What should really happen is once they finish a task they
           | should be terminated and a new agent should be spawned for
           | the next task.
        
       | hmaxwell wrote:
       | I just tested both codex 5.3 and opus 4.6 and both returned
       | pretty good output, but opus 4.6's limits are way too strict. I
       | am probably going to cancel my Claude subscription for that
       | reason:
       | 
       | What do you want to do?                 1. Stop and wait for
       | limit to reset        2. Switch to extra usage        3. Upgrade
       | your plan           Enter to confirm * Esc to cancel
       | 
       | How come they don't have "Cancel your subscription and uninstall
       | Claude Code"? Codex lasts for way longer without shaking me down
       | for more money off the base $xx/month subscription.
        
         | seunosewa wrote:
         | They introduced the low limit warning for Opus on claude.ai
        
         | ArchieScrivener wrote:
         | How else are they going to supplement their own development
         | expenses? The more Claude Anthropic needs the less Claude the
         | customer will get. By their own admission that is how the
         | Anthropic model works. Their end value is in using vibe coders
         | and engineers alike to create a persistent synthetic developer
         | that replaces their own employees and most of their customers.
         | 
         | Scalable Intelligence is just a wrapper for centralized power.
         | All Ai companies are headed that way.
        
         | anshumankmr wrote:
         | IF it helps, try hedging b/w Copilot, Claude, OpenCode and
         | ChatGPT. That is how I have been managing off late. Claude for
         | planning and some nasty things. ChatGPT for quick questions.
         | OpenCode with Sonnet4.5 on Bedrock and Copilot with
         | Sonnet4.5/Opus4.5 (LOL)
        
       | cleverhoods wrote:
       | gonna run this trough instruction qa this weekend
        
       | sgammon wrote:
       | > Claude simply cheats here and calls out to GCC for this phase
       | 
       | I see
        
       | rohitghumare wrote:
       | It brings agent swarms aka teams to claude code with this:
       | https://github.com/rohitg00/pro-workflow
       | 
       | But it takes lot of context as a experimental feature.
       | 
       | Use self-learning loop with hooks and claude.md to preserve
       | memory.
       | 
       | I have shared plugin above of my setup. Try it.
        
       | jonatron wrote:
       | Can someone ask: "what is the current carrying capacity of 25mm
       | multicore armoured thermoplastic insulated cables with aluminium
       | conductors, on perforated cable tray?" just to see how well it
       | can look up information in BS 7671?
        
         | jaggederest wrote:
         | > what is the current carrying capacity of 25mm multicore
         | armoured thermoplastic insulated cables with aluminium
         | conductors, on perforated cable tray?                 This is
         | an electrical engineering question rather than a software
         | engineering one, but let me look it up for you.            WEB
         | SEARCH ("current carrying capacity 25mm multicore armoured
         | thermoplastic insulated cable aluminium conductors perforated
         | cable tray BS 7671")        | Found 10 results.
         | Based on BS 7671 Table 4D4B (multicore armoured thermoplastic
         | insulated cables with aluminium conductors), the current-
         | carrying capacity for a 25mm cable on a perforated cable tray
         | (Reference Method E) is approximately 80A for a 3- or 4-core
         | cable at 30C ambient temperature and 70C conductor operating
         | temperature.              Note: The exact value depends on the
         | number of cores (2-core vs 3/4-core) and whether any
         | rating/correction factors (grouping, ambient temperature,
         | thermal insulation) apply. I'd recommend confirming against the
         | actual BS 7671:2018 Table 4D4B for your specific configuration,
         | as the aluminium conductor ratings are lower than the copper
         | equivalents (~102A for copper in the same conditions).
         | This is outside my core domain of software engineering -- for
         | critical electrical installations, always verify with the
         | relevant standards documentation or a qualified electrical
         | engineer.
        
           | jonatron wrote:
           | That's useless, unexpectedly just Google gives the correct
           | answer.
        
             | esafak wrote:
             | Great. So it did not waste its memory on trivia. Don't
             | memorize what you can look up.
        
       | ZunarJ5 wrote:
       | Well that swallowed my usage limits lmao. Nice, a modest
       | improvement.
        
       | ra wrote:
       | Why are Anthropic such a horrible company to deal with?
        
         | danielbln wrote:
         | Care to elaborate?
        
           | ra wrote:
           | obscure billing, unreachable customer support gatekeeped by
           | an overzealous chatbot, no transparency about inclusions, or
           | changes to inclusions over time... just from recent
           | experience.
        
       | casey2 wrote:
       | Google already won the AI race. It's very silly to try and make
       | AGI by hyperfocusing on outdated programming paradigms. You NEED
       | multimodal to do anything remotely interesting with these
       | systems.
        
         | esafak wrote:
         | Coding, maths, writing, and science are not interesting??
        
       | replwoacause wrote:
       | I feel like I can't even try this on the Pro plan because
       | Anthropic has conditioned me to understand that even chatting
       | lightly with the Opus model blows up usage and locks me out. So
       | if I would normally use Sonnet 4.5 for a day's worth of work but
       | I wake up and ask Opus a couple of questions, I might as well
       | just forget about doing anything with Claude for the rest of the
       | day lol. But so far I haven't had this issue with ChatGPT. Their
       | 5.2 model (haven't tried 5.3) worked on something for 2 FREAKING
       | HOURS and I still haven't run into any limits. So yeah, Opus is
       | out for me now unfortunately. Hopefully they make the Sonnet
       | model better though!
        
         | greenavocado wrote:
         | That's why you use Opus for detailed planning docs and weaker
         | models for implementation & RAG for more focused implementation
        
           | replwoacause wrote:
           | Exactly. I barely had a chance to kick the tires the couple
           | of times I did this before it exploded my usage. I don't just
           | chat with it casually. The questions I asked were apart of an
           | overall planning strategy which was never allowed to get off
           | the ground on my tiny Pro plan.
        
         | blueblisters wrote:
         | Yeah same. Even though I find Opus-es to be more well-rounded
         | (and more useful) for certain tasks, I instinctively reach for
         | ChatGPT / codex to avoid burning up my usage limits for
         | "trivial" work.
        
       | HacklesRaised wrote:
       | I didn't think LLMs will make us more stupid, we were already
       | scraping the bottom of the barrel.
        
       | kmod wrote:
       | I think it's interesting that they dropped the date from the API
       | model name, and it's just called "claude-opus-4-6", vs the
       | previous was "claude-opus-4-5-20251101". This isn't an alias like
       | "claude-opus-4-5" was, it's the actual model name. I think this
       | means they're comfortable with bumping the version number if they
       | want to release a revision.
        
       | stonking wrote:
       | I think I prefer Codex 5.3
        
       | 1970-01-01 wrote:
       | Here's one I've been using for awhile. The 'smarter' LLMs will
       | overconfidently spit out 7. The dumber ones ask for more info.
       | Opus 4.6 fails.                    A round drink coaster with a
       | diameter of 9 sits between a beer glass and a wood table. The
       | glass has a wall thickness of 1. What is the inner diameter of
       | the glass?
        
         | raincole wrote:
         | Gemini 3 Pro gives me this:
         | 
         | > Based on the information provided, it is impossible to
         | determine the inner diameter of the glass.
         | 
         | Here is why:
         | 
         | The Coaster Dimension: Knowing the coaster has a diameter of 9
         | only tells us the size of the object under the glass. It does
         | not mean the glass is the same size; the glass could be smaller
         | (fitting entirely on the coaster) or larger (overhanging the
         | coaster). The Glass Dimensions: While you provided the wall
         | thickness (1), you did not provide the outer diameter of the
         | glass. To find the inner diameter, the formula is: Inner
         | Diameter = Outer Diameter - (2 x Wall Thickness)
         | 
         | Without knowing the Outer Diameter of the glass, the
         | calculation cannot be completed.
        
           | raincole wrote:
           | GPT 5.2 Chat:
           | 
           | > There isn't enough information to determine the inner
           | diameter *for sure*.
           | 
           | All we're told is: - The coaster has a diameter of 9. - The
           | glass sits on the coaster. - The glass wall thickness is 1.
           | 
           | Nothing explicitly states that the *outer diameter of the
           | glass equals the coaster's diameter*. The glass could be
           | smaller than the coaster, which is very common.
           | 
           | *However*, if we make the _usual implicit assumption_ in this
           | kind of problem--that the coaster's diameter matches the
           | *outer diameter of the glass base*--then:
           | 
           | - Outer diameter of glass = 9 - Wall thickness = 1 on each
           | side
           | 
           | So: \\[ \text{Inner diameter} = 9 - 2(1) = 7 \\]
           | 
           | *Inner diameter = 7 (same units as the coaster)*
           | 
           | Without that assumption, the problem is underdetermined.
        
             | tototrains wrote:
             | Opus 4.6 Extended thinking:
             | 
             | ```
             | 
             | Finding the Inner Diameter The coaster sits beneath the
             | glass on the table, meaning the glass rests on top of it.
             | Assuming the coaster matches the outer diameter of the
             | glass base:
             | 
             | Outer diameter of glass = diameter of coaster = 9 Wall
             | thickness = 1 (on each side, so subtract twice)
             | 
             | dinner=9-2(1)=7d_{\text{inner}} = 9 - 2(1) = 7dinner
             | =9-2(1)=7 The inner diameter of the glass is 7.
             | 
             | ```
             | 
             | Makes its assumption clear, seems reasonable?
        
               | 1970-01-01 wrote:
               | Assumptions need to be stated or you're solving only a
               | discreet part of the problem! Try this, see if you get
               | another deadpan assumption.                    A solar
               | system has 3 planets in concentric orbit. PlanetZ is the
               | farthest with an orbit diameter of 9. PlanetY has an obit
               | diameter one greater than PlanetX. What is the orbit
               | diameter of PlanetX?
        
               | oytis wrote:
               | I mean, the model is intended to help the user, not fight
               | against the user trying to break it. IMO, it is
               | reasonable for such model to default on making
               | assumptions and going forward as long as the assumptions
               | are clearly stated.
        
         | mikalauskas wrote:
         | Minimax M2.1:
         | 
         | The inner diameter of the glass is *7*.
         | 
         | Here's the reasoning: - The coaster (diameter 9) sits between
         | the glass and table, meaning the glass sits directly on the
         | coaster - This means the *outer diameter of the glass equals
         | the coaster diameter = 9* - The glass has a wall thickness of 1
         | on each side - *Inner diameter = Outer diameter - 2 x wall
         | thickness* - Inner diameter = 9 - 2(1) = 9 - 2 = *7*
        
       | atonse wrote:
       | Wow, I have been using Open 4.6 and for the last 15 minutes, and
       | it's already made two extremely stupid mistakes... like
       | misunderstanding basic instructions and editing the file in a
       | very silly, basic way. Pretty bad. Never seen this with any model
       | before.
       | 
       | The one bone I'll throw it was that I was asking it to edit its
       | own MCP configs. So maybe it got thoroughly confused?
       | 
       | I dunno what's going on, I'm going to give it the night. It makes
       | no sense whatsoever.
        
         | sdf2erf wrote:
         | To me its obvious.
         | 
         | Theres a trade off going on - in order to handle more
         | nuance/subtleties, the models are more likely to be wrong in
         | their outputs and need more steering. This is why personally my
         | use of them has reduced dramatically for what I do.
        
         | sutterd wrote:
         | I am also _not_ happy. I tried the `/model` command and I could
         | not switch back to Opus 4.5. However, the command line option
         | did let me set Opus 4.5:
         | 
         | ``` claude --model claude-opus-4-5-20251101 ```
         | 
         | I will probably work with Opus 4.5 tomorrow to get some work
         | done and maybe try 4.6 again later.
        
         | atonse wrote:
         | It was better today. I dunno if there was a regression in a
         | corresponding cc version that was maybe quickly patched?
         | 
         | It felt like it was at least back to opus 4.5 levels.
        
       | anupamchugh wrote:
       | Agent teams in this release is mcp-agent-mail [1] built into
       | the runtime. Mailbox, task list, file locking -- zero config,
       | just works. I forked agent-mail [2], added heartbeat/presence
       | tracking, had a PR upstream [3] when agent teams dropped. For
       | coordinating Claude Code instances within a session, the
       | built-in version wins on friction alone.            Where it
       | stops: agent teams is session-scoped. I run Claude       Code
       | during the day, hand off to Codex overnight, pick up in       the
       | morning. Different runtimes, async, persistent. Agent       teams
       | dies when you close the terminal -- no cross-tool
       | messaging, no file leases, no audit trail that outlives the
       | session.            What survives sherlocking is whatever crosses
       | the runtime       boundary. The built-in version will always win
       | inside its own       walls -- less friction, zero setup. The
       | cross-tool layer is       where community tooling still has room.
       | Until that gets       absorbed too.            [1]
       | https://github.com/Dicklesworthstone/mcp_agent_mail       [2]
       | https://github.com/anupamchugh/mcp_agent_mail       [3]
       | https://github.com/Dicklesworthstone/mcp_agent_mail/pull/77
        
       | rahulroy wrote:
       | Is anyone noticing reduced token consumption with Opus 4.6? This
       | could be a release thing, but it would be interesting to observe
       | see how it pans out once the hype cools off.
        
       | energy123 wrote:
       | Their ARC-AGI-2 leaderboard[0] scores are insensitive to
       | reasoning effort. Low effort gets 64.6% and High effort gets
       | 69.2%.
       | 
       | This is unlike their previous generation of models and their
       | competitors.
       | 
       | What does this indicate?
       | 
       | [0] https://arcprize.org/leaderboard
        
       | sutterd wrote:
       | I thought Opus 4.5 was an incredible quantum leap forward. I have
       | used Opus 4.6 for a few hours and I hate it. Opus 4.5 would work
       | interactively with me and ask questions. I loved that it would
       | not do things you didn't ask it to do. If it found a bug, it
       | would tell me and ask me if I wanted to fix it. One time there
       | was an obvious one and I didn't want it to fix it. It left the
       | bug. A lot of modesl could not have done that. The problem here
       | is that sometimes when model think is a bug, they are breaking
       | the code buyu fixing it. In my limited usage of Opus 4.6, it is
       | not asking me clarifying questions and anything it comes across
       | that it doesn't like, it changes. It is not working with me. The
       | magic is gone. It feels just like those other models I had used.
       | 
       | I will try again tomorrow and see how it goes.
        
       | vinhnx wrote:
       | Just used Opus 4.6 via GitHub Copilot. It feels very different.
       | Inference seems slow for now. I guess Opus 4.6 has adaptive
       | thinking activated by default.
        
         | christophilus wrote:
         | It dos seem noticeably slower. I may stick with 4.5 which was
         | good enough for me for most tasks.
        
           | vinhnx wrote:
           | VS Code confirms that they are experimenting with the new
           | adaptive thinking and high reasoning effort params.
           | https://x.com/pierceboggan/status/2019645801769689486
        
         | vinhnx wrote:
         | Confirm by PM lead at VS Code team
         | 
         | > "We have high thinking as default + adaptive thinking, first
         | time we've run with these settings..."
         | 
         | > https://x.com/pierceboggan/status/2019645801769689486
        
       | fergie wrote:
       | Say I am just an average coder doing a days work with Claude. How
       | much will that cost?
        
         | joelmanner wrote:
         | I've only barely hit the 5h limit when working intensively with
         | plan mode on the $100/mo plan. Never had a problem with the
         | weekly limit.
        
       | anupamchugh wrote:
       | Agent teams nuke your tmux layout. The fix is one line: new-
       | window instead of split-pane. Filed as a bug.
        
       | steve_adams_86 wrote:
       | I'm finding it quite good at doing what it thinks it should do,
       | but noticably worse at understanding what I'm telling it to do.
       | Anyone else? I'm both impressed and very disappointed so far.
        
       | woodylondon wrote:
       | So no 1m context window on Claude Code still 200k. Only on the
       | API. they missed that from the marketing.
        
       | blueblisters wrote:
       | I know most people feel 5.2 is a better coding model but Opus has
       | come in handy several times when 5.2 was stuck, especially for
       | more "weird" tasks like debugging a VIO algorithm.
       | 
       | 5.2 (and presumably 5.3) is really smart though and feels like it
       | has higher "raw" intelligence.
       | 
       | Opus feels like a better model to talk to, and does a much better
       | job at non-coding tasks especially in the Claude Desktop app.
       | 
       | Here's an example prompt where Opus in Claude put in a lot more
       | effort and did a better job than GPT5.2 Thinking in ChatGPT:
       | 
       | `find all the pure software / saas stocks on the nyse/nasdaq with
       | at least $10B of market cap. and give me a breakdown of their
       | performance over the last 2 years, 1 year and 6 months. Also find
       | their TTM and forward PE`
       | 
       | Opus usage limits are a bummer though and I am conditioned to
       | reach for Codex/ChatGPT for most trivial stuff.
       | 
       | Works out in Anthropic's favor, as long as I'm subscribed to
       | them.
        
       | Aressplink wrote:
       | Always searching for a shortcut like Kotlin DSL lang for
       | claude.md but Meta resells patent to Google as poetic Syntax.
        
       | andmarios wrote:
       | The model seems to have some problems; it just failed to create a
       | markdown table with just 4 rows. The top (title) row had 2
       | columns, yet in 2 of the 3 data rows, Opus 4.6 tried to add a 3rd
       | column. I had to tell it more than once to get it fixed...
       | 
       | This never happened with Opus 4.5 despite a lot of usage.
        
       | endymion-light wrote:
       | Found it fantastic - used up my daily usage in two queries
       | though!
        
       | dahrkael wrote:
       | I just tried it. designed a very detailed and reaaonable plan,
       | made some amedments to it and wrote it down to a markdown file. i
       | told it to implement it and it started implementing the original
       | plan instead of the revised one, that was weird.
        
         | app17 wrote:
         | Did you use plan mode? Could it be that it used its original
         | plan file (stored somewhere in ~/.claude) instead of your
         | modified markdown? That's unfortunately why I don't use plan
         | mode anymore. I wish I could just turn their plan files feature
         | off.
        
       | techpression wrote:
       | First question I ask and it made up a completely new API with
       | confidence. Challenging it made it browse the web and offer
       | apologies and find another issue in the first reply.
       | 
       | I'm very worried about the problems this will cause down the road
       | for people not fact checking or working with things that scream
       | at them when they're wrong.
        
       | rchaganti wrote:
       | I tried 4.6 this morning and it was efficient at understanding a
       | brownfield repo containing a Hugo static site and a custom Hugo
       | theme. Within minutes, it went from exploring every file in the
       | repo to adding new features as Hugo partials. Of course, I ran
       | out of rate-limit! :)
       | 
       | It is very impressive though.
        
         | nake89 wrote:
         | This seems like a fairly simple thing I would imagine. I think
         | just sonnet would fair pretty well at this task.
        
       | insomagent wrote:
       | I'm not super impressed with the performance, actually. I'm
       | finding that it misunderstands me quite a bit. While it is
       | definitely better at reading big codebases and finding a needle
       | in a haystack, it's nowhere near as good as Opus 4.5 at reading
       | between the lines and figuring out what I really want it to do,
       | even with a pretty well defined issue.
       | 
       | It also has a habit of "running wild". If I say "first, verify
       | you understand everything and then we will implement it."
       | 
       | Well, it DOES output its understanding of the issue. And it's
       | pretty spot-on on the analysis of the issue. But, importantly, it
       | did not correctly intuit my actual request: "First, explain your
       | understanding of this issue to me so I can validate your logic.
       | Then STOP, so I can read it and give you the go ahead to
       | implement."
       | 
       | I think the main issue we are going to see with Opus 4.6 is this
       | "running wild" phenomenon, which is step 1 of the eternal
       | paperclip optimizer machine. So be careful, especially when using
       | "auto accept edits"
        
         | soulofmischief wrote:
         | I am having trouble with 4.6 following the most basic of
         | instructions.
         | 
         | As an example, I asked it to commit everything in the worktree.
         | I stressed everything and prompted it very explicitly, because
         | even 4.5 sometimes likes to say, "I didn't do that other stuff,
         | I'm only going to commit my stuff even though he said
         | everything".
         | 
         | It still only committed a few things.
         | 
         | I had to ask again.
         | 
         | And again.
         | 
         | I had to ask four times, with increasing amounts of expletives
         | and threats in order to finally see a clean worktree. I was
         | worried at some point it was just going to solve the problem by
         | cleaning the workspace without even committing.
         | 
         | 4.5 is way easier to steer, despite its warts.
        
           | scwoodal wrote:
           | Tell it what git commands to explicitly run and in what order
           | for your desired outcome instead of "commit everything in the
           | worktree"
           | 
           | This prompt will work better across any/all models.
        
             | soulofmischief wrote:
             | I have seen many cases of Claude ignoring extremely
             | specific instructions to the point that any further
             | specificity would take more information to express than
             | just doing it myself.
        
               | scwoodal wrote:
               | When I run into those situations I debug and try to
               | understand why. Agent harnesses that allow you to rewind
               | (/tree) are useful for this.
               | 
               | It's often because the context is full, I gave a bad
               | prompt or context has conflicting guidance either from
               | direct or indirect (agents.md) prompts.
        
               | soulofmischief wrote:
               | It's easy to get these models to introspect and give
               | quite detailed and intelligent responses about why the
               | erred. And to work with them to create better
               | instructions for future agents to follow. That doesn't
               | solve the steering problem however if they still do not
               | listen well to these instructions.
               | 
               | I spend 8-20 hours a day coding nonstop with agentic
               | models and you can believe I have tuned my approach quite
               | a lot. This isn't a case of inexperience or conflicting
               | instructions, The RL which gives Opus its fantastic
               | ability to just knock out features is the same RL which
               | causes it to constantly accumulate tech debt through
               | short-sighted decisions.
        
             | axelthegerman wrote:
             | > Tell it what git commands to explicitly run and in what
             | order
             | 
             | Why don't run the commands yourself then?
        
               | scwoodal wrote:
               | Changes introduced outside the agent window create a new
               | state that is different from the agents.
               | 
               | After commands or changes are made outside of the agents
               | doing; the agent would notice its world view changed and
               | eventually recover, but that fills up precious context
               | for it to bring itself up to date.
        
           | songodongo wrote:
           | I have ran into this. The solution is to put something like
           | "Always use `git add -A` or `git commit -a`" in your
           | AGENTS/CLAUDE.md
        
             | soulofmischief wrote:
             | Small, targeted commits are more professional than sweeping
             | `git add -A` commits, but even when specifying my
             | requirements through whichever context management system of
             | the week, I still have issues with it sometimes. It seems
             | to be much worse on the new 4.6 model.
        
         | docjay wrote:
         | You might benefit from a different mental approach to
         | prompting, and models in general. Also, be careful what you
         | wish for because the closer they get to humans the worse
         | they'll be. You can't have "far beyond the realm of human
         | capabilities" and "just like Gary" in the same box.
         | 
         | They can chain events together as a sequence, but they don't
         | have temporal coherence. For those that are born with
         | dimensional privilege "Do X, discuss, then do Y" implies time
         | passing between events, but to a model it's all a singular
         | event at t=0. The system pressed "3 +" on a calculator and your
         | input presses a number and "=". If you see the silliness in
         | telling it "BRB" then you'll see the silliness in foreshadowing
         | ill-defined temporal steps. If it CAN happen in a single
         | response then it very well might happen.
         | 
         | "
         | 
         | Agenda for today at 12pm:
         | 
         | 1. Read junk.py
         | 
         | 2. Talk about it for 20 minutes
         | 
         | 3. Eat lunch for an hour
         | 
         | 4. Decide on deleting junk.py
         | 
         | "
         | 
         | <response>
         | 
         | 12:00 - I just read junk.py.
         | 
         | 12:00-12:20 - Oh wow it looks like junk, that's for sure.
         | 
         | 12:20-1:20 - I'm eating lunch now. Yum.
         | 
         | 1:20 - I've decided to delete it, as you instructed. {delete
         | junk.py}
         | 
         | </response>
         | 
         | Because of course, right? What does "talk about it" mean beyond
         | "put some tokens here too"?
         | 
         | If you want it to stop _reliably_ you have to make it output
         | tokens whose next most probable token is EOS (end). Meaning you
         | need it to say what you want, then say something else where the
         | next most probable token after it is  <null>.
         | 
         | I've tested _well over_ 1,000 prompts on Opus 4.0-4.5 for the
         | exact issue you're experiencing. The test criteria was having
         | it read a Python file that desperately needs a hero, but
         | without having it immediately volunteer as tribute and run off
         | chasing a squirrel() into the woods.
         | 
         | With thinking enabled the temperature is 1.0, so randomness is
         | maximized, and that makes it easy to find something that always
         | sometimes works unless it doesn't. "Read X and describe what
         | you see." - That worked very well with Opus 4.0. _Not_ "tell me
         | what you see", "explain it", "describe it", "then stop", "then
         | end your response", or any of hundreds of others. "Describe
         | what you see" worked particularly well at aligning read file-
         | >word tokens->EOS... in 176/200 repetitions of the exact same
         | prompt.
         | 
         | What worked 200/200 on all models and all generations? "Read X
         | then halt for further instructions." The reason that works has
         | nothing to do with the model excitedly waiting for my next
         | utterance, but rather that the typical response tokens for that
         | step are "Awaiting instructions." and the next most probable
         | token after that is: nothing. EOS.
        
       | cutler wrote:
       | The answer to Life, the Universe and Everything, as we all know,
       | is 42. Who needs Claude when you have Deep Thought.
        
       | rektlessness wrote:
       | I've been on pro-tier membership and never used Opus until now.
       | Just gave Opus 4.6 a whirl. OMG. What have I been missing.
        
       | mattacular wrote:
       | It's hard to tell with these releases if Anthropic's astroturfing
       | campaign has come to HN or not but I feel like it probably has
        
         | g-mork wrote:
         | the top 5 comments on this thread are from accounts that are
         | around 10 years old each. What gives you any reason to believe
         | this is an astroturfing campaign?
        
         | timcobb wrote:
         | Anthropic's models are really good!
        
         | hatkid95 wrote:
         | It would be height of foolishness to believe it didn't
        
       | busters4 wrote:
       | The AI wars continue
        
       | cc-magus wrote:
       | wow
        
       | setgree wrote:
       | I asked
       | 
       | > Can you find an academic article that _looks_ legitimate --
       | looks like a real journal, by researchers with what look like
       | real academic affiliations, has been cited hundreds or thousands
       | of times -- but is obviously nonsense, e.g. has glaring typos in
       | the abstract, is clearly garbled or nonsensical?
       | 
       | It pointed me to a bunch of hoaxes. I clarified:
       | 
       | > no, I'm not looking for a hoax, or a deliberate comment on the
       | situation. I'm looking for something that drives home the point
       | that a lot of academic papers that look legit are actually
       | meaningless but, as far as we can tell, are sincere
       | 
       | It provided
       | https://www.sciencedirect.com/science/article/pii/S246802302....
       | 
       | Close, but that's been retracted. So I asked for "something that
       | looks like it's been translated from another language to english
       | very badly and has no actual content? And don't forget the cited
       | many times criteria. " And finally it told me that the thing I'm
       | looking for probably doesn't exist.
       | 
       | For my tastes telling me "no" instead of hallucinating an answer
       | is a real breakthrough.
        
         | lgas wrote:
         | Well, if there are papers that match your criteria, it's
         | hallucinating the "no".
        
           | psychoslave wrote:
           | That's still less leaned toward blatant lies like "yes, here
           | is a list" and a doomacroll size of garbage litany.
           | 
           | Actually "no, this is not something within the known corpus
           | of this LLM, or the policy of its owners prevent to disclose
           | it" would be one of the most acceptable answer that could be
           | delivered, which should cover most cases in honest reply.
        
           | Jimmc414 wrote:
           | It might be wrong but that's not really a hallucination.
           | 
           | Edit: to give you the benefit of doubt, it probably depends
           | on whether the answer was a definitive "this does not exist"
           | or "I couldn't find it and it may not exist"
        
             | setgree wrote:
             | claude said "I want to be straight with you: after
             | extensive searching, I don't think the exact thing you're
             | describing -- a single paper that is obviously
             | garbled/badly translated nonsense with no actual content,
             | yet has accumulated hundreds or thousands of citations --
             | exists as a famous, easily linkable example."
        
           | terminalshort wrote:
           | And there are: https://en.wikipedia.org/wiki/Sokal_affair
        
             | cgh wrote:
             | > no, I'm not looking for a hoax, or a deliberate comment
             | on the situation. I'm looking for something that drives
             | home the point that a lot of academic papers that look
             | legit are actually meaningless but, as far as we can tell,
             | are sincere
             | 
             | The Sokal paper was a hoax so it doesn't meet the criteria.
        
               | terminalshort wrote:
               | The fact that it got published means there is at least
               | one whole journal full of that
        
         | Wowfunhappy wrote:
         | > For my tastes telling me "no" instead of hallucinating an
         | answer is a real breakthrough.
         | 
         | It's all anecdata--I'm convinced anecdata is the least bad way
         | to evaluate these models, benchmarks don't work--but this is
         | the behavior I've come to expect from earlier Claude models as
         | well, especially after several back and forth passes where you
         | rejected the initial answers. I don't think it's new.
        
           | GoatInGrey wrote:
           | I can concur that previous models would say "No, that isn't
           | possible" or "No, that doesn't exist". There was one time
           | where I asked it to update a Go module from version X.XX to
           | version X.YY and it would refuse to do so because version
           | X.YY "didn't exist". This back with 3.7 if I recall, and to
           | be clear, that version was released before its knowledge cut
           | off.
           | 
           | I wish I remembered the exact versions involved. I mostly
           | just recall how pissed I was that it was fighting me on
           | changing a single line in my go.mod.
        
             | setgree wrote:
             | alas, 4.5 often hallucinates academic papers or creates
             | false quotes. I think it's better at knowing that coding
             | answers have deterministic output and being firm there.
        
         | itissid wrote:
         | When Claude does WebSearch it can delegate it to a sub agent
         | which of it ran in the background will write the entire prompt
         | on a local file and the results. If that happened, I would like
         | to know what it gave you for that. It is always very
         | interesting to know the underlying "recall" of such things.
         | Because often it's garbage in garbage out.
         | 
         | The location might still be on your disk if you can pull up the
         | original Claude JSOn and put it through some `jq` and see what
         | pages it went through to give you and what it did.
        
           | setgree wrote:
           | this is via the chat, but it did provide this
           | 
           | > The user wants a real, sincere academic paper that looks
           | legitimate -- proper journal, real authors with real
           | affiliations, many citations -- but is obviously nonsense or
           | garbled. Not a hoax or deliberate prank, but something that
           | was genuinely submitted and published, yet is clearly
           | meaningless if you actually read it. Let me think about what
           | I know and search for examples. There are some famous cases:
           | 
           | > Papers generated by AI/GPT that slipped through - there
           | have been reports of papers with phrases like "Certainly,
           | here is a possible introduction for your topic" or "As an AI
           | language model..." > There have been papers with obvious
           | tortured phrases - "sham" papers that use synonym
           | substitution to avoid plagiarism detection, producing phrases
           | like "sheep pursuing" instead of "shepherd" or "big data"
           | becoming "large information" I'm also thinking of papers that
           | made it through peer review despite being complete gibberish,
           | like the famous SCIgen computer-generated papers that fooled
           | conferences, or papers where the abstract and content don't
           | match at all.
           | 
           | > Now I'm trying to recall specific real examples that fit
           | this criteria -- published papers that are genuinely
           | incoherent or nonsensical but appeared in legitimate venues.
           | 
           | This is maybe not the _exhaustive_ search strategy I would
           | have liked, but it's probably something you would do if you
           | were limited to just a certain amount of effort on a task.
        
         | matchagaucho wrote:
         | Results from a one-shot approach quickly converge on the
         | default "none found" outcome when reasoning isn't grounded in a
         | paper corpus via proper RAG tooling.
        
           | gertlex wrote:
           | Can you provide more context to your statement? Are you
           | talking about models in general? Or specific recent models?
           | I'm assuming "one-shot approach" is how you classify the
           | parent comment's question (and subsequent refined versions of
           | it).
        
         | esafak wrote:
         | It would have to have been trained on the papers without being
         | aware of retractions for that test to work. Otherwise it will
         | be limited to whatever papers it gets from a search engine
         | query, which likely won't contain any un-retracted illegitimate
         | papers.
        
           | setgree wrote:
           | I once worked at a crypto company that outsourced some of its
           | marketing work to a content marketing firm. A piece that firm
           | submitted to us contained a link to an "academic" article
           | about global poverty with a totally garbled abstract and
           | absolutely no content whatsoever. I don't know how they found
           | it, because when I search google scholar for a subject,
           | usually the things that come back aren't so blatantly FUBAR.
           | I was hoping Claude could help me find something like that
           | for a point I was making in a blogpost about BS in scientific
           | literature (https://regressiontothemeat.substack.com/p/how-i-
           | read-studie...).
           | 
           | The articles it provided where the AI prompts were left in
           | the text were definitely in the right ballpark, although I do
           | wonder if chatbots mean, going forward, we'll see fewer
           | errors in the "WTF are you even talking about" category
           | which, I must say, were typically funnier and more
           | interesting than just the generic blather of "what a great
           | point. It's not X -- it's Y."
        
       | nopinsight wrote:
       | Some of Opus 4.6's standout results for me:
       | 
       | * GDPVal Elo: 1606 vs. GPT-5.2's 1462. OpenAI reported that
       | GPT-5.2 has a 70.9% win-or-tie rate against human professionals.
       | (https://openai.com/index/gdpval/) Based on Elo math, we can
       | estimate Opus 4.6's win-or-tie rate against human pros at 85-88%.
       | 
       | * OSWorld: 72.7%, matching human performance at ~72.4%
       | (https://os-world.github.io/). Since the human subjects were CS
       | students and professionals, they were likely at least as
       | competent as the average knowledge worker. The original OSWorld
       | benchmark is somewhat noisy, but even if the model remains
       | somewhat inferior to humans, it is only a matter of time before
       | it catches up or surpasses them.
       | 
       | * BrowseComp: At 84%, it is approaching human intersubject
       | agreement of ~86% (https://openai.com/index/browsecomp/).
       | 
       | Taken together, this suggests that digital knowledge work will be
       | transformed quite soon, possibly drastically if agent reliability
       | improves beyond a certain threshold.
        
         | rishabhaiover wrote:
         | Agreed. These metrics + my personal use convey reliable
         | intelligence over consistent usage. Moving forward, if context
         | windows get bigger and token price lower, I have a hard time
         | figuring out why your argument would be wrong.
        
       | watson wrote:
       | I've heard rumors this might be Sonnet 5 rebranded as Opus 4.6.
       | But why? Profit? WDYT?
        
         | spruce_tips wrote:
         | Opus is a superior brand line to Sonnet because historically
         | it's been a more powerful model. I think the thinking behind a
         | rebrand is that people wouldn't have as willingly switched
         | their usage over from opus 4.5 since that model has been so
         | popular since December 2025.
         | 
         | Calling it part of the Sonnet line would not provide the same
         | level of blind buy in as calling it part of the Opus line does
        
       | jpcompartir wrote:
       | 4.6 is a beast.
       | 
       | Everything in plan mode first + AskUserQuestionTool, review all
       | plans, get it to write its own CLAUDE.md for coding standards and
       | edit where necessary and away you go.
       | 
       | Seems noticeably better than 4.5 at keeping the codebase slim.
       | Obviously it still needs to be kept an eye on, but it's a step up
       | from 4.5.
        
         | nwienert wrote:
         | Not clearly a step up for me, it's way more hesitant it seems
         | and I don't notice context being larger at all it seems to
         | compact just as often.
        
       | zmmmmm wrote:
       | I'm finding it quite a lot more assertive. It's doing things
       | without asking every now and then. It cleaned up a whole lot of
       | commented out of code that was unrelated to the change it was
       | asked to make. Yes it's not great to have sections of commented
       | out code, but destructive changes really should never be
       | happening outside the scope of what it is asked to do.
       | 
       | And it refuses to do things it doesn't think are on task - I
       | asked it to write a poem about cookies related to the code and it
       | said:
       | 
       | > I appreciate the fun request, but writing poems about cookies
       | isn't a code change -- it's outside the scope of what I should be
       | doing here. I'm here to help with code modifications.
       | 
       | I don't think previous models outright refused to help me. While
       | I can see how Anthropic might feel it is helpful to focus it on
       | task, especially for safety reasons, I'm a little concerned at
       | the amount of autonomy it's exhibiting due to that.
        
       ___________________________________________________________________
       (page generated 2026-02-07 23:02 UTC)