[HN Gopher] 2025: The Year in LLMs
       ___________________________________________________________________
        
       2025: The Year in LLMs
        
       Author : simonw
       Score  : 838 points
       Date   : 2025-12-31 23:54 UTC (23 hours ago)
        
 (HTM) web link (simonwillison.net)
 (TXT) w3m dump (simonwillison.net)
        
       | AndyNemmity wrote:
       | These are excellent every year, thank you for all the wonderful
       | work you do.
        
         | tkgally wrote:
         | Same here. Simon is one of the main reasons I've been able to
         | (sort of) keep up with developments in AI.
         | 
         | I look forward to learning from his blog posts and HN comments
         | in the year ahead, too.
        
           | password4321 wrote:
           | Don't forget you can pay Simon to keep up with less!
           | 
           | > _At the end of every month I send out a much shorter
           | newsletter to anyone who sponsors me for $10 or more on
           | GitHub_
           | 
           | https://simonwillison.net/about/#monthly
        
       | waldrews wrote:
       | Remember, back in the day, when a year of progress was like, oh,
       | they voted to add some syntactic sugar to Java...
        
         | throwup238 wrote:
         | _> they voted to add some syntactic sugar to Java..._
         | 
         | I remember when we just wanted to rewrite everything in Rust.
         | 
         | Those were the simpler times, when crypto bros seemed like the
         | worst venture capitalism could conjure.
        
           | OGEnthusiast wrote:
           | Crypto bros in hindsight were so much less dangerous than AI
           | bros. At least they weren't trying to construct data centers
           | in rural America or prop up artificial stocks like $NVDA.
        
             | quaintpartridge wrote:
             | They were, just not as many.
             | https://www.wired.com/story/the-worlds-biggest-bitcoin-
             | mine-...
        
             | SauntSolaire wrote:
             | Instead they were building crypto mining warehouses in
             | rural America and propping up artificial currencies like
             | BTC.
        
               | ryandrake wrote:
               | Crazy how the two most hyped and funded technologies of
               | the decade were: energy wasting fake money for criminals
               | and energy wasting plagiarism machines.
        
             | zahlman wrote:
             | Speaking of which, we never found out the details (strike
             | price/expiration) of Michael Burry's puts, did we? It seems
             | he could have made bank if he'd waited one more month...
        
               | kamranjon wrote:
               | I think they expire in March 2026 if the NVIDIA stock
               | drops to $140 a share? Something close to that I think.
        
             | mgfist wrote:
             | It's funny how people complain about the rust belt dying
             | and factories leaving rural communities and so on, then
             | when someone wants to build something that can provide jobs
             | and tax revenue, everyone complains.
        
               | lostlogin wrote:
               | I've heard about the risk of AI leading to job losses and
               | wealth concentration.
               | 
               | I haven't heard about new businesses, job creation and
               | growth in former industrial towns. What have I missed?
        
               | jakeydus wrote:
               | How many people are employed at the average data center?
               | A few dozen? Versus a steel mill, that's nothing. A
               | chicken plant in Nebraska closed down this last month.
               | 3200 people lost their jobs. You think Meta will fill it
               | with GPUs and the whole town will have jobs again?
        
               | scotty79 wrote:
               | Many more are employed while building it. And they will
               | never stop building. It's modern version of rail. But
               | instead of distances it will cover the area.
        
               | uxcolumbo wrote:
               | Will local folks get those jobs to build the data center?
               | 
               | And if so, what happens to those builders once the data
               | center is built?
        
               | techpression wrote:
               | As if any taxes will be paid to the areas affected, and
               | add to that the billions in taxes used to subsidize
               | everything before a single cent is a net positive.
        
         | nrhrjrjrjtntbt wrote:
         | More like 6 different new nosql databases and js frameworks.
        
           | dotancohen wrote:
           | A Wordpress zero day and Linux not on the desktop. Netcraft
           | confirms it.
        
         | crystal_revenge wrote:
         | That must have been a _long_ time back. Having lived through
         | the time when web pages were served through CGI and mobile
         | phones only existed in movies, when SVMs where the _new
         | hotness_ in ML and people would write about how weird NNs were,
         | I feel like I 've seen _a lot_ more concrete progress in the
         | last few decades than this year.
         | 
         | This year honestly feels quite stagnant. LLMs are _literally_
         | technology that can only reproduce the past. They 're cool, but
         | they were _way_ cooler 4 years ago. We 've taken big ideas like
         | "agents" and "reinforcement learning" and basically stripped
         | them of all meaning in order to claim progress.
         | 
         | I mean, do you remember Geoffrey Hinton's RBM talk at Google in
         | 2010? [0] That was _absolutely insane_ for anyone keeping up
         | with that field. By the mid-twenty teens RBMs were _already_
         | outdated. I remember when everyone was implementing flavors of
         | RNNs and LSTMs. Karpathy 's character 2015 RNN project was
         | _insane_ [1].
         | 
         | This comment makes me wonder if part of the hype around LLMs is
         | just that a lot of software people simply weren't paying
         | attention to the absolutely mind-blowing progress we've seen in
         | this field for the last 20 years. But even ignoring ML, the
         | world's of web development and mobile application development
         | have gone through incredible progress over the last decade and
         | a half. I remember a time when JavaScript books would have a
         | section warning that you should _never_ use JS for anything
         | critical to the application. Then there 's the work in theorem
         | provers over the last decade... If you remember when syntactic
         | sugar was progress, either you remember _way_ further back than
         | I do, or you weren 't paying attention to what was happening in
         | the larger computing world.
         | 
         | 0. https://www.youtube.com/watch?v=VdIURAu1-aU
         | 
         | 1. https://karpathy.github.io/2015/05/21/rnn-effectiveness/
        
           | handoflixue wrote:
           | > LLMs are literally technology that can only reproduce the
           | past.
           | 
           | Funny, I've used them to create my own personalized text
           | editor, perfectly tailored to what I actually want. I'm
           | pretty sure that didn't exist before.
           | 
           | It's wild to me how many people who talk about LLM apparently
           | haven't learned how to use them for even very basic tasks
           | like this! No wonder you think they're not that powerful, if
           | you don't even know basic stuff like this. You really owe it
           | to yourself to try them out.
        
             | crystal_revenge wrote:
             | > You really owe it to yourself to try them out.
             | 
             | I've worked at multiple AI startups in lead AI Engineering
             | roles, both working on deploying user facing LLM products
             | and working on the research end of LLMs. I've done
             | collaborative projects and demos with a pretty wide range
             | of big names in this space (but don't want to doxx myself
             | _too_ aggressively), have had my LLM work cited on HN
             | multiple times, have LLM based github projects with
             | hundreds of stars, appeared on a few podcasts talking about
             | AI etc.
             | 
             | This gets to the point I was making. I'm starting to
             | realize that part of the disconnect between my opinions on
             | the state of the field and others is that many people
             | _haven 't_ really been paying much attention.
             | 
             | I can see if recent LLMs are your first intro to the state
             | of the field, it must feel incredible.
        
               | CamperBob2 wrote:
               | That's all very impressive, to be sure. But are you sure
               | you're getting the point? As of 2025, LLMs are now very
               | good at writing new code, creating new imagery, and
               | writing original text. They continue to improve at a
               | remarkable rate. They are helping their users create
               | things that didn't exist before. Additionally, they are
               | now very good at searching and utilizing web resources
               | that didn't exist at training time.
               | 
               | So it is absurdly incorrect to say "they can only
               | reproduce the past." Only someone who _hasn 't_ been
               | paying attention (as you put it) would say such a thing.
        
               | crystal_revenge wrote:
               | I think the confusion is people's misunderstanding of
               | what 'new code' and 'new imagery' mean. Yes, LLMs can
               | generate a specific CRUD webapp that hasn't existed
               | before but only based on interpolating between the
               | history of existing CRUD webapps. I mean traditional
               | Markov Chains can also produce 'new' text in the sense
               | that "this _exact_ text " hasn't been seen before, but
               | nobody would argue that traditional Markov Chains aren't
               | constrained by "only producing the past".
               | 
               | This is even more clear in the case of diffusion models
               | (which I personally love using, and have spent a lot of
               | time researching). All of the "new" images created by
               | even the most advanced diffusion models are fundamentally
               | remixing past information. This is really obvious to
               | anyone who has played around with these extensively
               | because they really can't produce truly novel concepts.
               | New concepts can be added by things like fine-tuning or
               | use of LoRAs, but _fundamentally_ you 're still just
               | remixing the past.
               | 
               | LLMs are always doing some form of interpolation between
               | different points in the past. Yes they can create a "new"
               | SQL query, but it's just remixing from the SQL queries
               | that have existed prior. This still makes them _very_
               | useful because _a lot_ of engineering work, including
               | writing a custom text editor, involve remixing existing
               | engineering work. If you could have stack-overflowed your
               | way to an answer in the past, an LLM will be much
               | superior. In fact, the phrase  "CRUD" largely exists to
               | point out that most webapps are fundamentally _the same_.
               | 
               | A great example of this limitation in practice is the
               | work that Terry Tao is doing with LLMs. One of the
               | largest challenges in automated theorem proving is
               | translating human proofs into the language of a theorem
               | prover (often Lean these days). The challenge is that
               | there is not very much Lean code currently available to
               | LLMs (especially with the necessary context of the
               | accompanying NL proof), so they struggle to correctly
               | translate. Most of the research in this area is around
               | improving LLM's representation of the mapping from human
               | proofs to Lean proofs (btw, I personally feel like LLMs
               | _do_ have a reasonably good chance of providing major
               | improvements in the space of formal theorem proving, in
               | conjunction with languages like Lean, because the
               | translation process is the biggest blocker to progress).
               | 
               | When you say:
               | 
               | > So it is absurdly incorrect to say "they can only
               | reproduce the past."
               | 
               | It's pretty clear you don't have a solid background in
               | generative models, because this is fundamentally what
               | they do: model an existing probability distribution and
               | draw samples from that. LLMs are doing this for a
               | _massive_ amount of human text, which is why they do
               | produce some impressive and _useful_ results, but this is
               | also a fundamental limitation.
               | 
               | But a world where we used LLMs for the majority of work,
               | would be a world with no fundamental breakthroughs. If
               | you've read _The Three Body Problem_ , it's very much
               | like living in the world where scientific progress is
               | impeded by sophons. In that world there is still some
               | progress (especially with abundant energy), but it
               | remains fundamentally and deeply limited.
        
               | throwaway7783 wrote:
               | Would you say that LLMs can discover patterns hitherto
               | unknown? It would still be generating from the past, but
               | patterns/connections not made before.
        
               | PeterHolzwarth wrote:
               | Just an innocent bystander here, so forgive me, but I
               | think the flack you are getting is because you appear to
               | be responding to claims that these tools will reinvent
               | everything and introduce a new halcyon age of creation -
               | when, at least on hacker news, and definitely in this
               | thread, no one is really making such claims.
               | 
               | Put another way, and I hate to throw in the now over-used
               | phrase, but I feel you may be responding to a strawman
               | that doesn't much appear in the article or the discussion
               | here: "Because these tools don't achieve a god-like level
               | of novel perfection that no one is really promising here,
               | I dismiss all this sorta crap."
               | 
               | Especially when I think you are also admitting that the
               | technology is a fairly useful tool on its own merits - a
               | stance which I believe represents the bulk of the
               | feelings that supporters of the tech here on HN are
               | describing.
               | 
               | I apologize if you feel I am putting unrepresentative
               | words in your mouth, but this is the reading I am taking
               | away from your comments.
        
               | signatoremo wrote:
               | Lot of impressive points. They are also irrelevant. The
               | majority of people also only extrapolate from the
               | knowledge they acquired in the past. That's why there is
               | the concept of inventor, someone who comes up with new
               | ideas. Many new inventions are also based on existing
               | ideas. Is that the reason to dismiss those achievements?
               | 
               | Do you only take LLM seriously if it can be another
               | Einstein?
               | 
               | > But a world where we used LLMs for the majority of
               | work, would be a world with no fundamental breakthroughs.
               | 
               | What do you consider recent fundamental breakthroughs?
               | 
               | Even if you are right, human can continue to work on hard
               | problems while letting LLM handle the majority of
               | derivative work
        
               | uxcolumbo wrote:
               | How do human brains create something novel and what will
               | it take for AIs to do the same?
        
               | threethirtytwo wrote:
               | > It's pretty clear you don't have a solid background in
               | generative models, because this is fundamentally what
               | they do
               | 
               | You don't have a solid background. No one does. We
               | fundamentally don't understand LLMs, this is an industry
               | and academic opinion. Sure there are high level
               | perspectives and analogies we can apply to LLMs and
               | machine learning in general like probability
               | distributions, curve fitting or interpolations... but
               | those explanations are so high level that they can
               | essentially be applied to humans as well. At a lower
               | level we cannot describe what's going on. We have no idea
               | how to reconstruct the logic of how an LLM arrived at a
               | specific output from a specific input.
               | 
               | It is impossible to have any sort of deterministic
               | function, process or anything produce new information
               | from old information. This limitation is fundamental to
               | logic and math and thus it will limit human output as
               | well.
               | 
               | You can combine information you can transform information
               | you can lose information. But producing new information
               | from old information from deterministic intelligence is
               | fundamentally impossible in reality and therefore
               | fundamentally impossible for LLMs and humans. But note
               | the keyword: "deterministic"
               | 
               | New information can literally only arise through
               | stochastic processes. That's all you have in reality. We
               | know it's stochastic because determinism vs.
               | stochasticism are literally your only two viable options.
               | You have a bunch of inputs, the outputs derived from it
               | are either purely deterministic transformations or if you
               | want some new stuff from the input you must apply
               | randomness. That's it.
               | 
               | That's essentially what creativity is. There is literally
               | no other logical way to generate "new information".
               | Purely random is never really useful so "useful
               | information" arrives only after it is filtered and we use
               | past information to filter the stochastic output and
               | "select" something that's not wildly random. We also only
               | use randomness to perturb the output a little bit so it's
               | not too crazy.
               | 
               | In the end it's this selection process and stochastic
               | process combined that forms creativity. We know this is a
               | general aspect of how creativity works because there's
               | literally no other way to do it.
               | 
               | LLMs do have stochastic aspects to them so we know for a
               | fact it is generating new things and not just drawing on
               | the past. We know it can fit our definition of "creative"
               | and we can literally see it be creative in front of your
               | eyes.
               | 
               | You're ignoring what you see with your eyes and drawing
               | your conclusions from a model of LLMs that isn't fully
               | accurate. Or you're not fully tying the mechanisms of how
               | LLMs work with what creativity or generating new data
               | from past data is in actuality.
               | 
               | The fundamental limitation with LLMs is not that it can't
               | create new things. It's that the context window is too
               | small to create new things beyond that. Whatever it can
               | create it is limited to the possibilities within that
               | window and that sets a limitation on creativity.
               | 
               | What you see happening with LEAN can also be an issue
               | with the context window being too small. If we have an
               | LLM with a giant context window bigger than anything
               | before... and pass it all the necessary data to "learn"
               | and be "trained" on lean it can likely start to produce
               | new theorems without literally being "trained".
               | 
               | Actually I wouldn't call this a "fundamental" problem.
               | More fundamental is the aspect of hallucinations. The
               | fact that LLMs produce new information from past
               | information in the WRONG way. Literally making up
               | bullshit out of thin air. It's the opposite problem of
               | what you're describing. These things are too creative and
               | making up too much stuff.
               | 
               | We have hints that LLMs know the difference between
               | hallucinations and reality but coaxing it to communicate
               | that differentiation to us is limited.
        
               | oedemis wrote:
               | as architectures evolve, i think it can be that we learn
               | more "side effects".. back in 2020 openai researchers
               | said "GPT-3 is applied without any gradient updates or
               | fine-tuning" the model emerges at a certain level of
               | scale...
        
               | aoeusnth1 wrote:
               | > It's pretty clear you don't have a solid background in
               | generative models, because this is fundamentally what
               | they do: model an existing probability distribution and
               | draw samples from that.
               | 
               | After post-training, this is definitively NOT what an LLM
               | does.
        
               | weatherlite wrote:
               | > So it is absurdly incorrect to say "they can only
               | reproduce the past."
               | 
               | Also , a shitton of what we do economically is
               | reproducing the past with slight tweaks and improvements.
               | We all do very repetitive things and these tools cut the
               | time / personnel needed by a significant factor.
        
               | windexh8er wrote:
               | > They are helping their users create things that didn't
               | exist before.
               | 
               | That is a derived output. That isn't new as in: novel. It
               | may be unique but it is derived from training data. LLMs
               | legitimately cannot think and thus they cannot create in
               | that way.
        
               | Kerrick wrote:
               | That is a pedantic distinction. You can create something
               | that didn't exist by combining two things that did exist,
               | in a way of combining things that already existed. For
               | example, you could use a blender to combine almond butter
               | and sawdust. While this may not be "novel", and it may be
               | derived from existing materials and methods, you may
               | still lay claim to having created something that didn't
               | exist before.
               | 
               | For a more practical example, creating bindings from
               | dynamic-language-A for a library in compiled-language-B
               | is a genuinely useful task, allowing you to create things
               | that didn't exist before. Those things are likely to
               | unlock great happiness and/or productivity, even if they
               | are derived from training data.
        
               | windexh8er wrote:
               | > That is a pedantic distinction. You can create
               | something that didn't exist by combining two things that
               | did exist, in a way of combining things that already
               | existed.
               | 
               | This is the definition of a derived product. Call it a
               | derivative work if we're being pedantic and, regardless,
               | is not any level of proof that LLMs "think".
        
               | threethirtytwo wrote:
               | Pedantic and not true. The LLM has stochastic processes
               | involved. Randomness. That's not old information. That's
               | newly generated stuff.
        
               | zingar wrote:
               | Could you give us an idea of what you're hoping for that
               | is not possible to derive from training data of the
               | entire internet and many (most?) published books?
        
               | techpression wrote:
               | This is the problem, the entire internet is a really bad
               | set of training data because it's extremely polluted.
               | 
               | Also the derived argument doesn't really hold, just
               | because you know about two things doesn't mean you'd be
               | able to come up with the third, it's actually very hard
               | most of the time and requires you to not do next token
               | prediction.
        
               | threethirtytwo wrote:
               | The emergent phenomenon is that the LLM can separate
               | truth from fiction when you give it a massive amount of
               | data. It can figure the world out just as we can figure
               | it out when we are as well inundated with bullshit data.
               | The pathways exist in the LLM but it won't necessarily
               | reveal that to you unless you tune it with RL.
        
               | ahtihn wrote:
               | > The emergent phenomenon is that the LLM can separate
               | truth from fiction when you give it a massive amount of
               | data.
               | 
               | I don't believe they can. LLMs have no concept of truth.
               | 
               | What's likely is that the "truth" for many subjects is
               | represented way more than fiction and when there is
               | objective truth it's consistently represented in similar
               | way. On the other hand there are many variations of
               | "fiction" for the same subject.
        
               | threethirtytwo wrote:
               | They can and we have definitive proof. When we tune LLM
               | models with reinforcement learning the models end up
               | hallucinating less and becoming more reliable. Basically
               | in a nut shell we reward the model when telling the truth
               | and punish it when it's not.
               | 
               | So think of it like this, to create the model we use
               | terabytes of data. Then we do RL which is probably less
               | than one percent of additional data involved in the
               | initial training.
               | 
               | The change in the model is that reliability is increased
               | and hallucinations are reduced at a far greater rate than
               | one percent. So much so that modern models can be used
               | for agentic tasks.
               | 
               | How can less than one percent of reinforcement training
               | get the model to tell the truth greater than one percent
               | of the time?
               | 
               | The answer is obvious. It ALREADY knew the truth. There's
               | no other logical way to explain this. The LLM in its
               | original state just predicts text but it doesn't care
               | about truth or the kind of answer you want. With a little
               | bit of reinforcement it suddenly does much better.
               | 
               | It's not a perfect process and reinforcement learning
               | often causes the model to be deceptive an not necessarily
               | tell the truth but it more gives an answer that may seem
               | like the truth or an answer that the trainer wants to
               | hear. In general though we can measurably see a
               | difference in truthfulness and reliability to an extent
               | far greater than the data involved in training and that
               | is logical proof it knows the difference.
               | 
               | Additionally while I say it knows the truth already this
               | is likely more of a blurry line. Even humans don't fully
               | know the truth so my claim here is that an LLM knows the
               | truth to a certain extent. It can be wildly off for
               | certain things but in general it knows and this "knowing"
               | has to be coaxed out of the model through RL.
               | 
               | Keep in mind the LLM is just auto trained on reams and
               | reams of data. That training is massive. Reinforcement
               | training is done on a human basis. A human must rate the
               | answers so it is significantly less.
        
               | habinero wrote:
               | > The answer is obvious. It ALREADY knew the truth.
               | There's no other logical way to explain this.
               | 
               | I can think of several offhand.
               | 
               | 1. The effect was never real, you've just convinced
               | yourself it is because you want it to be, ie you Clever
               | Hans'd yourself.
               | 
               | 2. The effect is an artifact of how you measure "truth"
               | and disappears outside that context ("It can be wildly
               | off for certain things")
               | 
               | 3. The effect was completely fabricated and is the result
               | of fraud.
               | 
               | If you want to convince me that "I threatened a
               | statistical model with a stick and it somehow got more
               | accurate, therefore it's both intelligent and lying" is
               | true, I need a lot less breathless overcredulity and a
               | lot more "I have actively tried to disprove this result,
               | here's what I found"
        
               | threethirtytwo wrote:
               | You asked for something concrete, so I'll anchor every
               | claim to either documented results or directly observable
               | training mechanics.
               | 
               | First, the claim that RLHF materially reduces
               | hallucinations and increases factual accuracy is not
               | anecdotal. It shows up quantitatively in benchmarks
               | designed to measure this exact thing, such as TruthfulQA,
               | Natural Questions, and fact verification datasets like
               | FEVER. Base models and RL-tuned models share the same
               | architecture and almost identical weights, yet the RL-
               | tuned versions score substantially higher. These
               | benchmarks are external to the reward model and can be
               | run independently.
               | 
               | Second, the reinforcement signal itself does not contain
               | factual information. This is a property of how RLHF
               | works. Human raters provide preference comparisons or
               | scores, and the reward model outputs a single scalar.
               | There are no facts, explanations, or world models being
               | injected. From an information perspective, this signal
               | has extremely low bandwidth compared to pretraining.
               | 
               | Third, the scale difference is documented by every group
               | that has published training details. Pretraining consumes
               | trillions of tokens. RLHF uses on the order of tens or
               | hundreds of thousands of human judgments. Even generous
               | estimates put it well under one percent of the total
               | training signal. This is not controversial.
               | 
               | Fourth, the improvement generalizes beyond the reward
               | distribution. RL-tuned models perform better on prompts,
               | domains, and benchmarks that were not part of the
               | preference data and are evaluated automatically rather
               | than by humans. If this were a Clever Hans effect or
               | evaluator bias, performance would collapse when the
               | reward model is not in the loop. It does not.
               | 
               | Fifth, the gains are not confined to a single definition
               | of "truth." They appear simultaneously in question
               | answering accuracy, contradiction detection, multi-step
               | reasoning, tool use success, and agent task completion
               | rates. These are different evaluation mechanisms. The
               | only common factor is that the model must internally
               | distinguish correct from incorrect world states.
               | 
               | Finally, reinforcement learning cannot plausibly inject
               | new factual structure at scale. This follows from
               | gradient dynamics. RLHF biases which internal activations
               | are favored, it does not have the capacity to encode
               | millions of correlated facts about the world when the
               | signal itself contains none of that information. This is
               | why the literature consistently frames RLHF as behavior
               | shaping or alignment, not knowledge acquisition.
               | 
               | Given those facts, the conclusion is not rhetorical. If a
               | tiny, low-bandwidth, non-factual signal produces large,
               | general improvements in factual reliability, then the
               | information enabling those improvements must already
               | exist in the pretrained model. Reinforcement learning is
               | selecting among latent representations, not creating
               | them.
               | 
               | You can object to calling this "knowing the truth," but
               | that's a semantic move, not a substantive one. A system
               | that internally represents distinctions that reliably
               | track true versus false statements across domains, and
               | can be biased to express those distinctions more
               | consistently, functionally encodes truth.
               | 
               | Your three alternatives don't survive contact with this.
               | Clever Hans fails because the effect generalizes.
               | Measurement artifact fails because multiple independent
               | metrics move together. Fraud fails because these results
               | are reproduced across competing labs, companies, and
               | open-source implementations.
               | 
               | If you think this is still wrong, the next step isn't
               | skepticism in the abstract. It's to name a concrete
               | alternative mechanism that is compatible with the
               | documented training process and observed generalization.
               | Without that, the position you're defending isn't
               | cautious, it's incoherent.
        
               | jama211 wrote:
               | Yeah you've lost me here I'm sorry. In the real world
               | humans work with AI tools to create new things. What
               | you're saying is the equivalent of "when a human writes a
               | book in English, because they use words and letters that
               | already exist and they already know they aren't creating
               | anything new".
        
               | nl wrote:
               | What does "think" mean?
               | 
               | Why is that kind of thinking required to create novel
               | works?
               | 
               | Randomness can create novelty.
               | 
               | Mistakes can be novel.
               | 
               | There are many ways to create novelty.
               | 
               | Also I think you might not know how LLMs are trained to
               | code. Pre-training gives them some idea of the syntax etc
               | but that only gets you to fancy autocomplete.
               | 
               | Modern LLMs are heavily trained using reinforcement data
               | which is custom task the labs pay people to do (or by
               | distilling another LLM which has had the process
               | performed on it).
        
               | closewith wrote:
               | By that definition, nearly all commercial software
               | development (and nearly all human output in general) is
               | derived output.
        
               | windexh8er wrote:
               | Wow.
               | 
               | You're using 'derived' to imply 'therefore equivalent.'
               | That's a category error. A cookbook is derived from food
               | culture. Does an LLM taste food? Can it think about how
               | good that cookie tastes?
               | 
               | A flight simulator is derived from aerodynamics - yet it
               | doesn't fly.
               | 
               | Likewise, text that resembles reasoning isn't the same
               | thing as a system that has beliefs, intentions, or
               | understanding. Humans do. LLMs don't.
               | 
               | Also... Ask an LLM what's the difference between a human
               | brain and an LLM. If an LLM could "think" it wouldn't
               | give you the answer it just did.
        
               | CamperBob2 wrote:
               | _Ask an LLM what 's the difference between a human brain
               | and an LLM. If an LLM could "think" it wouldn't give you
               | the answer it just did._
               | 
               | I imagine that sounded more profound when you wrote it
               | than it did just now, when I read it. Can you be a little
               | more specific, with regard to what features you would
               | expect to differ between LLM and human responses to such
               | a question?
               | 
               | Right now, LLM system prompts are strongly geared towards
               | _not_ claiming that they are humans or simulations of
               | humans. If your point is that a hypothetical  "thinking"
               | LLM would claim to be a human, that could certainly be
               | arranged with an appropriate system prompt. You wouldn't
               | know whether you were talking to an LLM or a human --
               | just as you don't now -- but nothing would be proved
               | either way. That's ultimately why the Turing test is a
               | poor metric.
        
               | closewith wrote:
               | You're arguing against a straw man. No one is claiming
               | LLMs have beliefs, intentions, or understanding. They
               | don't need them to be economically useful.
        
               | windexh8er wrote:
               | Oh yes, they are.
               | 
               | And beyond people claiming that LLMs are basically
               | sentient you have people like CamperBob2 who made this
               | wild claim:
               | 
               |  _" ""There's no such thing as people without language,
               | except for infants and those who are so mentally
               | incapacitated that the answer is self-evidently "No, they
               | cannot."
               | 
               | Language is the substrate of reason. It doesn't need to
               | be spoken or written, but it's a necessary and (as it
               | turns out) sufficient component of thought."""_
               | 
               | Let that sink. They literally think that there's no such
               | thing as people without language. Talk about a wild and
               | ignorant take on life in general!
        
               | ordersofmag wrote:
               | I will find this often-repeated argument compelling only
               | when someone can prove to me that the human mind works in
               | a way that isn't 'combining stuff it learned in the
               | past'.
               | 
               | 5 years ago a typical argument against AGI was that
               | computers would never be able to think because "real
               | thinking" involved mastery of language which was
               | something clearly beyond what computers would ever be
               | able to do. The implication was that there was some magic
               | sauce that human brains had that couldn't be replicated
               | in silicon (by us). That 'facility with language'
               | argument has clearly fallen apart over the last 3 years
               | and been replaced with what appears to be a different
               | magic sauce comprised of the phrases 'not really
               | thinking' and the whole 'just repeating what it's
               | heard/parrot' argument.
               | 
               | I don't think LLM's think or will reach AGI through
               | scaling and I'm skeptical we're particularly close to AGI
               | in any form. But I feel like it's a matter of incremental
               | steps. There isn't some magic chasm that needs to be
               | crossed. When we get there I think we will look back and
               | see that 'legitimately thinking' wasn't anything magic.
               | We'll look at AGI and instead of saying "isn't it amazing
               | computers can do this" we'll say "wow, was that all there
               | is to thinking like a human".
        
               | windexh8er wrote:
               | > 5 years ago a typical argument against AGI was that
               | computers would never be able to think because "real
               | thinking" involved mastery of language which was
               | something clearly beyond what computers would ever be
               | able to do.
               | 
               | Mastery of words is thinking? In that line of argument
               | then computers have been able to think for decades.
               | 
               | Humans don't think only in words. Our context, memory and
               | thoughts are processed and occur in ways we don't
               | understand, still.
               | 
               | There's a lot of great information out there describing
               | this [0][1]. Continuing to believe these tools are
               | thinking, however, is dangerous. I'd gather it has
               | something to do with logic: you can't see the process and
               | it's non-deterministic so it feels like thinking. ELIZA
               | tricked people. LLMs are no different.
               | 
               | [0] https://archive.is/FM4y8 [0]
               | https://www.theverge.com/ai-artificial-
               | intelligence/827820/l... [1]
               | https://www.raspberrypi.org/blog/secondary-school-maths-
               | show...
        
               | CamperBob2 wrote:
               | _Mastery of words is thinking?_
               | 
               | That's the crazy thing. Yes, in fact, it turns out that
               | language encodes and embodies reasoning. All you have to
               | do is pile up enough of it in a high-dimensional space,
               | use gradient descent to model its original structure, and
               | add some feedback in the form of RL. At that point,
               | reasoning is just a database problem, which we currently
               | attack with attention.
               | 
               | No one had the faintest clue. Even now, many people not
               | only don't understand what just happened, but they don't
               | think anything happened at all.
               | 
               | ELIZA, ROFL. How'd ELIZA do at the IMO last year?
        
               | meindnoch wrote:
               | So people without language cannot reason? I don't think
               | so.
        
               | CamperBob2 wrote:
               | There's no such thing as people without language, except
               | for infants and those who are so mentally incapacitated
               | that the answer is self-evidently "No, they cannot."
               | 
               | Language is the substrate of reason. It doesn't need to
               | be spoken or written, but it's a necessary and (as it
               | turns out) sufficient component of thought.
        
               | windexh8er wrote:
               | There are quite a few studies to refute this highly
               | ignorant comment. I'd suggest some reading [0].
               | 
               | From the abstract: _" Is thought possible without
               | language? Individuals with global aphasia, who have
               | almost no ability to understand or produce language,
               | provide a powerful opportunity to find out.
               | Astonishingly, despite their near-total loss of language,
               | these individuals are nonetheless able to add and
               | subtract, solve logic problems, think about another
               | person's thoughts, appreciate music, and successfully
               | navigate their environments. Further, neuroimaging
               | studies show that healthy adults strongly engage the
               | brain's language areas when they understand a sentence,
               | but not when they perform other nonlinguistic tasks like
               | arithmetic, storing information in working memory,
               | inhibiting prepotent responses, or listening to music.
               | Taken together, these two complementary lines of evidence
               | provide a clear answer to the classic question: many
               | aspects of thought engage distinct brain regions from,
               | and do not depend on, language."_
               | 
               | [0] https://pmc.ncbi.nlm.nih.gov/articles/PMC4874898/
        
               | arcatech wrote:
               | > I will find this often-repeated argument compelling
               | only when someone can prove to me that the human mind
               | works in a way that isn't 'combining stuff it learned in
               | the past'.
               | 
               | This is the definition of the word 'novel'.
        
               | handoflixue wrote:
               | Seriously, all that familiarity and you think an LLM
               | "literally" can't invent anything that didn't already
               | exist?
               | 
               | Like, I'm sorry, but you're just flat-out wrong and I've
               | got the proof sitting on my hard drive. I use this
               | supposedly impossible program daily.
        
               | bigyabai wrote:
               | FWIW, your "evidence" is a text editor. I'm glad you made
               | a tool that works for you, but the parent's point stands;
               | this is a 200-level course-curriculum homework
               | assignment. Tens of thousands of homemade editors exist,
               | in various states of disrepair and vain overengineering.
        
               | least wrote:
               | The difference between those is the person is actually
               | using this text editor that they built with the help of
               | LLMs. There's plenty of people creating novel scripts and
               | programs that can accommodate their own unique
               | specifications.
               | 
               | If a programmer creating their own software (or
               | contracting it out to a developer) would be a bespoke
               | suit and using software someone or some company created
               | without your input is an off the rack suit, I'd liken
               | these sorts of programs as semi-bespoke, or made to
               | measure.
               | 
               | "LLMs are literally technology that can only reproduce
               | the past" feels like an odd statement. I think the point
               | they're going for is that it's not thinking and so it's
               | not going to produce new ideas like a human would? But
               | literally no technology does that. That is all derived
               | from some human beings being particularly clever.
               | 
               | LLMs are tools. They can enable a human to create new
               | things because they are interfacing with a human to
               | facilitate it. It's merging the functional knowledge and
               | vision of a person and translating it into something
               | else.
        
               | resize2996 wrote:
               | compilers can only produce machine code. so unorginal.
        
               | windexh8er wrote:
               | Do you also think LLMs "think"?
               | 
               | From what you've described an LLM has not invented
               | anything. LLMs that can reason have a bit more slight of
               | hand but they're not coming up with new ideas outside of
               | the bounds of what a lot of words have encompassed in
               | both fiction and non.
               | 
               | Good for you that you've got a fun token of code that's
               | what you've always wanted, I guess. But this type of
               | fantasy take on LLMs seems to be more and more prevalent
               | as of late. A lot of people defending LLMs as if they're
               | owed something because they've built something or maybe
               | people are getting more and more attached to them from
               | the conversational angle. I'm not sure, but I've run
               | across more people in 2025 that are way too far in the
               | deep end of personifying their relationships with LLMs.
        
               | Kerrick wrote:
               | Hang on, you're now saying that if something has ever
               | been described in _fiction_ it doesn 't count as
               | invention? So if somebody literally developed a working
               | photon torpedo, that isn't new because "Star Trek Did
               | It"?
        
               | phatfish wrote:
               | Is there any danger an LLM is going to create a working
               | photo torpedo?
        
               | ben_w wrote:
               | Well, they can use tools, and tools includes physics
               | simulations, so if it is possible (and FWIW the tool-free
               | "intuition" of ChatGPT is "there will never be an age of
               | antimatter"), then why couldn't LLMs grind those tools to
               | get a solution?
        
               | windexh8er wrote:
               | You seem to be pretty far down the rabbit hole. How about
               | this... You task an LLM to create a photon torpedo. If it
               | can truly think then it should be able to provide you
               | with something tangible. When you've got that in hand let
               | us all know.
               | 
               | Back to the land of reality... Describing something in
               | fiction doesn't magically make it "not an invention".
               | Fiction can anticipate an idea, but invention is about
               | producing a working, testable implementation and usually
               | involves novel technical methods. "Star Trek did it" is
               | at most prior art for the concept, not a blueprint for
               | the mechanism. If you can't understand that differential
               | then maybe go ask an LLM.
        
               | Kerrick wrote:
               | I didn't say anything about an LLM. I said "somebody" not
               | "some predictive text engine."
        
               | ctxc wrote:
               | Some people cannot be convinced simply because their
               | expectation of "novel" is something that appears in an
               | Asimov novel.
               | 
               | I for one think your work is pretty cool - even though I
               | haven't seen it, using something you built everyday is a
               | claim not many can make!
        
               | 9rx wrote:
               | When a computer is able to invent things, we've achieved
               | AGI. Do you believe we are already in the AGI era, or is
               | the inventor in this case actually you?
        
               | threethirtytwo wrote:
               | Over half of HN still thinks it's a stochastic parrot and
               | that it's just a glorified google search.
               | 
               | The change hit us so fast a huge number of people don't
               | understand how capable it is yet.
               | 
               | Also it certainly doesn't help that it still
               | hallucinates. One mistake and it's enough to set someone
               | against LLMs. You really need to push through that
               | hallucinations are just the weak part of the process to
               | see the value.
        
               | CamperBob2 wrote:
               | The problem I see, over and over, is that people pose
               | poorly-formed questions to the free ChatGPT and Google
               | models, laugh at the resulting half-baked answers that
               | are often full of errors and hallucinations, and draw
               | conclusions about the technology as a whole.
               | 
               | Either that, or they tried it "last year" or "a while
               | back" and have no concept of how far things have gone in
               | the meantime.
               | 
               | It's like they wandered into a machine shop, cut off a
               | finger or two, and concluded that their grandpa's hammer
               | and hacksaw were all anyone ever needed.
        
               | habinero wrote:
               | No, frankly it's the difference between actual engineers
               | and hobbyists/amateurs/non-SWEs.
               | 
               | SWEs are trained to discard surface-level observations
               | and be adversarial. You can't just look at the happy
               | path, how does the system behave for edge cases? Where
               | does it break down and how? What are the failure modes?
               | 
               | The actual analogy to a machine shop would be to look at
               | whether the machines were adequate for their use case,
               | the building had enough reliable power to run and if
               | there were any safety issues.
               | 
               | It's easy to Clever Hans yourself and get snowed by what
               | _looks_ like sophisticated effort or flat out bullshit. I
               | had to gently tell a junior engineer that just because
               | the marketing claims something will work a certain way,
               | that doesn 't mean it will.
        
               | CamperBob2 wrote:
               | You sound pretty certain. There's often good money to be
               | made in taking the contrarian view, where you have
               | insights that the so-called "smart money" lacks. What are
               | some good investments to make in the extreme-bear case,
               | in which we're all just Clever Hans-ing ourselves as you
               | put it? Do you have skin in the game?
        
               | threethirtytwo wrote:
               | What you're describing is just competent engineering, and
               | it's already been applied to LLMs. People have been
               | adversarial. That's why we know so much about
               | hallucinations, jailbreaks, distribution shift failures,
               | and long-horizon breakdowns in the first place. If this
               | were hobbyist awe, none of those benchmarks or red-
               | teaming efforts would exist.
               | 
               | The key point you're missing is the type of failure.
               | Search systems fail by not retrieving. Parrots fail by
               | repeating. LLMs fail by producing internally coherent but
               | factually wrong world models. That failure mode only
               | exists if the system is actually modeling and reasoning,
               | imperfectly. You don't get that behavior from lookup or
               | regurgitation.
               | 
               | This shows up concretely in how errors scale. Ambiguity
               | and multi-step inference increase hallucinations.
               | Scaffolding, tools, and verification loops reduce them.
               | Step-by-step reasoning helps. Grounding helps. None of
               | that makes sense for a glorified Google search.
               | 
               | Hallucinations are a real weakness, but they're not
               | evidence of absence of capability. They're evidence of an
               | incomplete reasoning system operating without sufficient
               | constraints. Engineers don't dismiss CNC machines because
               | they crash bits. They map the envelope and design around
               | it. That's what's happening here.
               | 
               | Being skeptical of reliability in specific use cases is
               | reasonable. Concluding from those failure modes that this
               | is just Clever Hans is not adversarial engineering. It's
               | stopping one layer too early.
        
             | Greduan wrote:
             | Text editors in a thousand flavours has indeed already been
             | programmed though. I don't think you understood what op
             | meant.
             | 
             | Curious, does it perform at the limit of the hardware? Was
             | it programmed in a tools language (like C++, Rust, C, etc.)
             | or in a web tech?
        
               | zingar wrote:
               | What is the point that you believe would be demonstrated
               | by a new text editor running at the limit of hardware in
               | a compiled editor? Would that point apply to every other
               | text editor that exists already?
        
             | fmbb wrote:
             | Is your new text editor open source?
        
             | nsxwolf wrote:
             | The LLM didn't invent any new technology to do that,
             | though. You used the LLM to reorganize Lego building blocks
             | of knowledge into something new.
             | 
             | Without you, there was nothing.
        
           | waldrews wrote:
           | I'm being hyperbolic of course, but I'm a little dismissive
           | of the progress that happened since the days of BBS's and car
           | based cell phones - we just got more connectivity, more
           | capacity, more content, bigger/faster. Likewise, my attitude
           | toward machine learning before 2023 is a smug 'heh, these
           | computer scientists are doing undisciplined statistics at
           | scale, how nice for them.' Then all of a sudden the machines
           | woke up and started arguing with me, coherently, even about
           | niche topics I have a PhD in. I can appreciate in retrospect
           | how much of the machine learning progress ultimately went
           | into that, but, like fusion, the magic payoff was supposed to
           | be decades away and always remain decades away. This wasn't
           | supposed to happen in my lifetime. 2025 progress isn't the
           | 2023 shock, but this was the year LLM's-as-programmers (and
           | LLM's-as-mathematicians, and...) went from 'isn't that cute,
           | the machine is trying' to 'an expert with enough time would
           | make better choices than the machine did,' and that makes for
           | a different world. More so than, going from a Commodore Vic
           | 20 with 4k of RAM and a modem to the latest Macbook.
        
           | ako wrote:
           | > This year honestly feels quite stagnant. LLMs are literally
           | technology that can only reproduce the past.
           | 
           | Is this such a big limitation? Most jobs are basically people
           | trained on past knowledge applying it today. No need to
           | generate new knowledge.
           | 
           | And a lot of new knowledge is just combining 2 things from
           | the past in a new way.
        
           | HarHarVeryFunny wrote:
           | > LLMs are literally technology that can only reproduce the
           | past.
           | 
           | That's incorrect on many levels. They are drawing upon, and
           | reproducing, language patterns from "the past", but they are
           | combining those patterns in ways that may have never have
           | been seen before. They may not be truly creative, but they
           | are still capable of generating novel outputs.
           | 
           | > They're cool, but they were way cooler 4 years ago.
           | 
           | Maybe this year has been more about incremental progress with
           | LLMs than the shock/coolness factor of talking to an LLM for
           | the first time, but the utility of them, especially for
           | programming, has dramatically increased this year, really in
           | the last 6 months.
           | 
           | The improvement in "AI" image and video generation has also
           | been impressive, to the point now that fake videos on YouTube
           | can often only be identified as such by common sense rather
           | that the fact that they don't look real.
           | 
           | Incremental improvement can often be more impressive that
           | innovation, whose future importance can be hard to judge when
           | it first appears. How many people read "Attention is all you
           | need" in 2017 and thought "Wow! This is going to change the
           | world!". Not even the authors of the paper thought that.
        
         | odiroot wrote:
         | I'm very relieved we've moved away from rewriting everything in
         | Rust.
        
           | jll29 wrote:
           | There's no reason not to use Rust for LLM-generated code in
           | the longer term (other than lack of Rust code to learn from
           | in the shorter term).
           | 
           | The stricter typing of Rust would make sematic errors in
           | generated code come out more quickly than in e.g. Python
           | because using static typing the chances are that some of the
           | semantic errors are also type violations.
        
           | michaelcampbell wrote:
           | Have we though? I'm glad we're not shouting about it from the
           | rooftops like it's some magical "win" button as much, but TBH
           | the things I use routinely that HAVE been rewritten in rust
           | are generally much better. That could also just be because
           | they're newer and have the errors of the past to not repeat.
        
       | sanreau wrote:
       | > Vendor-independent options include GitHub Copilot CLI, Amp,
       | OpenHands CLI, and Pi
       | 
       | ...and the best of them all, OpenCode[1] :)
       | 
       | [1]: https://opencode.ai
        
         | simonw wrote:
         | Good call, I'll add that. I think I mentally scrambled it with
         | OpenHands.
        
           | the_mitsuhiko wrote:
           | Thanks for adding pi to it though :)
        
         | nineteen999 wrote:
         | How did I miss this until now! Thank you for sharing.
        
         | logicprog wrote:
         | I don't know why you're downloaded, OpenCode is by far the
         | best.
        
         | d4rkp4ttern wrote:
         | Can OpenCode be used with the Claude Max or ChatGPT Pro
         | subscriptions, i.e., without per-token API charges?
        
           | simonw wrote:
           | Apparently it does work with Claude Max:
           | https://opencode.ai/docs/providers/#anthropic
           | 
           | I don't see a similar option for ChatGPT Pro. Here's a closed
           | issue: https://github.com/sst/opencode/issues/704
        
             | williamstein wrote:
             | There's a plugin that evidently supports ChatGPT Pro with
             | Opencode: https://github.com/sst/opencode/issues/1686#issue
             | comment-349...
        
           | ewoodrich wrote:
           | Yes, I use it with a regular Claude Pro subscription. It also
           | supports using GitHub Copilot subscriptions as a backend.
        
       | the_mitsuhiko wrote:
       | > The (only?) year of MCP
       | 
       | I like to believe, but MCP is quickly turning into an enterprise
       | thing so I think it will stick around for good.
        
         | simonw wrote:
         | I think it will stick around, but I don't think it will have
         | another year where it's the hot thing it was back in January
         | through May.
        
           | Alex-Programs wrote:
           | I never quite got what was so "hot" about it. There seems to
           | be an entire parallel ecosystem of corporates that are just
           | begging to turn AI into PowerPoint slides so that they can
           | mould it into a shape that's familiar.
        
             | 9dev wrote:
             | One reason may be that it makes it a lot easier to open up
             | a product to AI. Instead of adding a bad ChatGPT UI clone
             | into your app, you inverse control and let external AI
             | tools interact with your application and its data, thus
             | giving your customers immediate benefits, while
             | simultaneously sating your investors/founders/managers
             | desire to somehow add AI.
        
         | nrhrjrjrjtntbt wrote:
         | MCP or skills? Can a skill negate the need for MCP. In addition
         | there was a YC startup who is looking at searching docs for
         | LLMs or similar. I think MCP may be less needed once you have
         | skills, openapi specs, and other things that LLMs can call
         | directly.
        
         | MitziMoto wrote:
         | MCP isn't going anywhere. Some developers can't seem to see
         | past their terminal or dev environment when it comes to MCP.
         | Skills, etc do not replace MCP and MCP is far more than just
         | documentation searching.
         | 
         | MCP is a great way for an LLM to connect to an external system
         | in a standardized way and immediately understand what tools it
         | has available, when and how to use them, what their inputs and
         | outputs are,etc.
         | 
         | For example, we built a custom MCP server for our CRM. Now our
         | voice and chat agents that run on elevenlabs infrastructure can
         | connect to our system with one endpoint, understand what
         | actions it can take, and what information it needs to collect
         | from the user to perform those actions.
         | 
         | I guess this could maybe be done with webhooks or an API spec
         | with a well crafted prompt? Or if eleven labs provided an
         | executable environment with tool calling? But at some point
         | you're just reinventing a lot of the functionality you get for
         | free from MCP, and all major LLMs seem to know how to use MCP
         | already.
        
           | simonw wrote:
           | Yeah, I don't think I was particularly clear in that section.
           | 
           | I don't think MCP is going to go away, but I do think it's
           | unlikely to ever achieve the level of excitement it had in
           | early 2025 again.
           | 
           | If you're not building inside a code execution environment
           | it's a very good option for plugging tools into LLMs,
           | especially across different systems that support the same
           | standard.
           | 
           | But code execution environments are so much more powerful and
           | flexible!
           | 
           | I expect that once we come up with a robust, inexpensive way
           | to run a little Bash environment - I'm still hoping
           | WebAssembly gets us there - there will be much less reason to
           | use MCP even outside of coding agent setups.
        
             | brabel wrote:
             | I disagree. MCP will remain the best way to do most things
             | for the same reason REST APIs are the main way to access
             | non local services: they provide a way to secure and audit
             | access to systems in a way that a coding environment
             | cannot. And you can authorize actions depending on the well
             | defined inputs and outputs. You can't do that using just a
             | bash script unless said script actually does SSO and calls
             | REST APIs but then you just have a worse MCP client without
             | any interoperability.
        
               | the_mitsuhiko wrote:
               | I find it very hard to pick winners and losers in this
               | environment where everything changes so quickly. Right
               | now a lot of people are using bash as a glue environment
               | for agents, even if they are not for developers.
        
         | cloudking wrote:
         | For connecting agents to third-party systems I prefer CLI
         | tools, less context bloat and faster. You can define the CLI
         | usage in your agent instructions. If the MCP you're using
         | doesn't exist as a CLI, build one with your agent.
        
       | npalli wrote:
       | Great summary of the year in LLMs. Is there a predictions (for
       | 2026) blogpost as well?
        
         | simonw wrote:
         | Given how badly my 2025 predictions aged I'm probably going to
         | sit that one out! https://simonwillison.net/2025/Jan/10/ai-
         | predictions/
        
           | DANmode wrote:
           | Don't be a bad sport, now!!
        
           | zahlman wrote:
           | Making predictions is useful even when they turn out very
           | wrong. Consider also giving confidence levels, so that you
           | can calibrate going forward.
        
             | jjude wrote:
             | I use predictions to prepare rather than to plan.
             | 
             | Planing depends on deterministic view of the future. I used
             | to plan (esp annual plans) until about 5 years. Now I scan
             | for trends and prepare myself for different scenarios that
             | can come in the future. Even if you get it approximately
             | right, you stand apart.
             | 
             | For tech trends, I read Simon, Benedict Evans, Mary Meeker
             | etc. Simon is in a better position make these predictions
             | than anyone else having closely analyzed these trends over
             | the last few years.
             | 
             | Here I wrote about my approach:
             | https://www.jjude.com/shape-the-future/
        
       | skydhash wrote:
       | [flagged]
        
         | MattRix wrote:
         | [flagged]
        
           | skydhash wrote:
           | Why do people assume negative critique is ignorance?
        
             | dmd wrote:
             | People denied that bicycles could possibly balance even as
             | others happily pedaled by. This is the same thing.
        
               | measurablefunc wrote:
               | Bicycles don't balance, the human on the bicycle is the
               | one doing the balancing.
        
               | dmd wrote:
               | Yes, that is the analogy I am making. People argued that
               | bicycles (a tool for humans to use) could not possibly
               | work - even as people were successfully using them.
        
               | measurablefunc wrote:
               | People use drugs as well but I'm not sure I'd call that
               | successful use of chemical compounds without further
               | context. There are many analogies one can apply here that
               | would be equally valid.
        
               | duchef wrote:
               | Bicycles (without a rider) do balance at sufficient speed
               | via a self steering and correction mechanism of the front
               | axle..
        
               | skydhash wrote:
               | Please tell me which one of the headings is not about
               | increased usage o LLMs and derived tools and is about
               | some improvement in the axes of reliability or or any
               | kind of usefulness.
               | 
               | Here is the changelog for OpenBSD 7.8:
               | 
               | https://www.openbsd.org/78.html
               | 
               | There's nothing here that says: We make it easier to use
               | it more of it. It's about using it better and fixing
               | underlying problems.
        
               | simonw wrote:
               | The coding agent heading. Claude Code and tools like it
               | represent a _huge_ improvement in what you can usefully
               | get done with LLMs.
               | 
               | Mistakes and hallucinations matter a whole lot less if a
               | reasoning LLM can try the code, see that it doesn't work
               | and fix the problem.
        
               | walt_grata wrote:
               | If it actually does that without an argument. I can't
               | believe I have to say that about a computer program
        
               | skydhash wrote:
               | > The coding agent heading. Claude Code and tools like it
               | represent a huge improvement in what you can usefully get
               | done with LLMs.
               | 
               | Does it? It's all prompt manipulation. Shell script are
               | powerful yes, but not really _huge_ improvement over
               | having a shell (REPL interface) to the system. And even
               | then a lot of programs just use syscalls or wrapper
               | libraries.
               | 
               | > can try the code, see that it doesn't work and fix the
               | problem.
               | 
               | Can you really say that does happens reliably?
        
               | simonw wrote:
               | Depends on what you mean by "reliably".
               | 
               | If you mean 100% correct all of the time then no.
               | 
               | If you mean correct often enough that you can expect it
               | to be a productive assistant that helps solve all sorts
               | of problems faster than you could solve them without it,
               | and which makes mistakes infrequently enough that you
               | waste less time fixing them than you would doing
               | everything by yourself then yes, it's plenty reliable
               | enough now.
        
               | dham wrote:
               | You're welcome to try the LLM's yourself and come up with
               | your own conclusions. By what you've posted it doesn't
               | look like you've tried the anything in the last 2 years.
               | Yes LLM's can be annoying, but there has been progress.
        
               | noodletheworld wrote:
               | I know it seems like forever ago, but claude code only
               | came out in 2025.
               | 
               | Its very difficult to argue the point that claude code:
               | 
               | 1) was a paradigm shift in terms of functionality,
               | despite, to be fair, at best, incremental improvements in
               | the underlying models.
               | 
               | 2) The results are an order of magnitude, I estimate,
               | better in terms of output.
               | 
               | I think its very fair to distill "AI progress 2025" to:
               | you can get better results (up to a point; better than
               | _raw output_ anyway; scaling to multiple agents has not
               | worked) without better models with clever tools and
               | loops. (...and video /image slop infests everything :p).
        
               | bandrami wrote:
               | Did more software ship in 2025 than in 2024? I'm still
               | looking for some actual indication of output here. I get
               | that people _feel_ more productive but the actual metrics
               | don 't seem to agree.
        
               | skydhash wrote:
               | I'm still waiting for the Linux drivers to be written
               | because of all the 20x improvements that AI hypers are
               | touting. I would even settle for Apple M3 and M4
               | computers to be supported by Asahi.
        
               | noodletheworld wrote:
               | I am not making any argument about productivity about
               | using AI vs. not using AI.
               | 
               | My point is purely that, _compared to 2024_ , the quality
               | of the code produced by LLM inference _agent systems_ is
               | better.
               | 
               | To say that 2025 was a nothing burger is objectively
               | incorrect.
               | 
               | Will it scale? Is it good enough to use professionally?
               | Is this like self driving cars where the best they ever
               | get is stuck with an odd shaped traffic cone? Is it
               | actually more productive?
               | 
               | Who knows?
               | 
               | Im just saying... LLM coding in 2024 sucked. 2025 was a
               | big year.
        
               | tehnub wrote:
               | People did?
        
               | rhubarbtree wrote:
               | It's possible this is correct.
               | 
               | It's also possible that people more experienced,
               | knowledgable and skilled than you can see fundamental
               | flaws in using LLMs for software engineering that you
               | cannot. I am not including myself in that category.
               | 
               | I'm personally honestly undecided. I've been coding for
               | over 30 years and know something like 25 languages. I've
               | taught programming to postgrad level, and built prototype
               | AI systems that foreshadowed LLMs, I've written
               | everything from embedded systems to enterprise, web,
               | mainframes, real time, physics simulation and research
               | software. I would consider myself an 7/10 or 8/10 coder.
               | 
               | A lot of folks I know are better coders. To put my
               | experience into context: one guy in my year at uni wrote
               | one of the world's most famous crypto systems; another
               | wrote large portions of some of the most successful games
               | of the last few decades. So I've grown up surrounded by
               | geniuses, basically, and whilst I've been lectured by
               | true greats I'm humble enough to recognise I don't bleed
               | code like they do. I'm just a dabbler. But it irks me
               | that a lot of folks using AI profess it's the future but
               | don't really know anything about coding compared to these
               | folks. Not to be a Luddite - they are the first people to
               | adopt new languages and techniques, but they also are
               | super sceptical about anything that smells remotely like
               | bullshit.
               | 
               | One of the most wise insights in coding is the
               | aphorism"beware the enthusiasm of the recently
               | converted." And I see that so much with AI. I've seen it
               | with compilers, with IDEs, paradigms, and languages.
               | 
               | I've been experimenting a lot with AI, and I've found it
               | fantastic for comprehending poor code written by others.
               | I've also found it great for bouncing ideas. And the code
               | it writes, beyond boiler plate, is hot garbage. It
               | doesn't properly reason, it can't design architecture, it
               | can't write code that is comprehensible to other
               | programmers, and treating it as a "black box to be
               | manipulated by AI" just leads to dead ends that can't be
               | escaped, terrible decisions that will take huge amounts
               | of expert coding time to undo, subtle bugs that AI can't
               | fix and are super hard to spot, and often you can't
               | understand their code enough to fix them, and security
               | nightmares.
               | 
               | Testing is insufficient for good code. Humans write code
               | in a way that is designed for general correctness. AI
               | does not, at least not yet.
               | 
               | I do think these problems can be solved. I think we
               | probably need automated reasoning systems, or else vastly
               | improved LLMs that border on automated reasoning much
               | like humans do. Could be a year. Could be a decade. But
               | right now these tools don't work well. Great for vibe
               | coding, prototyping, analysis, review, bouncing ideas.
        
               | CamperBob2 wrote:
               | _But right now these tools don't work well. Great for
               | vibe coding, prototyping, analysis, review, bouncing
               | ideas._
               | 
               | What are some of the models you've been working with?
        
               | blibble wrote:
               | people also said that selling jpegs of monkeys for
               | millions of dollars was a pump and dump scam, and would
               | collapse
               | 
               | they were right
        
               | sothatsit wrote:
               | JPEGs with no value other than fake scarcity is very
               | different to coding agents that people actively use to
               | ship real code.
        
             | kakapo5672 wrote:
             | Whenever someone tells me that AI is worthless, does
             | nothing, scam/slop etc, I ask them about their own AI
             | usage, and their general knowledge about what's going on.
             | 
             | Invariably they've never used AI, or at most very rarely.
             | (If they used AI beyond that, this would be admission that
             | it was useful at some level).
             | 
             | Therefore it's reasonable to assume that you are in that
             | boat. Now that might not be true in your case, who knows,
             | but it's definitely true on average.
        
               | snigsnog wrote:
               | It's not worthless, it's just not worldchanging as is
               | even in the fields where it's most useful, like
               | programming. If the trajectory changes and we reach AGI
               | then this changes too but right now it's just a way to
               | 
               | - fart out demos that you don't plan on maintaining, or
               | want to use as a starting place
               | 
               | - generate first-draft unit tests/documentation
               | 
               | - generate boilerplate without too much functionality
               | 
               | - refactor in a very well covered codebase
               | 
               | It's very useful for all of the above! But it doesn't
               | even replace a junior dev at my company in its current
               | state. It's too agreeable, makes subtle mistakes that it
               | can't permanently correct (GEMINI.md isn't a magic
               | bullet, telling it to not do something does not guarantee
               | that it won't do it again), and you as the developer
               | submitting LLM-generated code for review need to review
               | it closely before even putting it up (unless you feel
               | like offloading this to your team) to the point that it's
               | not that much faster than having written it yourself.
        
             | LewisVerstappen wrote:
             | because your "negative critique" is just idiotic and wrong
        
             | sothatsit wrote:
             | You did not make a negative critique. You completely
             | dismissed the value of coding agents on the basis that the
             | results are not predictable, which is both obvious and
             | doesn't matter in practice. Anyone who has given these
             | tools a chance will quickly realise that 1) they are
             | actually quite predictable in doing what you ask them to,
             | and 2) them being non-deterministic does not at all negate
             | their value. This is why people can immediately tell you
             | haven't used these tools, because your argument as to why
             | they're useless is so elementary.
        
           | dang wrote:
           | Please don't respond to a bad comment by breaking the site
           | guidelines yourself. That only makes things worse.
           | 
           | https://news.ycombinator.com/newsguidelines.html
        
         | senordevnyc wrote:
         | This comment is legitimately hilarious to me. I thought it was
         | satire at first. The list of what has happened in this field in
         | the last twelve months is _staggering_ to me, while you write
         | it off as essentially nothing.
         | 
         | Different strokes, but I'm getting so much more done and mostly
         | enjoying it. Can't wait to see what 2026 holds!
        
           | ronsor wrote:
           | People who dislike LLMs are generally insistent that they're
           | useless for everything and have infinitely negative value,
           | regardless of facts they're presented with.
           | 
           | Anyone that believes that they are completely useless is just
           | as deluded as anyone that believes they're going to bring an
           | AGI utopia next week.
        
         | n2d4 wrote:
         | This is extremely dismissive. Claude Code helps me make a
         | majority of changes to our codebase now, particularly small
         | ones, and is an insane efficiency boost. You may not have the
         | same experience for one reason or another, but plenty of devs
         | do, so "nothing happened" is absolutely wrong.
         | 
         | 2024 was a lot of talk, a lot of "AI could hypothetically do
         | this and that". 2025 was the year where it genuinely started to
         | enter people's workflows. Not everything we've been told would
         | happen has happened (I still make my own presentations and
         | write my own emails) but coding agents certainly have!
        
           | bandrami wrote:
           | Did you ship more in 2025 than in 2024?
        
             | wickedsight wrote:
             | I definitely did.
        
             | GCUMstlyHarmls wrote:
             | Shipping in 2025:
             | https://x.com/trq212/status/2001848726395269619
        
             | DANmode wrote:
             | I _definitely_ did.
             | 
             | Objectively 0->1 lots of backlog.
        
           | skydhash wrote:
           | And this is one of the _vague_ "AI helped me do more".
           | 
           | This is me touting for Emacs
           | 
           |  _Emacs was a great plus for me over the last year. The
           | integration with various tooling with comint (REPL
           | integration), compile (build or report tools), TUI (through
           | eat or ansi-term), gave me a unified experience through the
           | buffer paradigm of emacs. Using the same set of commands
           | boosted my editing process and the easy addition of new
           | commands make it easy to fit my development workflow to the
           | editor._
           | 
           | This is how easy it is to write a non-vague "tool X helped
           | me" and I'm not even an English native speaker.
        
             | n2d4 wrote:
             | That paragraph could be the truth, or it could be a lie.
             | Maybe Emacs really did make you more efficient, or you made
             | it all up, I don't know. Best I can do is trust you.
             | 
             | If you don't trust me, I can't conclusively convince you
             | that AI makes me more efficient, but if you want I'm happy
             | to hop on a screen-share and elaborate in what ways it has
             | boosted my workflow. I'm offering this because I'm also
             | curious what _your_ work looks like where AI cannot help at
             | all.
             | 
             | E-mail address is on my profile!
        
             | thunky wrote:
             | > This is how easy it is to write a non-vague "tool X
             | helped me" and I'm not even an English native speaker.
             | 
             | Your example is very vague.
             | 
             | See if you can spot the problem in my review of Excel in
             | your style:
             | 
             | "It's great and I like how it's formula paradigm gave me a
             | unified experience. It's table features boosted my science
             | workflows last year".
        
         | dang wrote:
         | Could you please stop posting dismissive, curmudgeonly
         | comments? It's not what this site is for, and destroys what it
         | is for.
         | 
         | We want _curious_ conversation here.
         | 
         | https://news.ycombinator.com/newsguidelines.html
        
           | Madmallard wrote:
           | His comment is far better than the rampant astroturfing from
           | stakeholders going on everywhere on this website that is
           | being mitigated not at all whatsoever. There is a wealth of
           | information present suggesting these things are so bad for
           | everyone in so many ways.
        
             | clawedcod wrote:
             | These people love generated content (like, they'll actually
             | read generated blog post word-for-word and not even be
             | angry; they'll skip a personal email for its machine
             | summary) and they can generate all the content they'd ever
             | want. If they want to take over HN this isn't a battle
             | we're going to win except with aggressive moderation, and
             | we know who feeds the mods.
             | 
             | HN isn't a place for thinking people any more (a long time
             | coming, but you could squint and pretend until recently).
             | Happy new year and adios, thanks for the 100s of accounts
             | dang. Double pinky swear I won't make another.
        
               | dang wrote:
               | Normally everyone who publicly declares they're done with
               | this site, will never make a new account again, etc.
               | etc., either has already made their next account or will
               | do so shortly. HN, for all its eternal decline generating
               | endless complaints, seems to be irresistible to this sort
               | of complainer.
        
             | dang wrote:
             | What are some specific links to the rampant astroturfing
             | that you feel is going on on this website and which
             | https://news.ycombinator.com/item?id=46450296 is better
             | than?
        
               | Madmallard wrote:
               | Let's see, going off of just top-level comments in this
               | thread alone:
               | 
               | didip, timonoko, mark_I_watson, icapybara, _pdp_,
               | agentifysh, sanreau,
               | 
               | There's no way to know if these are genuine thoughts or
               | incentivized compelled speech.
               | 
               | nativeit has a good way of putting it.
               | 
               | Your replies to "anonnon" make me less than hopeful for
               | the future of HN in regards to AI. Seems like this might
               | be trending in the direction of Reddit, where the
               | interests are basically all paid for and imposed rather
               | than being genuine and organic, and dissent is
               | aggressively shut out.
               | 
               | "Curious conversation" does not really apply when it is
               | compelled via monetary interest without any consideration
               | toward potentially serious side effects.
               | 
               | "At least when herding cats, you can be sure that if the
               | cats are hungry, they will try to get where the food is."
               | This part of the guy's comment is actually funny and apt.
               | Somehow that escaped you when you wrote your threat
               | reply. That makes me wonder how mind-controlled you are.
               | 
               | "yupyupyups" has a small summary of some of the
               | negatives, yet is being flagged. "techpression" similarly
               | does, though is a bit more negative in his remarks. Also
               | being flagged.
               | 
               | So the whole thread reads like this: 1.) talking about
               | benefits? bubble to the top 2.) criticize? Either
               | threatened by Dang or flagged to the bottom
               | 
               | Sounds a whole lot like compelled speech to me. Sounds a
               | whole lot like mind-control.
               | 
               | It's pretty sad to see really.
               | 
               | It might just be your rule system. I personally want to
               | see criticism. I don't have the sensitivity you have
               | toward personal attacks or what you "deem" personal
               | attacks when it is text on-screen. I don't care. I want
               | to see what useful information might come out of it. I
               | think your policing just makes everything worse to be
               | honest. The thread will just die out in a day anyway.
               | 
               | I think I have criticized it in the past and you or some
               | other staff said that it's a slippery slope toward
               | useless aggressive banter that derails topics, but I
               | don't know. I really don't agree with it. That's just my
               | life experience.
               | 
               | Reddit is kind of like this. And it's basically turned
               | into imposed topics rather than organic topics with
               | massive amounts of echo-chambering in each delusional
               | sub-reddit. Anything remotely against the grain is
               | harshly culled as soon as possible. You can only imagine
               | what the back-end looks like for that kind of thing.
               | Money being involved at many steps is guaranteed.
               | 
               | And yeah as another commenter pointed out, this one guy's
               | blog being at the top of hacker news every time is
               | potentially suspicious as well.
               | 
               | I think I originally came to this place more than Reddit
               | 10+ years ago because yeah it felt like people just
               | excited and curious about their tech topics and it didn't
               | feel like it was being rampantly policed or pushing a
               | political agenda etc. I guess I should just not
               | participate in these threads because the topic is tired
               | on me at this point.
               | 
               | Wait I just read your user page and this is actually
               | hilarious:
               | 
               | "Conflict is essential to human life, whether between
               | different aspects of oneself, between oneself and the
               | environment, between different individuals or between
               | different groups. It follows that the aim of healthy
               | living is not the direct elimination of conflict, which
               | is possible only by forcible suppression of one or other
               | of its antagonistic components, but the toleration of it
               | --the capacity to bear the tensions of doubt and of
               | unsatisfied need and the willingness to hold judgement in
               | suspense until finer and finer solutions can be
               | discovered which integrate more and more the claims of
               | both sides. It is the psychologist's job to make possible
               | the acceptance of such an idea so that the richness of
               | the varieties of experience, whether within the unit of
               | the single personality or in the wider unit of the group,
               | can come to expression."
               | 
               | Marion Milner, 'The Toleration of Conflict', Occupational
               | Psychology, 17, 1, January 1943
               | 
               | This made me immediately and uncontrollably guffaw.
        
       | aussieguy1234 wrote:
       | > The year of YOLO and the Normalization of Deviance #
       | 
       | On this including AI agents deleting home folders, I was able to
       | run agents in Firejail by isolating vscode (Most of my agents are
       | vscode based ones, like Kilo Code).
       | 
       | I wrote a little guide on how I did it
       | https://softwareengineeringstandard.com/2025/12/15/ai-agents...
       | 
       | Took a bit of tweaking, vscode crashing a bunch of times with not
       | being able to read its config files, but I got there in the end.
       | Now it can only write to my projects folder. All of my projects
       | are backed up in git.
        
         | NitpickLawyer wrote:
         | I have a bunch of tabs opened on this exact topic, so thank you
         | for sharing. So far I've been using devcontainers w/ vscode,
         | and mostly having a blast with it. It is a bit awkward since
         | some extensions need to be installed in the remote env, but
         | they seem to play nicely after you have it setup, and the keys
         | and stuff get populated so things like kilocode, cline, roo
         | work fine.
        
       | agentifysh wrote:
       | What an amazing progress in just short time. The future is
       | bright! Happy New Year y'all!
        
       | sho_hn wrote:
       | Not in this review: Also the record year in intelligent systems
       | aiding in and prompting human users into fatal self-harm.
       | 
       | Will 2026 fare better?
        
         | simonw wrote:
         | I really hope so.
         | 
         | The big labs are (mostly) investing a lot of resources into
         | reducing the chance their models will trigger self-harm and AI
         | psychosis and suchlike. See the GPT-4o retirement (and
         | resulting backlash) for an example of that.
         | 
         | But the number of users is exploding too. If they make things
         | 5x less likely to happen but sign up 10x more people it won't
         | be good on that front.
        
           | Nuzzerino wrote:
           | How does a model "trigger" self-harm? Surely it doesn't
           | catalyze the dissatisfaction with the human condition,
           | leading to it. There's no reliable data that can drive
           | meaningful improvement there, and so it is merely an
           | appeasement op.
           | 
           | Same thing with "psychosis", which is a manufactured moral
           | panic crisis.
           | 
           | If the AI companies really wanted to reduce actual self harm
           | and psychosis, maybe they'd stop prioritizing features that
           | lead to mass unemployment for certain professions. One of the
           | guys in the NYT article for AI psychosis had a successful
           | career before the economy went to shit. The LLM didn't create
           | those conditions, bad policies did.
           | 
           | It's time to stop parroting slurs like that.
        
         | measurablefunc wrote:
         | The people working on this stuff have convinced themselves
         | they're on a religious quest so it's not going to get better:
         | https://x.com/RobertFreundLaw/status/2006111090539687956
        
         | andai wrote:
         | Also essential self-fulfilment.
         | 
         | But that one doesn't make headlines ;)
        
           | sho_hn wrote:
           | Sure -- but that's fair game in engineering. I work on cars.
           | If we kill people with safety faults I expect it to make more
           | headlines than all the fun roadtrips.
           | 
           | What I find interesting with chat bots is that they're "web
           | apps" so to speak, but with safety engineering aspects that
           | type of developer is typically not exposed to or familiar
           | with.
        
             | simonw wrote:
             | One of the tough problems here is privacy. AI labs really
             | don't want to be in the habit of actively monitoring
             | people's conversations with their bots, but they also need
             | to prevent bad situations from arising and getting worse.
        
               | walt_grata wrote:
               | Until AI labs have the equivalent of an SLA for giving
               | accurate and helpful responses it don't get better.
               | They've not even able to measure if the agents work
               | correctly and consistently.
        
       | websiteapi wrote:
       | I'm curious how all of the progress will be seen if it does
       | indeed result in mass unemployment (but not eradication) of
       | professional software engineers.
        
         | simonw wrote:
         | I nearly added a section about that. I wanted to contrast the
         | thing where many companies are reducing junior engineering
         | hires with the thing where Cloudflare and Shopify are hiring
         | 1,000+ interns. I ran out of time and hadn't figured out a good
         | way to frame it though so I dropped it.
        
         | ori_b wrote:
         | My prediction: If we can successfully get rid of most software
         | engineers, we can get rid of most knowledge work. Given the
         | state of robotics, manual labor is likely to outlive
         | intellectual labor.
        
           | beardedwizard wrote:
           | "Given the state of robotics" reminds me a lot of what was
           | said about llms and image/video models over the past 3 years.
           | Considering how much llms improved, how long can robotics be
           | in this state?
           | 
           | I have to think 3 years from now we will be having the same
           | conversation about robots doing real physical labor.
           | 
           | "This is the worst they will ever be" feels more apt.
        
             | chii wrote:
             | but robotics had the means to do majority of the physical
             | labour already - it's just not worth the money to replace
             | humans, as human labour is cheap (and flexible - more than
             | robots).
             | 
             | With knowledge work being less high-paying, physical labour
             | supply should increase as well, which drops their price.
             | This means it's actually less likely that the advent of LLM
             | will make physical labour more automated.
        
             | Davidzheng wrote:
             | Robotics is coming FAST. Faster than LLM progress in my
             | opinion.
        
               | wh0knows wrote:
               | Curious if you have any links about the rapid progression
               | of robotics (as someone who is not educated on the
               | topic).
               | 
               | It was my feeling with robotics that the more challenging
               | aspect will be making them economically viable rather
               | than simply the challenge of the task itself.
        
               | beardedwizard wrote:
               | I mentioned military in my reply to the sibling comment -
               | that is the most ready example. What anduril and others
               | are doing today may be sloppy, but it's moving very
               | quickly.
        
               | throw1235435 wrote:
               | The question is how rapid the adoption is. The price of
               | failure in the real world is much higher ($$$,
               | environmental, physical risks) vs just
               | "rebuild/regenerate" in the digital realm.
        
               | beardedwizard wrote:
               | Military adoption is probably a decent proxy indicator -
               | and they are ready to hand the kill switch to autonomous
               | robots
        
               | throw1235435 wrote:
               | Maybe. There the cost of failure again is low. Its easier
               | to destroy than to create. Economic disruption to workers
               | will take a bit longer I think.
               | 
               | Don't get me wrong; I hope that we do see it in physical
               | work as well. There is more value to society there; and
               | consists of work that is risky and/or hard to do - and is
               | usually needed (food, shelter, etc). It also means that
               | the disruption is an "everyone" problem rather than
               | something that just affects those "intellectual" types.
        
           | BobbyJo wrote:
           | I would have agreed with this a few months ago, but something
           | Ive learned is that the ability to verify an LLMs output is
           | paramount to its value. In software, you can review its
           | output, add tests, on top of other adversarial techniques to
           | verify the output immediately after generation.
           | 
           | With most other knowledge work, I don't think that is the
           | case. Maybe actuarial or accounting work, but most knowledge
           | work exists at a cross section of function and taste, and the
           | latter isn't an automatically verifiable output.
        
             | throw1235435 wrote:
             | I also believe this - I think it will probably just disrupt
             | software engineering and any other digital medium with mass
             | internet publication (i.e. things RLVR can use). For the
             | short term future it seems to need a lot of data to train
             | on, and no other profession has posted the same amount of
             | verifiable material. The open source altruism has disrupted
             | the profession in the end; just not in the way people first
             | predicted. I don't think it will disrupt most knowledge
             | work for a number of reasons. Most knowledge professions
             | have "credentials' (i.e. gatekeeping) and they can see what
             | is happening to SWE's and are acting accordingly. I'm
             | hearing it firsthand at least locally in things like law,
             | even accounting, etc. Society will ironically respect these
             | professions more for doing so.
             | 
             | Any data, verifiability, rules of thumb, tests, etc are
             | being kept secret. You pay for the result, but don't know
             | the means.
        
               | coffeebeqn wrote:
               | I mean law and accounting usually have a "right" answer
               | that you can verify against. I can see a test data set
               | being built for most professions. I'm sure open source
               | helps with programming data but I doubt that's even the
               | majority of their training. If you have a company like
               | Google you could collect data on decades of software work
               | in all its dimensions from your workforce
        
               | District5524 wrote:
               | It's not about invalidating your conclusion, but I'm not
               | so sure about law having a right answer. At a very basic
               | level, like hypothetical conduct used in basic legal
               | training matrerials or MCQs, or in criminal/civil code
               | based situations in well-abstracting Roman law-based
               | jurisdictions, definitely. But the actual work, at least
               | for most lawyers is to build on many layers of such
               | abstractions to support your/client's viwepoint. And that
               | level is already about persuasion of other people, not
               | having the "right" legal argument or applying the most
               | correct case found. And this part is not documented well,
               | approaches changes a lot, even if law remains the same.
               | Think of family law or law of succession - does not
               | change much over centuries but every day, worldwide,
               | millions of people spend huge amounts of money and energy
               | on finding novel ways to turn those same paragraphs to
               | their advantage and put their "loved" ones and relatives
               | in a worse position.
        
               | throw1235435 wrote:
               | Not really. I used to think more general with the first
               | generation of LLM's but given all progress since o1 is RL
               | based I'm thinking most disruption will happen in open
               | productive domains and not closed domains. Speaking to
               | people in these professions they don't think SWE's have
               | any self respect and so in your example of law:
               | 
               | * Context is debatable/result isn't always clear: The way
               | to interpret that/argue your case is different (i.e. you
               | are paying for a service, not a product)
               | 
               | * Access to vast training data: Its very unlikely that
               | they will train you and give you data to their practice
               | especially as they are already in a union like
               | structure/accreditation. Its like paying for a binary (a
               | non-decompilable one) without source code (the result)
               | rather than the source and the validation the
               | practitioner used to get there.
               | 
               | * Variability of real world actors: There will be novel
               | interpretations that invalidate the previous one as new
               | context comes along.
               | 
               | * Velocity vs ability to make judgement: As a lawyer I
               | prefer to be paid higher for less velocity since it means
               | less judgement/less liability/less risk overall for
               | myself and the industry. Why would I change that even at
               | an individual level? Less problem of the commons here.
               | 
               | * Tolerance to failure is low: You can't iterate, get
               | feedback and try again until "the tests pass" in a court
               | room unlike "code on a text file". You need to have the
               | right argument the first time. AI/ML generally only works
               | where the end cost of failure is low (i.e can try again
               | and again to iron out error terms/hallucinations). Its
               | also why I'm skeptical AI will do much in the real
               | economy even with robots soon - failure has bigger
               | consequences in the real world ($$$, lives, etc).
               | 
               | * Self employment: There is no tension between say Google
               | shareholders and its employees as per your example -
               | especially for professions where you must trade in your
               | own name. Why would I disrupt myself? The cost I charge
               | is my profit.
               | 
               | TL;DR: Gatekeeping, changing context, and arms race
               | behavior between participants/clients. Unfortunately I do
               | think software, art, videos, translation, etc are unique
               | in that there's numerous examples online and has the
               | property "if I don't like it just re-roll" -> to me RLVR
               | isn't that efficient - it needs volumes of data to build
               | its view. Software sadly for us SWE's is the perfect
               | domain for this; and we as practitioners of it made it
               | that way through things like open source, TDD, etc and
               | giving it away free on public platforms in numerous
               | quantities.
        
           | JumpCrisscross wrote:
           | > _If we can successfully get rid of most software engineers,
           | we can get rid of most knowledge work_
           | 
           | Software, by its nature, is practically comprehensively
           | digitized, both in its code history as well as requirements.
        
           | 9dev wrote:
           | That's the deep irony of technology IMHO, that innovation
           | follows Conway's law on a meta layer: White collar workers
           | inevitably shaped high technology after themselves, and
           | instead of finally ridding humanity of hard physical labour--
           | as was the promise of the Industrial Revolution--we imitate
           | artists, scientists, and knowledge workers.
           | 
           | We can now use natural language to instruct computers
           | generate stock photos and illustrations that would take a
           | professional artist a few years ago, discover new molecule
           | shapes, beat the best Go players, build the code for entire
           | applications, or write documents of various shapes and
           | lengths--but painting a wall? An unsurmountable task that
           | requires a human to execute reliably, not even talking about
           | economics.
        
         | Madmallard wrote:
         | Why would it?
         | 
         | The ability to accurately describe what you want with all
         | constraints managed and with proactive design is the actual
         | skill. Not programming. The day PMs can do that and have LLMs
         | that can code to that, is the day software engineers en masse
         | will disappear. But that day is likely never.
         | 
         | The non-technical people I've ever worked for were hopelessly
         | terrible at attention to detail. They're hiring me primarily
         | for that anyway.
        
         | legulere wrote:
         | Even if it will make software engineering drastically more
         | productive, it's questionable that this will lead to
         | unemployment. Efficiency gains translate to lower prices.
         | Sometimes this leads to very few additional demand, as can be
         | seen with masses of typesetters that lost their jobs. Sometimes
         | this leads to a dramatically higher demand like you can see in
         | the classic Jevons paradox examples of coal and light bulbs. I
         | highly suspect software falls in the latter category
        
           | kingstnap wrote:
           | Software demand is philosophically limited by the question of
           | "What can your computer do for you?"
           | 
           | You can describe that somewhat formally as:
           | 
           | {What your computer can do} intersect {What you want done
           | (consciously or otherwise)}
           | 
           | Well a computer can technically calculate any computuable
           | task that fits in bounded memory, that is an enormous set so
           | its real limitations are its interfaces. In which case it can
           | send packets, make noises, and display images.
           | 
           | How many human desires are things that can be solved with
           | making noises, displaying images, and sending packets? Turns
           | out quite a few but its not everything.
           | 
           | Basically I'm saying we should hope more sorts of physical
           | interfaces come around (like VR and Robotics) so we cover
           | more human desires. Robotics is a really general physical
           | interface (like how ip packets are an extremely general
           | interface) so its pretty promising if it pans out.
           | 
           | Personally, I find it very hard to even articulate what
           | desires I have. I have this feeling that I might be
           | substantially happier if I was just sitting around a campfire
           | eating food and chatting with people instead of enjoying
           | whatever infinite stuff a super intelligent computer and
           | robots could do for me. At least some of the time.
        
         | fullstackchris wrote:
         | This overly discussed thesis is already laughable - decent LLMs
         | have been out for 3 years now and unemployment (using US as
         | example) is up around 1% over the same time frame - and even
         | attributing that small percentage change completely to AI is
         | also laughable
        
       | DrewADesign wrote:
       | _You're absolutely right! You astutely observed that 2025 was a
       | year with many LLMs and this was a selection of waypoints,
       | summarized in a helpful timeline._
       | 
       | That's what most non-tech-person's year in LLMs looked like.
       | 
       | Hopefully 2026 will be the year where companies realize that
       | implementing intrusive chatbots can't make better ::waving
       | hands:: ya know... _UX_ or whatever.
       | 
       | For some reason, they think its helpful to distractingly pop up
       | chat windows on their site because their customers need textual
       | kindergarten handholding to ... I don't know... find the ideal
       | pocket comb for their unique pocket/hair situation, or had an
       | unlikely question about that aerosol pan release spray that a
       | chatbot could actually answer. Well, my dog also thinks she's
       | helping me by attacking the vacuum when I'm trying to clean. Both
       | ideas are equally valid.
       | 
       | And spending a bazillion dollars implementing it doesn't mean
       | your customers won't hate it. And forcing your customers into
       | pathways they hate because of your sunk costs mindset means it
       | will never stop costing you more money than it makes.
       | 
       | I just hope companies start being honest with themselves about
       | whether or not these things are good, bad, or absolutely abysmal
       | for the customer experience and cut their losses when it makes
       | sense.
        
         | Night_Thastus wrote:
         | They need to be intrusive and shoved in your face. This way,
         | they can say they have a lot of people using them, which is a
         | good and useful metric.
        
         | ronsor wrote:
         | > For some reason, they think its helpful to distractingly pop
         | up chat windows on their site...
         | 
         | Companies have been doing this "live support" nonsense far
         | longer than LLMs have been popular.
        
           | DrewADesign wrote:
           | There was also source point pollution before the Industrial
           | Revolution. Useless, forced, irritating chat was _'nowhere
           | close'_ to as aggressive or pervasive as it is now. It used
           | to be a niche feature of some CRMs and now it's _everywhere._
           | 
           | I'm on LinkedIn Learning digging into something really
           | technical and practical and it's constantly pushing the chat
           | fly out with useless pre-populated prompts like "what are the
           | main takeaways from this video." And they moved their main
           | page search to a little icon on the title bar and sneakily
           | now what used to be the obvious, primary central search field
           | for years sends a prompt to their fucking chatbot.
        
         | zahlman wrote:
         | As much as I side with you on this one, I really don't think
         | this submission is the right place to rant about it.
        
         | fantasizr wrote:
         | I took the good with the bad: the ai assisted coding tools are
         | a multiplier, google ai overviews in search results are half
         | baked (at best) and often just factually wrong. AI was put in
         | the instagram search bar for no practical purpose etc.
        
       | techpression wrote:
       | Nothing about the severe impact on the environment, and the hand
       | waviness about water usage hurt to read. The referenced post was
       | missing every single point about the issue by making it global
       | instead of local. And as if data center buildouts are properly
       | planned and dimensioned for existing infrastructure...
       | 
       | Add to this that all the hardware is already old and the amount
       | of waste we're producing right now is mind boggling, and for
       | what, fun tools for the use of one?
       | 
       | I don't live in the US, but the amount of tax money being
       | siphoned to a few tech bros should have heads rolling and I
       | really don't want to see it happening in Europe.
       | 
       | But I guess we got a new version number on a few models and some
       | blown up benchmarks so that's good, oh and of course the svg
       | images we will never use for anything.
        
         | simonw wrote:
         | "Nothing about the severe impact on the environment"
         | 
         | I literally said:
         | 
         | "AI data centers continue to burn vast amounts of energy and
         | the arms race to build them continues to accelerate in a way
         | that feels unsustainable."
         | 
         | AND I linked to my coverage from last year, which is still true
         | today (hence why I felt no need to update it):
         | https://simonwillison.net/2024/Dec/31/llms-in-2024/#the-envi...
        
       | smileson2 wrote:
       | forgot to mention the first murder-suicide instigated by chatgpt
        
         | DANmode wrote:
         | These are _his_ highlights as a killer blogger,
         | 
         | not _AI's_ highlights.
         | 
         | Easy with the hot take.
        
       | didip wrote:
       | Indeed. I don't understand why Hacker News is so dismissive about
       | the coming of LLMs, maybe HN readers are going through 5 stages
       | of grief?
       | 
       | But LLM is certainly a game changer, I can see it delivering
       | impact bigger than the internet itself. Both require a lot of
       | investments.
        
         | cebert wrote:
         | Many people feel threatened by the rapid advancements in LLMs,
         | fearing that their skills may become obsolete, and in turn act
         | irrationally. To navigate this change effectively, we must keep
         | open minds, keep adaptable, and embrace continuous learning.
        
           | nickphx wrote:
           | rapid advancements in what? hallucinations..? FOMO marketing?
           | certainly nothing productive.
        
           | chii wrote:
           | > in turn act irrationally
           | 
           | it isn't irrational to act in self-interest. If LLM threatens
           | someone's livelihood, it matters not that it helps humanity
           | overall one bit - they will oppose it. I don't blame them.
           | But i also hope that they cannot succeed in opposing it.
        
             | Davidzheng wrote:
             | It's irrational to genuinely hold false beliefs about
             | capabilities of LLMs. But at this point I assume around
             | half of the skeptics are emotionally motivated anyway.
        
               | jdhsgsvsbzbd wrote:
               | As opposed to having skin in the game for llms and are
               | blind to their flaws???
               | 
               | I'd assume that around half of the optimists are
               | emotionally motivated this way.
        
           | rgoulter wrote:
           | Many comments discussing LLMs involve emotions, sure. :)
           | Including, obviously, comments in favour of LLMs.
           | 
           | But most discussion I see is vague and without specificity
           | and without nuance.
           | 
           | Recognising the shortcomings of LLMs makes comments praising
           | LLMs that much more believable; and recognising the benefits
           | of LLMs makes comments criticising LLMs more believable.
           | 
           | I'd completely believe anyone who says they've found the LLM
           | very helpful at greenfield frontend tasks, and I'd believe
           | someone who found the LLM unable to carry out subtle
           | refactors on an old codebase in a language that's not Python
           | or JavaScript.
        
           | reppap wrote:
           | I'm not threatened by LLMs taking my job as much as they are
           | taking away my sanity. Every time I tell someone no and they
           | come back to me with a "but copilot said.." it's followed by
           | something entirely incorrect it makes me want to
           | autodefenestrate.
        
             | callc wrote:
             | I am happy "autodefenestrate" is the first new word I
             | learned in 2026. Thank you.
             | 
             | Autodefenestrate - To eject or hurl oneself from a window,
             | especially lethally
        
         | snigsnog wrote:
         | The internet and smartphones were immediately useful in a
         | million different ways for almost every person. AI is not even
         | close to that level. Very to somewhat useful in some fields
         | (like programming) but the average person will easily be able
         | to go through their day without using AI.
         | 
         | The most wide-appeal possibility is people loving 100%-AI-slop
         | entertainment like that AI Instagram Reels product. Maybe I'm
         | just too disconnected with normies but I don't see this taking
         | off. Fun as a novelty like those Ring cam vids but I would
         | never spend all day watching AI generated media.
        
           | JumpCrisscross wrote:
           | > _AI is not even close to that level_
           | 
           | Kagi's Research Assistant is pretty damn useful, particularly
           | when I can have it poll different models. I remember when the
           | first iPhone lacked copy-paste. This feels similar.
           | 
           | (And I don't think we're heading towards AGI.)
        
           | SgtBastard wrote:
           | ... the internet was not immediately useful in a million
           | different ways for almost every person.
           | 
           | Even if you skip ARPAnet, you're forgetting the Gopher days
           | and even if you jump straight to WWW+email==the internet,
           | you're forgetting the mosaic days.
           | 
           | The applications that became useful to the masses emerged a
           | decade+ after the public internet and even then, it took 2+
           | decades to reach anything approaching saturation.
           | 
           | Your dismissal is not likely to age well, for similar
           | reasons.
        
             | chii wrote:
             | the "usefulness" excuse is irrelevant, and the claim that
             | phones/internet is "immediately useful" is just a post hoc
             | rationalization. It's basically trying to find a reasonable
             | reason why opposition to AI is valid, and is not in self-
             | interest.
             | 
             | The opposition to AI is from people who feel threatened by
             | it, because it either threatens their livelihood (or
             | family/friends'), and that they feel they are unable to
             | benefit from AI in the same way as they had internet/mobile
             | phones.
        
               | duchef wrote:
               | The usefulness of mobile phones was identifiable
               | immediately and it is absolutely not 'post hoc
               | rationalization'. The issue was the cost - once low cost
               | mobile telephones were produced they almost immediately
               | became ubiquitous (see nokia share price from the release
               | of the nokia 6110 onwards for example).
               | 
               | This barrier does not exist for current AI technologies
               | which are being given away free. Minor thought experiment
               | - just how radical would the uptake of mobile phones have
               | been if they were given away free?
        
               | jfyi wrote:
               | It's only low cost for general usage chat users. If you
               | are using it for anything beyond that, you are paying or
               | sitting in a long queue (likely both).
               | 
               | You may just be a little early to the renaissance. What
               | happens when the models we have today run on a mobile
               | device?
               | 
               | The nokia 6110 was released 15 years after the first
               | commercial cell phone.
        
               | duchef wrote:
               | Yes although even those people paying are likely still
               | being subsidized and not currently paying the full cost.
               | 
               | Interesting thought about current SOTA models running on
               | my mobile device. I've given it some thought and I don't
               | think it would change my life in any way. Can you suggest
               | some way that it would change yours?
        
               | qualifck wrote:
               | Eh, quite the contrary. A lot of anti AI people genuinely
               | wanted to use AI but run into the factual reality of the
               | limitations of the software. It's not that it's going to
               | take my job, it's that I was told it would redefine how I
               | do work and is exponentially improving only to find out
               | that it just kind of sucks and hasn't gotten much better
               | this year.
        
           | staticassertion wrote:
           | > Very to somewhat useful in some fields (like programming)
           | but the average person will easily be able to go through
           | their day without using AI.
           | 
           | I know a lot of "normal" people who have completely replaced
           | their search engine with AI. It's increasingly a staple for
           | people.
           | 
           | Smartphones were absolutely NOT immediately useful in a
           | million different ways for almost every person, that's total
           | revisionist history. I remember when the iPhone came out, it
           | was AT&T only, it did almost nothing useful. Smartphones were
           | a novelty for quite a while.
        
             | brabel wrote:
             | I agree with most points but as a tech enthusiast, I was
             | using a smart phone years before the iPhone, and I could
             | already use the internet, make video calls, email etc
             | around 2005. It was a small flip phone but it was not
             | uncommon for phones to do that already at that time, at
             | least in Australia and parts of Asia (a Singaporean friend
             | told me about the phone).
        
           | nen-nomad wrote:
           | ChatGPT has roughly 800 million weekly active users. Almost
           | everyone around me uses it daily. I think you are
           | underestimating the adoption.
        
             | throw1235435 wrote:
             | How many pay? And out of that how many are willing to pay
             | the amount to at least cover the inference costs (not loss
             | leading?)
             | 
             | Outside the verifiable domains I think the impact is more
             | assistance/augmentation than outright disruption (i.e. a
             | novelty which is still nice). A little tiny bit of value
             | sprinkled over a very large user base but each person
             | deriving little value overall.
             | 
             | Even as they use it as search it is at best an
             | incrementable improvement on what they used to do - not
             | life changing.
        
             | danielbln wrote:
             | Even my mom and aunts are using it frequently for all sorts
             | of things, and it took a long time for them to hop onto
             | internet and smartphones at first.
        
             | mrweasel wrote:
             | The adoption is just so weird to me. I cannot for the life
             | of me get LLM chatbot to work for me. Every time I try I
             | get into an argument with the stupid thing. They are still
             | wrong constantly, and when I'm wrong they won't correct me.
             | 
             | I have great faith in AI in e.g. medical equipment, or
             | otherwise as something built in, working on a single
             | problem in the background, but the chat interface is
             | terrible.
        
             | dragonwriter wrote:
             | "Almost everyone will use it at free or effectively
             | subsidized prices" and "It delivers utility which justifies
             | its variable costs + fixed costs amortized over useful
             | lifetime" are not the same thing, and its not clear how
             | much of the use is tied to novelty such that if new and
             | progressively more expensive to train releases at a regular
             | cadence dropped off, usage, even at subsidized prices,
             | would, too.
        
             | arctic-true wrote:
             | Usage plunges on the weekends and during the summer,
             | suggesting that a significant portion of users are students
             | using ChatGPT for free or at heavily subsidized rates to do
             | homework (i.e., extremely basic work that is
             | extraordinarily well-represented in the training data).
             | That usage will almost certainly never be monetizable, and
             | it suggests nothing about the trajectory of the
             | technology's capability or popularity. I suspect ChatGPT,
             | in particular, will see its usage slip considerably as the
             | education system (hopefully) adapts.
        
               | simonw wrote:
               | The summer slump was a thing in 2023 but apparently
               | didn't repeat in 2024:
               | https://www.similarweb.com/blog/insights/ai-news/chatgpt-
               | bea...
               | 
               | The weekend slumps could equally suggest people are using
               | it at work.
        
               | arctic-true wrote:
               | Interesting, thank you for that. I'd be curious to see
               | the data for 2025. I was basing my take off Google trends
               | data - the kind of person who goes to ChatGPT by googling
               | "chatGPT" seems to be using it less in the summer.
        
           | raincole wrote:
           | The early internet and smartphones (the Japanese ones, not
           | iPhone) were definitely not "immediately" adopted by the
           | mass, unlike LLM.
           | 
           | If "immediate" usefulness is the metric we measure, then the
           | internet and smartphones are pretty insignificant inventions
           | compared to LLM.
           | 
           | (of course it's not a meaningful metric, as there is no clear
           | line between a dumb phone and a smart phone, or a moderately
           | sized language model and a LLM)
        
           | fragmede wrote:
           | > The internet and smartphones were immediately useful in a
           | million different ways for almost every person. AI is not
           | even close to that level.
           | 
           | Those are some very rosy glasses you've got on there. The
           | nascent Internet took forever to catch on. It was for weird
           | nerds at universities and it'll never catch on, but here we
           | are.
        
           | what-the-grump wrote:
           | A year after the iPhone came out... it didn't have an App
           | Store, barely was able to play video, barely had enough power
           | to last a day. You just don't remember or were not around for
           | it.
           | 
           | A year after llms came out... are you kidding me?
           | 
           | Two years?
           | 
           | 10 years?
           | 
           | Today, by adding an MCP server to wrap the same API that's
           | been around forever for some system, makes the users of that
           | system prefer NLI over the gui almost immediately.
        
         | zvolsky wrote:
         | The idea of HN being dismissive of impactful technology is as
         | old as HN. And indeed, the crowd often appears stuck in the
         | past with hindsight. That said, HN discussions aren't
         | homogeneous, and as demonstrated by Karpathy in his recent
         | blogpost "Auto-grading decade-old Hacker News", at least some
         | commenters have impressive foresight:
         | https://karpathy.bearblog.dev/auto-grade-hn/
        
           | brabel wrote:
           | So exactly 10 years ago a lot of people believed that the
           | game Go would not be "conquered" by AI, but after just a few
           | months it was. People will always be skeptical of new things,
           | even people who are in tech, because many hyped things indeed
           | go nowhere... while it may look obvious in hindsight, it's
           | really hard to predict what will and what won't be
           | successful. On the LLM front I personally think it's
           | extremely foolish to still consider LLMs as going nowhere.
           | There's a lot more evidence today of the usefulness of LLMs
           | than there was of DeepMind being able to beat top human
           | players in Go 10 years ago.
        
         | crystal_revenge wrote:
         | > I don't understand why Hacker News is so dismissive about the
         | coming of LLMs
         | 
         | I find LLMs incredibly useful, but if you were following along
         | the last few years the promise was for "exponential progress"
         | with a teaser world destroying super intelligence.
         | 
         | We objectively are not on that path. There is no "coming of
         | LLMs". We might get some incremental improvement, but we're
         | very clearly seeing sigmoid progress.
         | 
         | I can't speak for everyone, but I'm tired of hyperbolic rants
         | that are unquestionably not justified (the nice thing about
         | exponential progress is you don't need to argue about it)
        
           | aoeusnth1 wrote:
           | We're very clearly seeing exponential progress - even above
           | trend, on METR, whose slope keeps getting revised to a higher
           | and higher estimate each time. Explain your perspective on
           | the objective evidence against exponential progress?
        
             | llmslave2 wrote:
             | Pretty neat how this exponential progress hasn't resulted
             | in exponential productivity. Perhaps you could explain your
             | perspective on that?
        
               | viraptor wrote:
               | Writing the code itself was never the main bottleneck.
               | Designing the bigger solution, figuring out tradeoffs,
               | taking to affected teams, etc. takes as much time as it
               | used to. But still, there's definitely a significant
               | improvement in code production part in many areas.
        
               | aoeusnth1 wrote:
               | It has! CLs/engineer increased by 10% this year.
               | 
               | LLMs from late 2024 were nearly worthless as coding
               | agents, so given they have quadrupled in capability since
               | then (exponential growth, btw), it's not surprising to
               | see a modestly positive impact on SWE work.
               | 
               | Also, I'm noticing you're not explaining yourself :)
        
               | llmslave2 wrote:
               | Hey, I'm not the OG commentator, why do I have to explain
               | myself! :)
               | 
               | When Fernando Alonso (best rookie btw) goes from 0-60 in
               | 2.4 seconds in his Aston Martin, is it reasonable to
               | assume he will near the speed of light in 20 seconds?
        
               | lopatin wrote:
               | > Hey, I'm not the OG commentator, why do I have to
               | explain myself! :)
               | 
               | The issue is that you're not acknowledging or replying to
               | people's explanations for _why_ they see this as
               | exponential growth. It's almost as if you skimmed through
               | the meat of the comment and then just re-phrased your
               | original idea.
               | 
               | > When Fernando Alonso (best rookie btw) goes from 0-60
               | in 2.4 seconds in his Aston Martin, is it reasonable to
               | assume he will near the speed of light in 20 seconds?
               | 
               | This comparison doesn't make sense because we know the
               | limits of cars but we don't yet know the limits of LLMs.
               | It's an open question. Whether or not an F1 engine can
               | make it the speed of light in 20 seconds is not an open
               | question.
        
               | llmslave2 wrote:
               | It's not in me to somehow disprove claims of exponential
               | growth when there isn't even evidence provided of it.
               | 
               | My point with the F1 comparison is to say that a short
               | period of rapid improvement doesn't imply exponential
               | growth and it's about as weird to expect that as it is
               | for an f1 car to reach the speed of light. It's possible
               | you know, the regulations are changing for next season -
               | if Leclerc sets a new lap record in Australia by .1 ms we
               | can just assume exponential improvements and surely
               | Ferrari will be lapping the rest of the field by the
               | summer right?
        
               | aoeusnth1 wrote:
               | There is already evidence provided of it! METR time
               | horizons is going up on an exponential trend. This is
               | literally the most famous AI benchmark and already
               | mentioned in this thread.
               | 
               | https://metr.org/blog/2025-03-19-measuring-ai-ability-to-
               | com...
               | 
               | https://metr.org/blog/2025-07-14-how-does-time-horizon-
               | vary-...
        
               | aoeusnth1 wrote:
               | If you're not going to explain yourself, at least stay on
               | topic. We're talking about exponential growth, so address
               | the points I'm making.
        
               | aoeusnth1 wrote:
               | I'm noticing you're not responding to my claim that
               | producivity has been impacted
        
               | Madmallard wrote:
               | LLMs a year ago were more able to do a complex project
               | I've repeatedly tried to do than they are now.
        
               | scotty79 wrote:
               | Try Antigravity with Gemini 3 Pro. Seems very capable to
               | me.
        
               | surajrmal wrote:
               | I think this is happening by raising the floor for job
               | roles which are largely boilerplate work. If you are on
               | the more skilled side or work in more original/ niche
               | areas, AI doesn't really help too much. I've only been
               | able to use AI effectively for scaling refactors, not
               | really much in feature development. It often just slows
               | me down when I try to use it. I don't see this changing
               | any time soon.
        
               | HPMOR wrote:
               | I think this is an open question still and very
               | interesting. Ilya discussed this on the Dwarkesh podcast.
               | But the capabilities of LLMs is clearly exponential and
               | perhaps super exponential. We went from something that
               | could string together incoherent text in 2022 to general
               | models helping people like Terrance Tao and Scott
               | Aaronson write new research papers. LLMs also beat IMO
               | and the ICPC. We have entered the John Henry era for
               | intellectual tasks...
        
               | llmslave2 wrote:
               | > But the capabilities of LLMs is clearly exponential and
               | perhaps super exponential
               | 
               | By what metric?
        
               | utopiah wrote:
               | BS metric... /s
        
               | tsimionescu wrote:
               | > LLMs also beat IMO and the ICPC
               | 
               | Very spurious claims, given that there was no effort made
               | to check whether the IMO or ICPC problems were in the
               | training set or not, or to quantify how far problems in
               | the training set were from the contest problems. IMO
               | problems are supposed to be unique, but since it's not at
               | the frontier of math research, there is no guarantee that
               | the same problem, or something very similar, was not
               | solved in some obscure manual.
        
               | mgfist wrote:
               | Because that requires adoption. Devs on hackernews are
               | already the most up to date folks in the industry and
               | even here adoption of LLMs is incredibly slow. And a lot
               | of the adoption that does happen is still with older tech
               | like ChatGPT or Cursor.
        
               | belmont_sup wrote:
               | What's the newer tech?
        
               | TeodorDyakov wrote:
               | Claude Code With Opus 4.5
        
               | scotty79 wrote:
               | How long before introduction of computers lead to
               | increases in average productivity? How long for the
               | internet? Business is just slow to figure out how to use
               | anything for its benefit, but it eventually gets there.
        
               | fmbb wrote:
               | > How long before introduction of computers lead to
               | increases in average productivity?
               | 
               | I think it never did. Still has not.
               | 
               | https://en.wikipedia.org/wiki/Productivity_paradox
        
               | spectralista wrote:
               | The best example is that even ATM machines didn't reduce
               | bank teller jobs.
               | 
               | Why? Because even the bank teller is doing more than
               | taking and depositing money.
               | 
               | IMO there is an ontological bias that pervades our modern
               | society that confuses the map for the territory and has a
               | highly distorted view of human existence through the lens
               | of engineering.
               | 
               | We don't see anything in this time series, because this
               | time series itself is meaningless nonsense that reflects
               | exactly this special kind of ontological stupidity:
               | 
               | https://fred.stlouisfed.org/series/PRS85006092
               | 
               | As if the sum of human interaction in an economy is some
               | kind of machine that we just need to engineer better
               | parts for and then sum the outputs.
               | 
               | Any non-careerist, thinking person that studies economics
               | would conclude we don't and will probably not have the
               | tools to properly study this subject in our lifetimes.
               | The high dimensional interaction of biology, entropy and
               | time. We have nothing. The career economist is
               | essentially forced to sing for their supper in a type of
               | time series theater. Then there is the method acting of
               | pretending to be surprised when some meaningless
               | reductionist aspect of human interaction isn't reflected
               | in the fake time series.
        
               | barrenko wrote:
               | Sir, we're in a modern economy, we don't ever _ever_ look
               | at productivity graphs (this is not to disparage LLMs,
               | just a comment on productivity in general)
        
           | viraptor wrote:
           | > exponential progress
           | 
           | First you need to define what it means. What's the metric?
           | Otherwise it's very much something you can argue about.
        
             | noodletheworld wrote:
             | > What's the metric?
             | 
             | Language model capability at generating text output.
             | 
             | The model progress this year has been a lot of:
             | 
             | - "We added multimodal"
             | 
             | - "We added a lot of _non AI_ tooling" (ie agents)
             | 
             | - "We put more compute into inference" (ie thinking mode)
             | 
             | So yes, there is still rapid progress, but these ^ make it
             | clear, at least to me, that next gen _models_ are
             | _significantly harder_ to build.
             | 
             | Simultaneously we see a distinct narrowing between players
             | (openai, deepseek, mistral, google, anthropic) in their
             | offerings.
             | 
             | Thats usually a signal that the rate of progress is
             | slowing.
             | 
             | Remind me what was so great about gpt 5? How about gpt4
             | from from gpt 3?
             | 
             | Do you even remember the releases? Yeah. I dont. I had to
             | look it up.
             | 
             | Just another model with more or less the same capabilities.
             | 
             | "Mixed reception"
             | 
             | That is not what exponential progress looks like, _by any
             | measure_.
             | 
             | The progress this year has been in the tooling around the
             | models, smaller faster models with similar capabilities.
             | Multimodal add ons that no one asked for, because its
             | easier to add image and audio processing than improve text
             | handling.
             | 
             | That may still be on a path to AGI, but it not an
             | _exponential_ path to it.
        
               | dragonwriter wrote:
               | > Language model capability at generating text output.
               | 
               | That's not a metric, that's a vague non-operationalized
               | concept, that could be operationalized into an infinite
               | number of different metrics. And an improvement that was
               | linear in one of those possible metrics would be
               | exponential in another one (well, actually, one that is
               | was linear in one would also be linear in an infinite
               | number of others, _as well as_ being exponential in an
               | infinite number of others.
               | 
               | That's why you have to define an actual metric, not
               | simply describe a vague concept of a kind of capacity of
               | interest, before you can meaningfully discuss whether
               | improvement is exponential. Because the answer is
               | necessarily entirely dependent on the specific
               | construction of the metric.
        
               | viraptor wrote:
               | > Language model capability at generating text output.
               | 
               | That's not a quantifiable sentence. Unless you put it in
               | numbers, anyone can argue exponential/not.
               | 
               | > next gen models are significantly harder to build.
               | 
               | That's not how we judge capability progress though.
               | 
               | > Remind me what was so great about gpt 5? How about gpt4
               | from from gpt 3?
               | 
               | > Do you even remember the releases?
               | 
               | At gpt 3 level we could generate some reasonable code
               | blocks / tiny features. (An example shown around at the
               | time was "explain what this function does" for a
               | "fib(n)") At gpt 4, we could build features and tiny
               | apps. At gpt 5, you can often one-shot build whole apps
               | from a vague description. The difference between them is
               | massive for coding capabilities. Sorry, but if you can't
               | remember that massive change... why are you making claims
               | about the progress in capabilities?
               | 
               | > Multimodal add ons that no one asked for
               | 
               | Not only does multimodal input training improve the model
               | overall, it's useful for (for example) feeding back
               | screenshots during development.
        
               | threethirtytwo wrote:
               | I don't think the path was ever exponential but your
               | claim here is almost as if the slow down hit an asymptote
               | like wall.
               | 
               | Most of the improvements are intangible. Can we truly say
               | how much more reliable the models are? We barely have
               | quantitative measurements on this so it's all vibes and
               | feels. We don't even have a baseline metric for what AGI
               | is and we invalidated the Turing test also based on vibes
               | and feels.
               | 
               | So my argument is that part of the slow down is in itself
               | an hallucination because the improvement is not actually
               | measurable or definable outside of vibes.
        
               | aoeusnth1 wrote:
               | > Language model capability at generating text output.
               | 
               | How would you put this on a graph?
        
             | scotty79 wrote:
             | Define it however you like. There's not a single chart you
             | can draw that even begins to look like a signoid.
        
             | nicbou wrote:
             | Time spent being human and enjoying life.
             | 
             | I can't point at many problems it has meaningfully solved
             | for me. I mean real problems , not tasks that I have to do
             | for my employer. It seems like it just made parts of my
             | existence more miserable, poisoned many of the things I
             | love, and generally made the future feel a lot less
             | certain.
        
           | scotty79 wrote:
           | > but we're very clearly seeing sigmoid progress.
           | 
           | Yeah, probably. But no chart actually shows it yet. For now
           | we are firmly in exponential zone of the signoid curve and
           | can't really tell if it's going to end in a year, decade or a
           | century.
        
             | utopiah wrote:
             | Doesn't even matter if the goal is extremely high. Talking
             | about exponential when we clearly see matching energy needs
             | proves there is no way we can maintain that pace without
             | radical (and thus unpredictable) improvements.
             | 
             | My own "feeling" is that it's definitely not exponential
             | but again, doesn't matter if it's unsustainable.
        
           | fullstackchris wrote:
           | I wrote an article complaining about the whole hype over a
           | year ago:
           | 
           | https://chrisfrewin.medium.com/why-llms-will-never-be-
           | agi-70...
           | 
           | Seems to be playing out that way.
        
           | senordevnyc wrote:
           | I've been reading this comment multiple times a week for the
           | last couple years. Constant assertions that we're starting to
           | hit limits, plateau, etc. But a cursory glance at where we
           | are today vs a year ago, let alone two years ago, makes it
           | wildly obvious that this is bullshit. The pace of improvement
           | of both models and tooling has been breathtaking. I could
           | give a shit whether you think it's "exponential", people like
           | you were dismissing all of this years ago, meanwhile I just
           | keep getting more and more productive.
        
             | qualifck wrote:
             | People keep saying stuff like this. That the improvements
             | are so obvious and breathtaking and astronomical and then I
             | go check out the frontier LLMs again and they're maybe a
             | tiny bit better than they were last year but I can't
             | actually be sure bcuz it's hard to tell.
             | 
             | sometimes it seems like people are just living in another
             | timeline.
        
           | aspenmartin wrote:
           | I'm not sure I understand: we are _objectively on that path_
           | -- we are increasing exponentially on a number of metrics
           | that may be imperfect but seem to paint a pretty consistent
           | picture. Scaling laws are exponential. METR's time horizon
           | benchmark is exponential. Lots of performance measures are
           | exponential, so why do you say we're objectively not on that
           | path?
           | 
           | > We might get some incremental improvement, but we're very
           | clearly seeing sigmoid progress.
           | 
           | again, if it is "very clear" can you point to some concrete
           | examples to illustrate what you mean?
           | 
           | > I can't speak for everyone, but I'm tired of hyperbolic
           | rants that are unquestionably not justified (the nice thing
           | about exponential progress is you don't need to argue about
           | it)
           | 
           | OK but what specifically do you have an issue with here?
        
         | Night_Thastus wrote:
         | LLMs hold _some_ real utility. But that real utility is buried
         | under a mountain of fake hype and over-promises to keep
         | shareholder value high.
         | 
         | LLMs have real limitations that aren't going away any time soon
         | - not until we move to a new technology fundamentally different
         | and separate from them - sharing almost nothing in common.
         | There's a lot of 'progress-washing' going on where people claim
         | that these shortfalls will magically disappear if we throw
         | enough data and compute at it when they clearly will not.
        
           | Gigachad wrote:
           | Pretty much. What actually exists is very impressive. But
           | what was promised and marketed has not been delivered.
        
             | rustystump wrote:
             | Markets never deliver. That isnt new, i do think llms are
             | not far off from google in terms of impact.
             | 
             | Search, as of today, is inferior to frontier models as a
             | product. However, best case still misses expected returns
             | by miles which is where the growsing comes from.
             | 
             | Generative art/ai is still up in the air for staying power
             | but id predict it isnt going away.
        
             | visarga wrote:
             | I think the missing ingredient is not something the LLMs
             | lack, but something we as developers don't do - we need to
             | constrain, channel, and guide agents by creating reactive
             | test environments around them. Not vibes, but hard tests,
             | they are the missing ingredient to coding agents. You can
             | even use AI to write most of these tests but the end result
             | depends on how well you structured your code to be
             | testable.
             | 
             | If you inherit 9000 tests from an existing project you can
             | vibe code a replacement on your phone in a holiday, like
             | Simon Willison's JustHTML port. We are moving from agents
             | semi-randomly flailing around to constraint satisfaction.
        
             | coffeebeqn wrote:
             | Yes and most of the investment has been kind of post-GPT4
             | betting that things will get exponentially more impressive
        
             | baq wrote:
             | I find opus 4.5 and gpt 5.2 mind blowing more often than I
             | find them dumb as rocks. I don't listen to or read any
             | marketing material, I just use the tools. I couldn't care
             | less about what the promises are, what I have now available
             | to me is fundamentally different from what I had in August
             | and it changed completely how I work.
        
         | probably_wrong wrote:
         | Speaking for myself: because if the hype were to be believed we
         | should have no relational databases when there's MongoDB, no
         | need for dollars when there's cryptocoins, all virtual goods
         | would be exclusively sold as NFTs, and we would be all driving
         | self-driving cars by now.
         | 
         | LLMs are being driven mostly by grifters trying to achieve a
         | monopoly before they run out of cash. Under those conditions I
         | find their promises hard to believe. I'll wait until they
         | either go broke or stop losing money left and right, and
         | whatever is left is probably actually useful.
        
           | simonw wrote:
           | The way I've been handling the deafening hype is to focus
           | exclusively on what the models that we have right now can do.
           | 
           | You'll note I don't mention AGI or future model releases in
           | my annual roundup at all. The closest I get to that is
           | expressing doubt that the METR chart will continue at the
           | same rate.
           | 
           | If you focus exclusively on what actually works the LLM space
           | is a whole lot more interesting and less frustrating.
        
             | magicalhippo wrote:
             | > focus exclusively on what the models that we have right
             | now can do
             | 
             | I'm just a casual user, but I've been doing the same and
             | have noticed the sharp improvements of the models we have
             | now vs a year ago. I have OpenAI Business subscription
             | through work, I signed up for Gemini at home after Gemini
             | 3, and I run local models on my GPU.
             | 
             | I just ask them various questions where I know the answer
             | well, or I can easily verify. Rewrite some code, factual
             | stuff etc. I compare and contrast by asking the same
             | question to different models.
             | 
             | AGI? Hell no. Very useful for some things? Hell yes.
        
         | asielen wrote:
         | It is an over correction because of all the empty promises of
         | LLMs. I use Claude and chatgpt daily at work and am amazed at
         | what they can do and how far they can come.
         | 
         | BUT when I hear my executive team talk and see demos of
         | "Agentforce" and every saas company becoming an AI company
         | promising the world, I have to roll my eyes.
         | 
         | The challenge I have with LLMs is they are great at creating
         | first draft shiny objects and the LLMs themselves over promise.
         | I am handed half baked work created by non technical people
         | that now I have to clean up. And they don't realize how much
         | work it is to take something from a 60% solution to a 100%
         | solution because it was so easy for them to get to the 60%.
         | 
         | Amazing, game changing tools in the right hands but also give
         | people false confidence.
         | 
         | Not that they are not also useful for non-technical people but
         | I have had to spend a ton of time explaining to copywriters on
         | the marketing team that they shouldn't paste their credentials
         | into the chat even if it tells them to and their vibe coded app
         | is a security nightmare.
        
           | semilin wrote:
           | This seems like the right take. The claims of the imminence
           | of AGI are exhausting and to me appear dissonant with
           | reality. I've tried gemini-cli and Claude Code and while
           | they're both genuinely quite impressive, they absolutely
           | suffer from a kind of prototype syndrome. While I could learn
           | to use these tools effectively for large-scale projects, I
           | still at present feel more comfortable writing such things by
           | hand.
           | 
           | The NVIDIA CEO says people should stop learning to code. Now
           | if LLMs will really end up as reliable as compilers, such
           | that they can write code that's better and faster than I can
           | 99% of the time, then he might be right. As things stand now,
           | that reality seems far-fetched. To claim that they're useless
           | because this reality has not yet been achieved would be
           | silly, but not more silly than claiming programming is a dead
           | art.
        
         | vunderba wrote:
         | _> I don 't understand why Hacker News is so dismissive about
         | the coming of LLMs._
         | 
         | Eh. I wouldn't be so quick to speak for the entirety of HN.
         | Several articles related to LLMs easily hit the front page
         | every single day, so clearly there are plenty of HN users
         | upvoting them.
         | 
         | I think you're just reading too much into what is more likely
         | classic HN cynicism and/or fatigue.
        
           | ewoodrich wrote:
           | Exactly. There was a stretch of 6 months or so right after
           | ChatGPT was released where approximately 50% of front page
           | posts at any given time were related to LLMs. And these days
           | every other Show HN is some kind of agentic dev tool and
           | Anthropic/OpenAI announcements routinely get 500+ comments in
           | a matter of hours.
        
           | utopiah wrote:
           | It's because both "side" tries to re-adjust.
           | 
           | When an "AI skeptic" sees a very positive AI comment, they
           | try to argue that it is indeed interesting but nowhere near
           | close to AI/AGI/ASI or whatever the hype at the moment uses.
           | 
           | When an "AI optimistic" sees a very negative AI comment, they
           | try to list all the amazing things they have done that they
           | were convinced was until then impossible.
        
         | viraptor wrote:
         | Based on quite a few comments recently, it also looks like many
         | have tried LLMs in the past, but haven't seriously revisited
         | either the modern or more expensive models. And I get it. Not
         | everyone wants to keep up to date every month, or burn cash on
         | experiments. But at the same time, people seem to have opinions
         | formed in 2024. (Especially if they talk about just
         | hallucinations and broken code - tell the agent to search for
         | docs and fix stuff) I'd really like to give them Opus 4.5 as an
         | agent to refresh their views. There's lots to complain about,
         | but the world has moved on significantly.
        
           | mirsadm wrote:
           | This has been the argument since day one. You just have to
           | try the latest model, that's where you went wrong. For the
           | record I use Claude Code quite a bit and I can't see much
           | meaningful improvements from the last few models. It is a
           | useful tool but it's shortcomings are very obvious.
        
           | techpression wrote:
           | Just last week Opus 4.5 decided that the way to fix a test
           | was to change the code so that everything else but the test
           | broke.
           | 
           | When people say "fix stuff" I always wonder if it actually
           | means fix, or just make it look like it works (which is
           | extremely common in software, LLM or not).
        
             | simonw wrote:
             | What did Opus do when you told it that it shouldn't have
             | done that?
        
               | layer8 wrote:
               | It apologized. ;)
        
             | viraptor wrote:
             | Sure, I get an occasional bad result from Opus - then I
             | revert and try again, or ask it for a fix. Even with a
             | couple of restarts, it's going to be faster than me on
             | average. (And that's ignoring the situations where I have
             | to restart myself)
             | 
             | Basically, you're saying it's not perfect. I don't think
             | anyone is claiming otherwise.
        
               | b3kart wrote:
               | The problem is it's imperfect in very unpredictable ways.
               | Meaning you always need to keep it on a short leash for
               | anything serious, which puts a limit on the productivity
               | boost. And that's fine, but does this match the level of
               | investment and expectations?
        
               | techpression wrote:
               | It's not about being perfect, it's about not being as
               | great as the marketing, and many proponents, claim.
               | 
               | The issue is that there's no common definition of
               | "fixed". "Make it run no matter what" is a more apt
               | description in my experience, which works to a point but
               | then becomes very painful.
        
             | baq wrote:
             | Nice. Did it realize the mistake and corrected it?
        
               | techpression wrote:
               | Nope, I did get a lot of fancy markdown with emojis
               | though so I guess that was a nice tradeoff.
               | 
               | In general, even with access to the entire code base
               | (which is very small), I find the inherent need in the
               | models to satisfy the prompter to be their biggest flaw
               | since it tends to constantly lead down this path. I often
               | have to correct over convoluted SQL too because my
               | problems are simple and the training data seems to favor
               | extremely advanced operations.
        
         | Madmallard wrote:
         | Have you tried using it for anything actually complicated?
         | 
         | Lol. It's worse than nothing at all.
        
           | lukaslalinsky wrote:
           | I think the split between vibe coding and AI-assisted coding
           | will only widen over time. If you ask LLMs to do something
           | complex, they will fail and you waste your time. If you work
           | with them as a peer, and you delegate tasks to them, they
           | will succeed and you save your time.
        
             | watwut wrote:
             | I work with leers by delegating complex task to them while
             | I do other complex tasks.
        
         | hapticmonkey wrote:
         | It's not the technology I'm dismissive about. It's the
         | economics.
         | 
         | 25 years ago I was optimistic about the internet, web sites,
         | video streaming, online social systems. All of that. Look at
         | what we have now. It was a fun ride until it all ended up
         | "enshitified". And it will happen to LLMs, too. Fool me once.
         | 
         | Some developer tools might survive in a useful state on
         | subscriptions. But soon enough the whole A.I. economy will
         | centralise into 2 or 3 major players extracting more and more
         | revenue over time until everyone is sick of them. In fact, this
         | process seems to be happening at a pretty high speed.
         | 
         | Once the users are captured, they'll orient the ad-spend market
         | around themselves. And _then_ they'll start taking advantage of
         | the advertisers.
         | 
         | I really hope it doesn't turn out this way. But it's hard to be
         | optimistic.
        
           | Al-Khwarizmi wrote:
           | Contrary to the case for the internet, there is a way out,
           | however - if local, open-source LLMs get good. I really hope
           | they do, because enshittification does seem unavoidable if we
           | depend on commercial offerings.
        
             | ndiddy wrote:
             | Well the "solution" for that will be the GPU vendors
             | focusing solely on B2B sales because it's more profitable,
             | therefore keeping GPUs out of the hands of average
             | consumers. There's leaks suggesting that nVidia will
             | gradually hike the prices of their 5090 cards from $2000 to
             | $5000 due to RAM price increases (
             | https://wccftech.com/geforce-rtx-5090-prices-to-soar-
             | to-5000... ). At that point, why even bother with the R&D
             | for newer consumer cards when you know that barely anyone
             | will be able to afford them?
        
         | tgv wrote:
         | The negatives outweigh the positives, if only because the
         | positives are so small. A bunch of coders making their lives
         | easier doesn't really matter, but pupils and students skipping
         | education does. As a meme said: you had better start eating
         | healthy, because your future doctor vibed his way through med
         | school.
        
         | phatfish wrote:
         | Maybe because the hype for an next gen search engine that can
         | also just make things up when you query it is a bit much?
        
         | jcims wrote:
         | It feels like there are several conversations happening that
         | sound the same but are actually quite different.
         | 
         | One of them is whether or not large models are useful and/or
         | becoming more useful over time. (To me, clearly the answer is
         | yes)
         | 
         | The other is whether or not they live up to the hype. (To me,
         | clearly the answer is no)
         | 
         | There are other skirmishes around capability for novelty, their
         | role in the economy, their impact on human cognition, if/when
         | AGI might happen and the overall impact to the largely tech-
         | oriented community on HN.
        
         | Atomic_Torrfisk wrote:
         | > HN readers are going through 5 stages of grief
         | 
         | So we are just irrational and sour?
        
         | claudiug wrote:
         | because lies. all the people involved in this, the one a C
         | title, tell us about how great is now.
        
       | syndacks wrote:
       | I can't get over the range of sentiment on LLMs. HN leans snake
       | oil, X leans "we're all cooked" --- can it possibly be both? How
       | do other folks make sense of this? I'm not asking for a side,
       | rather understanding the range. Does the range lead you to
       | believe X over Y?
        
         | zahlman wrote:
         | I'm not really convinced that anywhere leans heavily towards
         | anything; it depends which thread you're in etc.
         | 
         | It's polarizing because it represents a more radical shift in
         | expected workflows. Seeing that range of opinions doesn't
         | really give me a reason to update, no. I'm evaluating based on
         | what makes sense when I hear it.
        
         | thisoneisreal wrote:
         | My take (no more informed than anyone else's) is that the range
         | indicates this is a complex phenomenon that people are still
         | making sense of. My suspicion is that something like the
         | following is going on:
         | 
         | 1. LLMs can do some truly impressive things, like taking
         | natural language instructions and producing compiling,
         | functional code as output. This experience is what turns some
         | people into cheerleaders.
         | 
         | 2. Other engineers see that in real production systems, LLMs
         | lack sufficient background / domain knowledge to effectively
         | iterate. They also still produce output, but it's verbose and
         | essentially missing the point of a desired change.
         | 
         | 3. LLMs also can be used by people who are not knowledgeable to
         | "fake it," and produce huge amounts of output that is basically
         | besides-the-point bullshit. This makes those same senior folks
         | very, very resentful, because it wastes a huge amount of their
         | time. This isn't really the fault of the tool, but it's a
         | common way the tool gets used and so it gets tarnished by
         | association.
         | 
         | 4. There is a ridiculous amount of complexity in some of these
         | tools and workflows people are trying to invent, some of which
         | is of questionable value. So aside from the tools themselves
         | people are skeptical of the people trying to become thought
         | leaders in this space and the sort of wild hacks they're coming
         | up with.
         | 
         | 5. There are real macro questions about whether these tools can
         | be made economical to justify whatever value they do produce,
         | and broader questions about their net impact on society.
         | 
         | 6. Last but not least, these tools poke at the edges of
         | "intelligence," the crown jewel of our species and also a big
         | source of status for many people in the engineering community.
         | It's natural that we're a little sensitive about the prospect
         | of anything that might devalue or democratize the concept.
         | 
         | That's my take for what it's worth. It's a complex phenomenon
         | that touches all of these threads, so not only do you see a
         | bunch of different opinions, but the same person might feel
         | bullish about one aspect and bearish about another.
        
         | johnfn wrote:
         | I believe the spikiness in response is because AI itself is
         | spiky - it's incredibly good at some classes of tasks, and
         | remarkably poor at others. People who use it on the spikes are
         | genuinely amazed because of how good it is. This does nothing
         | but annoy the people who use it in the troughs, who become
         | increasingly annoyed that everyone seems to be losing their
         | mind over something that can't even do (whatever).
        
         | llmslave2 wrote:
         | Because there is a wide range of what people consider _good_.
         | If you look at that the people on X consider to be _good_ ,
         | it's not very surprising.
        
         | coffeefirst wrote:
         | Well, this is the internet. Arguing about everything is its
         | favorite pastime.
         | 
         | But generally yes, I think back to
         | Mongo/Node/metaverse/blockchain/IDEs/tablets and pretty much
         | everything has had its boosters and skeptics, this is just
         | more... intense.
         | 
         | Anyway I've decided to believe my own eyes. The crowds say a
         | lot of things. You can try most of it yourself and see what it
         | can and can't do. I make a point to compare notes with
         | competent people who also spent the time trying things. What's
         | interesting is most of their findings are _compatible with_
         | mine, including for folks who don 't work in tech.
         | 
         | Oh, and one thing is for sure: shoving this technology into
         | every single application imaginable is a good way to lose
         | friends and alienate users.
        
         | nstart wrote:
         | The problem with X is that so many people who have no
         | verifiable expertise are super loud in shouting "$INDUSTRY is
         | cooked!!" every time a new model releases. It's exhausting and
         | untrue. The kind of video generation we see might nail realism
         | but if you want to use it to create something meaningful which
         | involves solving a ton of problems and making difficult choices
         | in order to express an idea, you run into the walls of easy
         | work pretty quickly. It's insulting then for professionals to
         | see manga PFPs on X put some slop together and say "movie
         | industry is cooked!". It betrays a lack of understanding of
         | what it takes to make something good and it gives off a vibe of
         | "the loud ones are just trying to force this objectively meh-
         | by-default thing to happen".
         | 
         | The other day there was that dude loudly arguing about some
         | code they wrote/converted even after a woman with significant
         | expertise in the topic pointed out their errors.
         | 
         | Gen AI has its promise. But when you look at the lack of ethics
         | from the industry, the cacophony of voices of non experts
         | screaming "this time it's really doom", and the
         | weariness/wariness that set in during the crypto cycle, it's a
         | natural tendency that people are going to call snake oil.
         | 
         | That said, I think the more accurate representation here is
         | that HN as a whole is calling the hype snake oil. There's very
         | little question anymore about the tools being capable of
         | advanced things. But there is annoyance at proclamations of it
         | being beyond what it really is at the moment which is that it's
         | still at the stage of being an expertise+motivation multiplier
         | for deterministic areas of work. It's not replacing that facet
         | any time soon on its current trend (which could change wildly
         | in 2026). Not until it starts training itself I think. Could be
         | famous last words
        
           | senordevnyc wrote:
           | I'd put more faith in HN's proclamations if it hadn't widely
           | been wrong about AI in 2023, 2024, and now 2025. Watching the
           | tone shift here has been fascinating. As the saying goes, the
           | only thing moving faster than AI advances right now is the
           | speed at which HN haters move the goalposts...
        
             | habinero wrote:
             | Mmm. People who make AI their entire personality and brag
             | that other people are too stupid to see what they see and
             | soon they'll have to see the genius they're denying...does
             | not make me think "oh, wow, what have I missed in AI".
        
             | 3A2D50 wrote:
             | AI has risen the barrier to all but the top and is
             | threatening many peoples' livelihood. It has significantly
             | increase the cost of computer hardware and is projected to
             | increase the cost of electricity. I can definitely see why
             | there is a tone shift! I'm still rooting for AI in general.
             | Would love to see the end of a lot of diseases. I don't
             | think we humans can cure all disease on our own in any of
             | our lifetimes. Of course there all sorts of dystopian
             | consequences that may derive from AI fully comprehending
             | biology. I'm going to continue being naive and hope for the
             | best!
        
         | Madmallard wrote:
         | I use them daily and I actively lose progress on complex
         | problems and save time on simple problems.
        
         | PeterHolzwarth wrote:
         | I think it may be all summed up by Roy Amara's observation that
         | _" We tend to overestimate the effect of a technology in the
         | short run and underestimate the effect in the long run."_
        
           | ManuelKiessling wrote:
           | I think this is the most-fitting one-liner right now.
           | 
           | The arguments going back and forth in these threads are truly
           | a sight to behold. I don't want to lean to any one side, but
           | in 2025 I've begun to respond to everyone who still argues
           | that LLMs are only plagiarism machines, or are only better
           | autocompletes, or are only good at remixing the past: Yes,
           | correct!
           | 
           | And CPUs can only move zeros and ones.
           | 
           | This is likewise a very true statement. But look where having
           | 0s and 1s shuffled around has brought us.
           | 
           | The ripple effects of a machine doing something very simple
           | and near-meaningless, but doing it at high speed and again
           | and again without getting tired, cannot be underestimated.
           | 
           | At the same time, here is Nobel Laureate Robert Solow, who
           | famously, and at the time correctly, stated that "You can see
           | the computer age everywhere but in the productivity
           | statistics."
           | 
           | It took a while, but eventually, his statement became false.
        
           | legulere wrote:
           | The effects might be drastically different from what you
           | would expect though. We've seen this with machine learning/AI
           | again and again that what looks probable to work doesn't work
           | out and unexpected things work.
        
         | xboxnolifes wrote:
         | From my perspective, both show HN and Twitter's normal biases.
         | I view HN as generally leaning toward "new things suck, nothing
         | ever changes", and I view Twitter generally as "Things suck,
         | and everything is getting worse". Both of those align with
         | snake oil and we're all cooked.
        
         | sanderjd wrote:
         | As usual, somewhere in between!
        
         | sph wrote:
         | Truth lies in the middle. Yes LLM are an incredible piece of
         | technology, and yes we are cooked because once again
         | technologists and VC have no idea nor interest in understanding
         | the long-term societal ramifications of technology.
         | 
         | Now we are starting to agree that social media has had
         | disastrous effects that have not fully manifested yet, and in
         | the same breath we accept a piece of technology that promises
         | to replace large parts of society with machines controlled by a
         | few megacorps and we collectively shrug with "eh, we're gonna
         | be alright." I mean, until recently the stated goal was to
         | literally recreate advanced super-intelligence with the same
         | nonchalance one releases a new JavaScript framework unto the
         | world.
         | 
         | I find it utterly maddening how divorced STEM people have
         | become from philosophical and ethical concerns of their work. I
         | blame academia and the education system for creating this
         | massive blind spot, and it is most apparent in echo chambers
         | like HN that are mostly composed of Western-educated
         | programmers with a degree in computer science. At least on X
         | you get, among the lunatics, people that have read more than
         | just books on algorithms and startups.
        
         | senordevnyc wrote:
         | Because it turns out that HN is mostly made up of cranky
         | middle-aged conservatives (small c) who have largely defined
         | themselves around coding, and AI is an existential threat to
         | their core identity.
        
       | anonnon wrote:
       | Why do the mods allow Simon to spam HN with his blogposts and his
       | comments, which he often posts just for the sake of including a
       | link back to his blog? Seriously, go look at his post history and
       | see how often he includes a link to his blog, however
       | tangentially related, when he posts a comment. I actually flagged
       | this submission, which I never do, and encourage others to do
       | likewise.
        
         | simonw wrote:
         | Probably because my content gets a lot more upvotes than it
         | does flags.
         | 
         | If this post was by anyone _other_ than me would you have any
         | problems with its quality?
        
         | dang wrote:
         | He's one of the most valuable writers on LLMs, which are one of
         | the major topics at present. That's not spam.
        
           | anonnon wrote:
           | > He's one of the most valuable writers on LLMs
           | 
           | Is he, really? Most of his blog posts are little more than
           | opportunistic, buttressing commentary on someone else's blog
           | post or article, often with a bit of AI apologia sprinkled in
           | (for example, marginalizing people as paranoid for not taking
           | AI companies at their word that they aren't aggressively
           | scraping websites in violation of robots.txt, or exfiltrating
           | user data in AI-enbaled apps).
           | 
           | EDIT: and why must he link to his blog so often in his
           | comments? How is that not SEO/engagement farming? BTW dang, I
           | wasn't insinuating the mods were in league with him or
           | anything, just that, IMO, he's long past the point at which
           | good faith should no longer be assumed.
        
             | simonw wrote:
             | If you're not assuming good faith what _are_ you assuming
             | here? What 's my motivation?
             | 
             | "buttressing commentary on someone else's blog post"
             | 
             | That's how link blogs work. I wrote more about my approach
             | to that here: https://simonwillison.net/2024/Dec/22/link-
             | blog/
             | 
             | (And yes, there I go again linking to something I've
             | written from a comment. It's entirely relevant to the point
             | I am making here. That's why I have a blog - so I can put
             | useful information in one place.)
             | 
             | I'll also note that I don't ever share links to my link
             | blog posts on Hacker News myself - I don't think they're
             | the right format for a HN post. I can't help if other
             | people share them here:
             | https://news.ycombinator.com/from?site=simonwillison.net
        
               | anonnon wrote:
               | > What's my motivation?
               | 
               | Are you really going to insult my and others'
               | intelligence like this? Directly or indirectly, _your
               | motivation is money._ You already offer monthly
               | subscriptions to your blog, and you 're clearly trying to
               | build a monetizable brand for yourself as a leading
               | authority on AI, especially as it pertains to software
               | development.
        
               | simonw wrote:
               | If my motivation was money I would cash in on the
               | reputation I've already built and go and land a Silicon
               | Valley salaried job somewhere.
               | 
               | Sponsorship from my monthly newsletter doesn't come
               | close.
               | 
               | Seriously, do you have any idea how much money I'm
               | leaving on the table right now NOT having a real job in
               | this space?
               | 
               | Being a blogger is wildly financially irresponsible!
        
             | dang wrote:
             | Please stop.
        
               | th0ma5 wrote:
               | I think when a moderator keeps intervening like this it
               | really does mean that there's something wrong here. I
               | think people would be less mad if you just went ahead and
               | said that you have some kind of special arrangement here
               | with this influencer and post publicly that you like them
               | constantly spamming the site and letting their fans flood
               | the place with deflection and appeals for donations to
               | them. Even YouTube had to add a sponsored post
               | disclaimer.
        
               | dang wrote:
               | There's no special arrangement. The only issue is
               | clarifying what content is welcome vs. unwelcome on HN.
               | simonw's content is obviously welcome, and this ought to
               | be obvious.
               | 
               | > I think people would be less mad
               | 
               | People aren't mad about this. The vast majority of this
               | community values simonw's contributions, which are well
               | within the sweet spot for material on HN. That's why his
               | material gets upvoted, as minimaxir (no friend of
               | astroturfers) has pointed out elsewhere in this thread:
               | https://news.ycombinator.com/item?id=46451969.
        
           | rvz wrote:
           | It _is_ promotional spam.
           | 
           | But given the volume of LLM slop, it was kind of obvious and
           | known that even the moderators now have "favourites" over
           | guidelines.
           | 
           | > Please don't use HN primarily for promotion. It's ok to
           | post your own stuff part of the time, but the primary use of
           | the site should be for curiosity. [0]
           | 
           | The blog itself is clearly used as promotion all the time
           | when the original source(s) are buried deep in the post and
           | almost all of the links link back to his own posts.
           | 
           | This is now a first on HN and a new low for moderators and as
           | admitted have regular promotional favourites on the top of
           | HN.
           | 
           | [0] https://news.ycombinator.com/newsguidelines.html
        
             | minimaxir wrote:
             | The operative word there is "primarily". Simon comments on
             | a variety of topics and has far more interactions that
             | don't link to his blog than do.
             | 
             | Simon's posts are not engagement farming by any definition
             | of the term. He posts good content frequently which is then
             | upvoted by the Hacker News community, which should be the
             | ideal for a Hacker News contributor.
        
               | rvz wrote:
               | Except that the "content" that reaches the top is always
               | about AI / LLMs and nothing else and it is "all the
               | time". Any opportunity to comment, he will link back to
               | his own blog.
               | 
               | He even reposted the same link (which is about AI) with
               | one of his posts when the upvotes fell off and until the
               | second one reached the top, with the _intention_ of
               | promoting his own blog.
               | 
               | Let me simply prove my point to you on how predictable
               | this spam is.
               | 
               | He will do a blog post this month about this paper [0]
               | with an expert analysis by either someone else (or even
               | an LLM) with the primary intention of the blog being used
               | for self promotion with at least one link back to his own
               | blog.
               | 
               | > ...which is then upvoted by the Hacker News community
               | 
               | You don't know that. But what we do know is that even the
               | moderators now have "favourites". Anyone else would be
               | shot down for promotional spam.
               | 
               | [0] https://arxiv.org/abs/2512.24880
        
               | simonw wrote:
               | "He even reposted the same link (which is about AI) with
               | one of his posts when the upvotes fell off"
               | 
               | Where did I do that?
               | 
               | > He will do a blog post this month about this paper [0]
               | 
               | That paper you linked to is a perfect example of where my
               | approach can add value!
               | 
               | Did you read it? Do you understand what it saying? It is
               | _dense_.
               | 
               | I would love to read an evaluation of that paper by
               | someone who can rephrase the core ideas and conversations
               | into a couple of paragraphs that help me understand it,
               | and help me figure out if I should invest further effort
               | in learning more.
               | 
               | I have a whole tag on my blog for that kind of content
               | called paper-review:
               | https://simonwillison.net/tags/paper-review/ - it's my
               | version of the TikTok meme "I read X so you don't have
               | to".
               | 
               | Honestly, your problem doesn't seem to be with me so much
               | as it seems to be with the concept of _blogging in
               | general_.
        
               | th0ma5 wrote:
               | [flagged]
        
               | simonw wrote:
               | I had to paste that into a separate browser window (jwz
               | blocks Hacker News referral traffic) and I cannot figure
               | out how that story is relevant to this conversation. Did
               | you share the right link?
        
               | dang wrote:
               | You've posted over 40 replies hounding this one user whom
               | you seem to be fixated on. We've already asked you to
               | stop (https://news.ycombinator.com/item?id=44726957) but
               | you've continued:
               | 
               | https://news.ycombinator.com/item?id=46409736
               | 
               | https://news.ycombinator.com/item?id=46395646
               | 
               | https://news.ycombinator.com/item?id=46209386
               | 
               | This is obviously an abuse of HN, regardless of who
               | you're being aggressive towards. We ban accounts that
               | keep doing this. If you keep doing it, we will ban you,
               | so no more of this please.
        
         | firexcy wrote:
         | I appreciate his work for being more informative and organized
         | than average AI-related content. Without his blogging, it would
         | be a struggle to navigate the bombastic and narcissistic
         | Twitter/Reddit posts for AI updates. The barrier to entry for
         | AI reporting is so low that you just need to give a bit more
         | care to be distinguished, and he is getting the deserved
         | attention for doing exactly that in a systematical and
         | disciplined manner. (I do believe many on HN are more than
         | capable but not interested in doing the same.) Personally, I
         | sometimes find his posts more congratulatory or trivial than I
         | like, but I have learned to take what I want and ignore what I
         | don't.
        
       | vanderZwan wrote:
       | Speaking of new year and AI: my phone just suggested _" Happy
       | Birthday!"_ as the quick-reply to any _" Happy New Year!"_
       | notification I got in the last hours.
       | 
       | I'm not too worried about my job just yet.
        
         | pants2 wrote:
         | It won't help to point out the worst examples. You're not
         | competing with an outdated Apple LLM running on a phone. You're
         | competing with Anthropic frontier models running on a
         | multimillion dollar rack of servers.
        
           | vanderZwan wrote:
           | Sounds like I'm much more affordable with better ROI
        
         | gverrilla wrote:
         | This year I had a spotify and a youtube thing to "recall my
         | year", and it was abolute garbage (30% truth, to be exact). I
         | think they're doing it more like an exercise to build up
         | systems, infra, processes, people, etc - it's already clear
         | they don't actually care about users.
        
       | ogou wrote:
       | This is a good tooling survey of the past year. I have been
       | watching it as a developer re-entering the job market. The job
       | descriptions closely parallel the timeline used in the post.
       | That's bizarre to me because these approaches are changing so
       | fast. I see jobs for "Skill and Langchain experts with
       | production-grade 0>1 experience. Former founders preferred". That
       | is an expertise that is just a few months old and startups are
       | trying to build whole teams overnight with it. I'm sure January
       | and February will have job postings for whatever gets released
       | that week. It's all so many sand castles.
        
         | weatherlite wrote:
         | > Skill and Langchain experts with production-grade 0>1
         | experience.
         | 
         | Also , it's just normal backend work - calling a bunch of APIs.
         | What am I missing here?
        
           | walthamstow wrote:
           | Buzzwords.
        
           | XenophileJKO wrote:
           | That is like saying training tensorflow models is just
           | calling some APIs.
           | 
           | Actually making a system like this work seems easy, but isn't
           | really.
           | 
           | (Though with the CURRENT generation or two of models it has
           | gotten "pretty easy" I think. Before that, not so much.)
        
             | weatherlite wrote:
             | No idea about training tenserflow models - is it super
             | complex or is it just calling a couple of APIs ? Langchain
             | is literally calling an API. Maybe you need to get good
             | with prompting or whatever, but I don't see where the
             | complexity lies. Please let me know.
        
               | andy99 wrote:
               | Having used both Tensorflow (though I expect they mean
               | PyTorch which is way more popular, and I have also used)
               | and langchain, they are nothing alike.
               | 
               | They he ML frameworks are much closer to implementing the
               | mathematics of neural networks, with some abstractions
               | but much closer to the linear algebra level. It requires
               | an understanding of the underlying theory.
               | 
               | Langchain is a suite of convenience functions for
               | composing prompts to LLMs. I wouldn't consider there to
               | be some real domain knowledge one would need to use it.
               | There is a learning curve but it's about learning the
               | different components rather than learning a whole new
               | academic discipline.
        
               | HarHarVeryFunny wrote:
               | There's a big difference between building an ML framework
               | like Tensorflow or PyTorch (I built a Lua Torch-like one
               | in C++ myself) and just using it to build/train a model.
               | 
               | Building the model may range from very simple if you are
               | just recreating a standard architecture, or be a research
               | endeavor if you are designing something completely new.
               | 
               | The difficulty/complexity of then training the model
               | depends on what it is. For something simple like a CNN
               | for image recognition, it's really just a matter of
               | selecting a few hyperparameters and letting it rip. At
               | the other end of the spectrum you've got LLMs where
               | training (and coping with instabilities) is something of
               | a black art, with RL training completely different from
               | pre-training, and there is also the issue of
               | designing/discovering a pre/mid/post training curriculum.
               | 
               | But anyways, the actual training part can be very simple,
               | not requiring too much knowledge of what's going on under
               | the hood, depending on the model.
        
               | ogou wrote:
               | You're right, none of these new tools are disciplines.
               | They are vendor specific approaches that are very recent.
               | That's part of my overall point. Who is out there with 2+
               | years of very narrow tooling experience at another
               | company at a senior level and is available for a rando
               | startup (or desparate enterprise looking for bolt-on AI
               | features) at a fraction of the pay? Not many, I'm sure.
               | We can level up, do training, and maybe stand up a demo
               | project. But that won't satisfy an ATS scan. It's
               | unrealistic.
        
       | blutoot wrote:
       | I hope 2026 will be the year when software engineers and
       | recruiters will stop the obsession with leetcode and all other
       | forms of competitive programming bullshit
        
       | andrewinardeer wrote:
       | Thank you. Enjoyed this read.
       | 
       | AI slop videos will no doubt get longer and "more realistic" in
       | 2026.
       | 
       | I really hope social media companies plaster a prominent banner
       | over them which screams, "Likely/Made by AI" and give us the
       | option to automatically mute these videos from our timeline. That
       | would be the responsible thing to do. But I can't see Alphabet
       | doing that on YT, xAI doing that on X or Meta doing that on
       | FB/Insta as they all have skin in the video gen game.
        
         | sexy_seedbox wrote:
         | For image generation, it's already _too_ realistic with Z-Image
         | + Custom LoRas + SeedVR2 upscaling.
        
           | hooverd wrote:
           | I do think for the solution of say non-consensual pornography
           | the only solution is incredible violence against people
           | making it.
        
         | compass_copium wrote:
         | >I really hope social media companies plaster a prominent
         | banner over them which screams, "Likely/Made by AI" and give us
         | the option to automatically mute these videos from our
         | timeline.
         | 
         | They should just be deleted. They will not be, because they
         | clearly generate ad revenue.
        
         | cube00 wrote:
         | > social media companies plaster a prominent banner over them
         | 
         | Not going to happen as the social media companies realise they
         | can sell you the AI tools used to post slop back onto the
         | platform.
        
       | compass_copium wrote:
       | >I'm still holding hope that slop won't end up as bad a problem
       | as many people fear.
       | 
       | That's the pure, uncut copium. Meanwhile, in the real world,
       | search on major platforms is so slanted towards slop that people
       | need to specify that they want actual human music:
       | 
       | https://old.reddit.com/r/MusicRecommendations/comments/1pq4f...
        
       | apolloartemis wrote:
       | Thank you for your warning about the normalization of deviance.
       | Do you think there will be an AI agent software worm like
       | NotPetya which will cause a lot of economic damage?
        
         | simonw wrote:
         | I'm expecting something like a malicious prompt injection which
         | steals API keys and crypto wallets and uses additional tricks
         | to spread itself further.
         | 
         | Or targeted prompt injections - like spear phishing attacks -
         | against people with elevated privileges (think root sysadmins)
         | who are known to be using coding agents.
        
       | lukaslalinsky wrote:
       | Speaking of asynchronous agents, what do people use? Claude Code
       | for web is extremely limited, because you have no custom tools.
       | Claude Code in GitHub Actions is vastly more useful, due to the
       | custom environment, but ackward to use interactively. Are there
       | any good alternatives?
        
         | simonw wrote:
         | I use Claude Code for web with an environment allowing full
         | internet access, which means it can install extra tools as and
         | when it needs them. I don't run into limits with it very often.
        
         | jimmySixDOF wrote:
         | Pretty sure next year's wrapup will have "Year of the sub-
         | agent"
        
         | jes5199 wrote:
         | I'm running Claude Code in a tmux on a VPS, and I'm working on
         | setting up a meta-agent who can talk to me over text messages
        
           | absoluteunit1 wrote:
           | Hey - this sounds like really interesting set-up!
           | 
           | Would you be open to providing more details. Would love to
           | hear more, your workflows, etc.
        
         | fullstackchris wrote:
         | I just use a couple of custom MCP tools with the standard
         | claude desktop app:
         | 
         | https://chrisfrew.in/blog/two-of-my-favorite-mcp-tools-i-use...
         | 
         | IMO this is the best balance of getting agentic work done while
         | having immediate access to anything else you may need with your
         | development process.
        
         | ehsanu1 wrote:
         | What exactly do you mean by custom tools here? Just cli tools
         | accessible to the agent?
        
           | lukaslalinsky wrote:
           | Development environment needed to build and test the project.
        
       | lopatin wrote:
       | The "pelicans on a bike" challenge is pretty wide spread now. Are
       | we sure it's still not being trained on?
        
         | simonw wrote:
         | See https://simonwillison.net/2025/nov/13/training-for-
         | pelicans-... (also in the pelicans section of the post).
        
           | lopatin wrote:
           | > All I've ever wanted from life is a genuinely great SVG
           | vector illustration of a pelican riding a bicycle.
           | 
           | :)
        
       | Razengan wrote:
       | My experience with AI so far: It's still far from "butler" level
       | assistance for anything beyond simple tasks.
       | 
       | I posted about my failures to try to get them to review my bank
       | statements [0] and generally got gaslit about how I was doing it
       | wrong, that I if trust them to give them full access to my disk
       | and terminal, they could do it better.
       | 
       | But I mean, at that point, it's still more "manual intelligence"
       | than just telling someone what I want. A human could easily
       | understand it, but AI still takes a lot of wrangling and you
       | still need to think from the "AI's PoV" to get the good results.
       | 
       | [0] https://news.ycombinator.com/item?id=46374935
       | 
       | ----
       | 
       | But enough whining. I _want_ AI to get better so I can be lazier.
       | After trying them for a while, one feature that I think all
       | natural-language As need to have, would be the ability to mark
       | certain sentences as  "Do what I say" (aka Monkey's Paw) and "Do
       | what I mean", like how you wrap phrases in quotes on Google etc
       | to indicate a verbatim search.
       | 
       | So for example I could say "[[I was in Japan from the 5th to
       | 10th]], identify foreign currency transactions on my statement
       | with "POS" etc in the description" then the part in the [[]] (or
       | whatever other marker) would be literal, exactly as written, but
       | the rest of the text would be up to the AI's
       | interpretation/inference so it would also search for ATM
       | withdrawals etc.
       | 
       | Ideally, eventually we should be able to have multiple different
       | AI "personas" akin to different members of household staff: your
       | "chef" would know about your dietary preferences, your "maid"
       | would operate your Roomba, take care of your laundry, your
       | "accountant" would do accounty stuff.. and each of them would
       | only learn about that specific domain of your life: the chef
       | would pick up the times when you get hungry, but it won't know
       | about your finances, and so on. The current "Projects" paradigm
       | is not quite that yet.
        
       | ksec wrote:
       | All these improvement in a single year, 2025. While this may seem
       | obvious to those who follows along the AI / LLM news. It may be
       | worth pointing out again ChatGPT was introduced to us in November
       | 2022.
       | 
       | I still dont believe AGI, ASI or Whatever AI will take over human
       | in short period of time say 10 - 20 years. But it is hard to
       | argue against the value of current AI, which many of the vocal
       | critics on HN seems to have the opinion of. People are willing to
       | pay $200 per month, and it is getting $1B dollar runway
       | _already_.
       | 
       | Being more of a Hardware person, the most interesting part to me
       | is the funding of all the developments of latest hardware. I know
       | this is another topic HN hate because of the DRAM and NAND
       | pricing issue. But it is exciting to see this from a long term
       | view where the pricing are short term pain. Right now the
       | industry is asking, we have together over a trillion dollar to
       | spend on Capex over the next few years and will even borrow more
       | if it needs to be, when can you ship us 16A / 14A / 10A and 8A or
       | 5A, LPDDR6, Higher Capacity DRAM at lower power usage, better
       | packaging, higher speed PCIe or a jump to optical interconnect?
       | Every single part of the hardware stack are being fused with
       | money and demand. The last time we have this was Post-PC /
       | Smartphone era which drove the hardware industry forward for 10 -
       | 15 years. The current AI can at least push hardware for another 5
       | - 6 years while pulling forward tech that was initially 8 - 10
       | years away.
       | 
       | I so wished I brought some Nvidia stock. Again, I guess no one
       | knew AI would be as big as it is today, and it is only just
       | started.
        
         | coffeebeqn wrote:
         | Seems like Nvidia will be focusing on the super beefy GPUs and
         | leaving the consumer market to a smaller player
        
           | _s wrote:
           | AMD owns a lot of the consumer market already; handhelds,
           | consoles, desktop rigs and mobile ... they are not a small
           | player.
        
             | utopiah wrote:
             | They said "smaller" not small.
        
           | Flow wrote:
           | I don't get why Nvidia can't do both? Is it because of the
           | limited production capabilities of the factories?
        
             | ACCount37 wrote:
             | Yes. If you're bottlenecked on silicon and secondaries like
             | memory, why would you want to put more of those resources
             | into lower margin consumer products if you could use those
             | very resources to make and sell more high margin AI
             | accelerators instead?
             | 
             | From a business standpoint, it makes some sense to throttle
             | the gaming supply some. Not to the point of surrendering
             | the market to someone else probably, but to a measurable
             | degree.
        
               | ksec wrote:
               | We will have to wait and see but my bet is that Nvidia
               | will move to Leading Edge node N2 earlier now they have
               | the Margin to work with. Both Hopper and Blackwell were
               | too late in the design cycle. The AI hype and continue to
               | buy the latest and great leaving Gaming at a mainstream
               | node.
               | 
               | Nvidia using Mainstream node has always been the norm
               | considering most Fab capacity always goes to Mobile SoC
               | first. But I expect the internet / gamers will be angry
               | anyway because Nvidia does not provide them with the
               | latest and greatest.
               | 
               | In reality the extra R&D cost for designing with leading
               | edge will be amortised by all the AI order which give
               | Nvidia competitive advantage at the consumer level when
               | they compete. That is assuming there are competition
               | because most recent data have shown Nvidia owning 90%+ of
               | discreet market share, 9% for AMD and 1% for Intel.
        
         | utopiah wrote:
         | > All these improvement in a single year
         | 
         | > hard to argue against the value of current AI
         | 
         | > People are willing to pay $200 per month, and it is getting
         | $1B dollar runway already.
         | 
         | Those are 3 different things. There can be a LOT of fast and
         | significant improvements but still remain extremely far from
         | the actual goal, so far it looks like actually little progress.
         | 
         | People pay for a lot of things, including snake oil, so
         | convincing a lot of people to pay a bit is not in itself a
         | proof of value, especially when some people are basically
         | coerced into this, see how many companies changed their
         | "strategy" to mandating AI usage internally, or integration for
         | a captive audience e.g. Copilot.
         | 
         | Finally yes, $1B is a LOT of money for you and I... but for the
         | largest corporations it's actually not a lot. For reference
         | Google earned that in revenue... per day in 2023. Anyway that's
         | still a big number BUT it still has to be compared with, well
         | how much does OpenAI burn. I don't have any public number on
         | that but I believe the consensus is that it's a lot. So until
         | we know that number we can't talk about an actual runway.
        
           | aspenmartin wrote:
           | > People pay for a lot of things, including snake oil, so
           | convincing a lot of people to pay a bit is not in itself a
           | proof of value
           | 
           | But do you really believe e.g. Claude code is snake oil? I
           | pay $200 / month for Claude, which is something I would have
           | thought monumentally insane maybe 1-2 years ago (e.g. when
           | ChatGPT came out with their premium subscription price I
           | thought that seemed so out of touch). I don't think we would
           | be seeing the subscription rates and the retention numbers if
           | it really was snake oil.
           | 
           | > Finally yes, $1B is a LOT of money for you and I... but for
           | the largest corporations it's actually not a lot. For
           | reference Google earned that in revenue... per day in 2023.
           | Anyway that's still a big number BUT it still has to be
           | compared with, well how much does OpenAI burn. I don't have
           | any public number on that but I believe the consensus is that
           | it's a lot. So until we know that number we can't talk about
           | an actual runway.
           | 
           | this gets brought up a lot but I'm not sure I understand why
           | folks on a forum called YCombinator, a startup accelerator,
           | would make this sound like an obvious sign of charlatanism;
           | operating at a loss is nothing new and anthropic / openAI
           | strategy seems perfectly rational: they are scaling and
           | capturing market share, and TAM is insane.
        
         | chias wrote:
         | These are not all improvements. Listed:
         | 
         | * The year of YOLO and the Normalization of Deviance
         | 
         | * The year that Llama lost its way
         | 
         | * The year of alarmingly AI-enabled browsers
         | 
         | * The year of the lethal trifecta
         | 
         | * The year of slop
         | 
         | * The year that data centers got extremely unpopular
        
           | steveBK123 wrote:
           | > * The year that data centers got extremely unpopular
           | 
           | I was discussing the political angle with a friend recently.
           | I think Big Tech Bro / VC complex has done themselves a big
           | disservice by aligning so tightly with MAGA to the point AI
           | will be a political issue in 2026 & 2028.
           | 
           | Think about the message they've inadvertently created
           | themselves - AI is going to replace jobs, it's pushing
           | electric prices up, we need the government to bail us out AND
           | give us a regulatory light touch.
           | 
           | Super easy campaign for Dems - big tech trumpers are taking
           | your money, your jobs, causing inflation, and now they want
           | bailouts !!
        
           | mbesto wrote:
           | Said differently - the year we start to see all of the
           | externalities of a globally scaled hyped tech trend.
        
           | Y_Y wrote:
           | Not that YOLO, PJ Reddie released that in 2015
        
         | jillesvangurp wrote:
         | 2025 was the year of development tool using AI agents. I think
         | we'll shift attention to non development tool using AI agents.
         | Most business users are still stuck using chat gpt as some kind
         | of grand oracle that will write their email or powerpoint
         | slides. There are bits and pieces of mostly technology demo
         | level solutions but nothing that is widely used like AI coding
         | tools are so far. I don't think this is bottle necked on model
         | quality.
         | 
         | I don't need an AGI. I do need a secretary type agent that
         | deals with all the simple but yet laborious non technical tasks
         | that keep infringing on my quality engineering time. I'm CTO
         | for a small startup and the amount of non technical bullshit
         | that I need to deal with is enormous. Some examples of random
         | crap I deal with: figuring out contracts, their
         | meaning/implication to situations, and deciding on a course of
         | action; Customer offers, price calculations, scraping invoices
         | from emails and online SAAS accounts, formulating detailed
         | replies to customer requests, HR legal work, corporate
         | bureaucracy, financial planning, etc.
         | 
         | A lot of this stuff can be AI assisted (and we get a lot of
         | value out of ai tools for this) but context engineering is
         | taking up a non trivial amount of my time. Also most tools are
         | completely useless at modifying structured documents.
         | Refactoring a big code base, no problem. Adding structured text
         | to an existing structured document, hardest thing ever. The
         | state of the art here is an ff-ing sidebar that will suggest
         | you a markdown formatted text that you might copy/paste. Tool
         | quality is very primitive. And then you find yourself just
         | stripping all formatting and reformatting it manually. Because
         | the tools really suck at this.
        
           | arcatech wrote:
           | > Some examples of random crap I deal with: figuring out
           | contracts, their meaning/implication to situations, and
           | deciding on a course of action
           | 
           | This doesn't sound like bullshit you should hand off to an
           | AI. It sounds like stuff you would care about.
        
             | nrclark wrote:
             | Agree. Even asking it can anchor your thinking.
        
             | jillesvangurp wrote:
             | I do care about it; kind of my duty as a co-founder. Which
             | is why I'm spending double digit percentages of my time
             | doing this stuff. But I absolutely could use some tools to
             | cut down on a lot of the drudgery that is involved with
             | this. And me reading through 40 pages of dense legal German
             | isn't one of my strengths since I 1) do not speak German 2)
             | am not a lawyer and 3) am not necessarily deeply familiar
             | with all the bureaucracy, laws, etc.
             | 
             | But I can ask intelligent questions about that contract
             | from an LLM (in English) and shoot back and forth a few
             | things, come up with some kind of action plan, and then run
             | it by our laywers and other advisors.
             | 
             | That's not some kind of hypothetical thing. That's
             | something that happened multiple times in our company in
             | the last few months. LLMs are very empowering for dealing
             | with this sort of thing. You still need experts for some
             | stuff. But you can do a lot more yourself now. And as we've
             | found out, some of the "experts" that we relied on in the
             | past actually did a pretty shoddy job. A lot of this stuff
             | was about picking apart the mess they made and fixing it.
             | 
             | As soon as you start drafting contracts, it gets a lot
             | harder. I just went through a process like that as well. It
             | involves a lot of manual work that is basically about
             | formatting documents, drafting text, running pdfs and text
             | snippets through chat gpt for feedback, sparring,
             | criticism, etc. and iterating on that. This is not about
             | vibe coding some contract but making sure every letter of a
             | contract is right. That ultimately involves lawyers and
             | negotiating with other stakeholders but it helps if you
             | come prepared with a more or less ready to sign off on
             | document.
             | 
             | It's not about handing stuff off but about making LLMs work
             | for you. Just like with coding tools. I care about code
             | quality as well. But I still use the tools to save me a lot
             | of time.
        
               | simonw wrote:
               | One of the lessons I learned running a startup is that it
               | doesn't matter how good the professionals you hire are
               | for things like legal and accounting, you _still_ need to
               | put work in yourself.
               | 
               | Everyone makes mistakes and misses things, and as the co-
               | founder you have to care more about the details than
               | anyone else does.
               | 
               | I would have _loved_ to have weird-unreliable-paralegal-
               | Claude available back when I was doing that!
        
           | topaztee wrote:
           | `Also most tools are completely useless at modifying
           | structured documents`
           | 
           | we built a tool for this for the life science space and are
           | opening it up to the general public very soon. Email me I can
           | give you access (topaz at vespper dot com)
        
         | pjc50 wrote:
         | Investing a trillion dollars for a revenue of a billion dollars
         | doesn't sound great yet.
        
           | steveBK123 wrote:
           | Indeed, its the old Uber playbook at nearly two extra orders
           | of magnitude.
           | 
           | It is a large enough number to simply run out of private
           | capital to consume before it turns cash flow positive.
           | 
           | Lots of things sell well if sold at such a loss. I'd take a
           | new Ferrari for $2500 if it was on offer.
        
             | derwiki wrote:
             | Uber's playbook worked for Uber
        
             | aoeusnth1 wrote:
             | You say that as if Uber's playbook didn't work. Try this:
             | https://www.google.com/finance/quote/UBER:NYSE
        
             | pjc50 wrote:
             | Did Uber actually do a lot of capital investment? They
             | don't own the cars, for example.
        
               | simonw wrote:
               | I believe they spent a huge amount of money on incentives
               | to help sign up drivers, and discounts to help attract
               | customers.
        
         | ACCount37 wrote:
         | Is the AI progress in 2025 an outstanding breakthrough? Not
         | really. It's impressive but incremental.
         | 
         | Still, the gap between the capabilities of a cutting edge LLM
         | and that of a human is only this wide. There are only this many
         | increments it takes to cross it.
        
         | wpietri wrote:
         | This is not a great argument:
         | 
         | > But it is hard to argue against the value of current AI [...]
         | it is getting $1B dollar runway already.
         | 
         | The psychic services industry makes over $2 billion a year in
         | the US [1], with about a quarter of the population being actual
         | believers. [2].
         | 
         | [1] The https://www.ibisworld.com/united-
         | states/industry/psychic-ser...
         | 
         | [2] https://news.gallup.com/poll/692738/paranormal-phenomena-
         | met...
        
           | apexalpha wrote:
           | What if these provide actual value through placebo-effect?
        
             | recursive wrote:
             | You talking about psychics or LLMs?
        
               | grosswait wrote:
               | Yes
        
             | wpietri wrote:
             | I think we have different definitions of "actual value".
             | But even if I pick the flaccid definition, that isn't proof
             | of value of the thing itself, but of any placebo. In which
             | case we can focus on the cheapest/least harmful placebo.
             | Or, better, solving the underlying problem that the placebo
             | "helps".
        
               | computably wrote:
               | I'll preface by saying I fully agree that psychics aren't
               | providing any non-placebo value to believers, although I
               | think it's fine to provide entertainment for non-
               | believers.
               | 
               | > Or, better, solving the underlying problem that the
               | placebo "helps".
               | 
               | The underlying problems are often a lack of a decent
               | education and a generally difficult/unsatisfying life.
               | Systemic issues which can't be meaningfully "solved"
               | without massive resources and political will.
        
               | jay_kyburz wrote:
               | Actually, I'd go one step further and say they are
               | harmful to everybody else.
               | 
               | It might just be my circles, but I've seen Carl Sagans
               | quote everywhere in the last couple of months.
               | 
               | ""Science is more than a body of knowledge; it is a way
               | of thinking. I have a foreboding of an America in my
               | children's or grandchildren's time--when the United
               | States is a service and information economy; when nearly
               | all the key manufacturing industries have slipped away to
               | other countries; when awesome technological powers are in
               | the hands of a very few, and no one representing the
               | public interest can even grasp the issues; when the
               | people have lost the ability to set their own agendas or
               | knowledgeably question those in authority; when,
               | clutching our crystals and nervously consulting our
               | horoscopes, our critical faculties in decline, unable to
               | distinguish between what feels good and what's true, we
               | slide, almost without noticing, back into superstition
               | and darkness.""
        
           | ctoth wrote:
           | 2022/2023: "It hallucinates, it's a toy, it's useless."
           | 
           | 2024/2025: "Okay, it works, but it produces security
           | vulnerabilities and makes junior devs lazy."
           | 
           | 2026 (Current): "It is literally the same thing as a psychic
           | scam."
           | 
           | Can we at least make predictions for 2027? What shall the
           | cope be then! Lemme go ask my psychic.
        
             | bopbopbop7 wrote:
             | 2022/2023: "Next year software engineering is dead"
             | 
             | 2024: "Now this time for real, software engineering is dead
             | in 6 months, AI CEO said so"
             | 
             | 2025: "I know a guy who knows a guy who built a startup
             | with an LLM in 3 hours, software engineering is dead next
             | year!"
             | 
             | What will be the cope for you this year?
        
               | aspenmartin wrote:
               | The cope + disappointment will be knowing that a large
               | population of HN users will paint a weird alternative
               | reality. There are a multitude of messages about AI that
               | are out there, some are highly detached from reality (on
               | the optimistic and pessimistic side). And then there is
               | the rational middle, professionals who see the obvious
               | value of coding agents in their workflow and use them
               | extensively (or figure out how to best leverage them to
               | get the most mileage). I don't see software engineering
               | being "dead" ever, but the nature of the job _has already
               | changed_ and will continue to change. Look at Sonnet 3.5
               | -> 3.7 -> 4.5 -> Opus 4.5; that was 17 months of
               | development and the leaps in performance are quite
               | impressive. You then have massive hardware buildouts and
               | improvements to stack + a ton of R&D + competition to
               | squeeze the juice out of the current paradigm (there are
               | 4 orders of magnitude of scaling left before we hit real
               | bottlenecks) and also push towards the next paradigm to
               | solve things like continual learning. Some folks have
               | opted not to use coding agents (and some folks like
               | yourself seem to revel in strawmanning people who point
               | out their demonstrable usefulness). Not using coding
               | agents in Jan 2026 is defensible. It won't be defensible
               | for long.
        
               | bopbopbop7 wrote:
               | Please do provide some data for this "obvious value of
               | coding agents". Because right now the only thing obvious
               | is the increase in vulnerabilities, people claiming they
               | are 10x more productive but aren't shipping anything, and
               | some AI hype bloggers that fail to provide any
               | quantitative proof.
        
               | aspenmartin wrote:
               | Sure: at my MAANG company, where I watch the data closely
               | on adoption of CC and other internal coding agent tools,
               | most (significant) LOC are written by agents, and most
               | employees have adopted coding agents as WAU, and the
               | adoption rate is positively correlated with seniority.
               | 
               | Like a lot of things LLM related (Simon Willison's
               | pelican test, researchers + product leaders implementing
               | AI features) I also heavily "vibe" check the capabilities
               | myself on real work tasks. The fact of the matter is I am
               | able to dramatically speed up my work. It may be actually
               | writing production code + helping me review it, or it may
               | be tasks like: write me a script to diagnose this bug I
               | have, or build me a streamlit dashboard to analyze +
               | visualize this ad hoc data instead of me taking 1 hour to
               | make visualizations + munge data in a notebook.
               | 
               | > people claiming they are 10x more productive but aren't
               | shipping anything, and some AI hype bloggers that fail to
               | provide any quantitative proof.
               | 
               | what would satisfy you here? I feel you are strawmanning
               | a bit by picking the most hyperbolic statements and then
               | blanketing that on everyone else.
               | 
               | My workflow is now:
               | 
               | - Write code exclusively with Claude
               | 
               | - Review the code myself + use Claude as a sort of review
               | assistant to help me understand decisions about parts of
               | the code I'm confused about
               | 
               | - Provide feedback to Claude to change / steer it away or
               | towards approaches
               | 
               | - Give up when Claude is hopelessly lost
               | 
               | It takes a bit to get the hang of the right balance but
               | in my personal experience (which I doubt you will take
               | seriously but nevertheless): it is quite the game changer
               | and that's coming from someone who would have laughed at
               | the idea of a $200 coding agent subscription 1 year ago
        
               | bopbopbop7 wrote:
               | Anecdotes don't prove anything, ones without any metrics,
               | and especially at MAANG where AI use is strongly
               | incentivized.
               | 
               | Evidence is peer reviewed research, or at least something
               | with metrics. Like the METR study that shows that
               | experienced engineers often got slower on real tasks with
               | AI tools, even though they thought they were faster.
        
               | aspenmartin wrote:
               | That's why I gave you data! METR study was 16 people
               | using Sonnet 3.5/3.7. Data I'm talking about is 10s of
               | thousands of people and is much more up to date.
               | 
               | Some counter examples to METR that are in the literature
               | but I'll just say: "rigor" here is very difficult
               | (including METR) because outcomes are high dimensional
               | and nuanced, or ecological validity is an issue. It's
               | hard to have any approach that someone wouldn't be able
               | to dismiss due to some issue they have with the
               | methodology. The sources below also have methodological
               | problems just like METR
               | 
               | https://arxiv.org/pdf/2302.06590 -- 55% faster
               | implementing HTTP server in javascript with copilot (in
               | 2023!) but this is a single task and not really
               | representative.
               | 
               | https://demirermert.github.io/Papers/Demirer_AI_productiv
               | ity... -- "Though each experiment is noisy, when data is
               | combined across three experiments and 4,867 developers,
               | our analysis reveals a 26.08% increase (SE: 10.3%) in
               | completed tasks among developers using the AI tool.
               | Notably, less experienced developers had higher adoption
               | rates and greater productivity gains." (but e.g.
               | "completed tasks" as the outcome measure is of course
               | problematic)
               | 
               | To me, internal company measures for large tech companies
               | will be most reliable -- they are easiest to track and
               | measure, the scale is large enough, and the talent + task
               | pool is diverse (junior -> senior, different product
               | areas, different types of tasks). But then outcome
               | measures are always a problem...commits per developer per
               | month? LOC? task completion time? all of them are highly
               | problematic, especially because its reasonable to expect
               | AI tools would change the bias and variance of the proxy
               | so its never clear if you're measuring the change in
               | "style" or the change in the underlying latent measure of
               | productivity you care about
        
               | bopbopbop7 wrote:
               | To be fair, I'll take a non-biased 16 person study over
               | "internal measures" from a MAANG company that burned 100s
               | of billions on AI with no ROI.
        
               | insin wrote:
               | > - Give up when Claude is hopelessly lost
               | 
               | You love to see "Maybe completely waste my time" as part
               | of the normal flow for a productivity tool
        
               | nsxwolf wrote:
               | The nature of my job has always been fighting red tape,
               | process, and stake holders to deploy very small units of
               | code to production. AI really did not help with much of
               | that for me in 2025.
               | 
               | I'd imagine I'm not the only one who has a similar
               | situation. Until all those people and processes can be
               | swept away in favor of letting LLMS YOLO everything into
               | production, I don't see how that changes.
        
               | aspenmartin wrote:
               | No I think that's extremely correct. I work at a MAANG
               | where we have the resources to hook up custom internal
               | LLMs and agents to actually deal with that but that is
               | unique to an org of our scale.
        
         | Atomic_Torrfisk wrote:
         | > People are willing to pay $200 per month
         | 
         | Some people are of course, but how many?
         | 
         | > ... People are willing to pay $200 per month
         | 
         | This is just low-key hype. Careful with your portfolio...
        
         | HumblyTossed wrote:
         | It's a great tool, but right now it's only being used to feed
         | the greed.
         | 
         | >> Again, I guess no one knew AI would be as big as it is
         | today, and it is only just started.
         | 
         | People have been saying similar about self driving cars for
         | years now. "AI" is another one of those expensive ideas that
         | we'll get 85% of the way there and then to get the other 15%
         | will be way more expensive than anyone will want to pay for.
         | It's already happening - HW prices and electricity - people are
         | starting to ask, "if I put more $ into this machine, when am I
         | actually going to start getting money out?" The "true
         | believers" are like, soon! But people are right to be hugely
         | skeptical.
        
           | jliptzin wrote:
           | There are some things it's really great at. For example,
           | handling a css layout. If we have to spend trillions of
           | dollars and get nothing else out of it other than being able
           | to vertically center a <div> without wrestling with css and
           | wanting to smash the keyboard in the process, it will all
           | have been worth it.
        
           | aspenmartin wrote:
           | I agree -- skepticism is totally healthy. And there are so
           | many great ways to poke holes in the true underlying
           | narratives (not the headlines that people seem to pull from).
           | E.g. evaluation science is a wasteland (not for wont of very
           | smart people trying very hard to get them right). How do we
           | tackle the power requirements in a way that is sustainable?
           | Etc. etc.
           | 
           | But stuff like this im not sure I understand:
           | 
           | > It's a great tool, but right now it's only being used to
           | feed the greed.
           | 
           | if its a great tool, then how is it _only_ being used to
           | "feed the greed" and what do you mean by that?
           | 
           | Also I think folks are quick to make analogies to other
           | points in history: "AI is like the dot com boom we're going
           | to crash and burn" and "AI is like {self driving cars,
           | crypto, etc} and the promises will all be broken, its all
           | hype" but this removes the nuance: all of these things are
           | extremely different with very specific dynamics that in
           | _some_ ways may be similar but in many crucial and important
           | ways are completely different.
        
         | belter wrote:
         | >> But it is hard to argue against the value of current AI,
         | which many of the vocal critics on HN seems to have the opinion
         | of.
         | 
         | What is the concrete business case? Can anyone point to a
         | revenue producing company using AI in production, and where AI
         | is a material driver of profits?
         | 
         | Tool vendors don't count. I'm not interested in how much money
         | is being made selling shovels...show me a miner who actually
         | struck gold please.
        
         | layer8 wrote:
         | > Every single part of the hardware stack are being fused with
         | money and demand. The last time we have this was Post-PC /
         | Smartphone era which drove the hardware industry forward for 10
         | - 15 years. The current AI can at least push hardware for
         | another 5 - 6 years while pulling forward tech that was
         | initially 8 - 10 years away.
         | 
         | It's very unclear how much end-consumer hardware and DIY
         | builders will benefit from that, as opposed to server-grade
         | hardware that only makes sense for the enterprise marker. It
         | could have the opposite effect, like hardware manufacturers
         | leaving the consumer market (as in the case of Micron), because
         | there's just not that much money in it.
        
       | mrheosuper wrote:
       | I'm not against AI/LLM(in fact, i am quite supportive to it). But
       | one of my biggest fear is overusing AI. We may introduce some
       | tool that only "AI/LLM" can resonably do(Like tool with weird,
       | convoluted UI/UX, syntax) and no one against it because AI/LLM
       | can use/interact.
       | 
       | Then genAI, It's become more and more difficult to tell which is
       | AI and which is not, and AI is in everywhere. I dont know what to
       | think about it. "If you can't tell, does it matter ?"
        
         | netdur wrote:
         | i think the concern about software shifting toward ai design
         | ignores that the web hasn't been human-first for a long time.
         | most traffic is already machine to machine, like crawlers and
         | ci pipelines. we've tolerated systems that are barely legible
         | for years. anyone who has grepped through android studio logs
         | knows that human readability is usually a tertiary goal at
         | best. ai interacting with complex systems is just an evolution
         | of the glue code we've always written.
         | 
         | as for who made it, utility usually matters more than where it
         | came from. i used an agent for an oss changelog recently and it
         | picked up things i'd forgotten while structuring the narrative
         | better than i could. the intent and code were mine, but the ai
         | acted as a high fidelity compressor. the risk isn't ai being
         | everywhere. it's the atrophy of judgment where we stop using it
         | to support decisions and start using it to outsource thinking.
        
       | jama211 wrote:
       | The difference between the performance of models between 2024 and
       | 2025 has been so stark, that graph really shows it. There are
       | still many people on these forums who seem to think AI's produce
       | terrible code unless ultra supervised, and I can't help but
       | suspect some of them tried it a little while ago and just don't
       | understand how different it is now compared to even quite
       | recently.
        
         | Madmallard wrote:
         | I used Gemini Pro, Claude Pro yesterday a couple of dozen times
         | and basically have been daily.
         | 
         | I have a project to convert my multiplayer XNA game from C# to
         | Javascript and to add networking to the game-play using LLMs.
         | 
         | They are far worse at it now than they were a year ago. They
         | actually implemented the requirements (Though inaccurately) to
         | the best of their ability a year ago. Especially Gemini.
         | 
         | Now they don't even come remotely close to implementing just
         | the basic requirements.
         | 
         | The thing is, I'm giving them the entirety of the C# source
         | code and spelling out what they should do.
        
           | simonw wrote:
           | Weird. I would expect Gemini 3 Pro and Claude Opus 4.5 to run
           | rings around Gemini 1.5 Pro and Claude Sonnet 3.5.
           | 
           | How are you running them - regular chat interface or do you
           | have them setup with Claude Code or Gemini CLI?
        
             | Madmallard wrote:
             | Using the chat interface primarily with various prompting
             | strategies.
             | 
             | I am considering making a thread where I compel others to
             | attempt to get what I'm trying to get out of it and show me
             | their work.
             | 
             | The game is only around 25000-30000 LOC in C#.
        
       | andai wrote:
       | Re: yolo mode
       | 
       | I looked into docker and then realized the problem I'm actually
       | trying to solve was solved in like 1970 with users and
       | permissions.
       | 
       | I just made a agent user limited to its own home folder, and
       | added my user to its group. Then I run Claude code etc as the
       | agent user.
       | 
       | So it can only read write /home/agent, and it cannot read or
       | write my files.
       | 
       | I add myself to agent group so I can read/write the agent files.
       | 
       | I run into permission issues sometimes but, it's pretty smooth
       | for the most part.
       | 
       | Oh also I gave it root to a $3 VPS. It's so nice having a
       | sysadmin! :) That part definitely feels a bit deviant though!
        
         | jillesvangurp wrote:
         | I use a qemu vm for running codex cli in yolo mode and use
         | simple ssh based git operations for getting code in and out of
         | there. Works great. And you can also do fun things like let it
         | loose on multiple git projects in one prompt. The vm can run
         | docker as well which helps with containerized tests and other
         | more complicated things. One thing I've started to observe is
         | that you spend more time waiting for tool execution than for
         | model inference. So having a fast local vm is better than a
         | slower remote one.
        
         | some_developer wrote:
         | Docker in docker, with opencode.
         | 
         | Opencode plus some scripts on host and in its container works
         | well to run yolo and only see what it needs (via mounting). Has
         | git tools but can't push etc. is thought how to run tests with
         | the special container-in-container setup.
         | 
         | Including pre-configured MCPs, skills, etc.
         | 
         | The best part is that it just works for everyone on the team,
         | big plus.
        
         | knicholes wrote:
         | cgroups and namespaces
        
         | staeff777 wrote:
         | I really like this idea and just tried some steps for myself.
         | create user with homedir: sudo useradd -m agent add myself to
         | agent group: sudo usermod -a -G agent $USER
         | 
         | Allow agent group to agent home dir: sudo chmod -R 770
         | /home/agent
         | 
         | Start a new shell with the group (or login/logoff): newgrp
         | agent Now you should be able to change into the agent home.
         | 
         | Allow your user to sudo as agent: echo "$USER ALL=(agent)
         | NOPASSWD: ALL" |sudo tee -a /etc/sudoers.d/$USER-as-agent now
         | you can start your agent using sudo: sudo -u agent your_agent
         | 
         | works nice.
        
         | andai wrote:
         | Re: yolo mode
         | 
         | https://markdownpastebin.com/?id=1ef97add6ba9404b900929ee195...
         | 
         | My notes from back when I set this up! Includes instructions
         | for using a GUI file explorer as the agent user. As well as
         | setting up a systemd service to fix the permissions
         | automatically.
         | 
         | (And a nice trick which shows you which GUI apps are running as
         | which user...)
         | 
         | However, most of these are just workarounds for the permission
         | issue I kept running into, which is that Claude Code would for
         | some reason create files with incorrect permissions so that I
         | couldn't read or write those files from my normal account.
         | 
         | If someone knows how to fix that, or if someone at Anthropic is
         | reading, then most of this Rube Goldberg machine becomes
         | unnecessary :)
        
       | yupyupyups wrote:
       | Let's talk about the societal cost these models have had on us
       | including their high energy cost and the proliferation of auto-
       | generated slop media used to milk ad revenue, scam people, SEO
       | farm, do propaganda or automate trolling. What about these big
       | corporations collecting an astronomical amount of debt to hoard
       | DRAM and NAND in a way that has crippled the PC market within
       | weeks? And what are they going to do next, put a few dollars in
       | Trump's pocket so that they can rob/loot the US population
       | through bailouts? Who gets to keep all the hardware I wonder?
       | 
       | Nvidia, Samsung, SK Hynix and some other voltures I forgot to
       | mention are making serious bank right now.
        
       | ashishgupta2209 wrote:
       | 2026: The Year of Robots, note it for next year
        
       | fullstackchris wrote:
       | > The reason I think MCP may be a one-year wonder is the
       | stratospheric growth of coding agents. It appears that the best
       | possible tool for any situation is Bash--if your agent can run
       | arbitrary shell commands, it can do anything that can be done by
       | typing commands into a terminal.
       | 
       | I push back strongly from this. In the case of the solo, one-
       | machine coder, this is likely the case - if you're exposing
       | workflows or fixed tools to customers / collegues / the web at
       | large via API or similar, then MCP is still the best way to
       | expose it IMO.
       | 
       | Think about a GitHub or Jira MCP server - commandline alone they
       | are sure to make mistakes with REST requests, API schema etc.
       | With MCP the proper known commands are already baked in. Remember
       | always that LLMs will be better with natural language than code.
        
         | simonw wrote:
         | The solution to that is Anthropic's Skills.
         | 
         | Create a folder called skills/how-to-use-jira
         | 
         | Add several Bash scripts with the right curl commands to
         | perform specific actions
         | 
         | Add a SKILL.md file with some instructions in how to use those
         | scripts
         | 
         | You've effectively flattened that MCP server into some Markdown
         | and Bash, only the thing you have now is more flexible (the
         | coding agent can adapt those examples to cover new things you
         | hadn't thought to tell it) and much more context-efficient (it
         | only reads the Markdown the first time you ask it to do
         | something with JIRA).
        
           | aflukasz wrote:
           | But that moves the burden of maintenance from the provider of
           | the service to its users (and/or partially to intermediary in
           | form of "skills registry" of sorts, which apparently is a
           | thing now).
           | 
           | So maybe a hybrid approach would make more sense? Something
           | like /.well-known/skills/README.md exposed and owned by the
           | providers?
           | 
           | That is assuming that the whole idea of "skills" makes sense
           | in practice.
        
             | simonw wrote:
             | Yeah that's true, skill distribution isn't a solved problem
             | yet - MCPs have a URL, which is a great way of making them
             | available for people to start using without extra steps.
        
       | rr808 wrote:
       | What happened to Devin? 2024 it was a leading contender now it
       | isn't even included in the big list of coding agents.
        
         | fullstackchris wrote:
         | Wasn't it basically revealed as a scam? I remember some article
         | about their fancy demo video being sped up / unfairly cut and
         | sliced etc.
        
         | monkeydust wrote:
         | https://cognition.ai/blog/devin-annual-performance-review-20...
        
         | ColinEberhardt wrote:
         | It's still around, and tends to be adopted by big enterprises.
         | It's generally a decent product, but is facing a lot of equally
         | powerful competition and is very expensive.
        
         | simonw wrote:
         | To be honest that's more because I've never tried it myself, so
         | it isn't really on my radar.
         | 
         | I don't hear much buzz about it from the people I pay attention
         | to. I should still give it a go though.
        
       | Gud wrote:
       | What about self hosting?
        
         | simonw wrote:
         | I talked about that in this section
         | https://simonwillison.net/2025/Dec/31/the-year-in-llms/#the-...
         | - and touched on it a bit in the section about Chinese AI labs:
         | https://simonwillison.net/2025/Dec/31/the-year-in-llms/#the-...
        
       | politelemon wrote:
       | > The problem is that the big cloud models got better too--
       | including those open weight models that, while freely available,
       | were far too large (100B+) to run on my laptop.
       | 
       | The actual, notable progress will be models that can run
       | reasonably well on commodity, everyday hardware that the average
       | user has. From more accessibility will come greater usefulness.
       | Right now the way I see it, having to upgrade specs on a machine
       | to run local models keeps it in a niche hobbyist bubble.
        
       | timonoko wrote:
       | OpenSCAD-coding has improved significantly on all models. Now
       | syntax is always right and they understand the concept of
       | _negative space_.
       | 
       | Only problem is that they don't see connection between form and
       | function. They may make teapot perfectly but don't understand
       | that this form is supposed to contain liquid.
        
       | mark_l_watson wrote:
       | Thanks Simon, great writeup.
       | 
       | It has been an amazing year, especially around tooling (search,
       | code analysis, etc.) and surprisingly capable smaller models.
        
       | huqedato wrote:
       | I completely disagree with the idea that 2025 "The (only?) year
       | of MCP." In fact, I believe every year in the foreseeable future
       | will belong to MCP. It is here to stay. MCP was the best
       | (rational, scalable, predictable) thing since LLM madness broke
       | loose.
        
       | mmcnl wrote:
       | Let's hope 2026 will also have interesting innovations not
       | related to AI or LLMs.
        
         | spicyusername wrote:
         | 2025 had plenty of those, they just didn't get as many news
         | headlines.
         | 
         | One of the difficult things of modernity is that it's easy to
         | confuse what you hear about a lot with what is real.
         | 
         | One of the great things about modernity is that progress
         | continues, whether we know about it or not.
        
       | _pdp_ wrote:
       | With everything that we have done so far (our company) I believe
       | by end of 2026 our software will be self improving all the time.
       | 
       | And no it is not AI slop and we don't vibe code. There are a lot
       | of practical aspects of running software and maintaining /
       | improving code that can be done well with AI if you have the
       | right setup. It is hard to formulate what "right" looks like at
       | this stage as we are still iterating on this as well.
       | 
       | However, in our own experiments we can clearly see dramatic
       | increases in automation. I mean we have agents working overnight
       | as we sleep and this is not even pushing the limits. We are now
       | wrapping major changes that will allows us to run AI agents all
       | the time as long as we can afford them.
       | 
       | I can even see most of these materialising in Q1 2026.
       | 
       | Fun times.
        
         | papacj657 wrote:
         | What exactly are your agents doing overnight? I often hear
         | folks talk about their agents running for long periods of time
         | but rarely talk about the outcomes they're driving from those
         | agents.
        
           | _pdp_ wrote:
           | We have a lot of grunt work scheduled overnight like finding
           | bugs, creating tests where we don't have good coverage or
           | where we can improve, integrations, documentation work, etc.
           | 
           | Not everything gets accepted. There is a lot of work that is
           | discarded and much more pending verification and acceptance.
           | 
           | Frankly, and I hope I don't come as alarmist (judge for
           | yourself from my previous comments on Hn and Reddit) we
           | cannot keep up with the output! And a lot of it is actually
           | good and we should incorporate it even partially.
           | 
           | At the moment we are figuring out how to make things more
           | autonomous while we have the safety and guardrails in place.
           | 
           | The biggest issue I see at this stage is how to make sense of
           | it all as I do not believe we have the understanding of what
           | is happening - just the general notion of it.
           | 
           | I truly believe that we will reach the point where ideas
           | matter more than execution, which what I would expect to be
           | the case with more advanced and better applied AI.
        
       | asgR1t wrote:
       | Most LLMs got worse in 2025. Only addicts and the type of
       | computer gamer that feels drawn to complex setups, gamification
       | and does not care about the end result will feel positive about
       | the grift.
       | 
       | 2025: The Year in Open Source? Nothing, all resources were tied
       | up to debunk a couple of Python web developers who pose as the
       | ultimate experts in LLMs.
        
         | simonw wrote:
         | In what way did they get worse?
         | 
         | I made you a dashboard of my 2025 writing about open-source
         | that didn't include AI:
         | https://simonwillison.net/dashboard/posts-with-tags-in-a-yea...
        
       | icapybara wrote:
       | It was the year of Claude Code
        
       | nativeit wrote:
       | Between the people with invested and/conflicting interests, and
       | the hordes of dogmatic zealots, I find discussions about AI to be
       | the least productive or reliably informed on HN.
        
         | simonw wrote:
         | Honestly this thread was pretty disappointing. Many of the
         | comments here could have been attached to _any_ post about LLMs
         | in the past year or so.
        
       | ck2 wrote:
       | as I was clicking "gee I hope there's the year of pelicans riding
       | bicycles"
       | 
       | left satisfied, lol
        
       | losvedir wrote:
       | I predict 2026 will be the year of the first AI Agent "worm" (or
       | virus?). Kind of like the Morris worm running amok as an
       | experiment gone wrong, I think we will sometime soon have someone
       | set up an AI agent whose core loop is to try to propagate itself,
       | either as an experiment or just for the lulz.
       | 
       | The actual Agent payload would be very small, likely just a few
       | hundred line harness plus system prompt. It's just a question of
       | whether the agent will be skilled enough to find vulnerabilities
       | to propagate. The interesting thing about an AI worm is that it
       | can use different tricks on different hosts as it explores its
       | own environment.
       | 
       | If a pure agent worm isn't capable enough, I could see someone
       | embedding it on top of a more traditional virus. The normal virus
       | would propagate as usual, but it would also run an agent to
       | explore the system for things to extract or attack, and to find
       | easy additional targets on the same internal network.
       | 
       | A main difference here is that the agents have to call out to a
       | big SotA model somewhere. I imagine the first worm will simply
       | use Opus or ChatGPT with an acquired key, and part of it will be
       | trying to identify (or generate) new keys as it spreads.
       | 
       | Ultimately, I think this worm will be shut down by the model
       | vendor, but it will have to have made a big enough splash
       | beforehand to catch their attention and create a team to identify
       | and block keys making certain kinds of requests.
       | 
       | I'd hope OpenAI, Anthropic, etc have a team and process in place
       | already to identify suspicious keys, eg, those used from a huge
       | variety of IPs, but I wouldn't be surprised if this were low on
       | their list of priorities (until something like this hits).
        
       ___________________________________________________________________
       (page generated 2026-01-01 23:01 UTC)