[HN Gopher] 2025: The Year in LLMs
___________________________________________________________________
2025: The Year in LLMs
Author : simonw
Score : 838 points
Date : 2025-12-31 23:54 UTC (23 hours ago)
(HTM) web link (simonwillison.net)
(TXT) w3m dump (simonwillison.net)
| AndyNemmity wrote:
| These are excellent every year, thank you for all the wonderful
| work you do.
| tkgally wrote:
| Same here. Simon is one of the main reasons I've been able to
| (sort of) keep up with developments in AI.
|
| I look forward to learning from his blog posts and HN comments
| in the year ahead, too.
| password4321 wrote:
| Don't forget you can pay Simon to keep up with less!
|
| > _At the end of every month I send out a much shorter
| newsletter to anyone who sponsors me for $10 or more on
| GitHub_
|
| https://simonwillison.net/about/#monthly
| waldrews wrote:
| Remember, back in the day, when a year of progress was like, oh,
| they voted to add some syntactic sugar to Java...
| throwup238 wrote:
| _> they voted to add some syntactic sugar to Java..._
|
| I remember when we just wanted to rewrite everything in Rust.
|
| Those were the simpler times, when crypto bros seemed like the
| worst venture capitalism could conjure.
| OGEnthusiast wrote:
| Crypto bros in hindsight were so much less dangerous than AI
| bros. At least they weren't trying to construct data centers
| in rural America or prop up artificial stocks like $NVDA.
| quaintpartridge wrote:
| They were, just not as many.
| https://www.wired.com/story/the-worlds-biggest-bitcoin-
| mine-...
| SauntSolaire wrote:
| Instead they were building crypto mining warehouses in
| rural America and propping up artificial currencies like
| BTC.
| ryandrake wrote:
| Crazy how the two most hyped and funded technologies of
| the decade were: energy wasting fake money for criminals
| and energy wasting plagiarism machines.
| zahlman wrote:
| Speaking of which, we never found out the details (strike
| price/expiration) of Michael Burry's puts, did we? It seems
| he could have made bank if he'd waited one more month...
| kamranjon wrote:
| I think they expire in March 2026 if the NVIDIA stock
| drops to $140 a share? Something close to that I think.
| mgfist wrote:
| It's funny how people complain about the rust belt dying
| and factories leaving rural communities and so on, then
| when someone wants to build something that can provide jobs
| and tax revenue, everyone complains.
| lostlogin wrote:
| I've heard about the risk of AI leading to job losses and
| wealth concentration.
|
| I haven't heard about new businesses, job creation and
| growth in former industrial towns. What have I missed?
| jakeydus wrote:
| How many people are employed at the average data center?
| A few dozen? Versus a steel mill, that's nothing. A
| chicken plant in Nebraska closed down this last month.
| 3200 people lost their jobs. You think Meta will fill it
| with GPUs and the whole town will have jobs again?
| scotty79 wrote:
| Many more are employed while building it. And they will
| never stop building. It's modern version of rail. But
| instead of distances it will cover the area.
| uxcolumbo wrote:
| Will local folks get those jobs to build the data center?
|
| And if so, what happens to those builders once the data
| center is built?
| techpression wrote:
| As if any taxes will be paid to the areas affected, and
| add to that the billions in taxes used to subsidize
| everything before a single cent is a net positive.
| nrhrjrjrjtntbt wrote:
| More like 6 different new nosql databases and js frameworks.
| dotancohen wrote:
| A Wordpress zero day and Linux not on the desktop. Netcraft
| confirms it.
| crystal_revenge wrote:
| That must have been a _long_ time back. Having lived through
| the time when web pages were served through CGI and mobile
| phones only existed in movies, when SVMs where the _new
| hotness_ in ML and people would write about how weird NNs were,
| I feel like I 've seen _a lot_ more concrete progress in the
| last few decades than this year.
|
| This year honestly feels quite stagnant. LLMs are _literally_
| technology that can only reproduce the past. They 're cool, but
| they were _way_ cooler 4 years ago. We 've taken big ideas like
| "agents" and "reinforcement learning" and basically stripped
| them of all meaning in order to claim progress.
|
| I mean, do you remember Geoffrey Hinton's RBM talk at Google in
| 2010? [0] That was _absolutely insane_ for anyone keeping up
| with that field. By the mid-twenty teens RBMs were _already_
| outdated. I remember when everyone was implementing flavors of
| RNNs and LSTMs. Karpathy 's character 2015 RNN project was
| _insane_ [1].
|
| This comment makes me wonder if part of the hype around LLMs is
| just that a lot of software people simply weren't paying
| attention to the absolutely mind-blowing progress we've seen in
| this field for the last 20 years. But even ignoring ML, the
| world's of web development and mobile application development
| have gone through incredible progress over the last decade and
| a half. I remember a time when JavaScript books would have a
| section warning that you should _never_ use JS for anything
| critical to the application. Then there 's the work in theorem
| provers over the last decade... If you remember when syntactic
| sugar was progress, either you remember _way_ further back than
| I do, or you weren 't paying attention to what was happening in
| the larger computing world.
|
| 0. https://www.youtube.com/watch?v=VdIURAu1-aU
|
| 1. https://karpathy.github.io/2015/05/21/rnn-effectiveness/
| handoflixue wrote:
| > LLMs are literally technology that can only reproduce the
| past.
|
| Funny, I've used them to create my own personalized text
| editor, perfectly tailored to what I actually want. I'm
| pretty sure that didn't exist before.
|
| It's wild to me how many people who talk about LLM apparently
| haven't learned how to use them for even very basic tasks
| like this! No wonder you think they're not that powerful, if
| you don't even know basic stuff like this. You really owe it
| to yourself to try them out.
| crystal_revenge wrote:
| > You really owe it to yourself to try them out.
|
| I've worked at multiple AI startups in lead AI Engineering
| roles, both working on deploying user facing LLM products
| and working on the research end of LLMs. I've done
| collaborative projects and demos with a pretty wide range
| of big names in this space (but don't want to doxx myself
| _too_ aggressively), have had my LLM work cited on HN
| multiple times, have LLM based github projects with
| hundreds of stars, appeared on a few podcasts talking about
| AI etc.
|
| This gets to the point I was making. I'm starting to
| realize that part of the disconnect between my opinions on
| the state of the field and others is that many people
| _haven 't_ really been paying much attention.
|
| I can see if recent LLMs are your first intro to the state
| of the field, it must feel incredible.
| CamperBob2 wrote:
| That's all very impressive, to be sure. But are you sure
| you're getting the point? As of 2025, LLMs are now very
| good at writing new code, creating new imagery, and
| writing original text. They continue to improve at a
| remarkable rate. They are helping their users create
| things that didn't exist before. Additionally, they are
| now very good at searching and utilizing web resources
| that didn't exist at training time.
|
| So it is absurdly incorrect to say "they can only
| reproduce the past." Only someone who _hasn 't_ been
| paying attention (as you put it) would say such a thing.
| crystal_revenge wrote:
| I think the confusion is people's misunderstanding of
| what 'new code' and 'new imagery' mean. Yes, LLMs can
| generate a specific CRUD webapp that hasn't existed
| before but only based on interpolating between the
| history of existing CRUD webapps. I mean traditional
| Markov Chains can also produce 'new' text in the sense
| that "this _exact_ text " hasn't been seen before, but
| nobody would argue that traditional Markov Chains aren't
| constrained by "only producing the past".
|
| This is even more clear in the case of diffusion models
| (which I personally love using, and have spent a lot of
| time researching). All of the "new" images created by
| even the most advanced diffusion models are fundamentally
| remixing past information. This is really obvious to
| anyone who has played around with these extensively
| because they really can't produce truly novel concepts.
| New concepts can be added by things like fine-tuning or
| use of LoRAs, but _fundamentally_ you 're still just
| remixing the past.
|
| LLMs are always doing some form of interpolation between
| different points in the past. Yes they can create a "new"
| SQL query, but it's just remixing from the SQL queries
| that have existed prior. This still makes them _very_
| useful because _a lot_ of engineering work, including
| writing a custom text editor, involve remixing existing
| engineering work. If you could have stack-overflowed your
| way to an answer in the past, an LLM will be much
| superior. In fact, the phrase "CRUD" largely exists to
| point out that most webapps are fundamentally _the same_.
|
| A great example of this limitation in practice is the
| work that Terry Tao is doing with LLMs. One of the
| largest challenges in automated theorem proving is
| translating human proofs into the language of a theorem
| prover (often Lean these days). The challenge is that
| there is not very much Lean code currently available to
| LLMs (especially with the necessary context of the
| accompanying NL proof), so they struggle to correctly
| translate. Most of the research in this area is around
| improving LLM's representation of the mapping from human
| proofs to Lean proofs (btw, I personally feel like LLMs
| _do_ have a reasonably good chance of providing major
| improvements in the space of formal theorem proving, in
| conjunction with languages like Lean, because the
| translation process is the biggest blocker to progress).
|
| When you say:
|
| > So it is absurdly incorrect to say "they can only
| reproduce the past."
|
| It's pretty clear you don't have a solid background in
| generative models, because this is fundamentally what
| they do: model an existing probability distribution and
| draw samples from that. LLMs are doing this for a
| _massive_ amount of human text, which is why they do
| produce some impressive and _useful_ results, but this is
| also a fundamental limitation.
|
| But a world where we used LLMs for the majority of work,
| would be a world with no fundamental breakthroughs. If
| you've read _The Three Body Problem_ , it's very much
| like living in the world where scientific progress is
| impeded by sophons. In that world there is still some
| progress (especially with abundant energy), but it
| remains fundamentally and deeply limited.
| throwaway7783 wrote:
| Would you say that LLMs can discover patterns hitherto
| unknown? It would still be generating from the past, but
| patterns/connections not made before.
| PeterHolzwarth wrote:
| Just an innocent bystander here, so forgive me, but I
| think the flack you are getting is because you appear to
| be responding to claims that these tools will reinvent
| everything and introduce a new halcyon age of creation -
| when, at least on hacker news, and definitely in this
| thread, no one is really making such claims.
|
| Put another way, and I hate to throw in the now over-used
| phrase, but I feel you may be responding to a strawman
| that doesn't much appear in the article or the discussion
| here: "Because these tools don't achieve a god-like level
| of novel perfection that no one is really promising here,
| I dismiss all this sorta crap."
|
| Especially when I think you are also admitting that the
| technology is a fairly useful tool on its own merits - a
| stance which I believe represents the bulk of the
| feelings that supporters of the tech here on HN are
| describing.
|
| I apologize if you feel I am putting unrepresentative
| words in your mouth, but this is the reading I am taking
| away from your comments.
| signatoremo wrote:
| Lot of impressive points. They are also irrelevant. The
| majority of people also only extrapolate from the
| knowledge they acquired in the past. That's why there is
| the concept of inventor, someone who comes up with new
| ideas. Many new inventions are also based on existing
| ideas. Is that the reason to dismiss those achievements?
|
| Do you only take LLM seriously if it can be another
| Einstein?
|
| > But a world where we used LLMs for the majority of
| work, would be a world with no fundamental breakthroughs.
|
| What do you consider recent fundamental breakthroughs?
|
| Even if you are right, human can continue to work on hard
| problems while letting LLM handle the majority of
| derivative work
| uxcolumbo wrote:
| How do human brains create something novel and what will
| it take for AIs to do the same?
| threethirtytwo wrote:
| > It's pretty clear you don't have a solid background in
| generative models, because this is fundamentally what
| they do
|
| You don't have a solid background. No one does. We
| fundamentally don't understand LLMs, this is an industry
| and academic opinion. Sure there are high level
| perspectives and analogies we can apply to LLMs and
| machine learning in general like probability
| distributions, curve fitting or interpolations... but
| those explanations are so high level that they can
| essentially be applied to humans as well. At a lower
| level we cannot describe what's going on. We have no idea
| how to reconstruct the logic of how an LLM arrived at a
| specific output from a specific input.
|
| It is impossible to have any sort of deterministic
| function, process or anything produce new information
| from old information. This limitation is fundamental to
| logic and math and thus it will limit human output as
| well.
|
| You can combine information you can transform information
| you can lose information. But producing new information
| from old information from deterministic intelligence is
| fundamentally impossible in reality and therefore
| fundamentally impossible for LLMs and humans. But note
| the keyword: "deterministic"
|
| New information can literally only arise through
| stochastic processes. That's all you have in reality. We
| know it's stochastic because determinism vs.
| stochasticism are literally your only two viable options.
| You have a bunch of inputs, the outputs derived from it
| are either purely deterministic transformations or if you
| want some new stuff from the input you must apply
| randomness. That's it.
|
| That's essentially what creativity is. There is literally
| no other logical way to generate "new information".
| Purely random is never really useful so "useful
| information" arrives only after it is filtered and we use
| past information to filter the stochastic output and
| "select" something that's not wildly random. We also only
| use randomness to perturb the output a little bit so it's
| not too crazy.
|
| In the end it's this selection process and stochastic
| process combined that forms creativity. We know this is a
| general aspect of how creativity works because there's
| literally no other way to do it.
|
| LLMs do have stochastic aspects to them so we know for a
| fact it is generating new things and not just drawing on
| the past. We know it can fit our definition of "creative"
| and we can literally see it be creative in front of your
| eyes.
|
| You're ignoring what you see with your eyes and drawing
| your conclusions from a model of LLMs that isn't fully
| accurate. Or you're not fully tying the mechanisms of how
| LLMs work with what creativity or generating new data
| from past data is in actuality.
|
| The fundamental limitation with LLMs is not that it can't
| create new things. It's that the context window is too
| small to create new things beyond that. Whatever it can
| create it is limited to the possibilities within that
| window and that sets a limitation on creativity.
|
| What you see happening with LEAN can also be an issue
| with the context window being too small. If we have an
| LLM with a giant context window bigger than anything
| before... and pass it all the necessary data to "learn"
| and be "trained" on lean it can likely start to produce
| new theorems without literally being "trained".
|
| Actually I wouldn't call this a "fundamental" problem.
| More fundamental is the aspect of hallucinations. The
| fact that LLMs produce new information from past
| information in the WRONG way. Literally making up
| bullshit out of thin air. It's the opposite problem of
| what you're describing. These things are too creative and
| making up too much stuff.
|
| We have hints that LLMs know the difference between
| hallucinations and reality but coaxing it to communicate
| that differentiation to us is limited.
| oedemis wrote:
| as architectures evolve, i think it can be that we learn
| more "side effects".. back in 2020 openai researchers
| said "GPT-3 is applied without any gradient updates or
| fine-tuning" the model emerges at a certain level of
| scale...
| aoeusnth1 wrote:
| > It's pretty clear you don't have a solid background in
| generative models, because this is fundamentally what
| they do: model an existing probability distribution and
| draw samples from that.
|
| After post-training, this is definitively NOT what an LLM
| does.
| weatherlite wrote:
| > So it is absurdly incorrect to say "they can only
| reproduce the past."
|
| Also , a shitton of what we do economically is
| reproducing the past with slight tweaks and improvements.
| We all do very repetitive things and these tools cut the
| time / personnel needed by a significant factor.
| windexh8er wrote:
| > They are helping their users create things that didn't
| exist before.
|
| That is a derived output. That isn't new as in: novel. It
| may be unique but it is derived from training data. LLMs
| legitimately cannot think and thus they cannot create in
| that way.
| Kerrick wrote:
| That is a pedantic distinction. You can create something
| that didn't exist by combining two things that did exist,
| in a way of combining things that already existed. For
| example, you could use a blender to combine almond butter
| and sawdust. While this may not be "novel", and it may be
| derived from existing materials and methods, you may
| still lay claim to having created something that didn't
| exist before.
|
| For a more practical example, creating bindings from
| dynamic-language-A for a library in compiled-language-B
| is a genuinely useful task, allowing you to create things
| that didn't exist before. Those things are likely to
| unlock great happiness and/or productivity, even if they
| are derived from training data.
| windexh8er wrote:
| > That is a pedantic distinction. You can create
| something that didn't exist by combining two things that
| did exist, in a way of combining things that already
| existed.
|
| This is the definition of a derived product. Call it a
| derivative work if we're being pedantic and, regardless,
| is not any level of proof that LLMs "think".
| threethirtytwo wrote:
| Pedantic and not true. The LLM has stochastic processes
| involved. Randomness. That's not old information. That's
| newly generated stuff.
| zingar wrote:
| Could you give us an idea of what you're hoping for that
| is not possible to derive from training data of the
| entire internet and many (most?) published books?
| techpression wrote:
| This is the problem, the entire internet is a really bad
| set of training data because it's extremely polluted.
|
| Also the derived argument doesn't really hold, just
| because you know about two things doesn't mean you'd be
| able to come up with the third, it's actually very hard
| most of the time and requires you to not do next token
| prediction.
| threethirtytwo wrote:
| The emergent phenomenon is that the LLM can separate
| truth from fiction when you give it a massive amount of
| data. It can figure the world out just as we can figure
| it out when we are as well inundated with bullshit data.
| The pathways exist in the LLM but it won't necessarily
| reveal that to you unless you tune it with RL.
| ahtihn wrote:
| > The emergent phenomenon is that the LLM can separate
| truth from fiction when you give it a massive amount of
| data.
|
| I don't believe they can. LLMs have no concept of truth.
|
| What's likely is that the "truth" for many subjects is
| represented way more than fiction and when there is
| objective truth it's consistently represented in similar
| way. On the other hand there are many variations of
| "fiction" for the same subject.
| threethirtytwo wrote:
| They can and we have definitive proof. When we tune LLM
| models with reinforcement learning the models end up
| hallucinating less and becoming more reliable. Basically
| in a nut shell we reward the model when telling the truth
| and punish it when it's not.
|
| So think of it like this, to create the model we use
| terabytes of data. Then we do RL which is probably less
| than one percent of additional data involved in the
| initial training.
|
| The change in the model is that reliability is increased
| and hallucinations are reduced at a far greater rate than
| one percent. So much so that modern models can be used
| for agentic tasks.
|
| How can less than one percent of reinforcement training
| get the model to tell the truth greater than one percent
| of the time?
|
| The answer is obvious. It ALREADY knew the truth. There's
| no other logical way to explain this. The LLM in its
| original state just predicts text but it doesn't care
| about truth or the kind of answer you want. With a little
| bit of reinforcement it suddenly does much better.
|
| It's not a perfect process and reinforcement learning
| often causes the model to be deceptive an not necessarily
| tell the truth but it more gives an answer that may seem
| like the truth or an answer that the trainer wants to
| hear. In general though we can measurably see a
| difference in truthfulness and reliability to an extent
| far greater than the data involved in training and that
| is logical proof it knows the difference.
|
| Additionally while I say it knows the truth already this
| is likely more of a blurry line. Even humans don't fully
| know the truth so my claim here is that an LLM knows the
| truth to a certain extent. It can be wildly off for
| certain things but in general it knows and this "knowing"
| has to be coaxed out of the model through RL.
|
| Keep in mind the LLM is just auto trained on reams and
| reams of data. That training is massive. Reinforcement
| training is done on a human basis. A human must rate the
| answers so it is significantly less.
| habinero wrote:
| > The answer is obvious. It ALREADY knew the truth.
| There's no other logical way to explain this.
|
| I can think of several offhand.
|
| 1. The effect was never real, you've just convinced
| yourself it is because you want it to be, ie you Clever
| Hans'd yourself.
|
| 2. The effect is an artifact of how you measure "truth"
| and disappears outside that context ("It can be wildly
| off for certain things")
|
| 3. The effect was completely fabricated and is the result
| of fraud.
|
| If you want to convince me that "I threatened a
| statistical model with a stick and it somehow got more
| accurate, therefore it's both intelligent and lying" is
| true, I need a lot less breathless overcredulity and a
| lot more "I have actively tried to disprove this result,
| here's what I found"
| threethirtytwo wrote:
| You asked for something concrete, so I'll anchor every
| claim to either documented results or directly observable
| training mechanics.
|
| First, the claim that RLHF materially reduces
| hallucinations and increases factual accuracy is not
| anecdotal. It shows up quantitatively in benchmarks
| designed to measure this exact thing, such as TruthfulQA,
| Natural Questions, and fact verification datasets like
| FEVER. Base models and RL-tuned models share the same
| architecture and almost identical weights, yet the RL-
| tuned versions score substantially higher. These
| benchmarks are external to the reward model and can be
| run independently.
|
| Second, the reinforcement signal itself does not contain
| factual information. This is a property of how RLHF
| works. Human raters provide preference comparisons or
| scores, and the reward model outputs a single scalar.
| There are no facts, explanations, or world models being
| injected. From an information perspective, this signal
| has extremely low bandwidth compared to pretraining.
|
| Third, the scale difference is documented by every group
| that has published training details. Pretraining consumes
| trillions of tokens. RLHF uses on the order of tens or
| hundreds of thousands of human judgments. Even generous
| estimates put it well under one percent of the total
| training signal. This is not controversial.
|
| Fourth, the improvement generalizes beyond the reward
| distribution. RL-tuned models perform better on prompts,
| domains, and benchmarks that were not part of the
| preference data and are evaluated automatically rather
| than by humans. If this were a Clever Hans effect or
| evaluator bias, performance would collapse when the
| reward model is not in the loop. It does not.
|
| Fifth, the gains are not confined to a single definition
| of "truth." They appear simultaneously in question
| answering accuracy, contradiction detection, multi-step
| reasoning, tool use success, and agent task completion
| rates. These are different evaluation mechanisms. The
| only common factor is that the model must internally
| distinguish correct from incorrect world states.
|
| Finally, reinforcement learning cannot plausibly inject
| new factual structure at scale. This follows from
| gradient dynamics. RLHF biases which internal activations
| are favored, it does not have the capacity to encode
| millions of correlated facts about the world when the
| signal itself contains none of that information. This is
| why the literature consistently frames RLHF as behavior
| shaping or alignment, not knowledge acquisition.
|
| Given those facts, the conclusion is not rhetorical. If a
| tiny, low-bandwidth, non-factual signal produces large,
| general improvements in factual reliability, then the
| information enabling those improvements must already
| exist in the pretrained model. Reinforcement learning is
| selecting among latent representations, not creating
| them.
|
| You can object to calling this "knowing the truth," but
| that's a semantic move, not a substantive one. A system
| that internally represents distinctions that reliably
| track true versus false statements across domains, and
| can be biased to express those distinctions more
| consistently, functionally encodes truth.
|
| Your three alternatives don't survive contact with this.
| Clever Hans fails because the effect generalizes.
| Measurement artifact fails because multiple independent
| metrics move together. Fraud fails because these results
| are reproduced across competing labs, companies, and
| open-source implementations.
|
| If you think this is still wrong, the next step isn't
| skepticism in the abstract. It's to name a concrete
| alternative mechanism that is compatible with the
| documented training process and observed generalization.
| Without that, the position you're defending isn't
| cautious, it's incoherent.
| jama211 wrote:
| Yeah you've lost me here I'm sorry. In the real world
| humans work with AI tools to create new things. What
| you're saying is the equivalent of "when a human writes a
| book in English, because they use words and letters that
| already exist and they already know they aren't creating
| anything new".
| nl wrote:
| What does "think" mean?
|
| Why is that kind of thinking required to create novel
| works?
|
| Randomness can create novelty.
|
| Mistakes can be novel.
|
| There are many ways to create novelty.
|
| Also I think you might not know how LLMs are trained to
| code. Pre-training gives them some idea of the syntax etc
| but that only gets you to fancy autocomplete.
|
| Modern LLMs are heavily trained using reinforcement data
| which is custom task the labs pay people to do (or by
| distilling another LLM which has had the process
| performed on it).
| closewith wrote:
| By that definition, nearly all commercial software
| development (and nearly all human output in general) is
| derived output.
| windexh8er wrote:
| Wow.
|
| You're using 'derived' to imply 'therefore equivalent.'
| That's a category error. A cookbook is derived from food
| culture. Does an LLM taste food? Can it think about how
| good that cookie tastes?
|
| A flight simulator is derived from aerodynamics - yet it
| doesn't fly.
|
| Likewise, text that resembles reasoning isn't the same
| thing as a system that has beliefs, intentions, or
| understanding. Humans do. LLMs don't.
|
| Also... Ask an LLM what's the difference between a human
| brain and an LLM. If an LLM could "think" it wouldn't
| give you the answer it just did.
| CamperBob2 wrote:
| _Ask an LLM what 's the difference between a human brain
| and an LLM. If an LLM could "think" it wouldn't give you
| the answer it just did._
|
| I imagine that sounded more profound when you wrote it
| than it did just now, when I read it. Can you be a little
| more specific, with regard to what features you would
| expect to differ between LLM and human responses to such
| a question?
|
| Right now, LLM system prompts are strongly geared towards
| _not_ claiming that they are humans or simulations of
| humans. If your point is that a hypothetical "thinking"
| LLM would claim to be a human, that could certainly be
| arranged with an appropriate system prompt. You wouldn't
| know whether you were talking to an LLM or a human --
| just as you don't now -- but nothing would be proved
| either way. That's ultimately why the Turing test is a
| poor metric.
| closewith wrote:
| You're arguing against a straw man. No one is claiming
| LLMs have beliefs, intentions, or understanding. They
| don't need them to be economically useful.
| windexh8er wrote:
| Oh yes, they are.
|
| And beyond people claiming that LLMs are basically
| sentient you have people like CamperBob2 who made this
| wild claim:
|
| _" ""There's no such thing as people without language,
| except for infants and those who are so mentally
| incapacitated that the answer is self-evidently "No, they
| cannot."
|
| Language is the substrate of reason. It doesn't need to
| be spoken or written, but it's a necessary and (as it
| turns out) sufficient component of thought."""_
|
| Let that sink. They literally think that there's no such
| thing as people without language. Talk about a wild and
| ignorant take on life in general!
| ordersofmag wrote:
| I will find this often-repeated argument compelling only
| when someone can prove to me that the human mind works in
| a way that isn't 'combining stuff it learned in the
| past'.
|
| 5 years ago a typical argument against AGI was that
| computers would never be able to think because "real
| thinking" involved mastery of language which was
| something clearly beyond what computers would ever be
| able to do. The implication was that there was some magic
| sauce that human brains had that couldn't be replicated
| in silicon (by us). That 'facility with language'
| argument has clearly fallen apart over the last 3 years
| and been replaced with what appears to be a different
| magic sauce comprised of the phrases 'not really
| thinking' and the whole 'just repeating what it's
| heard/parrot' argument.
|
| I don't think LLM's think or will reach AGI through
| scaling and I'm skeptical we're particularly close to AGI
| in any form. But I feel like it's a matter of incremental
| steps. There isn't some magic chasm that needs to be
| crossed. When we get there I think we will look back and
| see that 'legitimately thinking' wasn't anything magic.
| We'll look at AGI and instead of saying "isn't it amazing
| computers can do this" we'll say "wow, was that all there
| is to thinking like a human".
| windexh8er wrote:
| > 5 years ago a typical argument against AGI was that
| computers would never be able to think because "real
| thinking" involved mastery of language which was
| something clearly beyond what computers would ever be
| able to do.
|
| Mastery of words is thinking? In that line of argument
| then computers have been able to think for decades.
|
| Humans don't think only in words. Our context, memory and
| thoughts are processed and occur in ways we don't
| understand, still.
|
| There's a lot of great information out there describing
| this [0][1]. Continuing to believe these tools are
| thinking, however, is dangerous. I'd gather it has
| something to do with logic: you can't see the process and
| it's non-deterministic so it feels like thinking. ELIZA
| tricked people. LLMs are no different.
|
| [0] https://archive.is/FM4y8 [0]
| https://www.theverge.com/ai-artificial-
| intelligence/827820/l... [1]
| https://www.raspberrypi.org/blog/secondary-school-maths-
| show...
| CamperBob2 wrote:
| _Mastery of words is thinking?_
|
| That's the crazy thing. Yes, in fact, it turns out that
| language encodes and embodies reasoning. All you have to
| do is pile up enough of it in a high-dimensional space,
| use gradient descent to model its original structure, and
| add some feedback in the form of RL. At that point,
| reasoning is just a database problem, which we currently
| attack with attention.
|
| No one had the faintest clue. Even now, many people not
| only don't understand what just happened, but they don't
| think anything happened at all.
|
| ELIZA, ROFL. How'd ELIZA do at the IMO last year?
| meindnoch wrote:
| So people without language cannot reason? I don't think
| so.
| CamperBob2 wrote:
| There's no such thing as people without language, except
| for infants and those who are so mentally incapacitated
| that the answer is self-evidently "No, they cannot."
|
| Language is the substrate of reason. It doesn't need to
| be spoken or written, but it's a necessary and (as it
| turns out) sufficient component of thought.
| windexh8er wrote:
| There are quite a few studies to refute this highly
| ignorant comment. I'd suggest some reading [0].
|
| From the abstract: _" Is thought possible without
| language? Individuals with global aphasia, who have
| almost no ability to understand or produce language,
| provide a powerful opportunity to find out.
| Astonishingly, despite their near-total loss of language,
| these individuals are nonetheless able to add and
| subtract, solve logic problems, think about another
| person's thoughts, appreciate music, and successfully
| navigate their environments. Further, neuroimaging
| studies show that healthy adults strongly engage the
| brain's language areas when they understand a sentence,
| but not when they perform other nonlinguistic tasks like
| arithmetic, storing information in working memory,
| inhibiting prepotent responses, or listening to music.
| Taken together, these two complementary lines of evidence
| provide a clear answer to the classic question: many
| aspects of thought engage distinct brain regions from,
| and do not depend on, language."_
|
| [0] https://pmc.ncbi.nlm.nih.gov/articles/PMC4874898/
| arcatech wrote:
| > I will find this often-repeated argument compelling
| only when someone can prove to me that the human mind
| works in a way that isn't 'combining stuff it learned in
| the past'.
|
| This is the definition of the word 'novel'.
| handoflixue wrote:
| Seriously, all that familiarity and you think an LLM
| "literally" can't invent anything that didn't already
| exist?
|
| Like, I'm sorry, but you're just flat-out wrong and I've
| got the proof sitting on my hard drive. I use this
| supposedly impossible program daily.
| bigyabai wrote:
| FWIW, your "evidence" is a text editor. I'm glad you made
| a tool that works for you, but the parent's point stands;
| this is a 200-level course-curriculum homework
| assignment. Tens of thousands of homemade editors exist,
| in various states of disrepair and vain overengineering.
| least wrote:
| The difference between those is the person is actually
| using this text editor that they built with the help of
| LLMs. There's plenty of people creating novel scripts and
| programs that can accommodate their own unique
| specifications.
|
| If a programmer creating their own software (or
| contracting it out to a developer) would be a bespoke
| suit and using software someone or some company created
| without your input is an off the rack suit, I'd liken
| these sorts of programs as semi-bespoke, or made to
| measure.
|
| "LLMs are literally technology that can only reproduce
| the past" feels like an odd statement. I think the point
| they're going for is that it's not thinking and so it's
| not going to produce new ideas like a human would? But
| literally no technology does that. That is all derived
| from some human beings being particularly clever.
|
| LLMs are tools. They can enable a human to create new
| things because they are interfacing with a human to
| facilitate it. It's merging the functional knowledge and
| vision of a person and translating it into something
| else.
| resize2996 wrote:
| compilers can only produce machine code. so unorginal.
| windexh8er wrote:
| Do you also think LLMs "think"?
|
| From what you've described an LLM has not invented
| anything. LLMs that can reason have a bit more slight of
| hand but they're not coming up with new ideas outside of
| the bounds of what a lot of words have encompassed in
| both fiction and non.
|
| Good for you that you've got a fun token of code that's
| what you've always wanted, I guess. But this type of
| fantasy take on LLMs seems to be more and more prevalent
| as of late. A lot of people defending LLMs as if they're
| owed something because they've built something or maybe
| people are getting more and more attached to them from
| the conversational angle. I'm not sure, but I've run
| across more people in 2025 that are way too far in the
| deep end of personifying their relationships with LLMs.
| Kerrick wrote:
| Hang on, you're now saying that if something has ever
| been described in _fiction_ it doesn 't count as
| invention? So if somebody literally developed a working
| photon torpedo, that isn't new because "Star Trek Did
| It"?
| phatfish wrote:
| Is there any danger an LLM is going to create a working
| photo torpedo?
| ben_w wrote:
| Well, they can use tools, and tools includes physics
| simulations, so if it is possible (and FWIW the tool-free
| "intuition" of ChatGPT is "there will never be an age of
| antimatter"), then why couldn't LLMs grind those tools to
| get a solution?
| windexh8er wrote:
| You seem to be pretty far down the rabbit hole. How about
| this... You task an LLM to create a photon torpedo. If it
| can truly think then it should be able to provide you
| with something tangible. When you've got that in hand let
| us all know.
|
| Back to the land of reality... Describing something in
| fiction doesn't magically make it "not an invention".
| Fiction can anticipate an idea, but invention is about
| producing a working, testable implementation and usually
| involves novel technical methods. "Star Trek did it" is
| at most prior art for the concept, not a blueprint for
| the mechanism. If you can't understand that differential
| then maybe go ask an LLM.
| Kerrick wrote:
| I didn't say anything about an LLM. I said "somebody" not
| "some predictive text engine."
| ctxc wrote:
| Some people cannot be convinced simply because their
| expectation of "novel" is something that appears in an
| Asimov novel.
|
| I for one think your work is pretty cool - even though I
| haven't seen it, using something you built everyday is a
| claim not many can make!
| 9rx wrote:
| When a computer is able to invent things, we've achieved
| AGI. Do you believe we are already in the AGI era, or is
| the inventor in this case actually you?
| threethirtytwo wrote:
| Over half of HN still thinks it's a stochastic parrot and
| that it's just a glorified google search.
|
| The change hit us so fast a huge number of people don't
| understand how capable it is yet.
|
| Also it certainly doesn't help that it still
| hallucinates. One mistake and it's enough to set someone
| against LLMs. You really need to push through that
| hallucinations are just the weak part of the process to
| see the value.
| CamperBob2 wrote:
| The problem I see, over and over, is that people pose
| poorly-formed questions to the free ChatGPT and Google
| models, laugh at the resulting half-baked answers that
| are often full of errors and hallucinations, and draw
| conclusions about the technology as a whole.
|
| Either that, or they tried it "last year" or "a while
| back" and have no concept of how far things have gone in
| the meantime.
|
| It's like they wandered into a machine shop, cut off a
| finger or two, and concluded that their grandpa's hammer
| and hacksaw were all anyone ever needed.
| habinero wrote:
| No, frankly it's the difference between actual engineers
| and hobbyists/amateurs/non-SWEs.
|
| SWEs are trained to discard surface-level observations
| and be adversarial. You can't just look at the happy
| path, how does the system behave for edge cases? Where
| does it break down and how? What are the failure modes?
|
| The actual analogy to a machine shop would be to look at
| whether the machines were adequate for their use case,
| the building had enough reliable power to run and if
| there were any safety issues.
|
| It's easy to Clever Hans yourself and get snowed by what
| _looks_ like sophisticated effort or flat out bullshit. I
| had to gently tell a junior engineer that just because
| the marketing claims something will work a certain way,
| that doesn 't mean it will.
| CamperBob2 wrote:
| You sound pretty certain. There's often good money to be
| made in taking the contrarian view, where you have
| insights that the so-called "smart money" lacks. What are
| some good investments to make in the extreme-bear case,
| in which we're all just Clever Hans-ing ourselves as you
| put it? Do you have skin in the game?
| threethirtytwo wrote:
| What you're describing is just competent engineering, and
| it's already been applied to LLMs. People have been
| adversarial. That's why we know so much about
| hallucinations, jailbreaks, distribution shift failures,
| and long-horizon breakdowns in the first place. If this
| were hobbyist awe, none of those benchmarks or red-
| teaming efforts would exist.
|
| The key point you're missing is the type of failure.
| Search systems fail by not retrieving. Parrots fail by
| repeating. LLMs fail by producing internally coherent but
| factually wrong world models. That failure mode only
| exists if the system is actually modeling and reasoning,
| imperfectly. You don't get that behavior from lookup or
| regurgitation.
|
| This shows up concretely in how errors scale. Ambiguity
| and multi-step inference increase hallucinations.
| Scaffolding, tools, and verification loops reduce them.
| Step-by-step reasoning helps. Grounding helps. None of
| that makes sense for a glorified Google search.
|
| Hallucinations are a real weakness, but they're not
| evidence of absence of capability. They're evidence of an
| incomplete reasoning system operating without sufficient
| constraints. Engineers don't dismiss CNC machines because
| they crash bits. They map the envelope and design around
| it. That's what's happening here.
|
| Being skeptical of reliability in specific use cases is
| reasonable. Concluding from those failure modes that this
| is just Clever Hans is not adversarial engineering. It's
| stopping one layer too early.
| Greduan wrote:
| Text editors in a thousand flavours has indeed already been
| programmed though. I don't think you understood what op
| meant.
|
| Curious, does it perform at the limit of the hardware? Was
| it programmed in a tools language (like C++, Rust, C, etc.)
| or in a web tech?
| zingar wrote:
| What is the point that you believe would be demonstrated
| by a new text editor running at the limit of hardware in
| a compiled editor? Would that point apply to every other
| text editor that exists already?
| fmbb wrote:
| Is your new text editor open source?
| nsxwolf wrote:
| The LLM didn't invent any new technology to do that,
| though. You used the LLM to reorganize Lego building blocks
| of knowledge into something new.
|
| Without you, there was nothing.
| waldrews wrote:
| I'm being hyperbolic of course, but I'm a little dismissive
| of the progress that happened since the days of BBS's and car
| based cell phones - we just got more connectivity, more
| capacity, more content, bigger/faster. Likewise, my attitude
| toward machine learning before 2023 is a smug 'heh, these
| computer scientists are doing undisciplined statistics at
| scale, how nice for them.' Then all of a sudden the machines
| woke up and started arguing with me, coherently, even about
| niche topics I have a PhD in. I can appreciate in retrospect
| how much of the machine learning progress ultimately went
| into that, but, like fusion, the magic payoff was supposed to
| be decades away and always remain decades away. This wasn't
| supposed to happen in my lifetime. 2025 progress isn't the
| 2023 shock, but this was the year LLM's-as-programmers (and
| LLM's-as-mathematicians, and...) went from 'isn't that cute,
| the machine is trying' to 'an expert with enough time would
| make better choices than the machine did,' and that makes for
| a different world. More so than, going from a Commodore Vic
| 20 with 4k of RAM and a modem to the latest Macbook.
| ako wrote:
| > This year honestly feels quite stagnant. LLMs are literally
| technology that can only reproduce the past.
|
| Is this such a big limitation? Most jobs are basically people
| trained on past knowledge applying it today. No need to
| generate new knowledge.
|
| And a lot of new knowledge is just combining 2 things from
| the past in a new way.
| HarHarVeryFunny wrote:
| > LLMs are literally technology that can only reproduce the
| past.
|
| That's incorrect on many levels. They are drawing upon, and
| reproducing, language patterns from "the past", but they are
| combining those patterns in ways that may have never have
| been seen before. They may not be truly creative, but they
| are still capable of generating novel outputs.
|
| > They're cool, but they were way cooler 4 years ago.
|
| Maybe this year has been more about incremental progress with
| LLMs than the shock/coolness factor of talking to an LLM for
| the first time, but the utility of them, especially for
| programming, has dramatically increased this year, really in
| the last 6 months.
|
| The improvement in "AI" image and video generation has also
| been impressive, to the point now that fake videos on YouTube
| can often only be identified as such by common sense rather
| that the fact that they don't look real.
|
| Incremental improvement can often be more impressive that
| innovation, whose future importance can be hard to judge when
| it first appears. How many people read "Attention is all you
| need" in 2017 and thought "Wow! This is going to change the
| world!". Not even the authors of the paper thought that.
| odiroot wrote:
| I'm very relieved we've moved away from rewriting everything in
| Rust.
| jll29 wrote:
| There's no reason not to use Rust for LLM-generated code in
| the longer term (other than lack of Rust code to learn from
| in the shorter term).
|
| The stricter typing of Rust would make sematic errors in
| generated code come out more quickly than in e.g. Python
| because using static typing the chances are that some of the
| semantic errors are also type violations.
| michaelcampbell wrote:
| Have we though? I'm glad we're not shouting about it from the
| rooftops like it's some magical "win" button as much, but TBH
| the things I use routinely that HAVE been rewritten in rust
| are generally much better. That could also just be because
| they're newer and have the errors of the past to not repeat.
| sanreau wrote:
| > Vendor-independent options include GitHub Copilot CLI, Amp,
| OpenHands CLI, and Pi
|
| ...and the best of them all, OpenCode[1] :)
|
| [1]: https://opencode.ai
| simonw wrote:
| Good call, I'll add that. I think I mentally scrambled it with
| OpenHands.
| the_mitsuhiko wrote:
| Thanks for adding pi to it though :)
| nineteen999 wrote:
| How did I miss this until now! Thank you for sharing.
| logicprog wrote:
| I don't know why you're downloaded, OpenCode is by far the
| best.
| d4rkp4ttern wrote:
| Can OpenCode be used with the Claude Max or ChatGPT Pro
| subscriptions, i.e., without per-token API charges?
| simonw wrote:
| Apparently it does work with Claude Max:
| https://opencode.ai/docs/providers/#anthropic
|
| I don't see a similar option for ChatGPT Pro. Here's a closed
| issue: https://github.com/sst/opencode/issues/704
| williamstein wrote:
| There's a plugin that evidently supports ChatGPT Pro with
| Opencode: https://github.com/sst/opencode/issues/1686#issue
| comment-349...
| ewoodrich wrote:
| Yes, I use it with a regular Claude Pro subscription. It also
| supports using GitHub Copilot subscriptions as a backend.
| the_mitsuhiko wrote:
| > The (only?) year of MCP
|
| I like to believe, but MCP is quickly turning into an enterprise
| thing so I think it will stick around for good.
| simonw wrote:
| I think it will stick around, but I don't think it will have
| another year where it's the hot thing it was back in January
| through May.
| Alex-Programs wrote:
| I never quite got what was so "hot" about it. There seems to
| be an entire parallel ecosystem of corporates that are just
| begging to turn AI into PowerPoint slides so that they can
| mould it into a shape that's familiar.
| 9dev wrote:
| One reason may be that it makes it a lot easier to open up
| a product to AI. Instead of adding a bad ChatGPT UI clone
| into your app, you inverse control and let external AI
| tools interact with your application and its data, thus
| giving your customers immediate benefits, while
| simultaneously sating your investors/founders/managers
| desire to somehow add AI.
| nrhrjrjrjtntbt wrote:
| MCP or skills? Can a skill negate the need for MCP. In addition
| there was a YC startup who is looking at searching docs for
| LLMs or similar. I think MCP may be less needed once you have
| skills, openapi specs, and other things that LLMs can call
| directly.
| MitziMoto wrote:
| MCP isn't going anywhere. Some developers can't seem to see
| past their terminal or dev environment when it comes to MCP.
| Skills, etc do not replace MCP and MCP is far more than just
| documentation searching.
|
| MCP is a great way for an LLM to connect to an external system
| in a standardized way and immediately understand what tools it
| has available, when and how to use them, what their inputs and
| outputs are,etc.
|
| For example, we built a custom MCP server for our CRM. Now our
| voice and chat agents that run on elevenlabs infrastructure can
| connect to our system with one endpoint, understand what
| actions it can take, and what information it needs to collect
| from the user to perform those actions.
|
| I guess this could maybe be done with webhooks or an API spec
| with a well crafted prompt? Or if eleven labs provided an
| executable environment with tool calling? But at some point
| you're just reinventing a lot of the functionality you get for
| free from MCP, and all major LLMs seem to know how to use MCP
| already.
| simonw wrote:
| Yeah, I don't think I was particularly clear in that section.
|
| I don't think MCP is going to go away, but I do think it's
| unlikely to ever achieve the level of excitement it had in
| early 2025 again.
|
| If you're not building inside a code execution environment
| it's a very good option for plugging tools into LLMs,
| especially across different systems that support the same
| standard.
|
| But code execution environments are so much more powerful and
| flexible!
|
| I expect that once we come up with a robust, inexpensive way
| to run a little Bash environment - I'm still hoping
| WebAssembly gets us there - there will be much less reason to
| use MCP even outside of coding agent setups.
| brabel wrote:
| I disagree. MCP will remain the best way to do most things
| for the same reason REST APIs are the main way to access
| non local services: they provide a way to secure and audit
| access to systems in a way that a coding environment
| cannot. And you can authorize actions depending on the well
| defined inputs and outputs. You can't do that using just a
| bash script unless said script actually does SSO and calls
| REST APIs but then you just have a worse MCP client without
| any interoperability.
| the_mitsuhiko wrote:
| I find it very hard to pick winners and losers in this
| environment where everything changes so quickly. Right
| now a lot of people are using bash as a glue environment
| for agents, even if they are not for developers.
| cloudking wrote:
| For connecting agents to third-party systems I prefer CLI
| tools, less context bloat and faster. You can define the CLI
| usage in your agent instructions. If the MCP you're using
| doesn't exist as a CLI, build one with your agent.
| npalli wrote:
| Great summary of the year in LLMs. Is there a predictions (for
| 2026) blogpost as well?
| simonw wrote:
| Given how badly my 2025 predictions aged I'm probably going to
| sit that one out! https://simonwillison.net/2025/Jan/10/ai-
| predictions/
| DANmode wrote:
| Don't be a bad sport, now!!
| zahlman wrote:
| Making predictions is useful even when they turn out very
| wrong. Consider also giving confidence levels, so that you
| can calibrate going forward.
| jjude wrote:
| I use predictions to prepare rather than to plan.
|
| Planing depends on deterministic view of the future. I used
| to plan (esp annual plans) until about 5 years. Now I scan
| for trends and prepare myself for different scenarios that
| can come in the future. Even if you get it approximately
| right, you stand apart.
|
| For tech trends, I read Simon, Benedict Evans, Mary Meeker
| etc. Simon is in a better position make these predictions
| than anyone else having closely analyzed these trends over
| the last few years.
|
| Here I wrote about my approach:
| https://www.jjude.com/shape-the-future/
| skydhash wrote:
| [flagged]
| MattRix wrote:
| [flagged]
| skydhash wrote:
| Why do people assume negative critique is ignorance?
| dmd wrote:
| People denied that bicycles could possibly balance even as
| others happily pedaled by. This is the same thing.
| measurablefunc wrote:
| Bicycles don't balance, the human on the bicycle is the
| one doing the balancing.
| dmd wrote:
| Yes, that is the analogy I am making. People argued that
| bicycles (a tool for humans to use) could not possibly
| work - even as people were successfully using them.
| measurablefunc wrote:
| People use drugs as well but I'm not sure I'd call that
| successful use of chemical compounds without further
| context. There are many analogies one can apply here that
| would be equally valid.
| duchef wrote:
| Bicycles (without a rider) do balance at sufficient speed
| via a self steering and correction mechanism of the front
| axle..
| skydhash wrote:
| Please tell me which one of the headings is not about
| increased usage o LLMs and derived tools and is about
| some improvement in the axes of reliability or or any
| kind of usefulness.
|
| Here is the changelog for OpenBSD 7.8:
|
| https://www.openbsd.org/78.html
|
| There's nothing here that says: We make it easier to use
| it more of it. It's about using it better and fixing
| underlying problems.
| simonw wrote:
| The coding agent heading. Claude Code and tools like it
| represent a _huge_ improvement in what you can usefully
| get done with LLMs.
|
| Mistakes and hallucinations matter a whole lot less if a
| reasoning LLM can try the code, see that it doesn't work
| and fix the problem.
| walt_grata wrote:
| If it actually does that without an argument. I can't
| believe I have to say that about a computer program
| skydhash wrote:
| > The coding agent heading. Claude Code and tools like it
| represent a huge improvement in what you can usefully get
| done with LLMs.
|
| Does it? It's all prompt manipulation. Shell script are
| powerful yes, but not really _huge_ improvement over
| having a shell (REPL interface) to the system. And even
| then a lot of programs just use syscalls or wrapper
| libraries.
|
| > can try the code, see that it doesn't work and fix the
| problem.
|
| Can you really say that does happens reliably?
| simonw wrote:
| Depends on what you mean by "reliably".
|
| If you mean 100% correct all of the time then no.
|
| If you mean correct often enough that you can expect it
| to be a productive assistant that helps solve all sorts
| of problems faster than you could solve them without it,
| and which makes mistakes infrequently enough that you
| waste less time fixing them than you would doing
| everything by yourself then yes, it's plenty reliable
| enough now.
| dham wrote:
| You're welcome to try the LLM's yourself and come up with
| your own conclusions. By what you've posted it doesn't
| look like you've tried the anything in the last 2 years.
| Yes LLM's can be annoying, but there has been progress.
| noodletheworld wrote:
| I know it seems like forever ago, but claude code only
| came out in 2025.
|
| Its very difficult to argue the point that claude code:
|
| 1) was a paradigm shift in terms of functionality,
| despite, to be fair, at best, incremental improvements in
| the underlying models.
|
| 2) The results are an order of magnitude, I estimate,
| better in terms of output.
|
| I think its very fair to distill "AI progress 2025" to:
| you can get better results (up to a point; better than
| _raw output_ anyway; scaling to multiple agents has not
| worked) without better models with clever tools and
| loops. (...and video /image slop infests everything :p).
| bandrami wrote:
| Did more software ship in 2025 than in 2024? I'm still
| looking for some actual indication of output here. I get
| that people _feel_ more productive but the actual metrics
| don 't seem to agree.
| skydhash wrote:
| I'm still waiting for the Linux drivers to be written
| because of all the 20x improvements that AI hypers are
| touting. I would even settle for Apple M3 and M4
| computers to be supported by Asahi.
| noodletheworld wrote:
| I am not making any argument about productivity about
| using AI vs. not using AI.
|
| My point is purely that, _compared to 2024_ , the quality
| of the code produced by LLM inference _agent systems_ is
| better.
|
| To say that 2025 was a nothing burger is objectively
| incorrect.
|
| Will it scale? Is it good enough to use professionally?
| Is this like self driving cars where the best they ever
| get is stuck with an odd shaped traffic cone? Is it
| actually more productive?
|
| Who knows?
|
| Im just saying... LLM coding in 2024 sucked. 2025 was a
| big year.
| tehnub wrote:
| People did?
| rhubarbtree wrote:
| It's possible this is correct.
|
| It's also possible that people more experienced,
| knowledgable and skilled than you can see fundamental
| flaws in using LLMs for software engineering that you
| cannot. I am not including myself in that category.
|
| I'm personally honestly undecided. I've been coding for
| over 30 years and know something like 25 languages. I've
| taught programming to postgrad level, and built prototype
| AI systems that foreshadowed LLMs, I've written
| everything from embedded systems to enterprise, web,
| mainframes, real time, physics simulation and research
| software. I would consider myself an 7/10 or 8/10 coder.
|
| A lot of folks I know are better coders. To put my
| experience into context: one guy in my year at uni wrote
| one of the world's most famous crypto systems; another
| wrote large portions of some of the most successful games
| of the last few decades. So I've grown up surrounded by
| geniuses, basically, and whilst I've been lectured by
| true greats I'm humble enough to recognise I don't bleed
| code like they do. I'm just a dabbler. But it irks me
| that a lot of folks using AI profess it's the future but
| don't really know anything about coding compared to these
| folks. Not to be a Luddite - they are the first people to
| adopt new languages and techniques, but they also are
| super sceptical about anything that smells remotely like
| bullshit.
|
| One of the most wise insights in coding is the
| aphorism"beware the enthusiasm of the recently
| converted." And I see that so much with AI. I've seen it
| with compilers, with IDEs, paradigms, and languages.
|
| I've been experimenting a lot with AI, and I've found it
| fantastic for comprehending poor code written by others.
| I've also found it great for bouncing ideas. And the code
| it writes, beyond boiler plate, is hot garbage. It
| doesn't properly reason, it can't design architecture, it
| can't write code that is comprehensible to other
| programmers, and treating it as a "black box to be
| manipulated by AI" just leads to dead ends that can't be
| escaped, terrible decisions that will take huge amounts
| of expert coding time to undo, subtle bugs that AI can't
| fix and are super hard to spot, and often you can't
| understand their code enough to fix them, and security
| nightmares.
|
| Testing is insufficient for good code. Humans write code
| in a way that is designed for general correctness. AI
| does not, at least not yet.
|
| I do think these problems can be solved. I think we
| probably need automated reasoning systems, or else vastly
| improved LLMs that border on automated reasoning much
| like humans do. Could be a year. Could be a decade. But
| right now these tools don't work well. Great for vibe
| coding, prototyping, analysis, review, bouncing ideas.
| CamperBob2 wrote:
| _But right now these tools don't work well. Great for
| vibe coding, prototyping, analysis, review, bouncing
| ideas._
|
| What are some of the models you've been working with?
| blibble wrote:
| people also said that selling jpegs of monkeys for
| millions of dollars was a pump and dump scam, and would
| collapse
|
| they were right
| sothatsit wrote:
| JPEGs with no value other than fake scarcity is very
| different to coding agents that people actively use to
| ship real code.
| kakapo5672 wrote:
| Whenever someone tells me that AI is worthless, does
| nothing, scam/slop etc, I ask them about their own AI
| usage, and their general knowledge about what's going on.
|
| Invariably they've never used AI, or at most very rarely.
| (If they used AI beyond that, this would be admission that
| it was useful at some level).
|
| Therefore it's reasonable to assume that you are in that
| boat. Now that might not be true in your case, who knows,
| but it's definitely true on average.
| snigsnog wrote:
| It's not worthless, it's just not worldchanging as is
| even in the fields where it's most useful, like
| programming. If the trajectory changes and we reach AGI
| then this changes too but right now it's just a way to
|
| - fart out demos that you don't plan on maintaining, or
| want to use as a starting place
|
| - generate first-draft unit tests/documentation
|
| - generate boilerplate without too much functionality
|
| - refactor in a very well covered codebase
|
| It's very useful for all of the above! But it doesn't
| even replace a junior dev at my company in its current
| state. It's too agreeable, makes subtle mistakes that it
| can't permanently correct (GEMINI.md isn't a magic
| bullet, telling it to not do something does not guarantee
| that it won't do it again), and you as the developer
| submitting LLM-generated code for review need to review
| it closely before even putting it up (unless you feel
| like offloading this to your team) to the point that it's
| not that much faster than having written it yourself.
| LewisVerstappen wrote:
| because your "negative critique" is just idiotic and wrong
| sothatsit wrote:
| You did not make a negative critique. You completely
| dismissed the value of coding agents on the basis that the
| results are not predictable, which is both obvious and
| doesn't matter in practice. Anyone who has given these
| tools a chance will quickly realise that 1) they are
| actually quite predictable in doing what you ask them to,
| and 2) them being non-deterministic does not at all negate
| their value. This is why people can immediately tell you
| haven't used these tools, because your argument as to why
| they're useless is so elementary.
| dang wrote:
| Please don't respond to a bad comment by breaking the site
| guidelines yourself. That only makes things worse.
|
| https://news.ycombinator.com/newsguidelines.html
| senordevnyc wrote:
| This comment is legitimately hilarious to me. I thought it was
| satire at first. The list of what has happened in this field in
| the last twelve months is _staggering_ to me, while you write
| it off as essentially nothing.
|
| Different strokes, but I'm getting so much more done and mostly
| enjoying it. Can't wait to see what 2026 holds!
| ronsor wrote:
| People who dislike LLMs are generally insistent that they're
| useless for everything and have infinitely negative value,
| regardless of facts they're presented with.
|
| Anyone that believes that they are completely useless is just
| as deluded as anyone that believes they're going to bring an
| AGI utopia next week.
| n2d4 wrote:
| This is extremely dismissive. Claude Code helps me make a
| majority of changes to our codebase now, particularly small
| ones, and is an insane efficiency boost. You may not have the
| same experience for one reason or another, but plenty of devs
| do, so "nothing happened" is absolutely wrong.
|
| 2024 was a lot of talk, a lot of "AI could hypothetically do
| this and that". 2025 was the year where it genuinely started to
| enter people's workflows. Not everything we've been told would
| happen has happened (I still make my own presentations and
| write my own emails) but coding agents certainly have!
| bandrami wrote:
| Did you ship more in 2025 than in 2024?
| wickedsight wrote:
| I definitely did.
| GCUMstlyHarmls wrote:
| Shipping in 2025:
| https://x.com/trq212/status/2001848726395269619
| DANmode wrote:
| I _definitely_ did.
|
| Objectively 0->1 lots of backlog.
| skydhash wrote:
| And this is one of the _vague_ "AI helped me do more".
|
| This is me touting for Emacs
|
| _Emacs was a great plus for me over the last year. The
| integration with various tooling with comint (REPL
| integration), compile (build or report tools), TUI (through
| eat or ansi-term), gave me a unified experience through the
| buffer paradigm of emacs. Using the same set of commands
| boosted my editing process and the easy addition of new
| commands make it easy to fit my development workflow to the
| editor._
|
| This is how easy it is to write a non-vague "tool X helped
| me" and I'm not even an English native speaker.
| n2d4 wrote:
| That paragraph could be the truth, or it could be a lie.
| Maybe Emacs really did make you more efficient, or you made
| it all up, I don't know. Best I can do is trust you.
|
| If you don't trust me, I can't conclusively convince you
| that AI makes me more efficient, but if you want I'm happy
| to hop on a screen-share and elaborate in what ways it has
| boosted my workflow. I'm offering this because I'm also
| curious what _your_ work looks like where AI cannot help at
| all.
|
| E-mail address is on my profile!
| thunky wrote:
| > This is how easy it is to write a non-vague "tool X
| helped me" and I'm not even an English native speaker.
|
| Your example is very vague.
|
| See if you can spot the problem in my review of Excel in
| your style:
|
| "It's great and I like how it's formula paradigm gave me a
| unified experience. It's table features boosted my science
| workflows last year".
| dang wrote:
| Could you please stop posting dismissive, curmudgeonly
| comments? It's not what this site is for, and destroys what it
| is for.
|
| We want _curious_ conversation here.
|
| https://news.ycombinator.com/newsguidelines.html
| Madmallard wrote:
| His comment is far better than the rampant astroturfing from
| stakeholders going on everywhere on this website that is
| being mitigated not at all whatsoever. There is a wealth of
| information present suggesting these things are so bad for
| everyone in so many ways.
| clawedcod wrote:
| These people love generated content (like, they'll actually
| read generated blog post word-for-word and not even be
| angry; they'll skip a personal email for its machine
| summary) and they can generate all the content they'd ever
| want. If they want to take over HN this isn't a battle
| we're going to win except with aggressive moderation, and
| we know who feeds the mods.
|
| HN isn't a place for thinking people any more (a long time
| coming, but you could squint and pretend until recently).
| Happy new year and adios, thanks for the 100s of accounts
| dang. Double pinky swear I won't make another.
| dang wrote:
| Normally everyone who publicly declares they're done with
| this site, will never make a new account again, etc.
| etc., either has already made their next account or will
| do so shortly. HN, for all its eternal decline generating
| endless complaints, seems to be irresistible to this sort
| of complainer.
| dang wrote:
| What are some specific links to the rampant astroturfing
| that you feel is going on on this website and which
| https://news.ycombinator.com/item?id=46450296 is better
| than?
| Madmallard wrote:
| Let's see, going off of just top-level comments in this
| thread alone:
|
| didip, timonoko, mark_I_watson, icapybara, _pdp_,
| agentifysh, sanreau,
|
| There's no way to know if these are genuine thoughts or
| incentivized compelled speech.
|
| nativeit has a good way of putting it.
|
| Your replies to "anonnon" make me less than hopeful for
| the future of HN in regards to AI. Seems like this might
| be trending in the direction of Reddit, where the
| interests are basically all paid for and imposed rather
| than being genuine and organic, and dissent is
| aggressively shut out.
|
| "Curious conversation" does not really apply when it is
| compelled via monetary interest without any consideration
| toward potentially serious side effects.
|
| "At least when herding cats, you can be sure that if the
| cats are hungry, they will try to get where the food is."
| This part of the guy's comment is actually funny and apt.
| Somehow that escaped you when you wrote your threat
| reply. That makes me wonder how mind-controlled you are.
|
| "yupyupyups" has a small summary of some of the
| negatives, yet is being flagged. "techpression" similarly
| does, though is a bit more negative in his remarks. Also
| being flagged.
|
| So the whole thread reads like this: 1.) talking about
| benefits? bubble to the top 2.) criticize? Either
| threatened by Dang or flagged to the bottom
|
| Sounds a whole lot like compelled speech to me. Sounds a
| whole lot like mind-control.
|
| It's pretty sad to see really.
|
| It might just be your rule system. I personally want to
| see criticism. I don't have the sensitivity you have
| toward personal attacks or what you "deem" personal
| attacks when it is text on-screen. I don't care. I want
| to see what useful information might come out of it. I
| think your policing just makes everything worse to be
| honest. The thread will just die out in a day anyway.
|
| I think I have criticized it in the past and you or some
| other staff said that it's a slippery slope toward
| useless aggressive banter that derails topics, but I
| don't know. I really don't agree with it. That's just my
| life experience.
|
| Reddit is kind of like this. And it's basically turned
| into imposed topics rather than organic topics with
| massive amounts of echo-chambering in each delusional
| sub-reddit. Anything remotely against the grain is
| harshly culled as soon as possible. You can only imagine
| what the back-end looks like for that kind of thing.
| Money being involved at many steps is guaranteed.
|
| And yeah as another commenter pointed out, this one guy's
| blog being at the top of hacker news every time is
| potentially suspicious as well.
|
| I think I originally came to this place more than Reddit
| 10+ years ago because yeah it felt like people just
| excited and curious about their tech topics and it didn't
| feel like it was being rampantly policed or pushing a
| political agenda etc. I guess I should just not
| participate in these threads because the topic is tired
| on me at this point.
|
| Wait I just read your user page and this is actually
| hilarious:
|
| "Conflict is essential to human life, whether between
| different aspects of oneself, between oneself and the
| environment, between different individuals or between
| different groups. It follows that the aim of healthy
| living is not the direct elimination of conflict, which
| is possible only by forcible suppression of one or other
| of its antagonistic components, but the toleration of it
| --the capacity to bear the tensions of doubt and of
| unsatisfied need and the willingness to hold judgement in
| suspense until finer and finer solutions can be
| discovered which integrate more and more the claims of
| both sides. It is the psychologist's job to make possible
| the acceptance of such an idea so that the richness of
| the varieties of experience, whether within the unit of
| the single personality or in the wider unit of the group,
| can come to expression."
|
| Marion Milner, 'The Toleration of Conflict', Occupational
| Psychology, 17, 1, January 1943
|
| This made me immediately and uncontrollably guffaw.
| aussieguy1234 wrote:
| > The year of YOLO and the Normalization of Deviance #
|
| On this including AI agents deleting home folders, I was able to
| run agents in Firejail by isolating vscode (Most of my agents are
| vscode based ones, like Kilo Code).
|
| I wrote a little guide on how I did it
| https://softwareengineeringstandard.com/2025/12/15/ai-agents...
|
| Took a bit of tweaking, vscode crashing a bunch of times with not
| being able to read its config files, but I got there in the end.
| Now it can only write to my projects folder. All of my projects
| are backed up in git.
| NitpickLawyer wrote:
| I have a bunch of tabs opened on this exact topic, so thank you
| for sharing. So far I've been using devcontainers w/ vscode,
| and mostly having a blast with it. It is a bit awkward since
| some extensions need to be installed in the remote env, but
| they seem to play nicely after you have it setup, and the keys
| and stuff get populated so things like kilocode, cline, roo
| work fine.
| agentifysh wrote:
| What an amazing progress in just short time. The future is
| bright! Happy New Year y'all!
| sho_hn wrote:
| Not in this review: Also the record year in intelligent systems
| aiding in and prompting human users into fatal self-harm.
|
| Will 2026 fare better?
| simonw wrote:
| I really hope so.
|
| The big labs are (mostly) investing a lot of resources into
| reducing the chance their models will trigger self-harm and AI
| psychosis and suchlike. See the GPT-4o retirement (and
| resulting backlash) for an example of that.
|
| But the number of users is exploding too. If they make things
| 5x less likely to happen but sign up 10x more people it won't
| be good on that front.
| Nuzzerino wrote:
| How does a model "trigger" self-harm? Surely it doesn't
| catalyze the dissatisfaction with the human condition,
| leading to it. There's no reliable data that can drive
| meaningful improvement there, and so it is merely an
| appeasement op.
|
| Same thing with "psychosis", which is a manufactured moral
| panic crisis.
|
| If the AI companies really wanted to reduce actual self harm
| and psychosis, maybe they'd stop prioritizing features that
| lead to mass unemployment for certain professions. One of the
| guys in the NYT article for AI psychosis had a successful
| career before the economy went to shit. The LLM didn't create
| those conditions, bad policies did.
|
| It's time to stop parroting slurs like that.
| measurablefunc wrote:
| The people working on this stuff have convinced themselves
| they're on a religious quest so it's not going to get better:
| https://x.com/RobertFreundLaw/status/2006111090539687956
| andai wrote:
| Also essential self-fulfilment.
|
| But that one doesn't make headlines ;)
| sho_hn wrote:
| Sure -- but that's fair game in engineering. I work on cars.
| If we kill people with safety faults I expect it to make more
| headlines than all the fun roadtrips.
|
| What I find interesting with chat bots is that they're "web
| apps" so to speak, but with safety engineering aspects that
| type of developer is typically not exposed to or familiar
| with.
| simonw wrote:
| One of the tough problems here is privacy. AI labs really
| don't want to be in the habit of actively monitoring
| people's conversations with their bots, but they also need
| to prevent bad situations from arising and getting worse.
| walt_grata wrote:
| Until AI labs have the equivalent of an SLA for giving
| accurate and helpful responses it don't get better.
| They've not even able to measure if the agents work
| correctly and consistently.
| websiteapi wrote:
| I'm curious how all of the progress will be seen if it does
| indeed result in mass unemployment (but not eradication) of
| professional software engineers.
| simonw wrote:
| I nearly added a section about that. I wanted to contrast the
| thing where many companies are reducing junior engineering
| hires with the thing where Cloudflare and Shopify are hiring
| 1,000+ interns. I ran out of time and hadn't figured out a good
| way to frame it though so I dropped it.
| ori_b wrote:
| My prediction: If we can successfully get rid of most software
| engineers, we can get rid of most knowledge work. Given the
| state of robotics, manual labor is likely to outlive
| intellectual labor.
| beardedwizard wrote:
| "Given the state of robotics" reminds me a lot of what was
| said about llms and image/video models over the past 3 years.
| Considering how much llms improved, how long can robotics be
| in this state?
|
| I have to think 3 years from now we will be having the same
| conversation about robots doing real physical labor.
|
| "This is the worst they will ever be" feels more apt.
| chii wrote:
| but robotics had the means to do majority of the physical
| labour already - it's just not worth the money to replace
| humans, as human labour is cheap (and flexible - more than
| robots).
|
| With knowledge work being less high-paying, physical labour
| supply should increase as well, which drops their price.
| This means it's actually less likely that the advent of LLM
| will make physical labour more automated.
| Davidzheng wrote:
| Robotics is coming FAST. Faster than LLM progress in my
| opinion.
| wh0knows wrote:
| Curious if you have any links about the rapid progression
| of robotics (as someone who is not educated on the
| topic).
|
| It was my feeling with robotics that the more challenging
| aspect will be making them economically viable rather
| than simply the challenge of the task itself.
| beardedwizard wrote:
| I mentioned military in my reply to the sibling comment -
| that is the most ready example. What anduril and others
| are doing today may be sloppy, but it's moving very
| quickly.
| throw1235435 wrote:
| The question is how rapid the adoption is. The price of
| failure in the real world is much higher ($$$,
| environmental, physical risks) vs just
| "rebuild/regenerate" in the digital realm.
| beardedwizard wrote:
| Military adoption is probably a decent proxy indicator -
| and they are ready to hand the kill switch to autonomous
| robots
| throw1235435 wrote:
| Maybe. There the cost of failure again is low. Its easier
| to destroy than to create. Economic disruption to workers
| will take a bit longer I think.
|
| Don't get me wrong; I hope that we do see it in physical
| work as well. There is more value to society there; and
| consists of work that is risky and/or hard to do - and is
| usually needed (food, shelter, etc). It also means that
| the disruption is an "everyone" problem rather than
| something that just affects those "intellectual" types.
| BobbyJo wrote:
| I would have agreed with this a few months ago, but something
| Ive learned is that the ability to verify an LLMs output is
| paramount to its value. In software, you can review its
| output, add tests, on top of other adversarial techniques to
| verify the output immediately after generation.
|
| With most other knowledge work, I don't think that is the
| case. Maybe actuarial or accounting work, but most knowledge
| work exists at a cross section of function and taste, and the
| latter isn't an automatically verifiable output.
| throw1235435 wrote:
| I also believe this - I think it will probably just disrupt
| software engineering and any other digital medium with mass
| internet publication (i.e. things RLVR can use). For the
| short term future it seems to need a lot of data to train
| on, and no other profession has posted the same amount of
| verifiable material. The open source altruism has disrupted
| the profession in the end; just not in the way people first
| predicted. I don't think it will disrupt most knowledge
| work for a number of reasons. Most knowledge professions
| have "credentials' (i.e. gatekeeping) and they can see what
| is happening to SWE's and are acting accordingly. I'm
| hearing it firsthand at least locally in things like law,
| even accounting, etc. Society will ironically respect these
| professions more for doing so.
|
| Any data, verifiability, rules of thumb, tests, etc are
| being kept secret. You pay for the result, but don't know
| the means.
| coffeebeqn wrote:
| I mean law and accounting usually have a "right" answer
| that you can verify against. I can see a test data set
| being built for most professions. I'm sure open source
| helps with programming data but I doubt that's even the
| majority of their training. If you have a company like
| Google you could collect data on decades of software work
| in all its dimensions from your workforce
| District5524 wrote:
| It's not about invalidating your conclusion, but I'm not
| so sure about law having a right answer. At a very basic
| level, like hypothetical conduct used in basic legal
| training matrerials or MCQs, or in criminal/civil code
| based situations in well-abstracting Roman law-based
| jurisdictions, definitely. But the actual work, at least
| for most lawyers is to build on many layers of such
| abstractions to support your/client's viwepoint. And that
| level is already about persuasion of other people, not
| having the "right" legal argument or applying the most
| correct case found. And this part is not documented well,
| approaches changes a lot, even if law remains the same.
| Think of family law or law of succession - does not
| change much over centuries but every day, worldwide,
| millions of people spend huge amounts of money and energy
| on finding novel ways to turn those same paragraphs to
| their advantage and put their "loved" ones and relatives
| in a worse position.
| throw1235435 wrote:
| Not really. I used to think more general with the first
| generation of LLM's but given all progress since o1 is RL
| based I'm thinking most disruption will happen in open
| productive domains and not closed domains. Speaking to
| people in these professions they don't think SWE's have
| any self respect and so in your example of law:
|
| * Context is debatable/result isn't always clear: The way
| to interpret that/argue your case is different (i.e. you
| are paying for a service, not a product)
|
| * Access to vast training data: Its very unlikely that
| they will train you and give you data to their practice
| especially as they are already in a union like
| structure/accreditation. Its like paying for a binary (a
| non-decompilable one) without source code (the result)
| rather than the source and the validation the
| practitioner used to get there.
|
| * Variability of real world actors: There will be novel
| interpretations that invalidate the previous one as new
| context comes along.
|
| * Velocity vs ability to make judgement: As a lawyer I
| prefer to be paid higher for less velocity since it means
| less judgement/less liability/less risk overall for
| myself and the industry. Why would I change that even at
| an individual level? Less problem of the commons here.
|
| * Tolerance to failure is low: You can't iterate, get
| feedback and try again until "the tests pass" in a court
| room unlike "code on a text file". You need to have the
| right argument the first time. AI/ML generally only works
| where the end cost of failure is low (i.e can try again
| and again to iron out error terms/hallucinations). Its
| also why I'm skeptical AI will do much in the real
| economy even with robots soon - failure has bigger
| consequences in the real world ($$$, lives, etc).
|
| * Self employment: There is no tension between say Google
| shareholders and its employees as per your example -
| especially for professions where you must trade in your
| own name. Why would I disrupt myself? The cost I charge
| is my profit.
|
| TL;DR: Gatekeeping, changing context, and arms race
| behavior between participants/clients. Unfortunately I do
| think software, art, videos, translation, etc are unique
| in that there's numerous examples online and has the
| property "if I don't like it just re-roll" -> to me RLVR
| isn't that efficient - it needs volumes of data to build
| its view. Software sadly for us SWE's is the perfect
| domain for this; and we as practitioners of it made it
| that way through things like open source, TDD, etc and
| giving it away free on public platforms in numerous
| quantities.
| JumpCrisscross wrote:
| > _If we can successfully get rid of most software engineers,
| we can get rid of most knowledge work_
|
| Software, by its nature, is practically comprehensively
| digitized, both in its code history as well as requirements.
| 9dev wrote:
| That's the deep irony of technology IMHO, that innovation
| follows Conway's law on a meta layer: White collar workers
| inevitably shaped high technology after themselves, and
| instead of finally ridding humanity of hard physical labour--
| as was the promise of the Industrial Revolution--we imitate
| artists, scientists, and knowledge workers.
|
| We can now use natural language to instruct computers
| generate stock photos and illustrations that would take a
| professional artist a few years ago, discover new molecule
| shapes, beat the best Go players, build the code for entire
| applications, or write documents of various shapes and
| lengths--but painting a wall? An unsurmountable task that
| requires a human to execute reliably, not even talking about
| economics.
| Madmallard wrote:
| Why would it?
|
| The ability to accurately describe what you want with all
| constraints managed and with proactive design is the actual
| skill. Not programming. The day PMs can do that and have LLMs
| that can code to that, is the day software engineers en masse
| will disappear. But that day is likely never.
|
| The non-technical people I've ever worked for were hopelessly
| terrible at attention to detail. They're hiring me primarily
| for that anyway.
| legulere wrote:
| Even if it will make software engineering drastically more
| productive, it's questionable that this will lead to
| unemployment. Efficiency gains translate to lower prices.
| Sometimes this leads to very few additional demand, as can be
| seen with masses of typesetters that lost their jobs. Sometimes
| this leads to a dramatically higher demand like you can see in
| the classic Jevons paradox examples of coal and light bulbs. I
| highly suspect software falls in the latter category
| kingstnap wrote:
| Software demand is philosophically limited by the question of
| "What can your computer do for you?"
|
| You can describe that somewhat formally as:
|
| {What your computer can do} intersect {What you want done
| (consciously or otherwise)}
|
| Well a computer can technically calculate any computuable
| task that fits in bounded memory, that is an enormous set so
| its real limitations are its interfaces. In which case it can
| send packets, make noises, and display images.
|
| How many human desires are things that can be solved with
| making noises, displaying images, and sending packets? Turns
| out quite a few but its not everything.
|
| Basically I'm saying we should hope more sorts of physical
| interfaces come around (like VR and Robotics) so we cover
| more human desires. Robotics is a really general physical
| interface (like how ip packets are an extremely general
| interface) so its pretty promising if it pans out.
|
| Personally, I find it very hard to even articulate what
| desires I have. I have this feeling that I might be
| substantially happier if I was just sitting around a campfire
| eating food and chatting with people instead of enjoying
| whatever infinite stuff a super intelligent computer and
| robots could do for me. At least some of the time.
| fullstackchris wrote:
| This overly discussed thesis is already laughable - decent LLMs
| have been out for 3 years now and unemployment (using US as
| example) is up around 1% over the same time frame - and even
| attributing that small percentage change completely to AI is
| also laughable
| DrewADesign wrote:
| _You're absolutely right! You astutely observed that 2025 was a
| year with many LLMs and this was a selection of waypoints,
| summarized in a helpful timeline._
|
| That's what most non-tech-person's year in LLMs looked like.
|
| Hopefully 2026 will be the year where companies realize that
| implementing intrusive chatbots can't make better ::waving
| hands:: ya know... _UX_ or whatever.
|
| For some reason, they think its helpful to distractingly pop up
| chat windows on their site because their customers need textual
| kindergarten handholding to ... I don't know... find the ideal
| pocket comb for their unique pocket/hair situation, or had an
| unlikely question about that aerosol pan release spray that a
| chatbot could actually answer. Well, my dog also thinks she's
| helping me by attacking the vacuum when I'm trying to clean. Both
| ideas are equally valid.
|
| And spending a bazillion dollars implementing it doesn't mean
| your customers won't hate it. And forcing your customers into
| pathways they hate because of your sunk costs mindset means it
| will never stop costing you more money than it makes.
|
| I just hope companies start being honest with themselves about
| whether or not these things are good, bad, or absolutely abysmal
| for the customer experience and cut their losses when it makes
| sense.
| Night_Thastus wrote:
| They need to be intrusive and shoved in your face. This way,
| they can say they have a lot of people using them, which is a
| good and useful metric.
| ronsor wrote:
| > For some reason, they think its helpful to distractingly pop
| up chat windows on their site...
|
| Companies have been doing this "live support" nonsense far
| longer than LLMs have been popular.
| DrewADesign wrote:
| There was also source point pollution before the Industrial
| Revolution. Useless, forced, irritating chat was _'nowhere
| close'_ to as aggressive or pervasive as it is now. It used
| to be a niche feature of some CRMs and now it's _everywhere._
|
| I'm on LinkedIn Learning digging into something really
| technical and practical and it's constantly pushing the chat
| fly out with useless pre-populated prompts like "what are the
| main takeaways from this video." And they moved their main
| page search to a little icon on the title bar and sneakily
| now what used to be the obvious, primary central search field
| for years sends a prompt to their fucking chatbot.
| zahlman wrote:
| As much as I side with you on this one, I really don't think
| this submission is the right place to rant about it.
| fantasizr wrote:
| I took the good with the bad: the ai assisted coding tools are
| a multiplier, google ai overviews in search results are half
| baked (at best) and often just factually wrong. AI was put in
| the instagram search bar for no practical purpose etc.
| techpression wrote:
| Nothing about the severe impact on the environment, and the hand
| waviness about water usage hurt to read. The referenced post was
| missing every single point about the issue by making it global
| instead of local. And as if data center buildouts are properly
| planned and dimensioned for existing infrastructure...
|
| Add to this that all the hardware is already old and the amount
| of waste we're producing right now is mind boggling, and for
| what, fun tools for the use of one?
|
| I don't live in the US, but the amount of tax money being
| siphoned to a few tech bros should have heads rolling and I
| really don't want to see it happening in Europe.
|
| But I guess we got a new version number on a few models and some
| blown up benchmarks so that's good, oh and of course the svg
| images we will never use for anything.
| simonw wrote:
| "Nothing about the severe impact on the environment"
|
| I literally said:
|
| "AI data centers continue to burn vast amounts of energy and
| the arms race to build them continues to accelerate in a way
| that feels unsustainable."
|
| AND I linked to my coverage from last year, which is still true
| today (hence why I felt no need to update it):
| https://simonwillison.net/2024/Dec/31/llms-in-2024/#the-envi...
| smileson2 wrote:
| forgot to mention the first murder-suicide instigated by chatgpt
| DANmode wrote:
| These are _his_ highlights as a killer blogger,
|
| not _AI's_ highlights.
|
| Easy with the hot take.
| didip wrote:
| Indeed. I don't understand why Hacker News is so dismissive about
| the coming of LLMs, maybe HN readers are going through 5 stages
| of grief?
|
| But LLM is certainly a game changer, I can see it delivering
| impact bigger than the internet itself. Both require a lot of
| investments.
| cebert wrote:
| Many people feel threatened by the rapid advancements in LLMs,
| fearing that their skills may become obsolete, and in turn act
| irrationally. To navigate this change effectively, we must keep
| open minds, keep adaptable, and embrace continuous learning.
| nickphx wrote:
| rapid advancements in what? hallucinations..? FOMO marketing?
| certainly nothing productive.
| chii wrote:
| > in turn act irrationally
|
| it isn't irrational to act in self-interest. If LLM threatens
| someone's livelihood, it matters not that it helps humanity
| overall one bit - they will oppose it. I don't blame them.
| But i also hope that they cannot succeed in opposing it.
| Davidzheng wrote:
| It's irrational to genuinely hold false beliefs about
| capabilities of LLMs. But at this point I assume around
| half of the skeptics are emotionally motivated anyway.
| jdhsgsvsbzbd wrote:
| As opposed to having skin in the game for llms and are
| blind to their flaws???
|
| I'd assume that around half of the optimists are
| emotionally motivated this way.
| rgoulter wrote:
| Many comments discussing LLMs involve emotions, sure. :)
| Including, obviously, comments in favour of LLMs.
|
| But most discussion I see is vague and without specificity
| and without nuance.
|
| Recognising the shortcomings of LLMs makes comments praising
| LLMs that much more believable; and recognising the benefits
| of LLMs makes comments criticising LLMs more believable.
|
| I'd completely believe anyone who says they've found the LLM
| very helpful at greenfield frontend tasks, and I'd believe
| someone who found the LLM unable to carry out subtle
| refactors on an old codebase in a language that's not Python
| or JavaScript.
| reppap wrote:
| I'm not threatened by LLMs taking my job as much as they are
| taking away my sanity. Every time I tell someone no and they
| come back to me with a "but copilot said.." it's followed by
| something entirely incorrect it makes me want to
| autodefenestrate.
| callc wrote:
| I am happy "autodefenestrate" is the first new word I
| learned in 2026. Thank you.
|
| Autodefenestrate - To eject or hurl oneself from a window,
| especially lethally
| snigsnog wrote:
| The internet and smartphones were immediately useful in a
| million different ways for almost every person. AI is not even
| close to that level. Very to somewhat useful in some fields
| (like programming) but the average person will easily be able
| to go through their day without using AI.
|
| The most wide-appeal possibility is people loving 100%-AI-slop
| entertainment like that AI Instagram Reels product. Maybe I'm
| just too disconnected with normies but I don't see this taking
| off. Fun as a novelty like those Ring cam vids but I would
| never spend all day watching AI generated media.
| JumpCrisscross wrote:
| > _AI is not even close to that level_
|
| Kagi's Research Assistant is pretty damn useful, particularly
| when I can have it poll different models. I remember when the
| first iPhone lacked copy-paste. This feels similar.
|
| (And I don't think we're heading towards AGI.)
| SgtBastard wrote:
| ... the internet was not immediately useful in a million
| different ways for almost every person.
|
| Even if you skip ARPAnet, you're forgetting the Gopher days
| and even if you jump straight to WWW+email==the internet,
| you're forgetting the mosaic days.
|
| The applications that became useful to the masses emerged a
| decade+ after the public internet and even then, it took 2+
| decades to reach anything approaching saturation.
|
| Your dismissal is not likely to age well, for similar
| reasons.
| chii wrote:
| the "usefulness" excuse is irrelevant, and the claim that
| phones/internet is "immediately useful" is just a post hoc
| rationalization. It's basically trying to find a reasonable
| reason why opposition to AI is valid, and is not in self-
| interest.
|
| The opposition to AI is from people who feel threatened by
| it, because it either threatens their livelihood (or
| family/friends'), and that they feel they are unable to
| benefit from AI in the same way as they had internet/mobile
| phones.
| duchef wrote:
| The usefulness of mobile phones was identifiable
| immediately and it is absolutely not 'post hoc
| rationalization'. The issue was the cost - once low cost
| mobile telephones were produced they almost immediately
| became ubiquitous (see nokia share price from the release
| of the nokia 6110 onwards for example).
|
| This barrier does not exist for current AI technologies
| which are being given away free. Minor thought experiment
| - just how radical would the uptake of mobile phones have
| been if they were given away free?
| jfyi wrote:
| It's only low cost for general usage chat users. If you
| are using it for anything beyond that, you are paying or
| sitting in a long queue (likely both).
|
| You may just be a little early to the renaissance. What
| happens when the models we have today run on a mobile
| device?
|
| The nokia 6110 was released 15 years after the first
| commercial cell phone.
| duchef wrote:
| Yes although even those people paying are likely still
| being subsidized and not currently paying the full cost.
|
| Interesting thought about current SOTA models running on
| my mobile device. I've given it some thought and I don't
| think it would change my life in any way. Can you suggest
| some way that it would change yours?
| qualifck wrote:
| Eh, quite the contrary. A lot of anti AI people genuinely
| wanted to use AI but run into the factual reality of the
| limitations of the software. It's not that it's going to
| take my job, it's that I was told it would redefine how I
| do work and is exponentially improving only to find out
| that it just kind of sucks and hasn't gotten much better
| this year.
| staticassertion wrote:
| > Very to somewhat useful in some fields (like programming)
| but the average person will easily be able to go through
| their day without using AI.
|
| I know a lot of "normal" people who have completely replaced
| their search engine with AI. It's increasingly a staple for
| people.
|
| Smartphones were absolutely NOT immediately useful in a
| million different ways for almost every person, that's total
| revisionist history. I remember when the iPhone came out, it
| was AT&T only, it did almost nothing useful. Smartphones were
| a novelty for quite a while.
| brabel wrote:
| I agree with most points but as a tech enthusiast, I was
| using a smart phone years before the iPhone, and I could
| already use the internet, make video calls, email etc
| around 2005. It was a small flip phone but it was not
| uncommon for phones to do that already at that time, at
| least in Australia and parts of Asia (a Singaporean friend
| told me about the phone).
| nen-nomad wrote:
| ChatGPT has roughly 800 million weekly active users. Almost
| everyone around me uses it daily. I think you are
| underestimating the adoption.
| throw1235435 wrote:
| How many pay? And out of that how many are willing to pay
| the amount to at least cover the inference costs (not loss
| leading?)
|
| Outside the verifiable domains I think the impact is more
| assistance/augmentation than outright disruption (i.e. a
| novelty which is still nice). A little tiny bit of value
| sprinkled over a very large user base but each person
| deriving little value overall.
|
| Even as they use it as search it is at best an
| incrementable improvement on what they used to do - not
| life changing.
| danielbln wrote:
| Even my mom and aunts are using it frequently for all sorts
| of things, and it took a long time for them to hop onto
| internet and smartphones at first.
| mrweasel wrote:
| The adoption is just so weird to me. I cannot for the life
| of me get LLM chatbot to work for me. Every time I try I
| get into an argument with the stupid thing. They are still
| wrong constantly, and when I'm wrong they won't correct me.
|
| I have great faith in AI in e.g. medical equipment, or
| otherwise as something built in, working on a single
| problem in the background, but the chat interface is
| terrible.
| dragonwriter wrote:
| "Almost everyone will use it at free or effectively
| subsidized prices" and "It delivers utility which justifies
| its variable costs + fixed costs amortized over useful
| lifetime" are not the same thing, and its not clear how
| much of the use is tied to novelty such that if new and
| progressively more expensive to train releases at a regular
| cadence dropped off, usage, even at subsidized prices,
| would, too.
| arctic-true wrote:
| Usage plunges on the weekends and during the summer,
| suggesting that a significant portion of users are students
| using ChatGPT for free or at heavily subsidized rates to do
| homework (i.e., extremely basic work that is
| extraordinarily well-represented in the training data).
| That usage will almost certainly never be monetizable, and
| it suggests nothing about the trajectory of the
| technology's capability or popularity. I suspect ChatGPT,
| in particular, will see its usage slip considerably as the
| education system (hopefully) adapts.
| simonw wrote:
| The summer slump was a thing in 2023 but apparently
| didn't repeat in 2024:
| https://www.similarweb.com/blog/insights/ai-news/chatgpt-
| bea...
|
| The weekend slumps could equally suggest people are using
| it at work.
| arctic-true wrote:
| Interesting, thank you for that. I'd be curious to see
| the data for 2025. I was basing my take off Google trends
| data - the kind of person who goes to ChatGPT by googling
| "chatGPT" seems to be using it less in the summer.
| raincole wrote:
| The early internet and smartphones (the Japanese ones, not
| iPhone) were definitely not "immediately" adopted by the
| mass, unlike LLM.
|
| If "immediate" usefulness is the metric we measure, then the
| internet and smartphones are pretty insignificant inventions
| compared to LLM.
|
| (of course it's not a meaningful metric, as there is no clear
| line between a dumb phone and a smart phone, or a moderately
| sized language model and a LLM)
| fragmede wrote:
| > The internet and smartphones were immediately useful in a
| million different ways for almost every person. AI is not
| even close to that level.
|
| Those are some very rosy glasses you've got on there. The
| nascent Internet took forever to catch on. It was for weird
| nerds at universities and it'll never catch on, but here we
| are.
| what-the-grump wrote:
| A year after the iPhone came out... it didn't have an App
| Store, barely was able to play video, barely had enough power
| to last a day. You just don't remember or were not around for
| it.
|
| A year after llms came out... are you kidding me?
|
| Two years?
|
| 10 years?
|
| Today, by adding an MCP server to wrap the same API that's
| been around forever for some system, makes the users of that
| system prefer NLI over the gui almost immediately.
| zvolsky wrote:
| The idea of HN being dismissive of impactful technology is as
| old as HN. And indeed, the crowd often appears stuck in the
| past with hindsight. That said, HN discussions aren't
| homogeneous, and as demonstrated by Karpathy in his recent
| blogpost "Auto-grading decade-old Hacker News", at least some
| commenters have impressive foresight:
| https://karpathy.bearblog.dev/auto-grade-hn/
| brabel wrote:
| So exactly 10 years ago a lot of people believed that the
| game Go would not be "conquered" by AI, but after just a few
| months it was. People will always be skeptical of new things,
| even people who are in tech, because many hyped things indeed
| go nowhere... while it may look obvious in hindsight, it's
| really hard to predict what will and what won't be
| successful. On the LLM front I personally think it's
| extremely foolish to still consider LLMs as going nowhere.
| There's a lot more evidence today of the usefulness of LLMs
| than there was of DeepMind being able to beat top human
| players in Go 10 years ago.
| crystal_revenge wrote:
| > I don't understand why Hacker News is so dismissive about the
| coming of LLMs
|
| I find LLMs incredibly useful, but if you were following along
| the last few years the promise was for "exponential progress"
| with a teaser world destroying super intelligence.
|
| We objectively are not on that path. There is no "coming of
| LLMs". We might get some incremental improvement, but we're
| very clearly seeing sigmoid progress.
|
| I can't speak for everyone, but I'm tired of hyperbolic rants
| that are unquestionably not justified (the nice thing about
| exponential progress is you don't need to argue about it)
| aoeusnth1 wrote:
| We're very clearly seeing exponential progress - even above
| trend, on METR, whose slope keeps getting revised to a higher
| and higher estimate each time. Explain your perspective on
| the objective evidence against exponential progress?
| llmslave2 wrote:
| Pretty neat how this exponential progress hasn't resulted
| in exponential productivity. Perhaps you could explain your
| perspective on that?
| viraptor wrote:
| Writing the code itself was never the main bottleneck.
| Designing the bigger solution, figuring out tradeoffs,
| taking to affected teams, etc. takes as much time as it
| used to. But still, there's definitely a significant
| improvement in code production part in many areas.
| aoeusnth1 wrote:
| It has! CLs/engineer increased by 10% this year.
|
| LLMs from late 2024 were nearly worthless as coding
| agents, so given they have quadrupled in capability since
| then (exponential growth, btw), it's not surprising to
| see a modestly positive impact on SWE work.
|
| Also, I'm noticing you're not explaining yourself :)
| llmslave2 wrote:
| Hey, I'm not the OG commentator, why do I have to explain
| myself! :)
|
| When Fernando Alonso (best rookie btw) goes from 0-60 in
| 2.4 seconds in his Aston Martin, is it reasonable to
| assume he will near the speed of light in 20 seconds?
| lopatin wrote:
| > Hey, I'm not the OG commentator, why do I have to
| explain myself! :)
|
| The issue is that you're not acknowledging or replying to
| people's explanations for _why_ they see this as
| exponential growth. It's almost as if you skimmed through
| the meat of the comment and then just re-phrased your
| original idea.
|
| > When Fernando Alonso (best rookie btw) goes from 0-60
| in 2.4 seconds in his Aston Martin, is it reasonable to
| assume he will near the speed of light in 20 seconds?
|
| This comparison doesn't make sense because we know the
| limits of cars but we don't yet know the limits of LLMs.
| It's an open question. Whether or not an F1 engine can
| make it the speed of light in 20 seconds is not an open
| question.
| llmslave2 wrote:
| It's not in me to somehow disprove claims of exponential
| growth when there isn't even evidence provided of it.
|
| My point with the F1 comparison is to say that a short
| period of rapid improvement doesn't imply exponential
| growth and it's about as weird to expect that as it is
| for an f1 car to reach the speed of light. It's possible
| you know, the regulations are changing for next season -
| if Leclerc sets a new lap record in Australia by .1 ms we
| can just assume exponential improvements and surely
| Ferrari will be lapping the rest of the field by the
| summer right?
| aoeusnth1 wrote:
| There is already evidence provided of it! METR time
| horizons is going up on an exponential trend. This is
| literally the most famous AI benchmark and already
| mentioned in this thread.
|
| https://metr.org/blog/2025-03-19-measuring-ai-ability-to-
| com...
|
| https://metr.org/blog/2025-07-14-how-does-time-horizon-
| vary-...
| aoeusnth1 wrote:
| If you're not going to explain yourself, at least stay on
| topic. We're talking about exponential growth, so address
| the points I'm making.
| aoeusnth1 wrote:
| I'm noticing you're not responding to my claim that
| producivity has been impacted
| Madmallard wrote:
| LLMs a year ago were more able to do a complex project
| I've repeatedly tried to do than they are now.
| scotty79 wrote:
| Try Antigravity with Gemini 3 Pro. Seems very capable to
| me.
| surajrmal wrote:
| I think this is happening by raising the floor for job
| roles which are largely boilerplate work. If you are on
| the more skilled side or work in more original/ niche
| areas, AI doesn't really help too much. I've only been
| able to use AI effectively for scaling refactors, not
| really much in feature development. It often just slows
| me down when I try to use it. I don't see this changing
| any time soon.
| HPMOR wrote:
| I think this is an open question still and very
| interesting. Ilya discussed this on the Dwarkesh podcast.
| But the capabilities of LLMs is clearly exponential and
| perhaps super exponential. We went from something that
| could string together incoherent text in 2022 to general
| models helping people like Terrance Tao and Scott
| Aaronson write new research papers. LLMs also beat IMO
| and the ICPC. We have entered the John Henry era for
| intellectual tasks...
| llmslave2 wrote:
| > But the capabilities of LLMs is clearly exponential and
| perhaps super exponential
|
| By what metric?
| utopiah wrote:
| BS metric... /s
| tsimionescu wrote:
| > LLMs also beat IMO and the ICPC
|
| Very spurious claims, given that there was no effort made
| to check whether the IMO or ICPC problems were in the
| training set or not, or to quantify how far problems in
| the training set were from the contest problems. IMO
| problems are supposed to be unique, but since it's not at
| the frontier of math research, there is no guarantee that
| the same problem, or something very similar, was not
| solved in some obscure manual.
| mgfist wrote:
| Because that requires adoption. Devs on hackernews are
| already the most up to date folks in the industry and
| even here adoption of LLMs is incredibly slow. And a lot
| of the adoption that does happen is still with older tech
| like ChatGPT or Cursor.
| belmont_sup wrote:
| What's the newer tech?
| TeodorDyakov wrote:
| Claude Code With Opus 4.5
| scotty79 wrote:
| How long before introduction of computers lead to
| increases in average productivity? How long for the
| internet? Business is just slow to figure out how to use
| anything for its benefit, but it eventually gets there.
| fmbb wrote:
| > How long before introduction of computers lead to
| increases in average productivity?
|
| I think it never did. Still has not.
|
| https://en.wikipedia.org/wiki/Productivity_paradox
| spectralista wrote:
| The best example is that even ATM machines didn't reduce
| bank teller jobs.
|
| Why? Because even the bank teller is doing more than
| taking and depositing money.
|
| IMO there is an ontological bias that pervades our modern
| society that confuses the map for the territory and has a
| highly distorted view of human existence through the lens
| of engineering.
|
| We don't see anything in this time series, because this
| time series itself is meaningless nonsense that reflects
| exactly this special kind of ontological stupidity:
|
| https://fred.stlouisfed.org/series/PRS85006092
|
| As if the sum of human interaction in an economy is some
| kind of machine that we just need to engineer better
| parts for and then sum the outputs.
|
| Any non-careerist, thinking person that studies economics
| would conclude we don't and will probably not have the
| tools to properly study this subject in our lifetimes.
| The high dimensional interaction of biology, entropy and
| time. We have nothing. The career economist is
| essentially forced to sing for their supper in a type of
| time series theater. Then there is the method acting of
| pretending to be surprised when some meaningless
| reductionist aspect of human interaction isn't reflected
| in the fake time series.
| barrenko wrote:
| Sir, we're in a modern economy, we don't ever _ever_ look
| at productivity graphs (this is not to disparage LLMs,
| just a comment on productivity in general)
| viraptor wrote:
| > exponential progress
|
| First you need to define what it means. What's the metric?
| Otherwise it's very much something you can argue about.
| noodletheworld wrote:
| > What's the metric?
|
| Language model capability at generating text output.
|
| The model progress this year has been a lot of:
|
| - "We added multimodal"
|
| - "We added a lot of _non AI_ tooling" (ie agents)
|
| - "We put more compute into inference" (ie thinking mode)
|
| So yes, there is still rapid progress, but these ^ make it
| clear, at least to me, that next gen _models_ are
| _significantly harder_ to build.
|
| Simultaneously we see a distinct narrowing between players
| (openai, deepseek, mistral, google, anthropic) in their
| offerings.
|
| Thats usually a signal that the rate of progress is
| slowing.
|
| Remind me what was so great about gpt 5? How about gpt4
| from from gpt 3?
|
| Do you even remember the releases? Yeah. I dont. I had to
| look it up.
|
| Just another model with more or less the same capabilities.
|
| "Mixed reception"
|
| That is not what exponential progress looks like, _by any
| measure_.
|
| The progress this year has been in the tooling around the
| models, smaller faster models with similar capabilities.
| Multimodal add ons that no one asked for, because its
| easier to add image and audio processing than improve text
| handling.
|
| That may still be on a path to AGI, but it not an
| _exponential_ path to it.
| dragonwriter wrote:
| > Language model capability at generating text output.
|
| That's not a metric, that's a vague non-operationalized
| concept, that could be operationalized into an infinite
| number of different metrics. And an improvement that was
| linear in one of those possible metrics would be
| exponential in another one (well, actually, one that is
| was linear in one would also be linear in an infinite
| number of others, _as well as_ being exponential in an
| infinite number of others.
|
| That's why you have to define an actual metric, not
| simply describe a vague concept of a kind of capacity of
| interest, before you can meaningfully discuss whether
| improvement is exponential. Because the answer is
| necessarily entirely dependent on the specific
| construction of the metric.
| viraptor wrote:
| > Language model capability at generating text output.
|
| That's not a quantifiable sentence. Unless you put it in
| numbers, anyone can argue exponential/not.
|
| > next gen models are significantly harder to build.
|
| That's not how we judge capability progress though.
|
| > Remind me what was so great about gpt 5? How about gpt4
| from from gpt 3?
|
| > Do you even remember the releases?
|
| At gpt 3 level we could generate some reasonable code
| blocks / tiny features. (An example shown around at the
| time was "explain what this function does" for a
| "fib(n)") At gpt 4, we could build features and tiny
| apps. At gpt 5, you can often one-shot build whole apps
| from a vague description. The difference between them is
| massive for coding capabilities. Sorry, but if you can't
| remember that massive change... why are you making claims
| about the progress in capabilities?
|
| > Multimodal add ons that no one asked for
|
| Not only does multimodal input training improve the model
| overall, it's useful for (for example) feeding back
| screenshots during development.
| threethirtytwo wrote:
| I don't think the path was ever exponential but your
| claim here is almost as if the slow down hit an asymptote
| like wall.
|
| Most of the improvements are intangible. Can we truly say
| how much more reliable the models are? We barely have
| quantitative measurements on this so it's all vibes and
| feels. We don't even have a baseline metric for what AGI
| is and we invalidated the Turing test also based on vibes
| and feels.
|
| So my argument is that part of the slow down is in itself
| an hallucination because the improvement is not actually
| measurable or definable outside of vibes.
| aoeusnth1 wrote:
| > Language model capability at generating text output.
|
| How would you put this on a graph?
| scotty79 wrote:
| Define it however you like. There's not a single chart you
| can draw that even begins to look like a signoid.
| nicbou wrote:
| Time spent being human and enjoying life.
|
| I can't point at many problems it has meaningfully solved
| for me. I mean real problems , not tasks that I have to do
| for my employer. It seems like it just made parts of my
| existence more miserable, poisoned many of the things I
| love, and generally made the future feel a lot less
| certain.
| scotty79 wrote:
| > but we're very clearly seeing sigmoid progress.
|
| Yeah, probably. But no chart actually shows it yet. For now
| we are firmly in exponential zone of the signoid curve and
| can't really tell if it's going to end in a year, decade or a
| century.
| utopiah wrote:
| Doesn't even matter if the goal is extremely high. Talking
| about exponential when we clearly see matching energy needs
| proves there is no way we can maintain that pace without
| radical (and thus unpredictable) improvements.
|
| My own "feeling" is that it's definitely not exponential
| but again, doesn't matter if it's unsustainable.
| fullstackchris wrote:
| I wrote an article complaining about the whole hype over a
| year ago:
|
| https://chrisfrewin.medium.com/why-llms-will-never-be-
| agi-70...
|
| Seems to be playing out that way.
| senordevnyc wrote:
| I've been reading this comment multiple times a week for the
| last couple years. Constant assertions that we're starting to
| hit limits, plateau, etc. But a cursory glance at where we
| are today vs a year ago, let alone two years ago, makes it
| wildly obvious that this is bullshit. The pace of improvement
| of both models and tooling has been breathtaking. I could
| give a shit whether you think it's "exponential", people like
| you were dismissing all of this years ago, meanwhile I just
| keep getting more and more productive.
| qualifck wrote:
| People keep saying stuff like this. That the improvements
| are so obvious and breathtaking and astronomical and then I
| go check out the frontier LLMs again and they're maybe a
| tiny bit better than they were last year but I can't
| actually be sure bcuz it's hard to tell.
|
| sometimes it seems like people are just living in another
| timeline.
| aspenmartin wrote:
| I'm not sure I understand: we are _objectively on that path_
| -- we are increasing exponentially on a number of metrics
| that may be imperfect but seem to paint a pretty consistent
| picture. Scaling laws are exponential. METR's time horizon
| benchmark is exponential. Lots of performance measures are
| exponential, so why do you say we're objectively not on that
| path?
|
| > We might get some incremental improvement, but we're very
| clearly seeing sigmoid progress.
|
| again, if it is "very clear" can you point to some concrete
| examples to illustrate what you mean?
|
| > I can't speak for everyone, but I'm tired of hyperbolic
| rants that are unquestionably not justified (the nice thing
| about exponential progress is you don't need to argue about
| it)
|
| OK but what specifically do you have an issue with here?
| Night_Thastus wrote:
| LLMs hold _some_ real utility. But that real utility is buried
| under a mountain of fake hype and over-promises to keep
| shareholder value high.
|
| LLMs have real limitations that aren't going away any time soon
| - not until we move to a new technology fundamentally different
| and separate from them - sharing almost nothing in common.
| There's a lot of 'progress-washing' going on where people claim
| that these shortfalls will magically disappear if we throw
| enough data and compute at it when they clearly will not.
| Gigachad wrote:
| Pretty much. What actually exists is very impressive. But
| what was promised and marketed has not been delivered.
| rustystump wrote:
| Markets never deliver. That isnt new, i do think llms are
| not far off from google in terms of impact.
|
| Search, as of today, is inferior to frontier models as a
| product. However, best case still misses expected returns
| by miles which is where the growsing comes from.
|
| Generative art/ai is still up in the air for staying power
| but id predict it isnt going away.
| visarga wrote:
| I think the missing ingredient is not something the LLMs
| lack, but something we as developers don't do - we need to
| constrain, channel, and guide agents by creating reactive
| test environments around them. Not vibes, but hard tests,
| they are the missing ingredient to coding agents. You can
| even use AI to write most of these tests but the end result
| depends on how well you structured your code to be
| testable.
|
| If you inherit 9000 tests from an existing project you can
| vibe code a replacement on your phone in a holiday, like
| Simon Willison's JustHTML port. We are moving from agents
| semi-randomly flailing around to constraint satisfaction.
| coffeebeqn wrote:
| Yes and most of the investment has been kind of post-GPT4
| betting that things will get exponentially more impressive
| baq wrote:
| I find opus 4.5 and gpt 5.2 mind blowing more often than I
| find them dumb as rocks. I don't listen to or read any
| marketing material, I just use the tools. I couldn't care
| less about what the promises are, what I have now available
| to me is fundamentally different from what I had in August
| and it changed completely how I work.
| probably_wrong wrote:
| Speaking for myself: because if the hype were to be believed we
| should have no relational databases when there's MongoDB, no
| need for dollars when there's cryptocoins, all virtual goods
| would be exclusively sold as NFTs, and we would be all driving
| self-driving cars by now.
|
| LLMs are being driven mostly by grifters trying to achieve a
| monopoly before they run out of cash. Under those conditions I
| find their promises hard to believe. I'll wait until they
| either go broke or stop losing money left and right, and
| whatever is left is probably actually useful.
| simonw wrote:
| The way I've been handling the deafening hype is to focus
| exclusively on what the models that we have right now can do.
|
| You'll note I don't mention AGI or future model releases in
| my annual roundup at all. The closest I get to that is
| expressing doubt that the METR chart will continue at the
| same rate.
|
| If you focus exclusively on what actually works the LLM space
| is a whole lot more interesting and less frustrating.
| magicalhippo wrote:
| > focus exclusively on what the models that we have right
| now can do
|
| I'm just a casual user, but I've been doing the same and
| have noticed the sharp improvements of the models we have
| now vs a year ago. I have OpenAI Business subscription
| through work, I signed up for Gemini at home after Gemini
| 3, and I run local models on my GPU.
|
| I just ask them various questions where I know the answer
| well, or I can easily verify. Rewrite some code, factual
| stuff etc. I compare and contrast by asking the same
| question to different models.
|
| AGI? Hell no. Very useful for some things? Hell yes.
| asielen wrote:
| It is an over correction because of all the empty promises of
| LLMs. I use Claude and chatgpt daily at work and am amazed at
| what they can do and how far they can come.
|
| BUT when I hear my executive team talk and see demos of
| "Agentforce" and every saas company becoming an AI company
| promising the world, I have to roll my eyes.
|
| The challenge I have with LLMs is they are great at creating
| first draft shiny objects and the LLMs themselves over promise.
| I am handed half baked work created by non technical people
| that now I have to clean up. And they don't realize how much
| work it is to take something from a 60% solution to a 100%
| solution because it was so easy for them to get to the 60%.
|
| Amazing, game changing tools in the right hands but also give
| people false confidence.
|
| Not that they are not also useful for non-technical people but
| I have had to spend a ton of time explaining to copywriters on
| the marketing team that they shouldn't paste their credentials
| into the chat even if it tells them to and their vibe coded app
| is a security nightmare.
| semilin wrote:
| This seems like the right take. The claims of the imminence
| of AGI are exhausting and to me appear dissonant with
| reality. I've tried gemini-cli and Claude Code and while
| they're both genuinely quite impressive, they absolutely
| suffer from a kind of prototype syndrome. While I could learn
| to use these tools effectively for large-scale projects, I
| still at present feel more comfortable writing such things by
| hand.
|
| The NVIDIA CEO says people should stop learning to code. Now
| if LLMs will really end up as reliable as compilers, such
| that they can write code that's better and faster than I can
| 99% of the time, then he might be right. As things stand now,
| that reality seems far-fetched. To claim that they're useless
| because this reality has not yet been achieved would be
| silly, but not more silly than claiming programming is a dead
| art.
| vunderba wrote:
| _> I don 't understand why Hacker News is so dismissive about
| the coming of LLMs._
|
| Eh. I wouldn't be so quick to speak for the entirety of HN.
| Several articles related to LLMs easily hit the front page
| every single day, so clearly there are plenty of HN users
| upvoting them.
|
| I think you're just reading too much into what is more likely
| classic HN cynicism and/or fatigue.
| ewoodrich wrote:
| Exactly. There was a stretch of 6 months or so right after
| ChatGPT was released where approximately 50% of front page
| posts at any given time were related to LLMs. And these days
| every other Show HN is some kind of agentic dev tool and
| Anthropic/OpenAI announcements routinely get 500+ comments in
| a matter of hours.
| utopiah wrote:
| It's because both "side" tries to re-adjust.
|
| When an "AI skeptic" sees a very positive AI comment, they
| try to argue that it is indeed interesting but nowhere near
| close to AI/AGI/ASI or whatever the hype at the moment uses.
|
| When an "AI optimistic" sees a very negative AI comment, they
| try to list all the amazing things they have done that they
| were convinced was until then impossible.
| viraptor wrote:
| Based on quite a few comments recently, it also looks like many
| have tried LLMs in the past, but haven't seriously revisited
| either the modern or more expensive models. And I get it. Not
| everyone wants to keep up to date every month, or burn cash on
| experiments. But at the same time, people seem to have opinions
| formed in 2024. (Especially if they talk about just
| hallucinations and broken code - tell the agent to search for
| docs and fix stuff) I'd really like to give them Opus 4.5 as an
| agent to refresh their views. There's lots to complain about,
| but the world has moved on significantly.
| mirsadm wrote:
| This has been the argument since day one. You just have to
| try the latest model, that's where you went wrong. For the
| record I use Claude Code quite a bit and I can't see much
| meaningful improvements from the last few models. It is a
| useful tool but it's shortcomings are very obvious.
| techpression wrote:
| Just last week Opus 4.5 decided that the way to fix a test
| was to change the code so that everything else but the test
| broke.
|
| When people say "fix stuff" I always wonder if it actually
| means fix, or just make it look like it works (which is
| extremely common in software, LLM or not).
| simonw wrote:
| What did Opus do when you told it that it shouldn't have
| done that?
| layer8 wrote:
| It apologized. ;)
| viraptor wrote:
| Sure, I get an occasional bad result from Opus - then I
| revert and try again, or ask it for a fix. Even with a
| couple of restarts, it's going to be faster than me on
| average. (And that's ignoring the situations where I have
| to restart myself)
|
| Basically, you're saying it's not perfect. I don't think
| anyone is claiming otherwise.
| b3kart wrote:
| The problem is it's imperfect in very unpredictable ways.
| Meaning you always need to keep it on a short leash for
| anything serious, which puts a limit on the productivity
| boost. And that's fine, but does this match the level of
| investment and expectations?
| techpression wrote:
| It's not about being perfect, it's about not being as
| great as the marketing, and many proponents, claim.
|
| The issue is that there's no common definition of
| "fixed". "Make it run no matter what" is a more apt
| description in my experience, which works to a point but
| then becomes very painful.
| baq wrote:
| Nice. Did it realize the mistake and corrected it?
| techpression wrote:
| Nope, I did get a lot of fancy markdown with emojis
| though so I guess that was a nice tradeoff.
|
| In general, even with access to the entire code base
| (which is very small), I find the inherent need in the
| models to satisfy the prompter to be their biggest flaw
| since it tends to constantly lead down this path. I often
| have to correct over convoluted SQL too because my
| problems are simple and the training data seems to favor
| extremely advanced operations.
| Madmallard wrote:
| Have you tried using it for anything actually complicated?
|
| Lol. It's worse than nothing at all.
| lukaslalinsky wrote:
| I think the split between vibe coding and AI-assisted coding
| will only widen over time. If you ask LLMs to do something
| complex, they will fail and you waste your time. If you work
| with them as a peer, and you delegate tasks to them, they
| will succeed and you save your time.
| watwut wrote:
| I work with leers by delegating complex task to them while
| I do other complex tasks.
| hapticmonkey wrote:
| It's not the technology I'm dismissive about. It's the
| economics.
|
| 25 years ago I was optimistic about the internet, web sites,
| video streaming, online social systems. All of that. Look at
| what we have now. It was a fun ride until it all ended up
| "enshitified". And it will happen to LLMs, too. Fool me once.
|
| Some developer tools might survive in a useful state on
| subscriptions. But soon enough the whole A.I. economy will
| centralise into 2 or 3 major players extracting more and more
| revenue over time until everyone is sick of them. In fact, this
| process seems to be happening at a pretty high speed.
|
| Once the users are captured, they'll orient the ad-spend market
| around themselves. And _then_ they'll start taking advantage of
| the advertisers.
|
| I really hope it doesn't turn out this way. But it's hard to be
| optimistic.
| Al-Khwarizmi wrote:
| Contrary to the case for the internet, there is a way out,
| however - if local, open-source LLMs get good. I really hope
| they do, because enshittification does seem unavoidable if we
| depend on commercial offerings.
| ndiddy wrote:
| Well the "solution" for that will be the GPU vendors
| focusing solely on B2B sales because it's more profitable,
| therefore keeping GPUs out of the hands of average
| consumers. There's leaks suggesting that nVidia will
| gradually hike the prices of their 5090 cards from $2000 to
| $5000 due to RAM price increases (
| https://wccftech.com/geforce-rtx-5090-prices-to-soar-
| to-5000... ). At that point, why even bother with the R&D
| for newer consumer cards when you know that barely anyone
| will be able to afford them?
| tgv wrote:
| The negatives outweigh the positives, if only because the
| positives are so small. A bunch of coders making their lives
| easier doesn't really matter, but pupils and students skipping
| education does. As a meme said: you had better start eating
| healthy, because your future doctor vibed his way through med
| school.
| phatfish wrote:
| Maybe because the hype for an next gen search engine that can
| also just make things up when you query it is a bit much?
| jcims wrote:
| It feels like there are several conversations happening that
| sound the same but are actually quite different.
|
| One of them is whether or not large models are useful and/or
| becoming more useful over time. (To me, clearly the answer is
| yes)
|
| The other is whether or not they live up to the hype. (To me,
| clearly the answer is no)
|
| There are other skirmishes around capability for novelty, their
| role in the economy, their impact on human cognition, if/when
| AGI might happen and the overall impact to the largely tech-
| oriented community on HN.
| Atomic_Torrfisk wrote:
| > HN readers are going through 5 stages of grief
|
| So we are just irrational and sour?
| claudiug wrote:
| because lies. all the people involved in this, the one a C
| title, tell us about how great is now.
| syndacks wrote:
| I can't get over the range of sentiment on LLMs. HN leans snake
| oil, X leans "we're all cooked" --- can it possibly be both? How
| do other folks make sense of this? I'm not asking for a side,
| rather understanding the range. Does the range lead you to
| believe X over Y?
| zahlman wrote:
| I'm not really convinced that anywhere leans heavily towards
| anything; it depends which thread you're in etc.
|
| It's polarizing because it represents a more radical shift in
| expected workflows. Seeing that range of opinions doesn't
| really give me a reason to update, no. I'm evaluating based on
| what makes sense when I hear it.
| thisoneisreal wrote:
| My take (no more informed than anyone else's) is that the range
| indicates this is a complex phenomenon that people are still
| making sense of. My suspicion is that something like the
| following is going on:
|
| 1. LLMs can do some truly impressive things, like taking
| natural language instructions and producing compiling,
| functional code as output. This experience is what turns some
| people into cheerleaders.
|
| 2. Other engineers see that in real production systems, LLMs
| lack sufficient background / domain knowledge to effectively
| iterate. They also still produce output, but it's verbose and
| essentially missing the point of a desired change.
|
| 3. LLMs also can be used by people who are not knowledgeable to
| "fake it," and produce huge amounts of output that is basically
| besides-the-point bullshit. This makes those same senior folks
| very, very resentful, because it wastes a huge amount of their
| time. This isn't really the fault of the tool, but it's a
| common way the tool gets used and so it gets tarnished by
| association.
|
| 4. There is a ridiculous amount of complexity in some of these
| tools and workflows people are trying to invent, some of which
| is of questionable value. So aside from the tools themselves
| people are skeptical of the people trying to become thought
| leaders in this space and the sort of wild hacks they're coming
| up with.
|
| 5. There are real macro questions about whether these tools can
| be made economical to justify whatever value they do produce,
| and broader questions about their net impact on society.
|
| 6. Last but not least, these tools poke at the edges of
| "intelligence," the crown jewel of our species and also a big
| source of status for many people in the engineering community.
| It's natural that we're a little sensitive about the prospect
| of anything that might devalue or democratize the concept.
|
| That's my take for what it's worth. It's a complex phenomenon
| that touches all of these threads, so not only do you see a
| bunch of different opinions, but the same person might feel
| bullish about one aspect and bearish about another.
| johnfn wrote:
| I believe the spikiness in response is because AI itself is
| spiky - it's incredibly good at some classes of tasks, and
| remarkably poor at others. People who use it on the spikes are
| genuinely amazed because of how good it is. This does nothing
| but annoy the people who use it in the troughs, who become
| increasingly annoyed that everyone seems to be losing their
| mind over something that can't even do (whatever).
| llmslave2 wrote:
| Because there is a wide range of what people consider _good_.
| If you look at that the people on X consider to be _good_ ,
| it's not very surprising.
| coffeefirst wrote:
| Well, this is the internet. Arguing about everything is its
| favorite pastime.
|
| But generally yes, I think back to
| Mongo/Node/metaverse/blockchain/IDEs/tablets and pretty much
| everything has had its boosters and skeptics, this is just
| more... intense.
|
| Anyway I've decided to believe my own eyes. The crowds say a
| lot of things. You can try most of it yourself and see what it
| can and can't do. I make a point to compare notes with
| competent people who also spent the time trying things. What's
| interesting is most of their findings are _compatible with_
| mine, including for folks who don 't work in tech.
|
| Oh, and one thing is for sure: shoving this technology into
| every single application imaginable is a good way to lose
| friends and alienate users.
| nstart wrote:
| The problem with X is that so many people who have no
| verifiable expertise are super loud in shouting "$INDUSTRY is
| cooked!!" every time a new model releases. It's exhausting and
| untrue. The kind of video generation we see might nail realism
| but if you want to use it to create something meaningful which
| involves solving a ton of problems and making difficult choices
| in order to express an idea, you run into the walls of easy
| work pretty quickly. It's insulting then for professionals to
| see manga PFPs on X put some slop together and say "movie
| industry is cooked!". It betrays a lack of understanding of
| what it takes to make something good and it gives off a vibe of
| "the loud ones are just trying to force this objectively meh-
| by-default thing to happen".
|
| The other day there was that dude loudly arguing about some
| code they wrote/converted even after a woman with significant
| expertise in the topic pointed out their errors.
|
| Gen AI has its promise. But when you look at the lack of ethics
| from the industry, the cacophony of voices of non experts
| screaming "this time it's really doom", and the
| weariness/wariness that set in during the crypto cycle, it's a
| natural tendency that people are going to call snake oil.
|
| That said, I think the more accurate representation here is
| that HN as a whole is calling the hype snake oil. There's very
| little question anymore about the tools being capable of
| advanced things. But there is annoyance at proclamations of it
| being beyond what it really is at the moment which is that it's
| still at the stage of being an expertise+motivation multiplier
| for deterministic areas of work. It's not replacing that facet
| any time soon on its current trend (which could change wildly
| in 2026). Not until it starts training itself I think. Could be
| famous last words
| senordevnyc wrote:
| I'd put more faith in HN's proclamations if it hadn't widely
| been wrong about AI in 2023, 2024, and now 2025. Watching the
| tone shift here has been fascinating. As the saying goes, the
| only thing moving faster than AI advances right now is the
| speed at which HN haters move the goalposts...
| habinero wrote:
| Mmm. People who make AI their entire personality and brag
| that other people are too stupid to see what they see and
| soon they'll have to see the genius they're denying...does
| not make me think "oh, wow, what have I missed in AI".
| 3A2D50 wrote:
| AI has risen the barrier to all but the top and is
| threatening many peoples' livelihood. It has significantly
| increase the cost of computer hardware and is projected to
| increase the cost of electricity. I can definitely see why
| there is a tone shift! I'm still rooting for AI in general.
| Would love to see the end of a lot of diseases. I don't
| think we humans can cure all disease on our own in any of
| our lifetimes. Of course there all sorts of dystopian
| consequences that may derive from AI fully comprehending
| biology. I'm going to continue being naive and hope for the
| best!
| Madmallard wrote:
| I use them daily and I actively lose progress on complex
| problems and save time on simple problems.
| PeterHolzwarth wrote:
| I think it may be all summed up by Roy Amara's observation that
| _" We tend to overestimate the effect of a technology in the
| short run and underestimate the effect in the long run."_
| ManuelKiessling wrote:
| I think this is the most-fitting one-liner right now.
|
| The arguments going back and forth in these threads are truly
| a sight to behold. I don't want to lean to any one side, but
| in 2025 I've begun to respond to everyone who still argues
| that LLMs are only plagiarism machines, or are only better
| autocompletes, or are only good at remixing the past: Yes,
| correct!
|
| And CPUs can only move zeros and ones.
|
| This is likewise a very true statement. But look where having
| 0s and 1s shuffled around has brought us.
|
| The ripple effects of a machine doing something very simple
| and near-meaningless, but doing it at high speed and again
| and again without getting tired, cannot be underestimated.
|
| At the same time, here is Nobel Laureate Robert Solow, who
| famously, and at the time correctly, stated that "You can see
| the computer age everywhere but in the productivity
| statistics."
|
| It took a while, but eventually, his statement became false.
| legulere wrote:
| The effects might be drastically different from what you
| would expect though. We've seen this with machine learning/AI
| again and again that what looks probable to work doesn't work
| out and unexpected things work.
| xboxnolifes wrote:
| From my perspective, both show HN and Twitter's normal biases.
| I view HN as generally leaning toward "new things suck, nothing
| ever changes", and I view Twitter generally as "Things suck,
| and everything is getting worse". Both of those align with
| snake oil and we're all cooked.
| sanderjd wrote:
| As usual, somewhere in between!
| sph wrote:
| Truth lies in the middle. Yes LLM are an incredible piece of
| technology, and yes we are cooked because once again
| technologists and VC have no idea nor interest in understanding
| the long-term societal ramifications of technology.
|
| Now we are starting to agree that social media has had
| disastrous effects that have not fully manifested yet, and in
| the same breath we accept a piece of technology that promises
| to replace large parts of society with machines controlled by a
| few megacorps and we collectively shrug with "eh, we're gonna
| be alright." I mean, until recently the stated goal was to
| literally recreate advanced super-intelligence with the same
| nonchalance one releases a new JavaScript framework unto the
| world.
|
| I find it utterly maddening how divorced STEM people have
| become from philosophical and ethical concerns of their work. I
| blame academia and the education system for creating this
| massive blind spot, and it is most apparent in echo chambers
| like HN that are mostly composed of Western-educated
| programmers with a degree in computer science. At least on X
| you get, among the lunatics, people that have read more than
| just books on algorithms and startups.
| senordevnyc wrote:
| Because it turns out that HN is mostly made up of cranky
| middle-aged conservatives (small c) who have largely defined
| themselves around coding, and AI is an existential threat to
| their core identity.
| anonnon wrote:
| Why do the mods allow Simon to spam HN with his blogposts and his
| comments, which he often posts just for the sake of including a
| link back to his blog? Seriously, go look at his post history and
| see how often he includes a link to his blog, however
| tangentially related, when he posts a comment. I actually flagged
| this submission, which I never do, and encourage others to do
| likewise.
| simonw wrote:
| Probably because my content gets a lot more upvotes than it
| does flags.
|
| If this post was by anyone _other_ than me would you have any
| problems with its quality?
| dang wrote:
| He's one of the most valuable writers on LLMs, which are one of
| the major topics at present. That's not spam.
| anonnon wrote:
| > He's one of the most valuable writers on LLMs
|
| Is he, really? Most of his blog posts are little more than
| opportunistic, buttressing commentary on someone else's blog
| post or article, often with a bit of AI apologia sprinkled in
| (for example, marginalizing people as paranoid for not taking
| AI companies at their word that they aren't aggressively
| scraping websites in violation of robots.txt, or exfiltrating
| user data in AI-enbaled apps).
|
| EDIT: and why must he link to his blog so often in his
| comments? How is that not SEO/engagement farming? BTW dang, I
| wasn't insinuating the mods were in league with him or
| anything, just that, IMO, he's long past the point at which
| good faith should no longer be assumed.
| simonw wrote:
| If you're not assuming good faith what _are_ you assuming
| here? What 's my motivation?
|
| "buttressing commentary on someone else's blog post"
|
| That's how link blogs work. I wrote more about my approach
| to that here: https://simonwillison.net/2024/Dec/22/link-
| blog/
|
| (And yes, there I go again linking to something I've
| written from a comment. It's entirely relevant to the point
| I am making here. That's why I have a blog - so I can put
| useful information in one place.)
|
| I'll also note that I don't ever share links to my link
| blog posts on Hacker News myself - I don't think they're
| the right format for a HN post. I can't help if other
| people share them here:
| https://news.ycombinator.com/from?site=simonwillison.net
| anonnon wrote:
| > What's my motivation?
|
| Are you really going to insult my and others'
| intelligence like this? Directly or indirectly, _your
| motivation is money._ You already offer monthly
| subscriptions to your blog, and you 're clearly trying to
| build a monetizable brand for yourself as a leading
| authority on AI, especially as it pertains to software
| development.
| simonw wrote:
| If my motivation was money I would cash in on the
| reputation I've already built and go and land a Silicon
| Valley salaried job somewhere.
|
| Sponsorship from my monthly newsletter doesn't come
| close.
|
| Seriously, do you have any idea how much money I'm
| leaving on the table right now NOT having a real job in
| this space?
|
| Being a blogger is wildly financially irresponsible!
| dang wrote:
| Please stop.
| th0ma5 wrote:
| I think when a moderator keeps intervening like this it
| really does mean that there's something wrong here. I
| think people would be less mad if you just went ahead and
| said that you have some kind of special arrangement here
| with this influencer and post publicly that you like them
| constantly spamming the site and letting their fans flood
| the place with deflection and appeals for donations to
| them. Even YouTube had to add a sponsored post
| disclaimer.
| dang wrote:
| There's no special arrangement. The only issue is
| clarifying what content is welcome vs. unwelcome on HN.
| simonw's content is obviously welcome, and this ought to
| be obvious.
|
| > I think people would be less mad
|
| People aren't mad about this. The vast majority of this
| community values simonw's contributions, which are well
| within the sweet spot for material on HN. That's why his
| material gets upvoted, as minimaxir (no friend of
| astroturfers) has pointed out elsewhere in this thread:
| https://news.ycombinator.com/item?id=46451969.
| rvz wrote:
| It _is_ promotional spam.
|
| But given the volume of LLM slop, it was kind of obvious and
| known that even the moderators now have "favourites" over
| guidelines.
|
| > Please don't use HN primarily for promotion. It's ok to
| post your own stuff part of the time, but the primary use of
| the site should be for curiosity. [0]
|
| The blog itself is clearly used as promotion all the time
| when the original source(s) are buried deep in the post and
| almost all of the links link back to his own posts.
|
| This is now a first on HN and a new low for moderators and as
| admitted have regular promotional favourites on the top of
| HN.
|
| [0] https://news.ycombinator.com/newsguidelines.html
| minimaxir wrote:
| The operative word there is "primarily". Simon comments on
| a variety of topics and has far more interactions that
| don't link to his blog than do.
|
| Simon's posts are not engagement farming by any definition
| of the term. He posts good content frequently which is then
| upvoted by the Hacker News community, which should be the
| ideal for a Hacker News contributor.
| rvz wrote:
| Except that the "content" that reaches the top is always
| about AI / LLMs and nothing else and it is "all the
| time". Any opportunity to comment, he will link back to
| his own blog.
|
| He even reposted the same link (which is about AI) with
| one of his posts when the upvotes fell off and until the
| second one reached the top, with the _intention_ of
| promoting his own blog.
|
| Let me simply prove my point to you on how predictable
| this spam is.
|
| He will do a blog post this month about this paper [0]
| with an expert analysis by either someone else (or even
| an LLM) with the primary intention of the blog being used
| for self promotion with at least one link back to his own
| blog.
|
| > ...which is then upvoted by the Hacker News community
|
| You don't know that. But what we do know is that even the
| moderators now have "favourites". Anyone else would be
| shot down for promotional spam.
|
| [0] https://arxiv.org/abs/2512.24880
| simonw wrote:
| "He even reposted the same link (which is about AI) with
| one of his posts when the upvotes fell off"
|
| Where did I do that?
|
| > He will do a blog post this month about this paper [0]
|
| That paper you linked to is a perfect example of where my
| approach can add value!
|
| Did you read it? Do you understand what it saying? It is
| _dense_.
|
| I would love to read an evaluation of that paper by
| someone who can rephrase the core ideas and conversations
| into a couple of paragraphs that help me understand it,
| and help me figure out if I should invest further effort
| in learning more.
|
| I have a whole tag on my blog for that kind of content
| called paper-review:
| https://simonwillison.net/tags/paper-review/ - it's my
| version of the TikTok meme "I read X so you don't have
| to".
|
| Honestly, your problem doesn't seem to be with me so much
| as it seems to be with the concept of _blogging in
| general_.
| th0ma5 wrote:
| [flagged]
| simonw wrote:
| I had to paste that into a separate browser window (jwz
| blocks Hacker News referral traffic) and I cannot figure
| out how that story is relevant to this conversation. Did
| you share the right link?
| dang wrote:
| You've posted over 40 replies hounding this one user whom
| you seem to be fixated on. We've already asked you to
| stop (https://news.ycombinator.com/item?id=44726957) but
| you've continued:
|
| https://news.ycombinator.com/item?id=46409736
|
| https://news.ycombinator.com/item?id=46395646
|
| https://news.ycombinator.com/item?id=46209386
|
| This is obviously an abuse of HN, regardless of who
| you're being aggressive towards. We ban accounts that
| keep doing this. If you keep doing it, we will ban you,
| so no more of this please.
| firexcy wrote:
| I appreciate his work for being more informative and organized
| than average AI-related content. Without his blogging, it would
| be a struggle to navigate the bombastic and narcissistic
| Twitter/Reddit posts for AI updates. The barrier to entry for
| AI reporting is so low that you just need to give a bit more
| care to be distinguished, and he is getting the deserved
| attention for doing exactly that in a systematical and
| disciplined manner. (I do believe many on HN are more than
| capable but not interested in doing the same.) Personally, I
| sometimes find his posts more congratulatory or trivial than I
| like, but I have learned to take what I want and ignore what I
| don't.
| vanderZwan wrote:
| Speaking of new year and AI: my phone just suggested _" Happy
| Birthday!"_ as the quick-reply to any _" Happy New Year!"_
| notification I got in the last hours.
|
| I'm not too worried about my job just yet.
| pants2 wrote:
| It won't help to point out the worst examples. You're not
| competing with an outdated Apple LLM running on a phone. You're
| competing with Anthropic frontier models running on a
| multimillion dollar rack of servers.
| vanderZwan wrote:
| Sounds like I'm much more affordable with better ROI
| gverrilla wrote:
| This year I had a spotify and a youtube thing to "recall my
| year", and it was abolute garbage (30% truth, to be exact). I
| think they're doing it more like an exercise to build up
| systems, infra, processes, people, etc - it's already clear
| they don't actually care about users.
| ogou wrote:
| This is a good tooling survey of the past year. I have been
| watching it as a developer re-entering the job market. The job
| descriptions closely parallel the timeline used in the post.
| That's bizarre to me because these approaches are changing so
| fast. I see jobs for "Skill and Langchain experts with
| production-grade 0>1 experience. Former founders preferred". That
| is an expertise that is just a few months old and startups are
| trying to build whole teams overnight with it. I'm sure January
| and February will have job postings for whatever gets released
| that week. It's all so many sand castles.
| weatherlite wrote:
| > Skill and Langchain experts with production-grade 0>1
| experience.
|
| Also , it's just normal backend work - calling a bunch of APIs.
| What am I missing here?
| walthamstow wrote:
| Buzzwords.
| XenophileJKO wrote:
| That is like saying training tensorflow models is just
| calling some APIs.
|
| Actually making a system like this work seems easy, but isn't
| really.
|
| (Though with the CURRENT generation or two of models it has
| gotten "pretty easy" I think. Before that, not so much.)
| weatherlite wrote:
| No idea about training tenserflow models - is it super
| complex or is it just calling a couple of APIs ? Langchain
| is literally calling an API. Maybe you need to get good
| with prompting or whatever, but I don't see where the
| complexity lies. Please let me know.
| andy99 wrote:
| Having used both Tensorflow (though I expect they mean
| PyTorch which is way more popular, and I have also used)
| and langchain, they are nothing alike.
|
| They he ML frameworks are much closer to implementing the
| mathematics of neural networks, with some abstractions
| but much closer to the linear algebra level. It requires
| an understanding of the underlying theory.
|
| Langchain is a suite of convenience functions for
| composing prompts to LLMs. I wouldn't consider there to
| be some real domain knowledge one would need to use it.
| There is a learning curve but it's about learning the
| different components rather than learning a whole new
| academic discipline.
| HarHarVeryFunny wrote:
| There's a big difference between building an ML framework
| like Tensorflow or PyTorch (I built a Lua Torch-like one
| in C++ myself) and just using it to build/train a model.
|
| Building the model may range from very simple if you are
| just recreating a standard architecture, or be a research
| endeavor if you are designing something completely new.
|
| The difficulty/complexity of then training the model
| depends on what it is. For something simple like a CNN
| for image recognition, it's really just a matter of
| selecting a few hyperparameters and letting it rip. At
| the other end of the spectrum you've got LLMs where
| training (and coping with instabilities) is something of
| a black art, with RL training completely different from
| pre-training, and there is also the issue of
| designing/discovering a pre/mid/post training curriculum.
|
| But anyways, the actual training part can be very simple,
| not requiring too much knowledge of what's going on under
| the hood, depending on the model.
| ogou wrote:
| You're right, none of these new tools are disciplines.
| They are vendor specific approaches that are very recent.
| That's part of my overall point. Who is out there with 2+
| years of very narrow tooling experience at another
| company at a senior level and is available for a rando
| startup (or desparate enterprise looking for bolt-on AI
| features) at a fraction of the pay? Not many, I'm sure.
| We can level up, do training, and maybe stand up a demo
| project. But that won't satisfy an ATS scan. It's
| unrealistic.
| blutoot wrote:
| I hope 2026 will be the year when software engineers and
| recruiters will stop the obsession with leetcode and all other
| forms of competitive programming bullshit
| andrewinardeer wrote:
| Thank you. Enjoyed this read.
|
| AI slop videos will no doubt get longer and "more realistic" in
| 2026.
|
| I really hope social media companies plaster a prominent banner
| over them which screams, "Likely/Made by AI" and give us the
| option to automatically mute these videos from our timeline. That
| would be the responsible thing to do. But I can't see Alphabet
| doing that on YT, xAI doing that on X or Meta doing that on
| FB/Insta as they all have skin in the video gen game.
| sexy_seedbox wrote:
| For image generation, it's already _too_ realistic with Z-Image
| + Custom LoRas + SeedVR2 upscaling.
| hooverd wrote:
| I do think for the solution of say non-consensual pornography
| the only solution is incredible violence against people
| making it.
| compass_copium wrote:
| >I really hope social media companies plaster a prominent
| banner over them which screams, "Likely/Made by AI" and give us
| the option to automatically mute these videos from our
| timeline.
|
| They should just be deleted. They will not be, because they
| clearly generate ad revenue.
| cube00 wrote:
| > social media companies plaster a prominent banner over them
|
| Not going to happen as the social media companies realise they
| can sell you the AI tools used to post slop back onto the
| platform.
| compass_copium wrote:
| >I'm still holding hope that slop won't end up as bad a problem
| as many people fear.
|
| That's the pure, uncut copium. Meanwhile, in the real world,
| search on major platforms is so slanted towards slop that people
| need to specify that they want actual human music:
|
| https://old.reddit.com/r/MusicRecommendations/comments/1pq4f...
| apolloartemis wrote:
| Thank you for your warning about the normalization of deviance.
| Do you think there will be an AI agent software worm like
| NotPetya which will cause a lot of economic damage?
| simonw wrote:
| I'm expecting something like a malicious prompt injection which
| steals API keys and crypto wallets and uses additional tricks
| to spread itself further.
|
| Or targeted prompt injections - like spear phishing attacks -
| against people with elevated privileges (think root sysadmins)
| who are known to be using coding agents.
| lukaslalinsky wrote:
| Speaking of asynchronous agents, what do people use? Claude Code
| for web is extremely limited, because you have no custom tools.
| Claude Code in GitHub Actions is vastly more useful, due to the
| custom environment, but ackward to use interactively. Are there
| any good alternatives?
| simonw wrote:
| I use Claude Code for web with an environment allowing full
| internet access, which means it can install extra tools as and
| when it needs them. I don't run into limits with it very often.
| jimmySixDOF wrote:
| Pretty sure next year's wrapup will have "Year of the sub-
| agent"
| jes5199 wrote:
| I'm running Claude Code in a tmux on a VPS, and I'm working on
| setting up a meta-agent who can talk to me over text messages
| absoluteunit1 wrote:
| Hey - this sounds like really interesting set-up!
|
| Would you be open to providing more details. Would love to
| hear more, your workflows, etc.
| fullstackchris wrote:
| I just use a couple of custom MCP tools with the standard
| claude desktop app:
|
| https://chrisfrew.in/blog/two-of-my-favorite-mcp-tools-i-use...
|
| IMO this is the best balance of getting agentic work done while
| having immediate access to anything else you may need with your
| development process.
| ehsanu1 wrote:
| What exactly do you mean by custom tools here? Just cli tools
| accessible to the agent?
| lukaslalinsky wrote:
| Development environment needed to build and test the project.
| lopatin wrote:
| The "pelicans on a bike" challenge is pretty wide spread now. Are
| we sure it's still not being trained on?
| simonw wrote:
| See https://simonwillison.net/2025/nov/13/training-for-
| pelicans-... (also in the pelicans section of the post).
| lopatin wrote:
| > All I've ever wanted from life is a genuinely great SVG
| vector illustration of a pelican riding a bicycle.
|
| :)
| Razengan wrote:
| My experience with AI so far: It's still far from "butler" level
| assistance for anything beyond simple tasks.
|
| I posted about my failures to try to get them to review my bank
| statements [0] and generally got gaslit about how I was doing it
| wrong, that I if trust them to give them full access to my disk
| and terminal, they could do it better.
|
| But I mean, at that point, it's still more "manual intelligence"
| than just telling someone what I want. A human could easily
| understand it, but AI still takes a lot of wrangling and you
| still need to think from the "AI's PoV" to get the good results.
|
| [0] https://news.ycombinator.com/item?id=46374935
|
| ----
|
| But enough whining. I _want_ AI to get better so I can be lazier.
| After trying them for a while, one feature that I think all
| natural-language As need to have, would be the ability to mark
| certain sentences as "Do what I say" (aka Monkey's Paw) and "Do
| what I mean", like how you wrap phrases in quotes on Google etc
| to indicate a verbatim search.
|
| So for example I could say "[[I was in Japan from the 5th to
| 10th]], identify foreign currency transactions on my statement
| with "POS" etc in the description" then the part in the [[]] (or
| whatever other marker) would be literal, exactly as written, but
| the rest of the text would be up to the AI's
| interpretation/inference so it would also search for ATM
| withdrawals etc.
|
| Ideally, eventually we should be able to have multiple different
| AI "personas" akin to different members of household staff: your
| "chef" would know about your dietary preferences, your "maid"
| would operate your Roomba, take care of your laundry, your
| "accountant" would do accounty stuff.. and each of them would
| only learn about that specific domain of your life: the chef
| would pick up the times when you get hungry, but it won't know
| about your finances, and so on. The current "Projects" paradigm
| is not quite that yet.
| ksec wrote:
| All these improvement in a single year, 2025. While this may seem
| obvious to those who follows along the AI / LLM news. It may be
| worth pointing out again ChatGPT was introduced to us in November
| 2022.
|
| I still dont believe AGI, ASI or Whatever AI will take over human
| in short period of time say 10 - 20 years. But it is hard to
| argue against the value of current AI, which many of the vocal
| critics on HN seems to have the opinion of. People are willing to
| pay $200 per month, and it is getting $1B dollar runway
| _already_.
|
| Being more of a Hardware person, the most interesting part to me
| is the funding of all the developments of latest hardware. I know
| this is another topic HN hate because of the DRAM and NAND
| pricing issue. But it is exciting to see this from a long term
| view where the pricing are short term pain. Right now the
| industry is asking, we have together over a trillion dollar to
| spend on Capex over the next few years and will even borrow more
| if it needs to be, when can you ship us 16A / 14A / 10A and 8A or
| 5A, LPDDR6, Higher Capacity DRAM at lower power usage, better
| packaging, higher speed PCIe or a jump to optical interconnect?
| Every single part of the hardware stack are being fused with
| money and demand. The last time we have this was Post-PC /
| Smartphone era which drove the hardware industry forward for 10 -
| 15 years. The current AI can at least push hardware for another 5
| - 6 years while pulling forward tech that was initially 8 - 10
| years away.
|
| I so wished I brought some Nvidia stock. Again, I guess no one
| knew AI would be as big as it is today, and it is only just
| started.
| coffeebeqn wrote:
| Seems like Nvidia will be focusing on the super beefy GPUs and
| leaving the consumer market to a smaller player
| _s wrote:
| AMD owns a lot of the consumer market already; handhelds,
| consoles, desktop rigs and mobile ... they are not a small
| player.
| utopiah wrote:
| They said "smaller" not small.
| Flow wrote:
| I don't get why Nvidia can't do both? Is it because of the
| limited production capabilities of the factories?
| ACCount37 wrote:
| Yes. If you're bottlenecked on silicon and secondaries like
| memory, why would you want to put more of those resources
| into lower margin consumer products if you could use those
| very resources to make and sell more high margin AI
| accelerators instead?
|
| From a business standpoint, it makes some sense to throttle
| the gaming supply some. Not to the point of surrendering
| the market to someone else probably, but to a measurable
| degree.
| ksec wrote:
| We will have to wait and see but my bet is that Nvidia
| will move to Leading Edge node N2 earlier now they have
| the Margin to work with. Both Hopper and Blackwell were
| too late in the design cycle. The AI hype and continue to
| buy the latest and great leaving Gaming at a mainstream
| node.
|
| Nvidia using Mainstream node has always been the norm
| considering most Fab capacity always goes to Mobile SoC
| first. But I expect the internet / gamers will be angry
| anyway because Nvidia does not provide them with the
| latest and greatest.
|
| In reality the extra R&D cost for designing with leading
| edge will be amortised by all the AI order which give
| Nvidia competitive advantage at the consumer level when
| they compete. That is assuming there are competition
| because most recent data have shown Nvidia owning 90%+ of
| discreet market share, 9% for AMD and 1% for Intel.
| utopiah wrote:
| > All these improvement in a single year
|
| > hard to argue against the value of current AI
|
| > People are willing to pay $200 per month, and it is getting
| $1B dollar runway already.
|
| Those are 3 different things. There can be a LOT of fast and
| significant improvements but still remain extremely far from
| the actual goal, so far it looks like actually little progress.
|
| People pay for a lot of things, including snake oil, so
| convincing a lot of people to pay a bit is not in itself a
| proof of value, especially when some people are basically
| coerced into this, see how many companies changed their
| "strategy" to mandating AI usage internally, or integration for
| a captive audience e.g. Copilot.
|
| Finally yes, $1B is a LOT of money for you and I... but for the
| largest corporations it's actually not a lot. For reference
| Google earned that in revenue... per day in 2023. Anyway that's
| still a big number BUT it still has to be compared with, well
| how much does OpenAI burn. I don't have any public number on
| that but I believe the consensus is that it's a lot. So until
| we know that number we can't talk about an actual runway.
| aspenmartin wrote:
| > People pay for a lot of things, including snake oil, so
| convincing a lot of people to pay a bit is not in itself a
| proof of value
|
| But do you really believe e.g. Claude code is snake oil? I
| pay $200 / month for Claude, which is something I would have
| thought monumentally insane maybe 1-2 years ago (e.g. when
| ChatGPT came out with their premium subscription price I
| thought that seemed so out of touch). I don't think we would
| be seeing the subscription rates and the retention numbers if
| it really was snake oil.
|
| > Finally yes, $1B is a LOT of money for you and I... but for
| the largest corporations it's actually not a lot. For
| reference Google earned that in revenue... per day in 2023.
| Anyway that's still a big number BUT it still has to be
| compared with, well how much does OpenAI burn. I don't have
| any public number on that but I believe the consensus is that
| it's a lot. So until we know that number we can't talk about
| an actual runway.
|
| this gets brought up a lot but I'm not sure I understand why
| folks on a forum called YCombinator, a startup accelerator,
| would make this sound like an obvious sign of charlatanism;
| operating at a loss is nothing new and anthropic / openAI
| strategy seems perfectly rational: they are scaling and
| capturing market share, and TAM is insane.
| chias wrote:
| These are not all improvements. Listed:
|
| * The year of YOLO and the Normalization of Deviance
|
| * The year that Llama lost its way
|
| * The year of alarmingly AI-enabled browsers
|
| * The year of the lethal trifecta
|
| * The year of slop
|
| * The year that data centers got extremely unpopular
| steveBK123 wrote:
| > * The year that data centers got extremely unpopular
|
| I was discussing the political angle with a friend recently.
| I think Big Tech Bro / VC complex has done themselves a big
| disservice by aligning so tightly with MAGA to the point AI
| will be a political issue in 2026 & 2028.
|
| Think about the message they've inadvertently created
| themselves - AI is going to replace jobs, it's pushing
| electric prices up, we need the government to bail us out AND
| give us a regulatory light touch.
|
| Super easy campaign for Dems - big tech trumpers are taking
| your money, your jobs, causing inflation, and now they want
| bailouts !!
| mbesto wrote:
| Said differently - the year we start to see all of the
| externalities of a globally scaled hyped tech trend.
| Y_Y wrote:
| Not that YOLO, PJ Reddie released that in 2015
| jillesvangurp wrote:
| 2025 was the year of development tool using AI agents. I think
| we'll shift attention to non development tool using AI agents.
| Most business users are still stuck using chat gpt as some kind
| of grand oracle that will write their email or powerpoint
| slides. There are bits and pieces of mostly technology demo
| level solutions but nothing that is widely used like AI coding
| tools are so far. I don't think this is bottle necked on model
| quality.
|
| I don't need an AGI. I do need a secretary type agent that
| deals with all the simple but yet laborious non technical tasks
| that keep infringing on my quality engineering time. I'm CTO
| for a small startup and the amount of non technical bullshit
| that I need to deal with is enormous. Some examples of random
| crap I deal with: figuring out contracts, their
| meaning/implication to situations, and deciding on a course of
| action; Customer offers, price calculations, scraping invoices
| from emails and online SAAS accounts, formulating detailed
| replies to customer requests, HR legal work, corporate
| bureaucracy, financial planning, etc.
|
| A lot of this stuff can be AI assisted (and we get a lot of
| value out of ai tools for this) but context engineering is
| taking up a non trivial amount of my time. Also most tools are
| completely useless at modifying structured documents.
| Refactoring a big code base, no problem. Adding structured text
| to an existing structured document, hardest thing ever. The
| state of the art here is an ff-ing sidebar that will suggest
| you a markdown formatted text that you might copy/paste. Tool
| quality is very primitive. And then you find yourself just
| stripping all formatting and reformatting it manually. Because
| the tools really suck at this.
| arcatech wrote:
| > Some examples of random crap I deal with: figuring out
| contracts, their meaning/implication to situations, and
| deciding on a course of action
|
| This doesn't sound like bullshit you should hand off to an
| AI. It sounds like stuff you would care about.
| nrclark wrote:
| Agree. Even asking it can anchor your thinking.
| jillesvangurp wrote:
| I do care about it; kind of my duty as a co-founder. Which
| is why I'm spending double digit percentages of my time
| doing this stuff. But I absolutely could use some tools to
| cut down on a lot of the drudgery that is involved with
| this. And me reading through 40 pages of dense legal German
| isn't one of my strengths since I 1) do not speak German 2)
| am not a lawyer and 3) am not necessarily deeply familiar
| with all the bureaucracy, laws, etc.
|
| But I can ask intelligent questions about that contract
| from an LLM (in English) and shoot back and forth a few
| things, come up with some kind of action plan, and then run
| it by our laywers and other advisors.
|
| That's not some kind of hypothetical thing. That's
| something that happened multiple times in our company in
| the last few months. LLMs are very empowering for dealing
| with this sort of thing. You still need experts for some
| stuff. But you can do a lot more yourself now. And as we've
| found out, some of the "experts" that we relied on in the
| past actually did a pretty shoddy job. A lot of this stuff
| was about picking apart the mess they made and fixing it.
|
| As soon as you start drafting contracts, it gets a lot
| harder. I just went through a process like that as well. It
| involves a lot of manual work that is basically about
| formatting documents, drafting text, running pdfs and text
| snippets through chat gpt for feedback, sparring,
| criticism, etc. and iterating on that. This is not about
| vibe coding some contract but making sure every letter of a
| contract is right. That ultimately involves lawyers and
| negotiating with other stakeholders but it helps if you
| come prepared with a more or less ready to sign off on
| document.
|
| It's not about handing stuff off but about making LLMs work
| for you. Just like with coding tools. I care about code
| quality as well. But I still use the tools to save me a lot
| of time.
| simonw wrote:
| One of the lessons I learned running a startup is that it
| doesn't matter how good the professionals you hire are
| for things like legal and accounting, you _still_ need to
| put work in yourself.
|
| Everyone makes mistakes and misses things, and as the co-
| founder you have to care more about the details than
| anyone else does.
|
| I would have _loved_ to have weird-unreliable-paralegal-
| Claude available back when I was doing that!
| topaztee wrote:
| `Also most tools are completely useless at modifying
| structured documents`
|
| we built a tool for this for the life science space and are
| opening it up to the general public very soon. Email me I can
| give you access (topaz at vespper dot com)
| pjc50 wrote:
| Investing a trillion dollars for a revenue of a billion dollars
| doesn't sound great yet.
| steveBK123 wrote:
| Indeed, its the old Uber playbook at nearly two extra orders
| of magnitude.
|
| It is a large enough number to simply run out of private
| capital to consume before it turns cash flow positive.
|
| Lots of things sell well if sold at such a loss. I'd take a
| new Ferrari for $2500 if it was on offer.
| derwiki wrote:
| Uber's playbook worked for Uber
| aoeusnth1 wrote:
| You say that as if Uber's playbook didn't work. Try this:
| https://www.google.com/finance/quote/UBER:NYSE
| pjc50 wrote:
| Did Uber actually do a lot of capital investment? They
| don't own the cars, for example.
| simonw wrote:
| I believe they spent a huge amount of money on incentives
| to help sign up drivers, and discounts to help attract
| customers.
| ACCount37 wrote:
| Is the AI progress in 2025 an outstanding breakthrough? Not
| really. It's impressive but incremental.
|
| Still, the gap between the capabilities of a cutting edge LLM
| and that of a human is only this wide. There are only this many
| increments it takes to cross it.
| wpietri wrote:
| This is not a great argument:
|
| > But it is hard to argue against the value of current AI [...]
| it is getting $1B dollar runway already.
|
| The psychic services industry makes over $2 billion a year in
| the US [1], with about a quarter of the population being actual
| believers. [2].
|
| [1] The https://www.ibisworld.com/united-
| states/industry/psychic-ser...
|
| [2] https://news.gallup.com/poll/692738/paranormal-phenomena-
| met...
| apexalpha wrote:
| What if these provide actual value through placebo-effect?
| recursive wrote:
| You talking about psychics or LLMs?
| grosswait wrote:
| Yes
| wpietri wrote:
| I think we have different definitions of "actual value".
| But even if I pick the flaccid definition, that isn't proof
| of value of the thing itself, but of any placebo. In which
| case we can focus on the cheapest/least harmful placebo.
| Or, better, solving the underlying problem that the placebo
| "helps".
| computably wrote:
| I'll preface by saying I fully agree that psychics aren't
| providing any non-placebo value to believers, although I
| think it's fine to provide entertainment for non-
| believers.
|
| > Or, better, solving the underlying problem that the
| placebo "helps".
|
| The underlying problems are often a lack of a decent
| education and a generally difficult/unsatisfying life.
| Systemic issues which can't be meaningfully "solved"
| without massive resources and political will.
| jay_kyburz wrote:
| Actually, I'd go one step further and say they are
| harmful to everybody else.
|
| It might just be my circles, but I've seen Carl Sagans
| quote everywhere in the last couple of months.
|
| ""Science is more than a body of knowledge; it is a way
| of thinking. I have a foreboding of an America in my
| children's or grandchildren's time--when the United
| States is a service and information economy; when nearly
| all the key manufacturing industries have slipped away to
| other countries; when awesome technological powers are in
| the hands of a very few, and no one representing the
| public interest can even grasp the issues; when the
| people have lost the ability to set their own agendas or
| knowledgeably question those in authority; when,
| clutching our crystals and nervously consulting our
| horoscopes, our critical faculties in decline, unable to
| distinguish between what feels good and what's true, we
| slide, almost without noticing, back into superstition
| and darkness.""
| ctoth wrote:
| 2022/2023: "It hallucinates, it's a toy, it's useless."
|
| 2024/2025: "Okay, it works, but it produces security
| vulnerabilities and makes junior devs lazy."
|
| 2026 (Current): "It is literally the same thing as a psychic
| scam."
|
| Can we at least make predictions for 2027? What shall the
| cope be then! Lemme go ask my psychic.
| bopbopbop7 wrote:
| 2022/2023: "Next year software engineering is dead"
|
| 2024: "Now this time for real, software engineering is dead
| in 6 months, AI CEO said so"
|
| 2025: "I know a guy who knows a guy who built a startup
| with an LLM in 3 hours, software engineering is dead next
| year!"
|
| What will be the cope for you this year?
| aspenmartin wrote:
| The cope + disappointment will be knowing that a large
| population of HN users will paint a weird alternative
| reality. There are a multitude of messages about AI that
| are out there, some are highly detached from reality (on
| the optimistic and pessimistic side). And then there is
| the rational middle, professionals who see the obvious
| value of coding agents in their workflow and use them
| extensively (or figure out how to best leverage them to
| get the most mileage). I don't see software engineering
| being "dead" ever, but the nature of the job _has already
| changed_ and will continue to change. Look at Sonnet 3.5
| -> 3.7 -> 4.5 -> Opus 4.5; that was 17 months of
| development and the leaps in performance are quite
| impressive. You then have massive hardware buildouts and
| improvements to stack + a ton of R&D + competition to
| squeeze the juice out of the current paradigm (there are
| 4 orders of magnitude of scaling left before we hit real
| bottlenecks) and also push towards the next paradigm to
| solve things like continual learning. Some folks have
| opted not to use coding agents (and some folks like
| yourself seem to revel in strawmanning people who point
| out their demonstrable usefulness). Not using coding
| agents in Jan 2026 is defensible. It won't be defensible
| for long.
| bopbopbop7 wrote:
| Please do provide some data for this "obvious value of
| coding agents". Because right now the only thing obvious
| is the increase in vulnerabilities, people claiming they
| are 10x more productive but aren't shipping anything, and
| some AI hype bloggers that fail to provide any
| quantitative proof.
| aspenmartin wrote:
| Sure: at my MAANG company, where I watch the data closely
| on adoption of CC and other internal coding agent tools,
| most (significant) LOC are written by agents, and most
| employees have adopted coding agents as WAU, and the
| adoption rate is positively correlated with seniority.
|
| Like a lot of things LLM related (Simon Willison's
| pelican test, researchers + product leaders implementing
| AI features) I also heavily "vibe" check the capabilities
| myself on real work tasks. The fact of the matter is I am
| able to dramatically speed up my work. It may be actually
| writing production code + helping me review it, or it may
| be tasks like: write me a script to diagnose this bug I
| have, or build me a streamlit dashboard to analyze +
| visualize this ad hoc data instead of me taking 1 hour to
| make visualizations + munge data in a notebook.
|
| > people claiming they are 10x more productive but aren't
| shipping anything, and some AI hype bloggers that fail to
| provide any quantitative proof.
|
| what would satisfy you here? I feel you are strawmanning
| a bit by picking the most hyperbolic statements and then
| blanketing that on everyone else.
|
| My workflow is now:
|
| - Write code exclusively with Claude
|
| - Review the code myself + use Claude as a sort of review
| assistant to help me understand decisions about parts of
| the code I'm confused about
|
| - Provide feedback to Claude to change / steer it away or
| towards approaches
|
| - Give up when Claude is hopelessly lost
|
| It takes a bit to get the hang of the right balance but
| in my personal experience (which I doubt you will take
| seriously but nevertheless): it is quite the game changer
| and that's coming from someone who would have laughed at
| the idea of a $200 coding agent subscription 1 year ago
| bopbopbop7 wrote:
| Anecdotes don't prove anything, ones without any metrics,
| and especially at MAANG where AI use is strongly
| incentivized.
|
| Evidence is peer reviewed research, or at least something
| with metrics. Like the METR study that shows that
| experienced engineers often got slower on real tasks with
| AI tools, even though they thought they were faster.
| aspenmartin wrote:
| That's why I gave you data! METR study was 16 people
| using Sonnet 3.5/3.7. Data I'm talking about is 10s of
| thousands of people and is much more up to date.
|
| Some counter examples to METR that are in the literature
| but I'll just say: "rigor" here is very difficult
| (including METR) because outcomes are high dimensional
| and nuanced, or ecological validity is an issue. It's
| hard to have any approach that someone wouldn't be able
| to dismiss due to some issue they have with the
| methodology. The sources below also have methodological
| problems just like METR
|
| https://arxiv.org/pdf/2302.06590 -- 55% faster
| implementing HTTP server in javascript with copilot (in
| 2023!) but this is a single task and not really
| representative.
|
| https://demirermert.github.io/Papers/Demirer_AI_productiv
| ity... -- "Though each experiment is noisy, when data is
| combined across three experiments and 4,867 developers,
| our analysis reveals a 26.08% increase (SE: 10.3%) in
| completed tasks among developers using the AI tool.
| Notably, less experienced developers had higher adoption
| rates and greater productivity gains." (but e.g.
| "completed tasks" as the outcome measure is of course
| problematic)
|
| To me, internal company measures for large tech companies
| will be most reliable -- they are easiest to track and
| measure, the scale is large enough, and the talent + task
| pool is diverse (junior -> senior, different product
| areas, different types of tasks). But then outcome
| measures are always a problem...commits per developer per
| month? LOC? task completion time? all of them are highly
| problematic, especially because its reasonable to expect
| AI tools would change the bias and variance of the proxy
| so its never clear if you're measuring the change in
| "style" or the change in the underlying latent measure of
| productivity you care about
| bopbopbop7 wrote:
| To be fair, I'll take a non-biased 16 person study over
| "internal measures" from a MAANG company that burned 100s
| of billions on AI with no ROI.
| insin wrote:
| > - Give up when Claude is hopelessly lost
|
| You love to see "Maybe completely waste my time" as part
| of the normal flow for a productivity tool
| nsxwolf wrote:
| The nature of my job has always been fighting red tape,
| process, and stake holders to deploy very small units of
| code to production. AI really did not help with much of
| that for me in 2025.
|
| I'd imagine I'm not the only one who has a similar
| situation. Until all those people and processes can be
| swept away in favor of letting LLMS YOLO everything into
| production, I don't see how that changes.
| aspenmartin wrote:
| No I think that's extremely correct. I work at a MAANG
| where we have the resources to hook up custom internal
| LLMs and agents to actually deal with that but that is
| unique to an org of our scale.
| Atomic_Torrfisk wrote:
| > People are willing to pay $200 per month
|
| Some people are of course, but how many?
|
| > ... People are willing to pay $200 per month
|
| This is just low-key hype. Careful with your portfolio...
| HumblyTossed wrote:
| It's a great tool, but right now it's only being used to feed
| the greed.
|
| >> Again, I guess no one knew AI would be as big as it is
| today, and it is only just started.
|
| People have been saying similar about self driving cars for
| years now. "AI" is another one of those expensive ideas that
| we'll get 85% of the way there and then to get the other 15%
| will be way more expensive than anyone will want to pay for.
| It's already happening - HW prices and electricity - people are
| starting to ask, "if I put more $ into this machine, when am I
| actually going to start getting money out?" The "true
| believers" are like, soon! But people are right to be hugely
| skeptical.
| jliptzin wrote:
| There are some things it's really great at. For example,
| handling a css layout. If we have to spend trillions of
| dollars and get nothing else out of it other than being able
| to vertically center a <div> without wrestling with css and
| wanting to smash the keyboard in the process, it will all
| have been worth it.
| aspenmartin wrote:
| I agree -- skepticism is totally healthy. And there are so
| many great ways to poke holes in the true underlying
| narratives (not the headlines that people seem to pull from).
| E.g. evaluation science is a wasteland (not for wont of very
| smart people trying very hard to get them right). How do we
| tackle the power requirements in a way that is sustainable?
| Etc. etc.
|
| But stuff like this im not sure I understand:
|
| > It's a great tool, but right now it's only being used to
| feed the greed.
|
| if its a great tool, then how is it _only_ being used to
| "feed the greed" and what do you mean by that?
|
| Also I think folks are quick to make analogies to other
| points in history: "AI is like the dot com boom we're going
| to crash and burn" and "AI is like {self driving cars,
| crypto, etc} and the promises will all be broken, its all
| hype" but this removes the nuance: all of these things are
| extremely different with very specific dynamics that in
| _some_ ways may be similar but in many crucial and important
| ways are completely different.
| belter wrote:
| >> But it is hard to argue against the value of current AI,
| which many of the vocal critics on HN seems to have the opinion
| of.
|
| What is the concrete business case? Can anyone point to a
| revenue producing company using AI in production, and where AI
| is a material driver of profits?
|
| Tool vendors don't count. I'm not interested in how much money
| is being made selling shovels...show me a miner who actually
| struck gold please.
| layer8 wrote:
| > Every single part of the hardware stack are being fused with
| money and demand. The last time we have this was Post-PC /
| Smartphone era which drove the hardware industry forward for 10
| - 15 years. The current AI can at least push hardware for
| another 5 - 6 years while pulling forward tech that was
| initially 8 - 10 years away.
|
| It's very unclear how much end-consumer hardware and DIY
| builders will benefit from that, as opposed to server-grade
| hardware that only makes sense for the enterprise marker. It
| could have the opposite effect, like hardware manufacturers
| leaving the consumer market (as in the case of Micron), because
| there's just not that much money in it.
| mrheosuper wrote:
| I'm not against AI/LLM(in fact, i am quite supportive to it). But
| one of my biggest fear is overusing AI. We may introduce some
| tool that only "AI/LLM" can resonably do(Like tool with weird,
| convoluted UI/UX, syntax) and no one against it because AI/LLM
| can use/interact.
|
| Then genAI, It's become more and more difficult to tell which is
| AI and which is not, and AI is in everywhere. I dont know what to
| think about it. "If you can't tell, does it matter ?"
| netdur wrote:
| i think the concern about software shifting toward ai design
| ignores that the web hasn't been human-first for a long time.
| most traffic is already machine to machine, like crawlers and
| ci pipelines. we've tolerated systems that are barely legible
| for years. anyone who has grepped through android studio logs
| knows that human readability is usually a tertiary goal at
| best. ai interacting with complex systems is just an evolution
| of the glue code we've always written.
|
| as for who made it, utility usually matters more than where it
| came from. i used an agent for an oss changelog recently and it
| picked up things i'd forgotten while structuring the narrative
| better than i could. the intent and code were mine, but the ai
| acted as a high fidelity compressor. the risk isn't ai being
| everywhere. it's the atrophy of judgment where we stop using it
| to support decisions and start using it to outsource thinking.
| jama211 wrote:
| The difference between the performance of models between 2024 and
| 2025 has been so stark, that graph really shows it. There are
| still many people on these forums who seem to think AI's produce
| terrible code unless ultra supervised, and I can't help but
| suspect some of them tried it a little while ago and just don't
| understand how different it is now compared to even quite
| recently.
| Madmallard wrote:
| I used Gemini Pro, Claude Pro yesterday a couple of dozen times
| and basically have been daily.
|
| I have a project to convert my multiplayer XNA game from C# to
| Javascript and to add networking to the game-play using LLMs.
|
| They are far worse at it now than they were a year ago. They
| actually implemented the requirements (Though inaccurately) to
| the best of their ability a year ago. Especially Gemini.
|
| Now they don't even come remotely close to implementing just
| the basic requirements.
|
| The thing is, I'm giving them the entirety of the C# source
| code and spelling out what they should do.
| simonw wrote:
| Weird. I would expect Gemini 3 Pro and Claude Opus 4.5 to run
| rings around Gemini 1.5 Pro and Claude Sonnet 3.5.
|
| How are you running them - regular chat interface or do you
| have them setup with Claude Code or Gemini CLI?
| Madmallard wrote:
| Using the chat interface primarily with various prompting
| strategies.
|
| I am considering making a thread where I compel others to
| attempt to get what I'm trying to get out of it and show me
| their work.
|
| The game is only around 25000-30000 LOC in C#.
| andai wrote:
| Re: yolo mode
|
| I looked into docker and then realized the problem I'm actually
| trying to solve was solved in like 1970 with users and
| permissions.
|
| I just made a agent user limited to its own home folder, and
| added my user to its group. Then I run Claude code etc as the
| agent user.
|
| So it can only read write /home/agent, and it cannot read or
| write my files.
|
| I add myself to agent group so I can read/write the agent files.
|
| I run into permission issues sometimes but, it's pretty smooth
| for the most part.
|
| Oh also I gave it root to a $3 VPS. It's so nice having a
| sysadmin! :) That part definitely feels a bit deviant though!
| jillesvangurp wrote:
| I use a qemu vm for running codex cli in yolo mode and use
| simple ssh based git operations for getting code in and out of
| there. Works great. And you can also do fun things like let it
| loose on multiple git projects in one prompt. The vm can run
| docker as well which helps with containerized tests and other
| more complicated things. One thing I've started to observe is
| that you spend more time waiting for tool execution than for
| model inference. So having a fast local vm is better than a
| slower remote one.
| some_developer wrote:
| Docker in docker, with opencode.
|
| Opencode plus some scripts on host and in its container works
| well to run yolo and only see what it needs (via mounting). Has
| git tools but can't push etc. is thought how to run tests with
| the special container-in-container setup.
|
| Including pre-configured MCPs, skills, etc.
|
| The best part is that it just works for everyone on the team,
| big plus.
| knicholes wrote:
| cgroups and namespaces
| staeff777 wrote:
| I really like this idea and just tried some steps for myself.
| create user with homedir: sudo useradd -m agent add myself to
| agent group: sudo usermod -a -G agent $USER
|
| Allow agent group to agent home dir: sudo chmod -R 770
| /home/agent
|
| Start a new shell with the group (or login/logoff): newgrp
| agent Now you should be able to change into the agent home.
|
| Allow your user to sudo as agent: echo "$USER ALL=(agent)
| NOPASSWD: ALL" |sudo tee -a /etc/sudoers.d/$USER-as-agent now
| you can start your agent using sudo: sudo -u agent your_agent
|
| works nice.
| andai wrote:
| Re: yolo mode
|
| https://markdownpastebin.com/?id=1ef97add6ba9404b900929ee195...
|
| My notes from back when I set this up! Includes instructions
| for using a GUI file explorer as the agent user. As well as
| setting up a systemd service to fix the permissions
| automatically.
|
| (And a nice trick which shows you which GUI apps are running as
| which user...)
|
| However, most of these are just workarounds for the permission
| issue I kept running into, which is that Claude Code would for
| some reason create files with incorrect permissions so that I
| couldn't read or write those files from my normal account.
|
| If someone knows how to fix that, or if someone at Anthropic is
| reading, then most of this Rube Goldberg machine becomes
| unnecessary :)
| yupyupyups wrote:
| Let's talk about the societal cost these models have had on us
| including their high energy cost and the proliferation of auto-
| generated slop media used to milk ad revenue, scam people, SEO
| farm, do propaganda or automate trolling. What about these big
| corporations collecting an astronomical amount of debt to hoard
| DRAM and NAND in a way that has crippled the PC market within
| weeks? And what are they going to do next, put a few dollars in
| Trump's pocket so that they can rob/loot the US population
| through bailouts? Who gets to keep all the hardware I wonder?
|
| Nvidia, Samsung, SK Hynix and some other voltures I forgot to
| mention are making serious bank right now.
| ashishgupta2209 wrote:
| 2026: The Year of Robots, note it for next year
| fullstackchris wrote:
| > The reason I think MCP may be a one-year wonder is the
| stratospheric growth of coding agents. It appears that the best
| possible tool for any situation is Bash--if your agent can run
| arbitrary shell commands, it can do anything that can be done by
| typing commands into a terminal.
|
| I push back strongly from this. In the case of the solo, one-
| machine coder, this is likely the case - if you're exposing
| workflows or fixed tools to customers / collegues / the web at
| large via API or similar, then MCP is still the best way to
| expose it IMO.
|
| Think about a GitHub or Jira MCP server - commandline alone they
| are sure to make mistakes with REST requests, API schema etc.
| With MCP the proper known commands are already baked in. Remember
| always that LLMs will be better with natural language than code.
| simonw wrote:
| The solution to that is Anthropic's Skills.
|
| Create a folder called skills/how-to-use-jira
|
| Add several Bash scripts with the right curl commands to
| perform specific actions
|
| Add a SKILL.md file with some instructions in how to use those
| scripts
|
| You've effectively flattened that MCP server into some Markdown
| and Bash, only the thing you have now is more flexible (the
| coding agent can adapt those examples to cover new things you
| hadn't thought to tell it) and much more context-efficient (it
| only reads the Markdown the first time you ask it to do
| something with JIRA).
| aflukasz wrote:
| But that moves the burden of maintenance from the provider of
| the service to its users (and/or partially to intermediary in
| form of "skills registry" of sorts, which apparently is a
| thing now).
|
| So maybe a hybrid approach would make more sense? Something
| like /.well-known/skills/README.md exposed and owned by the
| providers?
|
| That is assuming that the whole idea of "skills" makes sense
| in practice.
| simonw wrote:
| Yeah that's true, skill distribution isn't a solved problem
| yet - MCPs have a URL, which is a great way of making them
| available for people to start using without extra steps.
| rr808 wrote:
| What happened to Devin? 2024 it was a leading contender now it
| isn't even included in the big list of coding agents.
| fullstackchris wrote:
| Wasn't it basically revealed as a scam? I remember some article
| about their fancy demo video being sped up / unfairly cut and
| sliced etc.
| monkeydust wrote:
| https://cognition.ai/blog/devin-annual-performance-review-20...
| ColinEberhardt wrote:
| It's still around, and tends to be adopted by big enterprises.
| It's generally a decent product, but is facing a lot of equally
| powerful competition and is very expensive.
| simonw wrote:
| To be honest that's more because I've never tried it myself, so
| it isn't really on my radar.
|
| I don't hear much buzz about it from the people I pay attention
| to. I should still give it a go though.
| Gud wrote:
| What about self hosting?
| simonw wrote:
| I talked about that in this section
| https://simonwillison.net/2025/Dec/31/the-year-in-llms/#the-...
| - and touched on it a bit in the section about Chinese AI labs:
| https://simonwillison.net/2025/Dec/31/the-year-in-llms/#the-...
| politelemon wrote:
| > The problem is that the big cloud models got better too--
| including those open weight models that, while freely available,
| were far too large (100B+) to run on my laptop.
|
| The actual, notable progress will be models that can run
| reasonably well on commodity, everyday hardware that the average
| user has. From more accessibility will come greater usefulness.
| Right now the way I see it, having to upgrade specs on a machine
| to run local models keeps it in a niche hobbyist bubble.
| timonoko wrote:
| OpenSCAD-coding has improved significantly on all models. Now
| syntax is always right and they understand the concept of
| _negative space_.
|
| Only problem is that they don't see connection between form and
| function. They may make teapot perfectly but don't understand
| that this form is supposed to contain liquid.
| mark_l_watson wrote:
| Thanks Simon, great writeup.
|
| It has been an amazing year, especially around tooling (search,
| code analysis, etc.) and surprisingly capable smaller models.
| huqedato wrote:
| I completely disagree with the idea that 2025 "The (only?) year
| of MCP." In fact, I believe every year in the foreseeable future
| will belong to MCP. It is here to stay. MCP was the best
| (rational, scalable, predictable) thing since LLM madness broke
| loose.
| mmcnl wrote:
| Let's hope 2026 will also have interesting innovations not
| related to AI or LLMs.
| spicyusername wrote:
| 2025 had plenty of those, they just didn't get as many news
| headlines.
|
| One of the difficult things of modernity is that it's easy to
| confuse what you hear about a lot with what is real.
|
| One of the great things about modernity is that progress
| continues, whether we know about it or not.
| _pdp_ wrote:
| With everything that we have done so far (our company) I believe
| by end of 2026 our software will be self improving all the time.
|
| And no it is not AI slop and we don't vibe code. There are a lot
| of practical aspects of running software and maintaining /
| improving code that can be done well with AI if you have the
| right setup. It is hard to formulate what "right" looks like at
| this stage as we are still iterating on this as well.
|
| However, in our own experiments we can clearly see dramatic
| increases in automation. I mean we have agents working overnight
| as we sleep and this is not even pushing the limits. We are now
| wrapping major changes that will allows us to run AI agents all
| the time as long as we can afford them.
|
| I can even see most of these materialising in Q1 2026.
|
| Fun times.
| papacj657 wrote:
| What exactly are your agents doing overnight? I often hear
| folks talk about their agents running for long periods of time
| but rarely talk about the outcomes they're driving from those
| agents.
| _pdp_ wrote:
| We have a lot of grunt work scheduled overnight like finding
| bugs, creating tests where we don't have good coverage or
| where we can improve, integrations, documentation work, etc.
|
| Not everything gets accepted. There is a lot of work that is
| discarded and much more pending verification and acceptance.
|
| Frankly, and I hope I don't come as alarmist (judge for
| yourself from my previous comments on Hn and Reddit) we
| cannot keep up with the output! And a lot of it is actually
| good and we should incorporate it even partially.
|
| At the moment we are figuring out how to make things more
| autonomous while we have the safety and guardrails in place.
|
| The biggest issue I see at this stage is how to make sense of
| it all as I do not believe we have the understanding of what
| is happening - just the general notion of it.
|
| I truly believe that we will reach the point where ideas
| matter more than execution, which what I would expect to be
| the case with more advanced and better applied AI.
| asgR1t wrote:
| Most LLMs got worse in 2025. Only addicts and the type of
| computer gamer that feels drawn to complex setups, gamification
| and does not care about the end result will feel positive about
| the grift.
|
| 2025: The Year in Open Source? Nothing, all resources were tied
| up to debunk a couple of Python web developers who pose as the
| ultimate experts in LLMs.
| simonw wrote:
| In what way did they get worse?
|
| I made you a dashboard of my 2025 writing about open-source
| that didn't include AI:
| https://simonwillison.net/dashboard/posts-with-tags-in-a-yea...
| icapybara wrote:
| It was the year of Claude Code
| nativeit wrote:
| Between the people with invested and/conflicting interests, and
| the hordes of dogmatic zealots, I find discussions about AI to be
| the least productive or reliably informed on HN.
| simonw wrote:
| Honestly this thread was pretty disappointing. Many of the
| comments here could have been attached to _any_ post about LLMs
| in the past year or so.
| ck2 wrote:
| as I was clicking "gee I hope there's the year of pelicans riding
| bicycles"
|
| left satisfied, lol
| losvedir wrote:
| I predict 2026 will be the year of the first AI Agent "worm" (or
| virus?). Kind of like the Morris worm running amok as an
| experiment gone wrong, I think we will sometime soon have someone
| set up an AI agent whose core loop is to try to propagate itself,
| either as an experiment or just for the lulz.
|
| The actual Agent payload would be very small, likely just a few
| hundred line harness plus system prompt. It's just a question of
| whether the agent will be skilled enough to find vulnerabilities
| to propagate. The interesting thing about an AI worm is that it
| can use different tricks on different hosts as it explores its
| own environment.
|
| If a pure agent worm isn't capable enough, I could see someone
| embedding it on top of a more traditional virus. The normal virus
| would propagate as usual, but it would also run an agent to
| explore the system for things to extract or attack, and to find
| easy additional targets on the same internal network.
|
| A main difference here is that the agents have to call out to a
| big SotA model somewhere. I imagine the first worm will simply
| use Opus or ChatGPT with an acquired key, and part of it will be
| trying to identify (or generate) new keys as it spreads.
|
| Ultimately, I think this worm will be shut down by the model
| vendor, but it will have to have made a big enough splash
| beforehand to catch their attention and create a team to identify
| and block keys making certain kinds of requests.
|
| I'd hope OpenAI, Anthropic, etc have a team and process in place
| already to identify suspicious keys, eg, those used from a huge
| variety of IPs, but I wouldn't be surprised if this were low on
| their list of priorities (until something like this hits).
___________________________________________________________________
(page generated 2026-01-01 23:01 UTC)