[HN Gopher] Cloudlflare builds OAuth with Claude and publishes a...
___________________________________________________________________
Cloudlflare builds OAuth with Claude and publishes all the prompts
Author : gregorywegory
Score : 339 points
Date : 2025-06-02 14:24 UTC (8 hours ago)
(HTM) web link (github.com)
(TXT) w3m dump (github.com)
| gregorywegory wrote:
| From the readme: This library (including the schema
| documentation) was largely written with the help of Claude, the
| AI model by Anthropic. Claude's output was thoroughly reviewed by
| Cloudflare engineers with careful attention paid to security and
| compliance with standards. Many improvements were made on the
| initial output, mostly again by prompting Claude (and reviewing
| the results). Check out the commit history to see how Claude was
| prompted and what code it produced.
|
| "NOOOOOOOO!!!! You can't just use an LLM to write an auth
| library!"
|
| "haha gpus go brrr"
|
| In all seriousness, two months ago (January 2025), I (@kentonv)
| would have agreed. I was an AI skeptic. I thoughts LLMs were
| glorified Markov chain generators that didn't actually understand
| code and couldn't produce anything novel. I started this project
| on a lark, fully expecting the AI to produce terrible code for me
| to laugh at. And then, uh... the code actually looked pretty
| good. Not perfect, but I just told the AI to fix things, and it
| did. I was shocked.
|
| To emphasize, this is not "vibe coded". Every line was thoroughly
| reviewed and cross-referenced with relevant RFCs, by security
| experts with previous experience with those RFCs. I was trying to
| validate my skepticism. I ended up proving myself wrong.
|
| Again, please check out the commit history -- especially early
| commits -- to understand how this went.
| tonyhart7 wrote:
| same argument with me but only for claude
|
| another models feels like shit to use, but claude is good
| unshavedyak wrote:
| Yup. I'm more skeptic than pro-AI these days, but nonetheless
| i'm still trying to use AI in my workflows.
|
| I don't actually enjoy it, i generally find it difficult to use
| as i have more trouble explaining what i want than actually
| just doing it. However it seems clear that this is not going
| away and to some degree it's "the future". I suspect it's
| better to learn the new tools of my craft than to be caught
| unaware.
|
| With that said i still think we're in the infancy of actual
| tooling around this stuff though. I'm always interested to see
| novel UXs on this front.
| qsort wrote:
| Probably unrelated to the broader discussion, but I don't
| think the "skeptic vs pro-AI" distinction even makes that
| much sense.
|
| For example, I usually come off as being relatively skeptic
| within the HN crowd, but I'm actually pushing for _more_
| usage at work. This kind of "opinion arbitrage" is common
| with new technologies.
| diggan wrote:
| > but I don't think the "skeptic vs pro-AI" distinction
| even makes that much sense
|
| Tends to be like that with subjects once feelings get
| involved. Make any skepticism public, even if you don't
| feel strongly either way, and you get one side of
| extremists yelling at you about X. At the same time, say
| anything positive and you get the zealots from the other
| side yelling at you about Y.
|
| Us who tend to be not so extremist gets push back from both
| sides, either in the same conversations or in different
| places, while both see you as belonging to "the other side"
| while in reality you're just trying to take a somewhat
| balanced approach.
|
| These "us vs them" never made sense to me, for (almost) any
| topic. Truth usually sits somewhere around the middle, and
| a balanced approach seems to usually result in more
| benefits overall, at least personally for me.
| unshavedyak wrote:
| > but I don't think the "skeptic vs pro-AI" distinction
| even makes that much sense.
|
| Imo it does, because it frames the underlying assumptions
| around your comment. Ie there was some very pro-AI folks
| who think it's not just going to replace everything, but
| already is. That's an extreme example of course.
|
| I view it as valuable anytime there's extreme hype, party
| lines, etc. If you don't frame it yourself, others will and
| can misunderstand your comment when viewed through the
| wrong lens.
|
| Not a big deal of course, but neither is putting a
| qualifier on a comment.
| steveklabnik wrote:
| One recent post I read about improving the discourse (which
| I seem to have lost the link...) agrees, but in a different
| way: adding a "capable vs not" axis. that is, "I believe AI
| is good enough to replace humans, _and_ I am pro " is
| different than "I believe AI is good enough to replace
| humans, _and_ I am against " and while "I believe AI is not
| good enough to replace humans, and I am pro" is a weird
| position to take, "I believe AI is not good enough to
| replace humans, and I am against."
|
| These things are also not binary, they're a full grid of
| space.
| baq wrote:
| > "I believe AI is not good enough to replace humans, and
| I am pro" is a weird position to take
|
| Huh? The recipe how to be in this position is literally
| in the readme of the linked project. You don't even have
| to believe it, you just have to work it.
| steveklabnik wrote:
| I mean at the most extreme: that it can NEVER do so.
| Someone who holds this position would point to commits
| like https://news.ycombinator.com/item?id=44159659
| jakeydus wrote:
| > "I believe AI is not good enough to replace humans, and
| I am pro" is a weird position to take
|
| I think that's just the opinion of someone who doesn't
| think AI currently lives up to the hype but is optimistic
| about developing it further, not really that weird of a
| position in my opinion.
|
| Personally I'm moving more into the "I think AI is good
| enough to replace humans, and I am against" category.
| steveklabnik wrote:
| Yeah, I meant like, at the full extreme of "and it never
| will". Someone with the position you describe wouldn't be
| at the far end, but somewhere closer to the middle.
| immibis wrote:
| I believe compilers are not good enough to replace
| humans, and I am pro
| thewebguyd wrote:
| > I don't actually enjoy it, i generally find it difficult to
| use as i have more trouble explaining what i want than
| actually just doing it.
|
| This is my problem I run into quite frequently. I have more
| trouble trying to explain computing or architectural concepts
| in natural language to the AI than I do just coding the damn
| thing in the first place. There are many reasons we don't
| program in natural language, and this is one of them.
|
| I've never found natural language tools easier to use, in any
| iteration of them, and so I get no joy out of prompting AI.
| Outside of the increasingly excellent autocomplete, I find it
| actually slows me down to try and prompt "correctly."
| hattmall wrote:
| I guess for me the questions is, at what point do you feel it
| would be reasonable to this without the experts involved in
| your case?
|
| As an edit, after reading some of the prompts, what is the
| likelihood that a non-expert could even come up with those
| prompts?
|
| The really really interesting thing would be if an AI could
| actually generate the prompts.
| dkdcio wrote:
| Why do you need a non-expert? We built on layers of
| abstractions, AI will help you at whichever layer you're the
| "expert" at. Of course you'll need to understand low-level
| stuff to work on low-level code
|
| i.e. I might not use AI to build an OAuth library, but I
| might use AI to build a web app (which I am an expert at)
| that may use an OAuth library Cloudfare developed (which
| theya are experts at). Trying to make "anyone" code
| "anything" doesn't seem like the point to me
| nisegami wrote:
| GP is just quoting the readme, they aren't the author.
|
| My 2 cents:
|
| >I guess for me the questions is, at what point do you feel
| it would be reasonable to this without the experts involved
| in your case?
|
| No sooner and no later than we could say the same thing about
| a junior developer. In essence, if you can't validate the
| code produced by a LLM then you shouldn't really have been
| writing that code to begin with.
|
| >The really really interesting thing would be if an AI could
| actually generate the prompts.
|
| I think you've hit on something that is going underexplored
| right now in my opinion. Orchestration of AI agents, where a
| we have a high level planning agent delegating subtasks to
| more specialized agents to perform them and report back. I
| think an approach like that could help avoid context
| saturation for longer tasks. Cline / Aider / Roo Code / etc
| do something like this with architect mode vs coding mode but
| I think it can be generalized.
| kentonv wrote:
| (I'm the author of this library -- or, the guy who prompted
| the AI at least.)
|
| I absolutely would not vibe code an OAuth implementation! Or
| any other production code at Cloudflare. We've been using
| more AI internally, but made this rule very clear: the human
| engineer directing the AI must fully understand and take
| responsibility for any code which the AI has written.
|
| I do think vibe coding can be really useful in low-stakes
| environments, though. I vibe-coded an Android app to use as a
| baby monitor (it just streams audio from a Unifi camera in
| the kid's room). I had no previous Android experience, and it
| would have taken me weeks to learn without AI, but it only
| took a few hours with AI.
|
| I think we are in desperate need of _safe_ vibe coding
| environments where code runs in a sandbox with security
| policies that make it impossible to screw up. That would
| enable a whole lot of people to vibe-code personal apps for
| personal use cases. It happens I have some background
| building such platforms...
|
| But those guardrails only really make sense at the
| application level. At the systems level, I don't think this
| is possible. AI is not smart enough yet to build systems
| without serious bugs and security issues. So human experts
| are still going to be necessary for a while there.
| diggan wrote:
| > I think we are in desperate need of safe vibe coding
| environments where code runs in a sandbox with security
| policies that make it impossible to screw up.
|
| OpenAI's new Rust version of Codex might be of interest,
| haven't dived deeper into the codebase but seems they're
| thinking about sandboxing from the get-go: https://github.c
| om/openai/codex/blob/7896b1089dbf702dd079299...
| freedomben wrote:
| What tools did you use for the vibe coding an Android app?
| And was it able to do the UI stuff too?
|
| I've wanted to do this but am not sure how to get started.
| For example, should I generate a new app in Android Studio
| and then point Claude Code at it? Or can I ask Claude Code
| (or another agent) to start it from scratch? (in the past
| that did not work, but I'm curious if it's just a PEBKAC
| error)
| kentonv wrote:
| I used Claude Code. I actually just asked it what tools I
| needed for a CLI-driven build, and it told me what to
| install (or even installed it for me in some cases). I
| basically didn't read any documentation, just asked
| Claude what I should do.
| freedomben wrote:
| Amazing, thank you!
|
| Edit: Holy shit, in 30 minutes I used Claude code to make
| a simple PDF viewer app, and it totally works. I did have
| to prompt it through the process quite a bit, including
| correcting some obvious flubs, but I'm super impressed.
|
| I didn't even have to install the android dev tools
| because I asked it to generate a Dockerfile in which to
| do the build, and a simple script to copy the apk out
| when done :-D
| rangerelf wrote:
| > I guess for me the questions is, at what point do you feel
| it would be reasonable to this without the experts involved
| in your case?
|
| I don't know if it was the intent but these kind of questions
| bother me, the seem to hint at an agenda, "when can I have a
| farm of idiots with keyboards paid minimum wage churn out
| products indistinguishable from expertly designed
| applications".
|
| To me that's the danger of AI, not it's purported
| intelligence, but our manifested greed.
| hattmall wrote:
| Yeah, I mean that is definitely the intent of the question
| and it's absolutely one of the factors that's driving money
| into AI.
|
| Assisting competent engineers certainly has value, but it's
| not an easy calculation to assess that value compared to
| the actual non-subsidized cost of AI's current state.
|
| On the other hand having a farm of idiots, or even no
| idiots at all, just computers, churning out high quality
| applications is a completely different value proposition.
| mtlynch wrote:
| > _In all seriousness, two months ago (January 2025), I
| (@kentonv) would have agreed._
|
| I'm confused by "I (@kentonv)" means here because kentonv is a
| different user.[0] Are you saying this is your alt? Or is this
| a typo/misunderstanding?
|
| Edit: Figured out that most of your post is quoting the README.
| Consider using > and * characters to clarify.
|
| [0] https://news.ycombinator.com/user?id=kentonv
| kentonv wrote:
| He is quoting from the project readme. I wrote all this text.
| mdaniel wrote:
| Thanks for weighing in here
|
| If I might make a suggestion, based on how fast things
| change, even within a model family, you may benefit from
| saying Claude _what_. I was especially cognizant of this
| given the recent v4 release which (of course) hailed as the
| second coming. Regardless, you may want to update your
| readme to say
|
| It may also be _wildly_ out of scope for including in a
| project 's readme, but knowing which of the bazillions of
| coding tools you used would also help a tiny bit with this
| reproduction crises found in every single one of these
| style threads
| diggan wrote:
| > It may also be wildly out of scope for including in a
| project's readme
|
| The entire point of the repository seems to be to
| invalidate/validate the thesis if LLMs are good enough to
| be pair programmers right now. Removing it from the
| README makes no sense in that context.
| mdaniel wrote:
| I did consider that, but the repo isn't called "kentonv
| does a yolo" it's straight-up labeled as a provider
| library for CF workers under Cloudflare's brand
|
| Some hair splitting about whether including the Claude
| stanza is "full disclosure," or "AI advocacy," or just
| because it's cool
|
| Anyway, I mentioned the out of scope because if half the
| readme is about correct usage of the library, and half is
| about the sausage making, I'd be confused as a reader
| about whether this was designed to be for real or for
| funzies
| uludag wrote:
| I found it pretty strange to include in the readme as
| well. Like, imagine someone relied on fiverr or
| codementor.io to write this code. It'd be weird to say in
| the readme "I was fairly skeptical that I could get
| quality code written on Fiverr, but I tried it and it
| turns out it was pretty good!"
|
| My guess is there were some push to doing anything
| related to AI at the company. I feel a lot of companies
| are doing this these days.
| kentonv wrote:
| This library is a core component of our MCP framework,
| it's not just an experiment.
| kentonv wrote:
| I believe it's important to say when AI was used so
| heavily in building a library -- it would feel dishonest
| to me to claim I wrote it all myself. I also think it's
| just a pretty interesting thing to know about. So I think
| it belongs in the readme. (But I'm not making a moral
| judgment on what anyone else does.)
|
| It was almost entirely Claude Sonnet 3.7. I agree I
| should add the version to the readme.
| pera wrote:
| That's interesting. My experience with Sonnet 3.7 early
| this year was pretty poor: It simply couldn't reach the
| correct solution alone, even when explaining the issues
| explicitly. The proposed invalid solution was not too far
| from the correct one, so you could fix it manually if you
| knew what you were doing, but then the way the code was
| structured was not something that I would like to
| maintain in a real project. All this on top of the usual
| UX issues like hallucinated APIs. The experience
| refactoring was even worse.
|
| I guess your mileage is highly dependent on the domain of
| your problem? In my case was GIS by the way
| diggan wrote:
| It's a literal copy-paste from the README, I think it was
| supposed to be quoted but parent messed it up somehow.
|
| https://github.com/cloudflare/workers-oauth-
| provider/blob/fe...
| stego-tech wrote:
| On the one hand, I would expect LLMs to be able to crank out
| such code when prompted by skilled engineers who also
| understand prompting these tools correctly. OAuth isn't new,
| has tons of working examples to steal as training data from
| public projects, and in a variety of existing languages to suit
| most use cases or needs.
|
| On the other hand, where I remain a skeptic is this constant
| banging-on that somehow this will translate into entirely new
| things - research, materials science, economies, inventions,
| etc - because that requires learning "in real time" from
| information sources you're literally generating in that moment,
| not decades of Stack Overflow responses without context. That
| has been bandied about for years, with no evidence to show for
| it beyond specifically cherry-picked examples, often from
| highly-controlled environments.
|
| I never doubted that, with competent engineers, these tools
| could be used to generate "new" code from past datasets. What I
| continue to doubt is the utility of these tools given their
| immense costs, both environmentally and socially.
| TeMPOraL wrote:
| > _where I remain a skeptic is this constant banging-on that
| somehow this will translate into entirely new things -
| research, materials science, economies, inventions, etc -
| because that requires learning "in real time" from
| information sources you're literally generating in that
| moment, not decades of Stack Overflow responses without
| context._
|
| Personally I hope this will materialize, at the very least
| because there's plenty of discoveries to be made by cross-
| correlating discoveries already made; the necessary
| information _should_ be there, but reasoning capability (both
| that of the model and that added by orchestration) seems to
| be lacking. I 'm not sure if pure chat is the best way to
| access it, either. We need better, more hands-on tools to
| explore the latent spaces of LLMs.
| stego-tech wrote:
| I don't consider that "new" research, personally - because
| AI boosters don't consider that "new". The future they hype
| is one where these LLMs can magic up entirely new fields of
| research and study without human input, which isn't how
| these models are trained in the first place.
|
| That said, yes, it could be highly beneficial for
| identifying patterns in existing research that allows for
| new discoveries - provided we don't trust it blindly and
| actually validate it with science. Though I question its
| value to society in burning up fossil fuels, polluting the
| atmosphere, and draining freshwater supplies compared to
| doing the same work with Grad Students and Scientists with
| the associated societal feedback involved in said
| employment activities.
| TeMPOraL wrote:
| > _Though I question its value to society in burning up
| fossil fuels, polluting the atmosphere, and draining
| freshwater supplies compared to doing the same work with
| Grad Students and Scientists with the associated societal
| feedback involved in said employment activities._
|
| I'd imagine AI is much cheaper on that front than grad
| students, whether you count marginal contribution, or
| total costs of building and utilization. Humans are damn
| expensive and environmentally intensive to rear and keep
| around.
| stego-tech wrote:
| You really should read the papers and reporting coming
| out about the sheer cost of these AI models and their
| operation. It might seem significantly cheaper in the
| context of immediate impact, but those humans provide
| knock-on impacts that can _decrease_ their environmental
| impact (especially if done in concert), while the current
| crop of AI is content burning NatGas turbines and
| guzzling up groundwater just so a human isn't tasked with
| reading a full paragraph of information, or a white paper
| of important content - and that's the _most optimistic_
| view, at present.
|
| Evaluating a technology in a vacuum does not work when
| trying to assess its impact, and in that wider context I
| don't see the value-add of these models deployed at
| scale, especially when their marketing continues focusing
| on synthetic benchmarks and lofty future-hype instead of
| immediately practicable applications (like this one was).
| waynenilsen wrote:
| most engineering is glorified plumbing so as far as labour
| productivity goes, this should go a long way
| stego-tech wrote:
| I doubt it, for the simple reason that literal plumbers
| still make excellent money _because_ plumbing is ultimately
| bespoke output built on standards.
|
| Everyone wants to automate the (proverbial) plumbing, until
| shit spews everywhere and there's nobody to blame but
| yourself.
| realreality wrote:
| Plumbers make excellent money because regulations require
| licensed plumbers to do the work, and plumbing unions
| have a financial interest in limiting the number of
| plumbers.
|
| But anybody can do plumbing. It's not rocket science.
| stego-tech wrote:
| > regulations require licensed plumbers to do the work
|
| Regulations come about because of repeated failures that
| end up harming the public. Regulations aren't a dirty
| word, and aren't obstacles to be "disrupted" in most
| cases.
|
| > plumbing unions have a financial interest in limiting
| the number of plumbers
|
| Golly gee, it's almost as if - because we live in a
| society where everyone _must_ work in order to survive -
| that skilled professionals have a vested interest in
| ensuring only qualified candidates may join their ranks,
| to make it harder to depress wages below subsistence
| levels (the default behavior of unregulated capital).
|
| > But anybody can do plumbing. It's not rocket science.
|
| Oh wow, I had no idea I was qualified to design sewage
| infrastructure for my township just because I plumbed my
| Amazon bidet into the cold water line! Sure glad there's
| no regulations stopping me from becoming a licensed
| plumber since apparently that's all it takes to succeed!
|
| Sarcasm aside, your argument holds about as much
| substance as artificial sweetener: it _sounds_ informed
| and wise, but anyone with substantial experience in
| reality and collaborating with other people knows that
| all you're spewing is ignorance of the larger systems at
| work and their interplay.
| dragonwriter wrote:
| > Regulations come about because of repeated failures
| that end up harming the public.
|
| Sometimes, but see also the concepts of "iron triangles"
| and "regulatory capture".
| stego-tech wrote:
| You're not wrong (examples include the US FCC, ZA's
| Telekom, ye olde Standard Oil, vertical
| integrations...the list goes on, and even includes modern
| cloud services and AI tools, since the regulations they
| champion are often intended to block competitors with
| onerous compliance requirements), but in the context of
| the person I was replying to, they used "regulation" very
| much in the same context Uber/AirBnB and the SV
| Libertarian ilk decry "regulations".
|
| Regulations aren't a binary (exclusively good or
| exclusively bad), yet so many of the HN cohort have drank
| the "exclusively bad and everyone can be trusted to make
| good decisions forever" koolaid that seeks to dismantle
| regulations wholesale.
| realreality wrote:
| You're wasting your time fighting a straw man. I never
| said all regulations are bad.
|
| The question was why plumbers are expensive. I assert
| that it's not because plumbing is especially difficult.
| stego-tech wrote:
| > You're wasting your time fighting a straw man.
|
| Smartest thing you've said all day. Thanks for reminding
| me that trying to convince someone of something when they
| cannot be bothered to do research beyond first order
| impacts is a waste of my time.
| realreality wrote:
| Designing sewage infrastructure isn't rocket science,
| either. If citizens in your town needed to do it, they
| could figure it out, regardless of their credentials.
|
| Sometimes regulations come about to protect the public.
| Often, they're enacted to protect the profits of
| insurance companies, banks, and other influential
| industries. Don't be naive about "the systems at work and
| their interplay".
| trollbridge wrote:
| You can hire a non-union plumber. There isn't usually
| much of a price difference. Where I live, you can easily
| find a non-licenced plumber (called "moonlighting",
| usually done by apprentices of licenced plumbers). A lot
| of people prefer not to since you're on your own if
| something goes wrong.
|
| Plumbing requires skill, particularly for difficult jobs,
| and also requires advanced equipment to do such a job in
| a reasonable amount of time, such as special cameras to
| inspect a septic tank or drain line without having to
| actually cut into it.
| realreality wrote:
| Where I live, permits are only given to licensed
| plumbers, and all work on plumbing requires a permit
| (though I'm sure many people ignore the rule).
| btown wrote:
| It's said that much of research is data janitorial work, and
| from my experience that's not just limited to the machine
| learning space. Every research scientist wishes that they had
| an army of engineers to build bespoke tooling for their
| niche, so they could get back to trying ideas at the speed of
| thought rather than needing to spend a day writing utility
| functions for those tools and poring over tables to spot
| anomalies. Giving every researcher a priceless level of
| leverage is a tremendous social good.
|
| Of course, we won't be able to tell the real effects, now,
| because every longitudinal study of researchers will now be
| corrupted by the ongoing evisceration of academic research in
| the current environment. Vibe-coding won't be a net
| creativity gain to a researcher affected by vibe-immigration-
| policy, vibe-grant-availability, and vibe-firings, for all of
| which the unpredictability is a punitive design goal.
|
| Whether fear of LLMs taking jobs has contributed to a larger
| culture of fear and tribalism that has emboldened anti-
| intellectual movements worldwide, and what the attributable
| net effect on research and development will be... it's
| incredibly hard to quantify.
| stego-tech wrote:
| > Vibe-coding won't be a net creativity gain to a
| researcher affected by vibe-immigration-policy, vibe-grant-
| availability, and vibe-firings, for all of which the
| unpredictability is a punitive design goal.
|
| Quite literally this is what I'm trying to get at with my
| resistance to LLM adoption in the current environment.
| We're not using it to do hard work, we're throwing it
| everywhere in an intentional decision to dumb down more
| people and funnel resources and control into fewer hands.
|
| Current AI isn't democratizing _anything_ , it's just a
| shinier marketing ploy to get people to abandon skilled
| professions and leave the bulk of the populace only
| suitable for McJobs. The benefits of its use are seen by
| vanishingly few, while its harms felt by distressingly
| many.
|
| At present, it is a tool designed to improve existing
| neoliberal policies and wealth pumps by reducing the demand
| for skilled labor without properly compensating those
| affected by its use, nor allowing an exit from their walled
| gardens (because that is literally what all these XaaS AI
| firms are - walled gardens of pattern matchers masquerading
| as intelligence).
| prmph wrote:
| This is one of the best comments about the current AI
| hype.
|
| The elite really don't see why the proletariat should be
| interested in, or enjoy the dignity of, actual skill and
| quality.
|
| Hence the enshitification of everything, and now AI
| promises to commoditize everything into slop.
|
| Sad because it is the very deoth of society that has
| birthe
| Invictus0 wrote:
| John Rockefeller didn't sit down in a big chair, twirl
| his mustache, and invent AI to funnel money to the hands
| of the wealthy. This technology was created by
| researchers and has been mostly accessible to everyone
| for as long as it has been around.
|
| All technology has the effect of concentrating wealth,
| and anyone who insists on using their two hands to
| fashion things when machines exist that can do it better
| will always be relegated to the "artisan" bin as time
| rolls on.
| bethekidyouwant wrote:
| This could be a comment about the industrial revolution.
| cpursley wrote:
| That's one perspective, but it's wrong and typical
| gatekeeping (do you have a software degree by any
| chance?). People had the same attitude towards open
| source tooling and low code frameworks - god forbid
| someone not certified and ordained build a solution in
| something other than Java...
|
| AI code tools are allowing people to build things they
| couldn't before due to lack of skillset, time or budget.
| I've seen all sorts of problems solved by semi technical
| and even non-technical people. My brother for example
| built a thing with Microsoft copilot that helped automate
| more in his manufacturing facility (used to be paper).
|
| But yeah, keep yelling at that cloud - the rest of us
| will keep shipping cool things that we couldn't before,
| and faster.
| Workaccount2 wrote:
| >My brother for example built a thing with Microsoft
| copilot that helped automate more in his manufacturing
| facility (used to be paper).
|
| I have harped on this endlessly as a non-programmer
| working a non-tech job, with 7 "vibe-coded" programs now
| being used daily by people at my company.
|
| I am sorry, but the tech world is _completely_ missing
| the forest for the trees here. LLM 's are talked about
| purely as tools that were created to help devs. Some love
| them, some hate them, but pretty much all of them seem
| unaware that LLMs allow non-tech people to automate tasks
| with a computer _without_ having to go through a 3rd-
| party-created interface.
|
| So yea, maybe Claude is useless troubleshooting your
| cloud platform. But it certainly isn't useless in helping
| me forgo a cloud platform by setting up a simple local
| database to use instead.
| cpursley wrote:
| Yep, and it allows them to build POCs that they can pass
| to "real" devs in a way that was not possible before.
| Workaccount2 wrote:
| Real devs excel at writing software for hundreds,
| thousands, millions of users with fractal use cases and
| feature needs.
|
| LLMs excel at writing software for one or a handful of
| users with a very narrow but very well defined use cases.
|
| I don't need an LLM to write Excel.exe for keeping track
| of 20 employee's hours. A simple GUI on a SQLite database
| can easily do that.
| kentonv wrote:
| Yes!
|
| We're about to enter a world where everyone has their own
| custom software for their specific use cases. Each of
| these is relatively simple, yet they may replace
| something complex. Excel is complex because it needs to
| handle everyone's use cases, but for any one particular
| spreadsheet, you could pretty easily vibe-code a
| replacement that does that one spreadsheet's job better
| than Excel can.
|
| I've also found that vibe-coding a presentation as a
| React app is better than using Power Point.
| sitharus wrote:
| The problem isn't that people can quickly prototype an
| idea that they've had without contracting an expensive
| professional, I think this is great. This will give ideas
| that would never see the light of day a chance. Plus this
| gives a much better talking point if they do choose to
| get a professional onboard.
|
| The problem is that it's sold as a complete solution. Use
| the LLM and you'll get a fully working product. However
| if you're not an experienced programmer you won't know
| what's missing, if it's using outdated and insecure
| options, or is just badly written. This still needs a
| professional.
|
| The technology is great and it has real potential to
| change how things are made, but it's being marketed as
| something it isn't (yet).
| kentonv wrote:
| > However if you're not an experienced programmer you
| won't know what's missing, if it's using outdated and
| insecure options, or is just badly written. This still
| needs a professional.
|
| I think a lot of this could be solved by a platform that
| implements appropriate guardrails so that the application
| code literally cannot screw up the security. Not every
| conceivable type of software would fit in such a
| platform, but a lot of what people want to do to automate
| their day-to-day lives could.
| immibis wrote:
| You know the industry that will need a lot more
| professionals after this - cybersecurity.
| btown wrote:
| This is a bit stronger than my point, I should say. I do
| think that LLMs would have a net benefit to society, by
| way of their effects on research and innovation... if we
| could get our political houses in order such that we
| weren't negating those effects, and such that we were
| empowering small businesses and high-tech startups to
| build with the results of this innovation sustainably.
|
| And in a world where policy is horrid and the effects are
| mainly negated, things would be even worse if the
| remaining researchers lost AI as a tool. For better or
| for worse, fire has been shared with humanity, and we
| might as well cook.
| manquer wrote:
| > much of research is data janitorial work
|
| In applied research perhaps, Fundamental research is
| nothing like that in any field including ML.
| freehorse wrote:
| All experimental or empirical research is like that, is
| closer to the point.
| QuantumGood wrote:
| There's too much "close enough" in virtually all these
| discussions. LLM is not a hand grenade. It's important to
| keep in mind what LLMs and related tech can be relied upon
| do or assist with, can't be relied upon to do or assist
| with, and might be relied on to do in the future.
| diggan wrote:
| > On the other hand, where I remain a skeptic is this
| constant banging-on that somehow this will translate into
| entirely new things - research, materials science, economies,
| inventions, etc
|
| Does it even have to be able to do so? Just the ability to
| speed up exploration and validation based on what a human
| tells it to do is already enormously useful, depending on how
| much you can speed up those things, and how accurate it can
| be.
|
| Too slow or too inaccurate and it'll have a strong slowdown
| factor. But once some threshold been reached, where it makes
| either of those things faster, I'd probably consider the
| whole thing "overall useful". Nut of course that isn't the
| full picture and ignoring all the tradeoffs is kind of
| cheating, there are more things to consider too as you
| mention.
|
| I'm guessing we aren't quite over the threshold because it is
| still very young all things considered, although the
| ecosystem is already pretty big. I feel like generally things
| tend to grow beyond their usefulness initially, and we're at
| that stage right now, and people are shooting it all kind of
| directions to see what works or not.
| dingnuts wrote:
| > Just the ability to speed up exploration and validation
| based on what a human tells it to do is already enormously
| useful, depending on how much you can speed up those
| things, and how accurate it can be.
|
| The big question is: is it useful enough to justify the
| cost when the VC subsidies go away?
|
| My phone recently offered me Gemini "now for free" and I
| thought "free for now, you mean. I better not get used to
| that. They should be required to call it a free trial."
| diggan wrote:
| > The big question is: is it useful enough to justify the
| cost when the VC subsidies go away?
|
| I won't claim local LLMs as nearly as good as various top
| models behind paid subscriptions/APIs, but I'm certain
| I'd be able find a way (for me) of working with them well
| enough, if the entire paid/hosted ecosystem disappeared
| over night. Even with models released today.
|
| I think the VC subsidies probably "make stuff happen"
| faster, and without it we'd see slower progress, but I
| don't think 100% of the ecosystem would disappear even if
| 100% of VC funding disappeared. We're bound for another
| AI winter at one point, and some will surely survive even
| that :)
| jsnell wrote:
| Inference is actually quite cheap. Like, a highly
| competitive LLM can cost 1/25th of a search query. And it
| is not due to inference being subsidized by VC money.
|
| It's also getting cheaper all the time. Something like
| 1000x cheaper in the last two years at the same quality
| level, and there's not yet any sign of a plateau.
|
| So it'd be quite surprising if the only long-term
| business model turned out to be subscriptions.
| Denzel wrote:
| Can you link to any sources that support your claim?
| jsnell wrote:
| Sure. Here's something I'd written on the subject that
| I'd left lying in my drafts folder for a month, but I've
| now published just for you :)
|
| https://www.snellman.net/blog/archive/2025-06-02-llms-
| are-ch...
|
| It has links to public sources on the pricing of both
| LLMs and search, and explains why the low inference
| prices can't be due the inference being subsidized. (And
| while there are other possible explanations, it includes
| a calculator for what the compound impact of all of those
| possible explanations could be.)
| whilenot-dev wrote:
| Just had a quick glance, but I think I found something to
| add to the _Objection!_ -section of your post:
|
| Brave's Search API is 3$ CPM and includes Web search,
| Images, Videos, News, Goggles[0]. Anthropic's API is 10$
| CPM for Web search (and text only?), excluding any
| input/output tokens from your model of choice[1], that'd
| be an additional 15$ CPM, assuming 1KTok per request and
| Claude Sonnet 4 as a good model, so ~25$ CPM.
|
| So your default "Ratio (Search cost / LLM cost): 25.0x"
| seems to be more on the 0.12x side of things (Search cost
| / LLM cost). Mind you, I just flew over everything in 10
| mins and have no experience using either API.
|
| [0]: https://brave.com/search/api/
|
| [1]: https://www.anthropic.com/pricing#anthropic-api
| Denzel wrote:
| Thanks for sharing!
|
| It's worthwhile to note that https://github.com/deepseek-
| ai/open-infra-index/blob/main/20... shows cost vs.
| _theoretical income_. They don 't show 80% gross margins
| and there's probably a reason they don't share their
| actual gross margin.
|
| OpenAI is the easiest counterexample that proves
| inference is subsidized right now. They've taken $50B in
| investment; surpassed 400M WAUs
| (https://www.reuters.com/technology/artificial-
| intelligence/o...); lost $5B on $4B in revenue for 2024
| (https://finance.yahoo.com/news/openai-thinks-revenue-
| more-tr...); and project they won't be cash-flow positive
| until 2029.
|
| Prices would be significantly higher if OpenAI was priced
| for unit profitability right now.
|
| As for the mega-conglomerates (Google, Meta, Microsoft),
| GenAI is a loss leader to build platform power. GenAI
| doesn't need to be unit profitable, it just needs to
| attract and retain people on their platform, ie you need
| a Google Cloud account to use Gemini API.
| jsnell wrote:
| Thanks,
|
| I believe the API prices are not subsidized, and there's
| an entire section devoted to that. To recap:
|
| 1) pure compute providers (rather than companies
| providing both the model and the compute) can't really
| gain anything from subsidizing. That market is already
| commoditized and supply-limited.
|
| 2) there is no value to gaining paid API market share --
| the market share isn't sticky, and there's no benefit to
| just getting more usage since the terms of service for
| all the serious providers promise that the data won't be
| used for training.
|
| 3) we have data from a frontier lab on what the economics
| of their paid API inference are (but not the economics of
| other types of usage)
|
| So the API prices set a ceiling on what the actual cost
| of inference can be. And that ceiling is very low
| relative to the prices of a comparable (but not
| identical) non-AI product category.
|
| That's a very distinct case from free APIs and consumer
| products. The former is being given out for no cost in
| exchange for data, the latter for data and sticky market
| share. So unlike paid APIs, the incentives are there.
|
| But given the cost structure of paid APIs, we can tell
| that it would be trivial for the consumer products to be
| profitably monetized with ads. They've got a ton of
| users, and the way users interact with their main product
| would be almost perfect for advertising.
|
| The reason OpenAI is not making a profit isn't that
| inference is expensive. It's that they're choosing not to
| monetize like 95% of their users, despite the unit
| economics being very lucrative in principle. They're
| making a loss because for now they can, and for now the
| only goal of their consumer business is to maximize their
| growth and consumer mindshare.
|
| If OpenAI needed to make a profit, they would not raise
| their prices on things being paid for. They'd just need
| to extract a very modest revenue from their unpaid users.
| (It's 500M unpaid users. To make $5B/year in revenue from
| them, you'd need just a $1 ARPU. That's an order of
| magnitude below what's realistic. Hell, that's lower than
| the famously hard to monetize _Reddit 's_ global ARPU.)
| butlike wrote:
| So isn't the heuristic that if your job is easily
| digestible by an LLM, you're probably replaceable, but if
| the strong slowdown factor presents itself, you're probably
| doing novel work and have job security?
| lincoln20xx wrote:
| I have a non-zero number of industrial process patents under
| my belt. Allegedly, that means that I had ideas that had not
| previously been recorded. Once I wrote them down, paid some
| lawyers a bunch of money, and did some paperwork, I have the
| right to pay lawyers more money to make someone's life
| difficult if I think that someone ever tries to do something
| with the same thoughts, regardless of if they had those
| thoughts before, after, or independently of me.
|
| In my opinion, there is a very valid argument that the vast
| majority of things that are patented are not "new" things,
| because everything builds on something else that came before
| it.
|
| The things that are seen as "new" are not infrequently
| something where someone in field A sees something in field B,
| ponders it for a minute, and goes "hey, if we take that idea
| from field B, twist it clockwise a bit, and bolt it onto the
| other thing we already use, it would make our lives easier
| over in this nasty corner of field A." Congratulations! "New"
| idea, and the patent lawyers and finance wonks rejoice.
|
| LLMs may not be able to truly "invent" "new" things,
| depending on where you place those particular goalposts.
|
| However, even a year or two ago - well before Deep Research
| et al - they could be shockingly useful for drawing
| connections between disparate fields and applications. I was
| working through a "try to sort out the design space of a
| chemical process" type exercise, and decided to ask whichever
| GPT was available and free at the time about analogous
| applications and processes in various industries.
|
| After a bit of prodding it made some suggestions that I
| definitely could have come up on my own if I had the
| requisite domain knowledge, but would almost certainly never
| have managed on my own. It also caused me to make a
| connection between a few things that I don't think I would
| have stumbled upon otherwise.
|
| I checked with my chemist friends, and they said the
| resulting ideas were worth testing. After much iteration, one
| of the suggested compounds/approaches ended up generating the
| least bad result from that set of experiments.
|
| I've previously sketched out a framework for using these
| tools (combined with other similar machine
| learning/AI/simulation tools) to massively improve the energy
| consumption of industrial chemical processes. It seems to me
| that that type of application is one where the LLM's
| environmental cost could be very much offset by the advances
| it provides.
|
| The social cost is a completely different question though,
| and I think a very valid one. I also don't think our economic
| system is structured in such a way that the social costs will
| ever be mitigated.
|
| Where am I going with this? I'm not sure.
|
| Is there a "ghost in the machine"? I wouldn't place a bet on
| yes, at least not today. But I think that there is a fair bit
| of something there. Utility, if nothing else. They seem like
| a force multiplier to me, and I think that with proper
| guidance, that force multiplier could be applied to basic
| research, material science, economics, and "inventions".
|
| Right now, it does seem that it takes someone with a lot of
| knowledge about the specific area, process, or task to get
| really good results out of LLMs.
|
| Will that always be true? I don't know. I think there's at
| least one piece of the puzzle we don't have sorted out yet,
| and that the utility of the existing models/architectures
| will ride the s-curve up a bit longer but ultimately flatten
| out.
|
| I'm also wrong a LOT, so I wouldn't bet a shiny nickel on
| that.
| svara wrote:
| > On the other hand, where I remain a skeptic is this
| constant banging-on that somehow this will translate into
| entirely new things
|
| Really a lot of innovation, even at the very cutting edge, is
| about combining old things in new ways, and these are great
| productivity tools for this.
|
| I've been "vibe coding" quite a bit recently, and it's been
| going great. I still end up reading all the code and fixing
| issues by hand occasionally, but it does remove a lot of the
| grunt work of looking up simple things and typing out obvious
| code.
|
| It helps me spend more time designing and thinking about how
| things should work.
|
| It's easily a 2-3x productivity boost versus the old
| fashioned way of doing things, possibly more when you take
| into account that I also end up implementing extra bells and
| whistles that I would otherwise have been too lazy to add,
| but that come almost for free with LLMs.
|
| I don't think the stereotype of vibe coding, that is of
| coding without understanding what's going on, actually works
| though. I've seen the tools get stuck on issues they don't
| seem to be able to understand fully too often to believe
| that.
|
| I'm not worried at all that LLMs are going to take software
| engineering jobs soon. They're really just making engineers
| more powerful, maybe like going from low level languages to
| high level compiled ones. I don't think anyone was worried
| about the efficiency gains from that destroying jobs either.
|
| There's still a lot of domain knowledge that goes into using
| LLMs for coding effectively. I have some stories on this too
| but that'll be for another day...
| abalone wrote:
| I like to make a rough analogy with autonomous vehicles.
| There's a leveling system from 1 (old school cruise control)
| to 5 (full automation):
|
| * We achieved Level 2 autonomy first, which requires you to
| fully supervise and retain control of the vehicle and expect
| mistakes at any moment. So kind of neat but also can get you
| in big trouble if you don't supervise properly. Some people
| like it, some people don't see it as a net gain given the
| oversight required.
|
| ^ This is where Tesla "FSD beta" is at, and probably where
| LLM codegen tools are at today.
|
| * After many years we have achieved a degree of Level 4
| autonomy on well-trained routes albeit with occasional human
| intervention. This is where Waymo is at in certain cities.
| Level 4 means autonomy within specific but broad
| circumstances like a given area and weather conditions. While
| it is still somewhat early days it looks like we can
| generally trust these to operate safely and ask for help when
| they are not confident. Humans are not out of the loop.[1]
|
| ^ This is probably what where we can expect codegen to grow
| after many more years of training and refinement in specific
| domains. I.e. a lot of what CloudFlare engineers did with
| their prompt engineering tweaking was of this nature. Think
| of them as the employees driving the training vehicles around
| San Francisco for the past decade. And similarly, "L4
| codegen" needs to prioritize code safety which in part means
| ensuring humans can understand situations and step in to
| guide and debug when the tool gets stuck.
|
| * We are still nowhere close to Level 5 "drive anywhere and
| under any conditions a human can." And IMHO it's not clear we
| ever will based purely on the technology and methods that got
| us to L4. There are other brain mechanisms at work that need
| to be modeled.
|
| [1] https://www.cnbc.com/2023/11/06/cruise-confirms-
| robotaxis-re...
| pphysch wrote:
| That's a good analogy. OAuth libraries and integrations are
| like a highly-mapped California city. Just because you can
| drive a Waymo or coding agent there, doesn't mean you can
| drive it through the Rockies.
| abalone wrote:
| Note that even with OAuth it took, as of today, security
| engineers many iterations of review and prompt tweaking
| to get this result. We're still in the "mapping" phase.
| jauntywundrkind wrote:
| > _Again, please check out the commit history -- especially
| early commits -- to understand how this went._
|
| Direct link to earliest page of history:
| https://github.com/cloudflare/workers-oauth-provider/commits...
|
| A lot of very explicit & clear prompting, with direct
| directions to go. Some examples on the first page:
| https://github.com/cloudflare/workers-oauth-provider/commit/...
| https://github.com/cloudflare/workers-oauth-provider/commit/...
| teaearlgraycold wrote:
| > Every line was thoroughly reviewed and cross-referenced with
| relevant RFCs, by security experts with previous experience
| with those RFCs.
|
| This sounds like coding but slower
| TeMPOraL wrote:
| The point was validating a hypothesis. That is the validation
| part.
| kentonv wrote:
| I would say it ended up being much faster than had I written
| it by hand. It took a few days to produce this library -- it
| would almost certainly have taken me weeks to write it
| myself.
| noodletheworld wrote:
| If you had written it by hand would the verification
| process been as time consuming?
|
| i.e. overall _including_ the time spent verifying that it
| was correct, do you consider it a net win?
| kentonv wrote:
| I was already including my own time spent verifying the
| output, which I mostly did right away as the code was
| being generated (approving or rejecting each edit).
|
| And the separate security review would have been required
| either way.
|
| So yes, it saved time.
| apwell23 wrote:
| does reviewing have the same fidelity as writing ?
|
| reminded me of my university classes where i took my own
| notes vs studied someone else's notes. you can guess which
| one was superior.
| kentonv wrote:
| It's certainly true that my own recall of the code would
| be better if I had written it by hand.
|
| But I don't think the final code is all that far off from
| what I would have written.
| apwell23 wrote:
| yes there are plenty of examples of ppl writing tic-tac-toe or
| a flying simulator with llm all over youtube. what does that
| prove exactly? oauth is as routine as it gets.
| weinzierl wrote:
| _" I thoughts LLMs were glorified Markov chain generators"_
|
| _" the code actually looked pretty good. Not perfect, but I
| just told the AI to fix things, and it did. I was shocked."_
|
| These two views are by no means mutually exclusive. I find LLMs
| extremely useful and still believe they are glorified Markov
| generators.
|
| The take away should be that that is all you need and humans
| likely are nothing more than that.
| smallnix wrote:
| > humans likely are nothing more than that
|
| Relevant post: https://news.ycombinator.com/item?id=44089156
| Flemlo wrote:
| The way the input doesn't match the output should imply that
| it's not just statistics.
|
| As soon as compression happens, optimization happens which
| can lead to rules/learning of principles which got feed by
| statistics.
| immibis wrote:
| That's "just" more statistics though.
| kentonv wrote:
| I suppose it's all a continuum and we can each have different
| opinions on what the threshold for "glorified markov
| generator" is.
|
| But there have been many cases in my experience where the LLM
| could not possibly have been simply pattern-matching to
| something it had seen before. It really did "understand" the
| meaning of the code by any definition that makes sense to me.
| palata wrote:
| > It really did "understand" the meaning of the code by any
| definition that makes sense to me.
|
| I find it dangerous to say it "understands". People are
| fast to say it "is sentient by any definition that makes
| sense to them".
|
| Also, would we say that a compiler "understands" the
| meaning of the code?
| bufferoverflow wrote:
| > _I find LLMs extremely useful and still believe they are
| glorified Markov generators._
|
| Then you should be able to make a markov chain generator
| without deep neural nets, and it should be on the same level
| of performance as current LLMs.
|
| But we both know you can't.
| ronsor wrote:
| You can, but it will require far more memory than any
| computer has.
| varispeed wrote:
| The thing is you need to know what exactly LLM should create
| and you need to know what it is doing wrong and tell it to fix
| it. Meaning, if you don't already have skill to build something
| yourself, AI might not be as useful. Think of it as keyboard on
| steroids. Instead of typing literally what you want to see, you
| just describe it in detail and LLM decompresses that thought.
| bsder wrote:
| > Claude's output was thoroughly reviewed by Cloudflare
| engineers with careful attention paid to security and
| compliance with standards.
|
| So, for those of us who are not OAuth experts, don't have a
| team of security engineers on call, and are likely to fall into
| _all_ the security and compliance traps, how does this help?
|
| I don't need AI to write my shitty code. I need AI to _review
| and correct_ my shitty code.
| pier25 wrote:
| Did you really save time given that every line of code was
| "thoroughly reviewed"?
| caycep wrote:
| tbh I would find it annoying to have to go audit someone else
| (i.e. an LLM's) code...
|
| Also, maybe the humbling question is, maybe we humans aren't so
| exceptional if 90% of the sum of human knowledge can be
| predicted by next-word-prediction
| mmaunder wrote:
| Claude 4 in agent mode is incredible. Nothing compares. But you
| need to have a deep technical understanding of what you're
| building and how to split it into achievable milestones and
| build on each one. It also helps to provide it with URLs with
| specs, standards, protocols, RFCs etc that are applicable and
| then tell it what to use from the docs.
| csmpltn wrote:
| There are tens (if not hundreds) of thousands of OAuth
| libraries out there. Probably millions of relevant codebases on
| GitHub, Bitbucket, etc. Possibly millions of questions on
| StackOverflow, Reddit, Quora. Vast amounts of documentation
| across many products and websites. RFCs. All kinds of forums.
| Wikipedias...
|
| Why are you so surprised an LLM could regurgitate one back? I
| wouldn't celebrate this example as a noteworthy achievement...
| blibble wrote:
| > I thoughts LLMs were glorified Markov chain generators that
| didn't actually understand code and couldn't produce anything
| novel.
|
| so he's been convinced by it shitting out yet another
| javascript oauth library?
|
| this experiment proves nothing re: novelty
| ayuhito wrote:
| Good thing most of my tasks don't require novelty, just
| working code.
| chrisweekly wrote:
| mods: typo in title "CloudLflare"
| mdaniel wrote:
| There is no "@" system here, you are welcome to email
| hn@ycombinator.com or hope that we're still within the edit
| window for the title
| abroadwin wrote:
| Oh hey, looks like it's mostly Kenton Varda, who you may
| recognize from his LAN party house:
| https://news.ycombinator.com/item?id=42156977
| davidjfelix wrote:
| Or Cap'n'Proto, Protobuf, Cloudflare workers, Cloudflare
| Durable Objects. The LAN house is cool too.
| paxys wrote:
| This is exactly the direction I expect AI-assisted coding to go
| in. Not software engineers being kicked out and some business
| person pressing a few buttons to have a fully functional app (as
| is playing out in a lot of fantasies on LinkedIn & X), but rather
| experienced engineers using AI to generate bits of code and then
| meticulously reviewing and testing them.
|
| The million dollar (perhaps literally) question is - could
| @kentonv have written this library quicker by himself without any
| AI help?
| dkdcio wrote:
| > The million dollar (perhaps literally) question is - could
| @kentonv have written this library quicker by himself without
| any AI help?
|
| I *think* the answer to this is clearly no: or at least, given
| what we can accomplish today with the tools we have now, and
| that we are still collectively learning how to effectively use
| this, there's no way it won't be faster (with effective use) in
| another 3-6 months to fully-code new solutions with AI. I think
| it requires a lot of work: well-documented, well-structured
| codebases with fast built-in feedback loops (good linting/unit
| tests etc.), but we're heading there no
| bigstrat2003 wrote:
| > but rather experienced engineers using AI to generate bits of
| code and then meticulously testing and reviewing them.
|
| My problem is that (in my experience anyways) this is _slower_
| than me just writing the code myself. That 's why AI is not a
| useful tool right now. They only get it right sometimes so it
| winds up being easier to just do it yourself in the first
| place. As the saying goes: bad help is worse than no help at
| all, and AI is bad help right now.
| uludag wrote:
| I feel this is on point. So not only is there the time lost
| correcting and testing AI generated code, but there's also
| the mental model you build of the code when you write it
| yourself.
|
| Assuming you want a strong mental model of what the code does
| and how it works (which you'd use in conversations with
| stakeholders and architecture discussions for example),
| writing the code manually, with perhaps minor completion-like
| AI assistance, may be the optimal approach.
| JimDabell wrote:
| > My problem is that (in my experience anyways) this is
| _slower_ than me just writing the code myself.
|
| How much experience do you have writing code vs how much
| experience do you have prompting using AI though? You have to
| factor in that these tools are new and everybody is still
| figuring out how to use them effectively.
| belter wrote:
| The million-dollar question is not whether you can review at
| the speed the model is coding. It is whether you can trust
| review alone to catch everything.
|
| If a robot assembles cars at lightning speed... but
| occasionally misaligns a bolt, and your only safeguard is a
| visual inspection afterward, some defects will roll off the
| assembly line. Human coders prevent many bugs by thinking
| during assembly.
| chrisweekly wrote:
| THIS.
|
| IMHO more rigorous test automation (including fuzzing and
| related techniques) is needed. Actually that holds whether AI
| is involved or not, but probably more so if it is.
| pton_xd wrote:
| > Human coders prevent many bugs by thinking during assembly.
|
| I'm far from an AI true believer but come on -- human coders
| write bugs, tons and tons of bugs. According to Peopleware,
| software has "an average defect density of one to three
| defects per hundred lines of code"!
| gokhan wrote:
| > Not software engineers being kicked out ... but rather
| experienced engineers using AI to generate bits of code and
| then meticulously reviewing and testing them.
|
| But what if you only need 2 kentonv's instead of 20 at the end?
| Do you assume we'll find enough new tasks that will occupy the
| other 18? I think that's the question.
|
| And the author is implementing a fairly technical project in
| this case. How about routine LoB app development?
| paxys wrote:
| Increased productivity means increased opportuntity. There
| isn't going to be a time (at least not anytime soon) when we
| can all sit back and say "yup, we have accomplished
| everything there is to do with software and don't need more
| engineers".
| spiderice wrote:
| But there very well might be a time very soon where human's
| no longer offer economic value to the software engineering
| process. If you could (and currently you can't) pay an AI
| $10k/year to do what a human could do in a year, why would
| you pay the human 6 figures? Or even $20k?
|
| Nobody is claiming that human's won't have jobs simply
| because "we have accomplished everything this is to do".
| It's that humans will offer zero economic value compared to
| AI because AI gets so good and so cheap.
| paxys wrote:
| And there might be a giant asteroid that strikes the
| earth a few years down the line ending human
| civilization.
|
| If there is some magic $10k AI that can fully replace a
| $200k software engineer then I'd love to see it. Until
| that happens this entire discussion is science fiction.
| spiderice wrote:
| If experts were saying the astroid will hit earth in the
| next 5 years, would it still be science fiction?
|
| You acting like those two scenarios are the same is
| disingenuous. Fuck that.
| paxys wrote:
| Remove all the "experts" who have a major conflict of
| interest (running AI startups, selling AI courses,
| wanting to pump their company's stock price by
| associating with AI) and you'll find that very few actual
| experts in the field hold this view.
| TeMPOraL wrote:
| Yup, because it's a stupid view. Good enough AI is right
| here, right now, today; it's already impacting day-to-day
| work in the software industry. That one is blindingly
| obvious to anyone who actually bothers to look around.
| You don't need experts to tell you the water is wet. It
| takes something special to try and deny this.
|
| It may not manifest as job loss _yet_ , but the market
| response to changes is a whole other thing. For one, it's
| likely to first manifest as slowing down hiring relative
| to amount of projects being started and then released.
| Software is a growing market after all.
| lukeschlather wrote:
| Experts understand orbital mechanics pretty well. If
| experts say an asteroid in the next 5 years it's pretty
| similar to saying that a rock dropped from the top of a
| skyscraper will hit the ground. It happens billions of
| times every day, we know the cause and effect.
|
| With AI, there's no real expertise involved in saying
| "well, it was very stupid 5 years ago, now it's starting
| to seem smart, if we extrapolate it's going to be smarter
| than me in 5 years." But no one really knows what level
| of effort is required to make it smarter than me. No one
| is an expert in something that doesn't exist yet.
| TeMPOraL wrote:
| It's not. Consider that replacing the _only_ $200k
| software engineer on the project is different than
| replacing the third or tenth $200k software engineer on
| the project. To the extent AI is improving productivity
| of those engineers, it reduces the need for adding more
| engineers to that team. That may mean firing some of
| them, or just not hiring new ones (or fewer of them) as
| the project expands, as existing ones + AI can keep up
| with increased workload.
| nand_gate wrote:
| I'm biased but my money's on the end result of AI being
| fewer engineers per team but also teams as a concept
| becoming obsolete.
|
| Why keep legacy structures, with luxuries like POs or PMs
| if AI becomes powerful as you say - it'll just be 'one
| man startups' for better or worse.
|
| Any empire-building VP should probably fear the wishful
| AI future they're praying for!
| alastairr wrote:
| You don't need to completely replace a whole 200k
| engineer. You just need to increase each engineer's
| productivity sufficiently that you can reduce the total
| number of engineers in your company.
| hooverd wrote:
| You run into knowledge collapse because nobody is
| socially reproducing that knowledge.
| amanaplanacanal wrote:
| This seems an important thing that _somebody_ should be
| concerned about. How do we get the next generation of
| engineers? And how will they even be able to do the
| senior engineer work of validating the LLM output if they
| haven 't had the years of experience writing code
| themselves?
| lanthissa wrote:
| it doesn't even have to be that. software engineer used
| to be a medium pay job, theres no law of the universe
| that says it cant go back to that.
| nand_gate wrote:
| In this scenario who would be buying this product that
| offers 'zero economic value compared to AI because AI
| gets so good and so cheap'.
| thewebguyd wrote:
| > But what if you only need 2 kentonv's instead of 20 at the
| end? Do you assume we'll find enough new tasks that will
| occupy the other 18? I think that's the question.
|
| This is likely where all this will end up. I have doubts that
| AI will replace all engineers, but I have no doubt in my mind
| that we'll certainly need a lot less engineers.
|
| A not so dissimilar thing happened in the sysadmin world (my
| career) when everything transitioned from ClickOps to the
| cloud & Infrastructure as Code. Infrastructure that needed 10
| sysadmins to manage now only needed 1 or 2 infrastructure
| folks.
|
| The role still exists, but the quantity needed is drastically
| reduced. The work that I do now by myself would have needed
| an entire team before AWS/Ansible/Terraform, etc.
| mikeocool wrote:
| Though arguably cloud infra made it so that a lot more
| companies who never would have built out a data center or
| leased a chunk of space in one were spinning up some
| serious infra in AWS or Azure -- and thus hiring at least
| 1-2 devops engineers.
|
| Before the end of zero interest rate policy, all the
| sysadmins I knew who the made the transition to devops were
| never stuck looking for a job for long.
| achierius wrote:
| To be clear, the number of people employed as "SREs" or
| "production engineers" is actually far, far higher (at
| least an order of magnitude) than in the days before cloud
| became a thing. There are simply far more apps / companies
| / businesses / etc. who use cloud hosting than there ever
| were doing on-prem work.
| kentonv wrote:
| I think there's a huge huge space of software to build that
| isn't being touched today because it's not cost-effective
| to have an engineer build them.
|
| But if the time it takes an engineer to build any one thing
| goes down, now there are a lot more things that are cost
| effective.
|
| Consider niche use cases. Every company tends to have
| custom processes and workflows. Think about being an
| accountant at one company vs. another -- while a lot of the
| job is the same, there will always be parts that are
| significantly different. Those bespoke processes often
| involve manual labor because off-the-shelf accounting
| software cannot add custom features for every company.
|
| But what if it could? What if an engineer working with AI
| could knock out customer-specific features 10x as fast as
| they could in the past. Now it actually makes sense to
| build those features, to improve the productivity of each
| company's accounting department.
|
| It's hard to say if demand for engineers will go down or
| up. I'm not pretending to know for sure. But I can see a
| possibility that we actually have way more developers in
| coming years!
| int_19h wrote:
| It's interesting that you bring up accounting software as
| an example. In jurisdictions where legal requirements
| around it are a lot more specific than in e.g. US,
| accounting suites usually already come with a lot of
| customization hooks (up to and including full-fledged
| scripting DSLs), and there are software engineers and
| companies who specialize in using those to implement
| bespoke accounting requirements.
| kentonv wrote:
| I admit I have no specific knowledge of accounting and
| just meant to reference any random department that isn't
| engineering.
|
| (Though I think it's true of engineering too. We all have
| our own weird team-specific processes for code reviews
| and CI and deployments which could probably use better
| automation.)
|
| But even where lots of customization exists today (such
| as in engineering!), more is always possible. It's always
| just a question of whether the automation saves as much
| time as it took to build. If the automations can be built
| faster, then it makes sense to build more of them.
| hn_acc1 wrote:
| After 30+ years in the software field, and a user for
| 40+, having at times heavily customized my desktop or
| editor, for example - I've concluded that the best thing
| for most apps is for me to learn to use them with stock
| settings.
|
| Why? Inevitably, I changed positions / jobs / platforms,
| and all that effort was lost / inapplicable, and I had to
| relearn to use the stock settings anyway.
|
| Now, I understand that some companies have different
| setups, but it might just make more sense to change the
| company's accounting procedures (if possible) to conform
| to most accounting software defaults, rather than invest
| heavily in modifying the setup, unless you're a huge
| conglomerate and can keep people on staff. Why? Because
| someone, somewhere will have to maintain those changes.
| Sure, you can then hire someone else to update those
| changes - but guess what? Most likely, unless they open-
| source their changes, no LLM will have seen those
| changes, and even if they are allowed to fine-tune on it,
| they'll have seen exactly ONE instance of these changes.
| Odds they'll get everything right, AND the person using
| the LLM will recognize when it doesn't go right? Oh
| right, they invested in hundreds of unit tests to ensure
| everything works as expected even with changes, and I'm
| the tooth fairy..
| kentonv wrote:
| There are good arguments to just conform. But it is in
| fact true nevertheless that many companies and teams
| continue to choose bespoke workflows over standardized
| ones. So I guess there must be something driving that.
|
| I don't actually think this is going to take the form of
| LLMs implementing custom patches to off-the-shelf
| software. I think instead it's going to look like LLMs
| writing code that uses APIs offered by off-the-shelf
| software to script specific workflows.
| thewebguyd wrote:
| > I think there's a huge huge space of software to build
| that isn't being touched today because it's not cost-
| effective to have an engineer build them.
|
| That's definitely an interesting area, but I think we'll
| actually see (maybe) individual employees solving some of
| these problems on their own without involving IT/the dev
| team.
|
| We kind of see it already - a lot of these problem spaces
| are being solved with complex Excel workflows, crappy
| Access databases, etc. because the team needed their
| problem solved now, and resources couldn't be given to
| them.
|
| Maybe AI is the answer to that so that instead of
| building a house of cards on Excel, these non-tech teams
| can have something a little more robust.
|
| It's interesting you mentioned accounting, because that's
| the one department/area I see taking off and running with
| it the most. They are already the department that's
| effectively programming already with Excel workflows &
| DSLs in whatever ERP du jour.
|
| So it doesn't necessarily open up more dev jobs, but
| maybe fulfills the old the mantra of "everyone will
| become a programmer." and we see more advanced computing
| become a commodity thanks to AI - much like everyone can
| click their way through an office suite with little
| experience or training, everyone will be able to use AI
| to automate large chunks of their job or departmental
| processes.
| kentonv wrote:
| > I think we'll actually see (maybe) individual employees
| solving some of these problems on their own without
| involving IT/the dev team.
|
| I agree, but in my book, those employees are now
| developers. And so by that definition, there will be a
| lot more developers.
|
| Will we see more or fewer people whose primary job is
| software development? That's harder to answer. I do think
| we'll see a lot more consultant-type roles, with
| experienced software developers helping other people
| write their own personal automations.
| the_sleaze_ wrote:
| Banking allegedly runs on ancient cobalt cathedrals and
| mystical runes.
|
| Will AI be able translate all that into rust?
| simonw wrote:
| I guess I have trouble emphasizing with "But what if you only
| need 2 kentonv's instead of 20 at the end?" because I'm an
| open source oriented developer.
|
| What's open source for if not allowing 2 developers to
| achieve projects that previously would have taken 20?
| danans wrote:
| > Not software engineers being kicked out and some business
| person pressing a few buttons to have a fully functional app
| (as is playing out in a lot of fantasies on LinkedIn & X)
|
| The theory of enshittification says that "business person
| pressing a few buttons" approach will be pursued, even if it
| lowers quality, to save costs, at least until that approach
| undermines quality so much that it undermines the business
| model. However, nobody knows how much quality tradeoff
| tolerance is there to mine.
| stackskipton wrote:
| >experienced engineers using AI to generate bits of code and
| then meticulously reviewing and testing them
|
| And where are supposed to get experienced engineers if replaced
| all Jr Devs with AI? There is a ton of benefit from drudgery of
| writing classes even if seems like grunt work at the time.
| hooverd wrote:
| AI is great for undifferentiated heavy lifting and surfacing
| knowledge, but by the time I've made all the decisions, I can
| just write the code that matters myself there.
| kentonv wrote:
| It took me a few days to build the library with AI.
|
| I estimate it would have taken a few weeks, maybe months to
| write by hand.
|
| That said, this is a pretty ideal use case: implementing a
| well-known standard on a well-known platform with a clear API
| spec.
|
| In my attempts to make changes to the Workers Runtime itself
| using AI, I've generally not felt like it saved much time.
| Though, people who don't know the codebase as well as I do have
| reported it helped them a lot.
|
| I have found AI incredibly useful when I jump into _other
| people 's_ complex codebases, that I'm not familiar with. I now
| feel like I'm comfortable doing that, since AI can help me find
| my way around very quickly, whereas previously I generally
| shied away from jumping in and would instead try to get someone
| on the team to make whatever change I needed.
| philipwhiuk wrote:
| > Though, people who don't know the codebase as well as I do
| have reported it helped them a lot.
|
| My problem I guess is that maybe this is just Dunning-Kruger
| esq. When you don't know what you don't know you get the
| impression it's smart. When you do, you think it's rubbish.
|
| Like when you see a media report on a subject you know about
| and you see it's inaccurate but then somehow still trust the
| media on a subject you're a non-expert on.
| throwaway314155 wrote:
| I think most of this just amounts to the same old good
| developers vs. bad developers situation that we've been in
| for decades.
| giantrobot wrote:
| > Like when you see a media report on a subject you know
| about and you see it's inaccurate but then somehow still
| trust the media on a subject you're a non-expert on.
|
| Gell-Mann Amnesia https://en.m.wikipedia.org/wiki/Gell-
| Mann_amnesia_effect
| 9dev wrote:
| Funny thing. I have built something similar recently, that is
| a 2.1-compliant authorisation server in TypeScript[0]. I did
| it by hand, with some LLM help on the documentation. I think
| it took me about two weeks full time, give or take, and
| there's still work to do, especially on the testing side of
| things, so I would agree with your estimate.
|
| I'm going to take a very close look at your code base :)
|
| [0] https://github.com/colibri-
| hq/colibri/blob/next/packages/oau...
| upstairs-war wrote:
| Thanks kentonv. I picked up where you left off, supported
| with oauth2.1 rfc, and integrated ms oauth to our internal
| mcp server. Cool to have Claude be business aware
| srhtftw wrote:
| > It took me a few days to build the library with AI. ... > I
| estimate it would have taken a few weeks, maybe months to
| write by hand.
|
| I don't think this is a fair assessment give the summary of
| the commit history https://pastebin.com/bG0j2ube shows your
| work started on 2025-02-27 and started trailing off at
| 2025-03-20 as others joined in. Minor changes continue to
| present.
|
| > That said, this is a pretty ideal use case: implementing a
| well-known standard on a well-known platform with a clear API
| spec.
|
| Still, this allowed you to complete in a month what may have
| taken two. That's a remarkable feat considering the time and
| value of someone of your caliber.
| manquer wrote:
| Is it though?
|
| Would someone of author's caliber even be working on
| trivial slog item like Oauth2 implementation, if not for
| the novel development approach he wanted to attempt here ?
|
| For the kind of regular jobs a engineer typically is
| expected to do, would it give 100% productivity jump ?
| tkiolp4 wrote:
| Why is speed important in this context? If the code is
| published one week/month later, would that affect what exactly?
| It's open source.
| skybrian wrote:
| Looking at the commit history, there's a fair bit of manual
| intervention to fix bugs and remove unused code.
| qsort wrote:
| I think this is pretty cool, but it doesn't really move my priors
| that much. Looking at the commit history shows a lot of
| handholding even in pretty basic situations, but on the other
| hand they probably saved _a lot_ of time vs. doing everything
| manually.
| jes5199 wrote:
| I've been using Claude (via Cursor) on a greenfield project for
| the last couple months and my observation is:
|
| 1. I am much more productive/effective
|
| 2. It's way _more_ cognitively demanding than writing code the
| old-fashioned way
|
| 3. Even over this short timespan, the tools have improved
| significantly, amplifying both of the points above
| diggan wrote:
| > It's way more cognitively demanding than writing code the
| old-fashioned way
|
| How are you using it?
|
| I've been mainly doing "pair programming" with my own agent
| (using Devstral as of late) and find the reviewing much easier
| than it would been to literally type all of the code it
| produces, at least time wise.
|
| I've also tried vibe coding for a bit, and for that I'd agree
| with you, as you don't have any context if you end up wanting
| to review something. Basically, if the project was vibe coded
| from the beginning, it's much harder to get into the codebase.
|
| But when pair programming with the LLM, I already have a built
| up context, and understand how I want things to be and so on,
| so reviewing pair programmed code goes a lot faster than
| reviewing vibe coded code.
| jes5199 wrote:
| I've tried a bunch of things but now I'm mostly using Cursor
| in agent mode with Claude Sonnet 4, doing small-ish pull-
| request-sized prompts. I don't have to review code as
| carefully as I did with Claude 3.7
|
| but I'm finding the bottleneck now is architecture design. I
| end up having these long discussions with chatGPT-o3 about
| design patterns, sometimes days of thinking, and then
| relatively quick implementation sessions with Cursor
| SkyPuncher wrote:
| > 2. It's way more cognitively demanding than writing code the
| old-fashioned way
|
| Funnily, enough, I find the exact opposite. I feel so much
| relief that I don't have to waste time figuring out every,
| single detail. It frees me up to focus on architectural and
| higher level changes.
| jes5199 wrote:
| I guess what I mean is, I found the details sort of
| "mindless" before. Code that I could write in my sleep. Now I
| only have to do the thinky parts
| layer8 wrote:
| This means that you fully trust the LLM to get the details
| right.
| pton_xd wrote:
| This mirrors my experience and those I've talked to.
|
| LLM assisted coding is a way to get stuff done much faster, at
| a greatly increased mental cost / energy spent. Oddly enough.
| piker wrote:
| The small dopamine hits you get from "it compiles" are
| completely automated away, and you're forced to survive on
| the goal alone. The issues are necessarily complex and
| require thinking about how the LLM has gotten it subtly
| wrong.
|
| Painful, but effective?
| infinitebattery wrote:
| From this commit: https://github.com/cloudflare/workers-oauth-
| provider/commit/...
|
| ===
|
| "Fix Claude's bug manually. Claude had a bug in the previous
| commit. I prompted it multiple times to fix the bug but it kept
| doing the wrong thing.
|
| So this change is manually written by a human.
|
| I also extended the README to discuss the OAuth 2.1 spec
| problem."
|
| ===
|
| This is super relatable to my experience trying to use these AI
| tools. They can get halfway there and then struggle immensely.
| nisegami wrote:
| Same. But I personally find it a lot easier to do those bits at
| the end than to begin from a blank file/function, so it's a
| good match for me.
| SkyPuncher wrote:
| Same here. Sometimes you just need time to stew in the
| problem/solution space.
|
| LLMs let me be ultraproductive upfront then come in at the
| end to clean up when I have a full understanding.
| diggan wrote:
| > They can get halfway there and then struggle immensely.
|
| Restart the conversation from scratch. As soon as you get
| something incorrect, begin from the beginning.
|
| It seems to me like any mistake in a messages
| chain/conversation instantly poisons the output afterwards,
| even if you try to "correct" it.
|
| So if something was wrong at one point, you need to go back to
| the initial message, and adjust it to clarify the prompt enough
| so it doesn't make that same mistake again, and regenerate the
| conversation from there on.
| eikenberry wrote:
| I thought Claude still has a problem generating the same
| output for the same input? That you can't just rewind and
| rerun and get to the same point again.
| diggan wrote:
| > I thought Claude still has a problem generating the same
| output for the same input?
|
| I haven't used Anthropic's models/software in a long time
| (months, basically forever in AI ecosystem), so don't know
| exactly how it works now.
|
| But last time I used Claude, you could edit the first
| message, and then re-generate the assistants next message
| based on your edit. Most of the LLM interfaces has one or
| another way of doing this, I can't imagine they got rid of
| that feature.
|
| What I'm suggesting isn't to use the exact same input (the
| first message), but rather change it so you remove the
| chances of something incorrect happening later after that.
| throwaway314155 wrote:
| > can't just rewind and rerun and get to the same point
| again
|
| Why would you want to? The whole point of a retry is that
| your previous conversation attempt went poorly.
| eikenberry wrote:
| Good engineering? You want automated steps to be
| repeatable so you know your tweak to the previous
| conversation have the effect you desire. Though using an
| AI for coding is probably closer in spirit the the art of
| writing code than the engineering of writing code and art
| is pretty much unrepeatable by definition.
| throwaway314155 wrote:
| Fair enough. Use the respective API or Google Gemini
| which will let you set temperature to zero resulting in
| deterministic output barring FP errors accumulating when
| paired with non-standard GPU/TPU configurations. Likely
| not to differ by much in the vast majority of cases
| though.
| dingnuts wrote:
| Can you imagine if Excel worked like this? the formula put
| out the wrong result, so try again! It's like that scene from
| The Office where Michael has an accountant "run it again."
| It's farcical. They have created computers that are bad at
| math and I will never forgive them.
|
| Also, each try costs money! You're pulling the lever on a god
| damned slot machine!
|
| I will TRY AGAIN with the same prompt when I start getting a
| refund for my wasted money and time when the model outputs
| bullshit, otherwise this is all confirmation and sunk cost
| bias talking, I'm sure if it.
| diggan wrote:
| > Can you imagine if Excel worked like this?
|
| I mean, why would I imagine that? Who would want that? It's
| like the argument against legal marijuana, and someone
| replies "But would you like your pilot to be high when
| flying?!". Right tool for the right job, clearly when you
| want 100% certainty then LLMs aren't the tool for that.
| Just because they're useful for some things don't mean we
| have to replace everything with them.
|
| > Also, each try costs money!
|
| I guess you're using some paid API? Try a different way
| then. I mostly use the web UI from OpenAI, or Codex lately,
| or ran locally with my own agent using local weights,
| neither is "each try costs money" more than writing data to
| my SSD is costing me money.
|
| It's not a holy grail some people paint it, and not sure
| we're across the "productivity threshold"
| (https://news.ycombinator.com/item?id=44160664) yet, but
| it's worth trying it out probably before jumping to
| conclusions. But no one is forcing you either, YMMV and all
| that.
| int_19h wrote:
| Chatbot UIs really need better support for conversation
| branching all around. It's very handy to be able to just
| right-click on any random message in the conversation in LM
| Studio and say, "branch from here".
| diggan wrote:
| Maybe it's contrarian, maybe it's not, but I don't think
| Chat UIs are well suited for software
| engineering/programming at all, we need something
| completely different. Being able to branch conversations
| and such would be useful, but probably not for the way I do
| software. Besides, I'm rarely beyond 3 messages (1 system,
| 1 user, 1 assistant) in any usage of the chat UIs. Maybe
| it's more useful to people with different workflows.
| carlosjs23 wrote:
| AI Studio has this, I usually ask it to plan and I do some
| rounds of refining until the plan covers all my
| requirements, then I branch this conversation, a branch for
| each feature, none of the branches get polluted this way.
| viktorcode wrote:
| It can be done, but for my environment the sum of all prompts
| that I end up typing to get the right result ends up being
| longer than the actual code.
|
| So now I'm using LLMs as crapshoot machines for generating
| ideas which I then implement manually
| mysterydip wrote:
| This to me is why I think these tools don't have actual
| understanding, and are instead producing emergent output from
| pooling an incomprehensibly large set of pattern-recognized
| data.
| diggan wrote:
| > these tools don't have actual understanding, and are
| instead producing emergent output from pooling an
| incomprehensibly large set of pattern-recognized data
|
| I mean, bypassing the fact that "actual understanding"
| doesn't have any consensus about what it is, does it matter
| if it's "actual understanding" or "kind of understanding", or
| even "barely understanding", as long as it produces the
| results you expect?
| sceptic123 wrote:
| > as long as it produces the results you expect?
|
| But it's more the case of "until it doesn't produce the
| results you expect" and then what do you do?
| diggan wrote:
| > "until it doesn't produce the results you expect" and
| then what do you do?
|
| I'm not sure I understand what you mean. You're asking it
| to do something, and it doesn't do that?
| dingnuts wrote:
| if you give an LLM a spec with a new language and no
| examples, it can't write the new language.
|
| until someone does that, I think we've demonstrated that
| they do not have understanding or abstract thought. they
| NEED examples in a way humans do not.
| Powdering7082 wrote:
| https://openreview.net/pdf?id=GTHD2UnDIb
| mysterydip wrote:
| Interesting paper, thanks for sharing. I assume the
| effectiveness depends greatly on the syntax of the
| language to be learned (c-like, etc).
| seunosewa wrote:
| Then you teach it. Even humans don't always produce the
| results we expect.
| mysterydip wrote:
| No, I was not making a critique on its effectiveness at
| generating usable results. I was responding to what I've
| seen in several other articles here arguing towards
| anthropomorphism.
| krooj wrote:
| The comment in lines 163 - 172 make some claims that are
| outright false and/or highly A/S dependent, to the point where
| I question the validity of this post entirely. While it's
| possible that an A/S can be pseudo-generated based on lots of
| training data, each implementation makes very specific design
| choices: i.e.: Auth0's A/S allows for a notion of "leeway"
| within the scope of refresh token grant flows to account for
| network conditions, but other A/S implementations may be far
| more strict in this regard.
|
| My point being: assuming you have RFCs (which leave A LOT to
| the imagination) and some OSS implementations to train on, each
| implementation usually has too many highly specific choices
| made to safely assume an LLM would be able to cobble something
| together without an amount of oversight effort approaching
| simply writing the damned thing yourself.
| nicce wrote:
| I am waiting for studies whether we have just an illusion of
| production or these actually save man hours in the long term in
| creation of production-level systems.
| arendtio wrote:
| One way to mitigate the issue is to use tests or specifications
| and let the AI find a solution to the spec.
|
| A few months ago, solving such a spec riddle could take a
| while, and most of the time, the solutions that were produced
| by long run times were worse than the quick solutions. However,
| recently the models have become significantly better at solving
| such riddles, making it fun (depending on how well your use
| case can be put into specs).
|
| In my experience, sonnet 3.7 represented a significant step
| forward compared to sonnet 3.5 in this discipline, and Gemini
| 2.5 Pro was even more impressive. Sonnet 4 makes even fewer
| mistakes, but it is still necessary to guide the AI through
| sound software engineering practices (obtaining requirements,
| discovering technical solutions, designing architecture,
| writing user stories and specifications, and writing code) to
| achieve good results.
|
| Edit: And there is another trick: Provide good examples to the
| AI. Recently, I wanted to create an app with the OpenAI
| Realtime API and at first it failed miserably, but then I added
| the most important two pages of the documentation and one of
| the demo projects into my workspace and just like that it
| worked (even though fur my use-case the API calls had to be use
| quite differently).
| fxnn wrote:
| That's one thing where I love Golang. I just tell Aider to
| `/run go doc github.com/some/package`, and it includes the
| full signatures in the chat history.
|
| It's true: often enough AI struggles to use libraries, and
| doesn't remember the usage correctly. Simply adding the go
| doc fixed that often.
| thih9 wrote:
| Congrats and thanks for sharing, both the code and the story.
|
| Which Claude plan did you use? Was it enough or did you feel
| limited by the quotas?
| kentonv wrote:
| This was mostly Claude Code, which runs on API credits. I think
| I spent a two-digit number of dollars. The model was Sonnet 3.7
| (this was all a couple months ago, before Claude 4).
| declan_roberts wrote:
| Getting a "Too Many Requests" error is kind of hilarious given
| the company involved.
| rcastellotti wrote:
| same
| _tqr3 wrote:
| I've tried building a web app with LLMs before. Two of them went
| in circles--I'd ask them to fix an infinite loop, they'd remove
| the code for a feature; I'd ask them to add the feature back,
| they'd bring back the infinite loop, and so on. The third one
| kept losing context--after just 2-3 messages, it would rebuild
| the whole thing differently.
|
| They'll probably get better, but for now I can safely say I've
| spent more time building and tweaking prompts than getting
| helpful results.
| diggan wrote:
| Rather than doing that approach which eventually builds up to
| 10+ messages or more, iterate on your initial prompt and you'll
| see better results. So if the first prompt correctly fixed the
| infinite loop, but removed something else, instead of saying
| "Add that back again", change the initial prompt to include
| "Don't remove anything else than what's explicitly mentioned"
| or similar, and you'll either get exactly what you want, or
| some other issue. Then rinse and repeat until completed.
|
| Eventually you'll build up a somewhat reusable template you can
| use as a system prompt to guide it exactly how you want.
|
| Basically, you get what you ask for, nothing else and nothing
| more. If you're unclear, it'll produce unclear outputs, if you
| didn't mention something, it'll do whatever with that. You have
| to be really, really explicit about everything.
| horacemorace wrote:
| I did the same thing a few months ago with 4o. This stuff works
| fine if done with care.
| jonplackett wrote:
| This doesn't seem like much of a surprise that it's possible - if
| you are a security expert, you can make LLMs write secure code.
| c-linkage wrote:
| I very much appreciate the fact that the OP posted not just the
| code developed by AI but also posted the prompts.
|
| I have tried to develop some code (typically non-web-based code)
| with LLMs but never seem to get very far before the
| hallucinations kick in and drive me mad. Given how many other
| people claim to have success, I figure maybe I'm just not writing
| the prompts correctly.
|
| Getting a chance to see the prompts shows I'm not actually that
| far off.
|
| Perhaps the LLMs don't work great for me because the problems I'm
| working on a somewhat obscure (currently reverse engineering SAP
| ABAP code to make a .NET implementation on data hosted in
| Snowflake) and often quite novel (I'm sure there is an OpenAuth
| implementation on gitbub somewhere from which the LLM can crib).
| sceptic123 wrote:
| If you need to be an expert to use AI tools safely, what does
| that say about AI tools?
| dkdcio wrote:
| Genuinely curious what your point is? Do you know how to use a
| ventillator? A A timing gun? A tonometer? A keratometer? Can
| you use all of those in a "production" setting safely without
| expertise?
| Bjartr wrote:
| They didn't make a point, they asked a question. Sometimes
| people do still ask questions because they're interested in
| the answer.
| alanfranz wrote:
| Carefully reviewed greenfield project; I don't think this is
| astonishing, and I very much love they recorded the prompts.
|
| Question is: will this work for non-greenfield projects as well?
| Usually 95% of work in a lifetime is not greenfield.
|
| Or will we throw away more and more code as we go, since AI will
| rewrite it, and we'll probably introduce subtle bugs as we go?
| multimoon wrote:
| I think this reinforces that "vibecoding" is silly and won't
| survive. It still needed immensely skilled programmers to work
| with it and check its output, and fix several bugs it refused to
| fix.
|
| Like anything else it will be a tool to speed up a task, but
| never do the task on its own without supervision or someone who
| can already do the task themselves, since at a minimum they have
| to already understand how the service is to work. You might be
| able to get by to make things like a basic website, but tools
| have existed to autogenerate stuff like that for a decade.
| NitpickLawyer wrote:
| I don't think it does. Vibecoding is currently best suited for
| low-stakes stuff. Get a gui up, crud stuff, write an app for a
| silly one time use, etc. There's a ton of usage there. And it's
| putting that power in the hands of people that didn't have the
| capabilities before.
|
| This isn't vibecoding. This is LLM-assisted coding.
| subarctic wrote:
| I get the sense that "vibecoding" is used like a strawman these
| days, something people keep moving the goal posts on so they
| can keep saying it's silly. Getting an LLM to write code for
| you that mostly works with some tweaks is vibe coding, isn't
| it?
| ZiiS wrote:
| Shouldn't they really have asked it to read
| https://developers.cloudflare.com/workers/examples/protect-a...
| kentonv wrote:
| The secret token is hashed first, and it's the hash that is
| looked up in storage. In this arrangement, an attacker cannot
| use timing to determine the correct value byte-by-byte, because
| any change to the secret token is expected to randomize the
| whole hash. So, timing-safe equality is not needed.
|
| That said, if you have spotted a place in the code where you
| believe there is such a vulnerability, please do report it.
| Disclosure guidelines are at:
| https://github.com/cloudflare/workers-oauth-provider/blob/ma...
| ZiiS wrote:
| I am not confident enough in this area to to report a
| vunrability, the networking alone probably makes timing
| impractical. I thought it was now practical to generate known
| prefix Sha256, so some information could be extracted? Not
| enough to compromise but the function is right there.
| kentonv wrote:
| Learning a prefix of the hash doesn't really get you
| anywhere. The hash itself isn't a secret -- it could be
| published publicly without breaking the security model. You
| still need to derive a token that hashes to that value in
| full, and if you can do that then you've broken the hash
| algorithm by definition.
| ZiiS wrote:
| Yes I guess if you trust the hash implementation
| completly; I just favour a bit more defence in depth.
| DJBunnies wrote:
| I feel like well defined RFCs and standards are easily coded
| against, and I question the investment/value/time tradeoff here.
| These things happily regurgitate training data, but seriously
| struggle when they don't have a pool of perfect examples to pull
| from.
|
| When Claude can do something new, then I think it will be
| impressive.
|
| Otherwise it's just piecing together existing examples.
| kentonv wrote:
| I'm the author of this library! Or uhhh... the AI prompter, I
| guess...
|
| I'm also the lead engineer and initial creator of the Cloudflare
| Workers platform.
|
| --------------
|
| Plug: This library is used as part of the Workers MCP framework.
| MCP is a protocol that allows you to make APIs available directly
| to AI agents, so that you can ask the AI to do stuff and it'll
| call the APIs. If you want to build a remote MCP server, Workers
| is a great way to do it! See:
|
| https://blog.cloudflare.com/remote-model-context-protocol-se...
|
| https://developers.cloudflare.com/agents/guides/remote-mcp-s...
|
| --------------
|
| OK, personal commentary.
|
| As mentioned in the readme, I was a huge AI skeptic until this
| project. This changed my mind.
|
| I had also long been rather afraid of the coming future where I
| mostly review AI-written code. As the lead engineer on Cloudflare
| Workers since its inception, I do a LOT of code reviews of
| regular old human-generated code, and it's a slog. Writing code
| has always been the fun part of the job for me, and so delegating
| that to AI did not sound like what I wanted.
|
| But after actually trying it, I find it's quite different from
| reviewing human code. The biggest difference is the feedback loop
| is much shorter. I prompt the AI and it produces a result within
| seconds.
|
| My experience is that this actually makes it feels more like I am
| authoring the code. It feels similarly fun to writing code by
| hand, except that the AI is exceptionally good at boilerplate and
| test-writing, which are exactly the parts I find boring. So... I
| actually like it.
|
| With that said, there's definitely limits on what it can do. This
| OAuth library was a pretty perfect use case because it's a well-
| known standard implemented in a well-known language on a well-
| known platform, so I could pretty much just give it an API spec
| and it could do what a generative AI does: generate. On the other
| hand, I've so far found that AI is not very good at refactoring
| complex code. And a lot of my work on the Workers Runtime ends up
| being refactoring: any new feature requires a bunch of upfront
| refactoring to prepare the right abstractions. So I am still
| writing a lot of code by hand.
|
| I do have to say though: The LLM _understands_ code. I can 't
| deny it. It is not a "stochastic parrot", it is not just
| repeating things it has seen elsewhere. It looks at the code,
| understands what it means, explains it to me mostly correctly,
| and then applies my directions to change it.
| rethab wrote:
| Fancy! Why are the first twenty commits or so created in the
| same minute though? Surely you can't be that fast if you need
| to prompt for each commit
| kentonv wrote:
| That's weird! It must be due to a history rewrite I did later
| on to clean up the repo, removing some files that weren't
| really part of the project. I didn't realize when I first
| started the experiment that we'd actually end up releasing
| the code so I had to go back and clean it up later. I am
| surprised though that this messed up the timestamps --
| usually rebases retain timestamps. I think I used `git
| filter-branch`, though. Maybe that doesn't retain timestamps.
| euiq wrote:
| I know that `git rebase` changes the committer date while
| keeping the author date the same, so I'm assuming something
| similar happened here. For example, many of the early
| commits have this committer date with varying author dates:
| $ git show --format=fuller
| 3dafc8f5de6ffe46fb223a75a46a6bd848b6daf8 commit
| 3dafc8f5de6ffe46fb223a75a46a6bd848b6daf8 Author:
| Kenton Varda <kenton@cloudflare.com> AuthorDate:
| Thu Feb 27 17:15:37 2025 -0600 Commit: Kenton
| Varda <kenton@cloudflare.com> CommitDate: Tue Mar 4
| 14:48:59 2025 -0600 Add storage schema
| by Claude.
|
| GitHub uses the committer date for its history, which is
| annoying if you rebase frequently; I like to run a non-
| interactive `git rebase` with `--commmiter-date-is-author-
| date` in such cases.
| davidwu wrote:
| Thanks for so meticulously documenting the prompts you used and
| whether or not a commit was done manually or via the AI.
| revskill wrote:
| So who's the experts here ?
| ookblah wrote:
| AI critics always have to make strawmen arguments about how there
| has to be a human in the loop to "fix" things when that's never
| been the argument AI proponents ever make (at least those who
| deal with it day to day). This will only get better with time. AI
| can frequently one-shot throwaway scripts that I need get things
| done. For actual features I typically start and have it go thru
| the initial slog and then finish it off. You must be reviewing
| the entire time, but it takes a huge cognitive load off. You can
| rubber-duck debug with it.
|
| I do agree if you have no idea what you are doing or are still
| learning it could be a detriment, but like anything it's just a
| tool. I feel for junior devs and the future. Lazy coders get
| lazier, those who utilize them to the fullest extent get even
| better, just like with any tech.
| skydhash wrote:
| The one thing about concocting throwaway scripts yourself is
| the increased familiarity with the tooling you use. And you're
| not actually throwing away those scripts. I have random scripts
| laying around my file system (and my shell history) to check
| how I did a task in the past.
| bongodongobob wrote:
| I used to do that too. I find I don't really need to save
| anything less than 100 lines these days because I can just
| ask again when I need it.
| NitpickLawyer wrote:
| > increased familiarity with the tooling you use
|
| In general I agree, but sometimes you want something that you
| haven't done in years but vaguely remember.
|
| ~20 years ago I worked with ffmpeg and vlc extensively in an
| IPTV project. It took me months to RTFM, implement stuff,
| test and so on. Documentation was king, and really the only
| thing I could use. Old-school. But after that project I moved
| on.
|
| In 2018 I worked on a ML - CV project. I knew vlc / ffmpeg
| could do everything that I needed, but I had forgotten most
| of everything by then. So I googled/so/random-blogs, plus a
| bit of RTFM where things didn't match. But it still took a
| few days to cobble together the thing I needed.
|
| Now I just ask, and the perfect one-liner pops-up, I run it,
| check that it does what I need it to, and go on my merry way.
| Verification is much faster than context changing, searching,
| reading, understanding, testing it out, using a work-around
| for the features that ffmpeg supports but not that python
| wrapper, and so on.
| wooque wrote:
| Not surprised, this is perfect task for AI, boilerplaty code that
| implements something that is implemented 100 times. And it's
| small project, 1200 lines of pure code.
|
| I'm surprised I took them more than 2 days to do that with AI.
| vjerancrnjak wrote:
| I like how it is just 1 file.
|
| Wonder how well incremental editing works with such a big file. I
| keep pushing for 1 file implementations, yet people split it up
| into bazillion files because it works better with AI.
| freedomben wrote:
| On a meta-note, it's (seriously) kind of refreshing to see that
| other people make this same typo when trying to type Cloudflare.
| I also often write CLoudflare, Cloudlfare, and Cloudfare:
|
| > _Cloudlflare builds OAuth with Claude and publishes all the
| prompts_
| IncreasePosts wrote:
| Has this source been compared with other oauth libraries, to see
| if it is just license-violating some other open source code it
| was trained on?
| mehdibl wrote:
| That's great.
|
| But Claude don't allow yet to add APPS in their backend.
|
| Mainly only closed beta for integration.
|
| How you can configure an app to leverage correctly Oauth and have
| your own app secret ID/ Client ID!
| globular-toast wrote:
| Should I be impressed? Oauth already exists and there are
| countless libraries implementing it. Is it impressive that an LLM
| can regurgitate yet another one?
| EtienneK wrote:
| > _This is a TypeScript library that implements the provider side
| of the OAuth 2.1 protocol with PKCE support._
|
| What is the "provider" side? OAuth 2.1 has no definition of a
| "provider". Is this for Clients? Resource Servers? Authorization
| Server?
|
| Quickly skimming the rest of the README it seems this is for
| creating a mix of a Client and a Resource Server, but I could be
| mistaken.
|
| > _To emphasize, this is not "vibe coded". Every line was
| thoroughly reviewed and cross-referenced with relevant RFCs, by
| security experts with previous experience with those RFCs_
|
| Experience with the RFCs but have not been able to correctly name
| it.
| DaiPlusPlus wrote:
| > OAuth 2.1 has no definition of a "provider"
|
| Strictly speaking, yes. But speaking of IDPs more broadly, it's
| perfectly acceptable to refer to the authorisation-server as an
| auth-provider, especially in OIDC (which _is_ OAuth, with
| extensions) where it's explicitly called "OpenID provider" - so
| it's natural for anyone well-versed in both to cross
| terminology like that.
| kentonv wrote:
| This library helps implement both the resource server and
| authorization server. Most people understand these two things
| to be, collectively, the "provider" side of OAuth -- the
| service provider, who is providing an API that requires
| authorization. The intent when using this library is that you
| write one Worker that does both. This library has no use on the
| client side.
|
| This is intended for building lightweight services quickly.
| Historically there has been no real need for "lightweight"
| OAuth providers -- if you were big enough that people wanted to
| connect to you using OAuth, you were not lightweight. MCP has
| sort of changed that as the "big" side of an MCP interaction is
| the client side (the LLM provider), whereas lots of people want
| to create all kinds of little MCP servers to do all kinds of
| little things. But MCP specifies OAuth as the authentication
| mechanism. So now people need to be able to implement OAuth
| from the provider side easily.
|
| > Experience with the RFCs but have not been able to correctly
| name it.
|
| These docs are written for people building MCP servers, most of
| whom only know they want to expose an API to AIs and have never
| read OAuth RFCs. They do not know or care about the difference
| between an authorization server and a resource server.
| catigula wrote:
| When I want to spend a dollar or two, it's much faster to just
| instruct Claude on how to write my code and prompt/correct it
| than to write it myself.
|
| It feels probably similarly from going from dumb or semi-dumb
| text editor to an IDE.
| simonw wrote:
| The most clearly Claude-written commits are on the first page,
| this link should get you to them:
| https://github.com/cloudflare/workers-oauth-provider/commits...
| throwaway314155 wrote:
| "built OAuth" here means they "implemented OAuth for CloudFlare
| workers" FYI.
| rienbdj wrote:
| Why not use an existing OAuth library?
| keeda wrote:
| A number of comments point out that OAuth is a well known
| standard and wonder how AI would perform on less explored problem
| spaces. As it happens I have some experience there, which I wrote
| about in this long-ass post nobody ever read:
| https://www.linkedin.com/pulse/adventures-coding-ai-kunal-ka...
|
| It's now a year+ old and models have advanced radically, but most
| of the key points still hold, which I've summarized here. The
| post has way more details if you need. Many of these points have
| also been echoed by others like @simonw.
|
| Background:
|
| * The main project is specialized and "researchy" enough that
| there is no direct reference on the Internet. The core idea has
| been explored in academic literature, a couple of relevant
| proprietary products exist, but nobody is doing it the way I am.
|
| * It has the advantage of being greenfield, but the drawback of
| being highly "prototype-y", so some gnarly, hacky code and a ton
| of exploratory / one-off programs.
|
| * Caveat: my usage of AI is actually very limited compared to
| power users (not even on agents yet!), and the true potential is
| likely far greater than what I've described.
|
| Highlights:
|
| * At least 30% and maybe > 50% of the code is AI-generated. Not
| only are autocompletes frequent, I do a lot of "chat-oriented"
| and interactive "pair programming", so precise attribution is
| hard. It has written large, decently complicated chunks of code.
|
| * It does boilerplate extremely easily, but it also handles novel
| use-cases very well.
|
| * It can refactor existing code decently well, but probably
| because I'ver worked to keep my code highly modular and
| functional, which greatly limits what needs to be in the context
| (which I often manage manually.) Errors for even pretty
| complicated requests are rare, especially with newer models.
|
| Thoughts:
|
| * AI has let me be productive - and even innovate! - despite
| having limited prior background in the domains involved. The vast
| majority of all innovation comes from combining and applying
| well-known concepts in new ways. My workflow is basically a "try
| an approach -> analyze results -> synthesize new approach" loop,
| which generates a lot of such unique combinations, and the AI
| handles those just fine. As @kentonv says in the comments, there
| is no doubt in my mind that these models "understand" code, as
| opposed to being stochastic parrots. Arguments about what
| constitutes "reasoning" are essentially philosophical at this
| point.
|
| * While the technical ideas so far have come from me, AI now
| shows the potential to be inventive by itself. In a recent
| conversation ChatGPT reasoned out a novel algorithm and code for
| an atypical, vaguely-defined problem. (I could find no reference
| to either the problem or the solution online.) Unfortunately, it
| didn't work too well :-) I suspect, however, that if I go full
| agentic by giving it full access to the underlying data and
| letting it iterate, it might actually refine its idea until it
| works. The main hurdles right now are logistics and cost.
|
| * It took me months to become productive with AI, having to find
| a workflow AND code structure that works well _for me_. I don't
| think enough people have put in the effort to find out what works
| _for them_ , and so you get these polarized discussions online. I
| implore everyone, find a sufficiently interesting personal
| project and spend a few weekends coding with AI. You owe it to
| yourself, because 1) it's free and 2)...
|
| * Jobs are absolutely going to be impacted. Mostly entry-level
| and junior ones, but maybe even mid-level ones. Without AI, I
| would have needed a team of 3+ (including a domain expert) to do
| this work in the same time. All knowledge jobs rely on a mountain
| of donkey work, and the donkey is going the way of the dodo. The
| future will require people who uplevel themselves to the state of
| the art and push the envelope using these tools.
|
| * How we create AI-capable senior professionals without junior
| apprentices is going to be a critical question for many
| industries. My preliminary take is that motivated apprentices
| should voluntarily eschew all AI use until they achieve a
| reasonable level of proficiency.
| tveita wrote:
| Some examples of prompt exchanges that seem representative:
|
| https://claude-workerd-transcript.pages.dev/oauth-provider-t...
| ("Total cost: $6.45")!
|
| https://github.com/cloudflare/workers-oauth-provider/commit/...
|
| https://github.com/cloudflare/workers-oauth-provider/commit/...
|
| The first transcript includes the cost, would be interesting to
| know the ballpark of total Claude spend on this library so far.
|
| --
|
| This is opportune for me, as I've been looking for a description
| of AI workflows from people of some presumed competency. You'd
| think there would be many, but it's hard to find anything
| reliable amidst all the hype. Is anyone live coding anything but
| todo lists?
|
| antirez:
| https://antirez.com/news/144#:~:text=Yesterday%20I%20needed%...
|
| tptacek: https://news.ycombinator.com/item?id=44163292
| kentonv wrote:
| I didn't keep extract track but I'd estimate the total cost of
| Claude credits to build this library was somewhere around $50,
| which is pretty negligible compared to the time saved.
| Squeeeez wrote:
| So I randomly clicked on a commit, went through the changes, went
| WTF? a few times, expanded the diff window a few times, expanded
| some more... 2.5k+ lines method? If gate cascades? Looking at
| more of the PRs and I would not have accepted such code smells
| from humans. So, maybe it was reviewed, in some way. To me it
| doesn't look like the reviewer(s) truly had their critical
| thinking caps on, more like their cowboy hats. Yeehaa!
___________________________________________________________________
(page generated 2025-06-02 23:01 UTC)