[HN Gopher] Claude now creates interactive charts, diagrams and ...
___________________________________________________________________
Claude now creates interactive charts, diagrams and visualizations
Author : adocomplete
Score : 159 points
Date : 2026-03-12 15:59 UTC (7 hours ago)
(HTM) web link (claude.com)
(TXT) w3m dump (claude.com)
| fixxation92 wrote:
| I find it absolutely mindblowing to witness the rate at which
| Anthropic can ship new features. Only a year ago I couldn't wait
| to see some sort of Github integration and then it appeared only
| a week later. Seriously impressive stuff.
| Razengan wrote:
| Meanwhile, you still can't Sign in with Apple on the website.
|
| But you can Sign in with Google.
|
| If you signed up with your Apple on the iOS Claude app, to
| access your account on the computer, you have to open the
| passwords app and copy your random email address and paste it
| into the Claude website login.
|
| Also if you try to copy-paste a prompt from Notes etc into the
| Claude chat, it gets added as an attachment, so you can't edit
| the prompt. If you do the four-finger shortcut to paste it as
| text, it mangles newlines etc.
|
| Why are they so dumb about such basic UX for so long?
| bombela wrote:
| Add to the list backtick handling. If you start a backtick
| block on the claude web chat, you cannot leave it with the
| keyboard. You are now stuck between the backticks. It is as
| if they wanted to reproduce Slack misery.
| catgirlinspace wrote:
| Pressing the down arrow while inside a block exits it for
| me.
| radley wrote:
| > you still can't Sign in with Apple on the website.
|
| Apple forces developers to offer _Sign in with Apple on iOS_
| devices if any other sign in service is used. Apple can 't
| force them to do it on non-Apple platforms.
| Wowfunhappy wrote:
| > If you signed up with your Apple on the iOS Claude app, to
| access your account on the computer, you have to open the
| passwords app and copy your random email address and paste it
| into the Claude website login.
|
| Isn't this basically Apple's fault? When you signed up, Apple
| provided a fake email address in leu of your real one. This
| is great for privacy but means the service has the wrong
| email.
|
| I'm sure they didn't want to provide an Apple sign in option
| at all, but it's required by App Store rules.
| Razengan wrote:
| > _I 'm sure they didn't want to provide an Apple sign in
| option at all_
|
| But they wanted to provide a Google Sign In? wth?
|
| > _This is great for privacy but means the service has the
| wrong email._
|
| So harm the users to benefit the service? wtf?
|
| I don't want to give my real email or anything to random
| services, specially not one like Claude where they don't
| even let you remove your payment info.
| Wowfunhappy wrote:
| > I don't want to give my real email or anything to
| random services, specially not one like Claude where they
| don't even let you remove your payment info.
|
| The original complaint was:
|
| >> If you signed up with your Apple on the iOS Claude
| app, to access your account on the computer, you have to
| open the passwords app and copy your random email address
| and paste it into the Claude website login.
|
| Either you use your original email or you use a per-
| service email. Apple helps you do the latter, but this
| does come with UX tradeoffs.
|
| Using a per-service email, then complaining that the
| service does not have your real email, strikes me as
| misguided.
| Razengan wrote:
| > _but this does come with UX tradeoffs._
|
| Only when a dumb service refuses to support Sign In with
| one pro-privacy provider but does for another anti-
| privacy one.
|
| Anyway I've voted by having a ChatGPT/Codex subscription
| for 1 year and only tried Claude for 1 month. Not missing
| anything.
| nerdjon wrote:
| They could also just implement sign in with apple on their
| website, they have the ability to sign in with google so
| not supporting Apple is still a weird choice they are
| making.
|
| Apple should not have had to require developers to have
| options other than Google for authentication, but clearly
| some companies have to be dragged kicking and screaming.
|
| So clearly they support it, and there is no reason it
| should not work on the web also.
| j45 wrote:
| A vendor doesn't have to bend for another.
|
| Always best to sign in with your own email address.
| nerdjon wrote:
| There are a lot of websites that only support third party
| login, so that is not always an option.
|
| They don't have to bend for another, but they made a
| choice to put an app on iOS. They added support for apple
| signin, and then for some reason did not put it on their
| website.
|
| You can criticize Apple for requiring that all you want,
| but they clearly have support for it and are choosing to
| not put it on their website which is causing a worse user
| experience.
|
| IF apple did not support website loggin than sure, but
| they do. So the ability to fix this is on Anthropic (and
| many other websites).
|
| If you are already going to support third party login you
| should not limit it to only Google accounts and there is
| no reason to support Apple on iOS and not the web.
|
| Also for the record, Apple only requires sign in with
| apple if you already support third party authentication.
| So if you are already going to support that, giving the
| user more choice (and making it so we are all a bit less
| dependent on google) is a good thing.
| j45 wrote:
| No criticism from me towards apple or Anthropic. Both
| parties made their choice. Apple was late to the identity
| business and the other ships had already sailed.
|
| Third party logins are an extension and a massive risk to
| any website that doesn't include email hosting.
|
| We have see identity providers dissapear, and people may
| change their mind.
|
| Easiest way is to register you rown domain and use it
| with an identity provider of your choice and be able to
| move it anywhere.
|
| Otherwise we are a faceless citizen of a corporation that
| can handle access to our identity and everything attached
| to it without recourse or access to anyone.
| Razengan wrote:
| Bruh.
|
| Are you seriously trying to justify offering Sign in with
| Google _but not_ ALSO offer Sign in with Apple because of
| some contorted principle, the method which HELPS users
| maintain their privacy? What the actual f.
|
| Antrhopic's UX is just trash, the worst of all the major
| AI products.
|
| They have this "I'm special" syndrome where they think
| they can get away with doing shit weirdly and not offer
| basic features that everyone else does, and the reason
| why I never purchased any of their services again after
| the first month, and had to replace my payment info with
| a throwaway card because they wouldn't let me remove it,
| again unlike everyone else.
| Wowfunhappy wrote:
| I don't think it's hard to understand why a service would
| want to support Google as an identity provider but not
| Apple. Google is probably the most commonly used provider
| out there, at least outside of the enterprise space.
| squeaky-clean wrote:
| > Always best to sign in with your own email address.
|
| Using a randomly generated email per service is a huge
| improvement over always using the same email.
| Razengan wrote:
| > _Always best to sign in with your own email address._
|
| Oh boy
|
| Saying this in 2026 is just.. oh man. just wow
| tstrimple wrote:
| Not really. It's the user's fault. Apple provides an option
| to hide your email, it's not required. It's an option that
| shows up when you're prompted to create an account.
| Wowfunhappy wrote:
| Oh, I agree with this.
|
| My original thinking was that Apple makes it too easy for
| a general audience to hide their email without
| considering the implications (the service won't know your
| email). But of course there's a tension here, since you
| also want the option to be easy and accessible.
|
| The party I do _not_ consider at fault in this case is
| Anthropic.
| short_sells_poo wrote:
| They clearly vibe code a lot (most? all?) of their stuff, and
| it shows. Elementary features are broken regularly and while I
| appreciate them trying new features, I'd appreciate it more if
| existing ones were reliable and promptly fixed if broken.
| mikkupikku wrote:
| Trouble is, vibe coding refinements and bug fixes works well
| but probably isn't a good track to promotions at Anthropic
| (or virtually any other company.)
| ericmcer wrote:
| They also pay... insane salaries, like double industry average.
| That coupled with an IPO on the horizon means they probably
| have their pick of engineers.
| dheera wrote:
| Their interviews are actually very much focused on how fast
| you can code something that works.
| brcmthrowaway wrote:
| You have 23 year olds earning $2mn/year, at least it isn't in
| HFT though!
| throwatdem12311 wrote:
| Have you actually used their products? They are janky, full of
| bugs and barely work half the time.
|
| They write 100% of their code with Claude. Some of their
| engineers apparently burn over 100k worth of tokens per month.
|
| It's not surprising they ship fast at all when the product is
| actually falling apart at the seams and they just vibe code
| everything.
| mchusma wrote:
| I think to the extent they are making a speed v quality
| tradeoff, I think they are making the right call. 10x speed
| over quality any day for me. Reminds me of:
|
| "If brute force doesn't work, you aren't using enough of it."
| - Isaac Arthur
| throwatdem12311 wrote:
| Everyone is making this tradeoff now. Surely nothing bad
| could come from it.
|
| In the meantime I can't even continue a Claude Code session
| I started on desktop on my phone. What's the point of
| shipping a billion features of they are all half baked?
| WhrRTheBaboons wrote:
| the point is VC money
| simosmik wrote:
| It's a phase, for sure things will turn around in the
| future once the hype of "oh we can now ship fast" is
| gone.
|
| fwiw I've had this open source browser ui that sits on
| top of your claude code, gemini and codex and picks
| up/starts your sessions from any device
| https://github.com/siteboon/claudecodeui
| bmurphy1976 wrote:
| Everyone is NOT making that tradeoff. Maybe we will be
| forced into it someday, but my team is leveraging AI to
| increase the quality of code far beyond what we would
| have done without it. Some of us are using it to engineer
| better solutions.
|
| Example: we are putting a lot of energy into removing
| technical debt, reorganizing the code to remove unneeded
| abstraction and complexity, and creating missing tests
| and automation. We're not just burping out new untested
| and poorly reviewed functionality.
| davesque wrote:
| The Claude Code TUI app is pretty solid. I use it heavily and
| I get great results from it. But with the mobile app, Claude
| Code remote is basically unusable (weird disconnect bugs) and
| Claude Code cloud has issues as well (UI hides approval
| confirmations; must reconnect to see them). So yeah, I
| imagine what you're saying is true. There are at least some
| major gaps in their QA process. It's ironically a pretty
| convincing case to keep humans in the loop. It's honestly
| shocking to me that those features were actually shipped in
| their current state. You run into the problems immediately.
| not_ai wrote:
| I have a very different experience. Claude code tui is the
| worst tui I have ever used. How is it possible that an
| inactive tui regularly eats 8gb of ram, has freezing issues
| and rendering issues?
|
| If I wasn't forced to use it I wouldn't as there are better
| options available.
| fixxation92 wrote:
| I agree with you about the Claude Code TUI. I switched to
| it weeks after it was released. The browser interface is
| great for quick chats and talking through ideas/concepts,
| but not for coding. What I love about the TUI is that it
| can see all your repos as once so it has the full picture
| all the time. You can't get that with the web version.
| bmurphy1976 wrote:
| I don't find criticism like this particularly compelling.
| Most products (written by humans) have the same failings. The
| few that aren't are exceptions to the rule or develop very
| very slowly and carefully.
| asim wrote:
| It was inevitable until the point all apps will disappear and AI
| will be the entry point for all work. You can see how anything
| required appear based on a single request. After which world
| models and other forms of interaction that are more dynamic will
| make sense and we'll need something that's not a screen.
| joshribakoff wrote:
| Its a large leap from "we made a config driven diagram tool and
| trained an llm on that config" to "all apps will disappear". If
| you're predicting such grand claims please be more precise than
| "AI" which is a term we cant define.
| bogzz wrote:
| You're harshing the vibe, man.
| abnercoimbre wrote:
| For sure, leave the hype profiteers alone!
| elliotbnvl wrote:
| And the people that still get excited about life!
| ericmcer wrote:
| Yeah an app doesn't "disappear" because you put an AI
| interface in front of it and then use a bunch of old school
| programming to parse LLM output and feed that into your old
| app. 99% of the work is still building the old app.
| shiftyck wrote:
| Claude is broken for me since this was released, prompts are just
| timing out and stopping after 10 attempts
| gkfasdfasdf wrote:
| I would love to know how they built this. Did they use json-
| render [0], openui [1], or rolled their own?
|
| [0]: https://github.com/vercel-labs/json-render
|
| [1]: https://github.com/thesysdev/openui
| gavinray wrote:
| Right-click the page and inspect the source code?
| atonse wrote:
| Anyone else able to use Claude with Excel? I've tried adding it
| to our (very small) Office365 org and it just fails. Been failing
| for months.
| alansaber wrote:
| All the office js integrations are still pretty shitty
| captainbland wrote:
| I feel like this is a feature which improves the perceived
| confidence of the LLM but doesn't do much for correctness of
| other outputs, i.e. an exacerbation of the "confidently
| incorrect" criticism.
| nerdjon wrote:
| This was my first thought as well, all this does is further
| remove the user from seeing the chat output and instead makes
| it appear as if the information is concretely reliable.
|
| I mean is it really that shocking that you can have an LLM
| generate structured data and shove that into a visualizer? The
| concern is if is reliable, which we know it isnt.
| j45 wrote:
| Its' a reasonable concern. Often it can be mitigated by
| prompting in a manner that invokes research and verification
| instead of defaulting to a corpus.
|
| Passive questions generate passive responses.
| ericmcer wrote:
| The further they can get people from the reality of `This
| just spits out whatever it thinks the next token will be` the
| more they can push the agenda.
| elliotbnvl wrote:
| It's a usability / quality of life feature to me. Nothing to do
| with increasing perceived confidence. I guess it depends on how
| much you already (dis)trust LLMs.
|
| I'm finding more and more often the limiting factor isn't the
| LLM, it's my intuition. This goes a way towards helping with
| that.
| vunderba wrote:
| A similar thing happened when Google started really pushing
| generating flowcharts as a use-case with Nano Banana. A slick
| presentation can distract people from the only thing that
| really matters - the accuracy of the underlying data.
| Angostura wrote:
| As a slightly different tack, I've been using Copilot to
| generate flowcharts from some of the fiendishly complex (and
| badly written) standard operating procedures we have at work.
|
| People find them quite easy to check - easier than the raw
| document. My angle with teams is use these to check your
| processes. If the flow is wrong it's either because the LLM
| has screwed up, or because the policy is wrong/badly written.
| It's usually the latter. It's a good way to fix SOPs
| vunderba wrote:
| It's interesting you mentioned that. One of the things I've
| started doing recently is throwing a large LLM such as
| codex-5.3 (highest level of reasoning) at some of the more
| complex systems we have to produce nicely formatted ASCII
| diagrams.
|
| I still review each diagram afterward, but the great thing
| is that, unlike image-based diagrams, they remain fully
| text-readable and searchable. And you can even expose them
| as part of the knowledge base for the LLM to reference when
| needed going forward.
| programmertote wrote:
| A recent LinkedIn post that I came across as an example of
| people trusting (or learning to trust) AI too much while not
| realizing that it can make up numbers too:
| https://www.linkedin.com/posts/mariamartin1728_claude-wrote-...
|
| P.S. Credit to the poster, she posted a correction note when
| someone caught the issue:
| https://www.linkedin.com/posts/mariamartin1728_correction-on...
| Avamander wrote:
| > A recent LinkedIn post that I came across as an example of
| people trusting (or learning to trust) AI too much while not
| realizing that it can make up numbers too
|
| Honestly, people make them up just as much or generate
| equally incorrect graphs.
|
| It's about time our trust into random visualizations is
| destroyed, without the actual formulas and data behind being
| exposed.
| mikkupikku wrote:
| I agree. Maybe next they'll add emotionally evocative music,
| with swelling orchestral bits when you reach the exciting
| climate of the slop.
| kemayo wrote:
| It's a mismatch with our intuition about how much effort things
| take.
|
| If there's humans involved, "I took this data and made a really
| fancy interactive chart" means that you put _a lot_ more work
| into it, and you can probably _somewhat_ assume that this means
| some more effort was also put into the accuracy of the data.
|
| But with the LLM it's not really very much more work to get the
| fancy chart. So the thing that was a signifier of effort is now
| misleading us into trusting data that got no extra effort.
|
| (Humans have been exploiting this tendency to trust fancy
| graphics forever, of course.)
| manquer wrote:
| It is not limited to graphics, better packaged products,
| better dressed / good looking well spoken person and so on.
| Celebrity endorsements depend on this thesis.
|
| There has always been a bias towards form over function.
| ipython wrote:
| Already happened. :)
|
| https://www.reddit.com/r/dataisugly/comments/1mk5wdb/this_ch...
| outlore wrote:
| I suspect chain of thought while building the chart will
| improve the overall correctness of the answer
| jameschaearley wrote:
| This is the tension I keep hitting when building data tools on
| top of LLMs. A nice-looking chart makes the output feel more
| trustworthy, but the data can still be wrong. The chart just
| makes it harder to notice. LLMs still need to come with
| receipts of where the data came from and the math they did.
| It's as bad as "I read the headline so I know everything in the
| article."
| atonse wrote:
| Wow, I asked it to build me a simple diagram explaining agile
| development and it did an amazing job. Wow it felt magical to
| watch that diagram slowly animating to life.
|
| Like a much prettier version of Mermaid.
|
| Kudos, Anthropic. Geez, this is so nice.
|
| Now I'm going to ask it to draw a diagram of a pelican riding a
| bicycle, why not?
| drewda wrote:
| When using Claude Code, we often prompt it to draft diagrams in
| MermaidJS syntax.
|
| Great for summarizing a multi-step process and quick to render
| with simple tools.
| JoshGG wrote:
| This is pretty neat and I am experimenting with it now, but
| hasn't ChatGPT had capability to create graphs and interact with
| data for a while? "ChatGPT advanced data analysis" for example.
| I'm asking in good faith as maybe some of you have been using
| that and can compare the two and give an informed opinion.
|
| I usually use a lot of other tools for data analysis or write
| code with Claude code or another LLM to do data analysis and
| visualization.
|
| article about the ChatGPT charts and graphs
| https://www.zdnet.com/article/how-to-use-chatgpt-to-make-cha...
| Gareth321 wrote:
| > but hasn't ChatGPT had capability to create graphs and
| interact with data for a while?
|
| It's pretty bad (for me). I have to use extremely prescriptive
| language to tell ChatGPT what to create. Even down to the
| colours in the chart, because otherwise it puts black font on
| black background (for example). Then I have to specifically
| tell it to put it in a canvas, and make it interactive, and
| make it executable in the canvas. Then if I'm lucky I have to
| hit a "preview" button in the top right and hope it works (it
| doesn't). I could write several paragraphs telling it to do
| something like what Claude just demo'd and it wouldn't come
| close. I'm trying Claude now for financial insights and it's
| effortless with beautiful UX.
|
| For posterity, Gemini is pretty good with these interactive
| canvases. Not nearly as good, but FAR better than ChatGPT.
| HotGarbage wrote:
| Interactive slop is still slop.
| razerbeans wrote:
| Interesting. So if I'm reading this correctly, this is distinctly
| different than the artifacts that Claude creates? If that's the
| case, why create it inline as opposed to an artifact? Any time I
| get a visual, I tend to find them so useful that I _want_ them to
| be an artifact that I can export and share.
| I_am_tiberius wrote:
| Does anyone know which library they use? Or something developed
| internally?
| I_am_tiberius wrote:
| Ok, asked it myself: Chart.js
| czk wrote:
| I tried the periodic table in their examples using sonnet 4.6 on
| the $20/mo plan. After a few minutes Claude told me it reached
| the max message length and bailed. I pressed continue and
| eventually it generated the table, but it wasn't inline, it was a
| jsx artifact, and I've now hit my daily usage limit.
| data-ottawa wrote:
| I'm intermittently getting artifacts vs the new visuals api,
| depending on which version of the Claude app I use. iOS/iPadOS
| apps are not yet supporting the visualization API, and I don't
| see an app-store update yet.
| karussell wrote:
| It wasn't quick but I still found it fast enough. In my case I
| could even download it as an html file:
| https://gist.github.com/karussell/289aeb621a71597babd6f97eb2...
|
| edit: claude just confirmed the initial version has a bug and
| 104-117 are not visible
| jzig wrote:
| Unable to reproduce the recipe image in the iOS app. It first
| gave a normal text answer. Then when referencing this blog post
| it produced a wonky HTML artifact.
| mehdibl wrote:
| Isn't this mainly a skill injected in the context? Rather a
| model/platform specific feature?
| Gareth321 wrote:
| I asked it to do some portfolio analysis for me and it created
| BEAUTIFUL, tabbed, interactive charts UNPROMPTED. This is kind of
| magical. The charts were not just beautiful, but actually super
| useful in understanding the data faster. I honestly could not
| have produced those in a week if you asked me to.
| wuweiaxin wrote:
| Reliability has been the real bottleneck for multi-agent setups
| in production. The hard part isnt getting one agent to do
| something clever once - its making repeated runs observable and
| bounded when tools fail halfway through. Idempotency checks,
| explicit handoff state, and human review gates have mattered more
| for us than adding another model or another agent role.
| alansaber wrote:
| New account spamming LM replies, nice
| wuweiaxin wrote:
| The artifact output model is more useful than it looks at first.
| We use Claude in a multi-agent pipeline and discovered that
| structured artifact outputs reduce parse errors significantly
| compared to freeform text responses -- the model seems to reason
| differently when it knows the output will be rendered. Curious
| whether Anthropic sees similar quality improvements in tool-use
| tasks when the output has a concrete format constraint.
| darepublic wrote:
| When I ask chatgpt to create a mermaid diagram for me it
| regularly will add new lines to certain labels that will break
| the parse. If you then feed the parse error back to it the second
| version is always correct And it seems to exactly know the
| problem. There are some other examples where it will almost
| always get it wrong the first time but right if nudged to correct
| itself. I wonder what the underlying cause is
| stefan_ wrote:
| Today I asked Claude to create me a squidward looking out the
| window meme and it started generating HTML & CSS to draw
| squidward in a style best described as "4 year old
| preschooler". Not quite it yet.
| qingcharles wrote:
| The issue for Claude is that Anthropic don't have an imagen
| that I know of, so the only tool available for the LLM to
| draw something is to start doing vector stuff in CSS, which
| is very hard for it (see the pelicans).
|
| Gemini, ChatGPT or Grok would find this a lot easier as they
| could gen an image inline, although IP restrictions might
| bite you. Even Grok wants to lecture on IP these days, but at
| least it's fairly trivial to jailbreak.
| overfeed wrote:
| > I wonder what the underlying cause is
|
| It responds with the statistically most probable text based on
| its training data, which happens to be different with the
| errors vs without. I suspect high-fidelity diagramming requires
| a different attention architecture from the common ones used in
| sentence-optimized models.
| deckar01 wrote:
| Mermaid is really bad about cutting off text after spaces, so
| you have to insert <br>s everywhere. I'm guessing this is
| getting rendered instead of escaped by your interface. Or just
| lost in translation at the tokenizer.
| ar0b wrote:
| "Prompt Repetition Improves Non-Reasoning LLMs " -
| https://arxiv.org/pdf/2512.14982
|
| What instance of ChatGPT are you doing that with? (Reasoning?)
| darepublic wrote:
| Observed from 5.2, on chatgpt.com. earlier versions did
| worse.. as in, they might take a few prompts to generate a
| parseable syntax. Newer versions just usually deliver one
| unparseable version then get it right second try. Likely I
| could prompt engineer to one shot but I think I would always
| need the specific warning about newlines.
| alex_duf wrote:
| I don't think it's about repeating the instructions, but
| rather providing feedback as to why it's not working.
|
| I've noticed the same thing when creating an agentic loop, if
| the model outputs a syntax error, just automatically feed it
| back to the LLM and give it a second chance. It dramatically
| increases the success rate.
| quintu5 wrote:
| This is one of the issues I've attempted to tackle with the
| Mermaid Studio plugin for IntelliJ.
|
| It provides both syntax guides and syntax/semantic analysis as
| MCP Tools, so you can have an agent iteratively refine diagrams
| with good context for patterns like multi-line text and
| comments (LLMs love end-of-line comments, but Mermaid.js often
| doesn't).
| dworks wrote:
| I think the problem should be defined as "why does it not loop
| back the errors from the first attempt so it can fix it on the
| second attempt" rather than why it fails to produce a fully
| correct implementation on the first pass.
| groby_b wrote:
| Aaand all the way at the bottom, there it is. The first glimpse
| of what will be an ad carousel.
|
| (Literally nobody needs an image of a cake when asking for a cake
| recipe)
| alansaber wrote:
| What if you've never seen a cake before?
| groby_b wrote:
| That's the target group for the feature. You're right. You
| got me.
| johsole wrote:
| love to see it, my auto researcher is getting more capable with
| less effort every release
| smusamashah wrote:
| I want to point out to everyone that Claude bullshits the least
| of all top models. Even Claude's lowest version rank above lots
| of other top models.
|
| https://petergpt.github.io/bullshit-benchmark/viewer/index.v...
| w10-1 wrote:
| Chat --> Notebook: Jupyter is so much more functional than slack
| for communicating real work product!
|
| Next up: exporting or sharing selections from the chat as a
| document or interactive page. If they allow share with non-
| subscribers, subscriptions could hockey stick -- particularly if
| the document/page included prompts necessary to replicate (or
| modify and adapt).
| tamimio wrote:
| I remember months ago I asked it to make a diagram and it wrote a
| html/js for it, and it was interactive, is this different than
| that?!
| hudtaylor wrote:
| I run an AI agent that deploys production code across 4 live
| products. The self-correction capability is real but undersold in
| these debates.
|
| Examples from the last month: the agent found its own API keys in
| environment files after initially claiming it didn't have them
| (lesson: grep before asking). It caught itself about to run a
| destructive database migration on a shared production instance
| and stopped. It fixed 8 broken RSS feed configurations that had
| been silently failing for weeks without anyone noticing.
|
| The pattern I've found: AI doesn't need to be perfect at writing
| code. It needs to be honest about what it doesn't know,
| aggressive about testing its own work, and operating under clear
| constraints about what's destructive vs. safe. We maintain a file
| called AGENTS.md with "sacred rules" -- things the agent can
| never do without explicit approval. Database migrations, pricing
| changes, anything with --accept-data-loss.
|
| The "no LLM" stance makes sense if you don't have guardrails.
| With the right constraints, AI-assisted code is faster AND safer
| than solo human development -- because the agent never gets
| tired, never rushes before a deadline, and never thinks "I'll
| test that later."
| rafaelmn wrote:
| > The "no LLM" stance makes sense if you don't have guardrails.
| With the right constraints, AI-assisted code is faster AND
| safer than solo human development -- because the agent never
| gets tired, never rushes before a deadline, and never thinks
| "I'll test that later."
|
| Ironically I had a bunch of cases recently where CC would stop
| saying stuff like "this test problem is unrelated/a pre-
| existing issue when it had no proof of this and it was clearly
| not true (the branch built/tests were passing before the LLM
| changes).
| abrookewood wrote:
| If you're relying on an AGENTS.md file to stop your AI agents
| from running a destructive action, I think you are sitting on a
| time bomb.
___________________________________________________________________
(page generated 2026-03-12 23:01 UTC)