[HN Gopher] Speed at the cost of quality: Study of use of Cursor...
___________________________________________________________________
Speed at the cost of quality: Study of use of Cursor AI in open
source projects (2025)
Author : wek
Score : 82 points
Date : 2026-03-16 17:07 UTC (5 hours ago)
(HTM) web link (arxiv.org)
(TXT) w3m dump (arxiv.org)
| rfw300 wrote:
| Super interesting study. One curious thing I've noticed is that
| coding agents tend to increase the code complexity of a project,
| but simultaneously _massively reduce_ the cost of that code
| complexity.
|
| If a module becomes unsustainably complex, I can ask Claude
| questions about it, have it write tests and scripts that
| empirically demonstrate the code's behavior, and worse comes to
| worst, rip out that code entirely and replace it with something
| better in a fraction of the time it used to take.
|
| That's not to say complexity isn't bad anymore--the paper's
| findings on diminishing returns on velocity seem well-grounded
| and plausible. But while the newest (post-Nov. 2025) models often
| make inadvisable design decisions, they rarely do things that are
| outright wrong or hallucinated anymore. That makes them much more
| useful for cleaning up old messes.
| joshribakoff wrote:
| Bad code has real world consequences. Its not limited to having
| to rewrite it. The cost might also include sanctions, lost
| users, attrition, and other negative consequences you don't
| just measure in dev hours
| SR2Z wrote:
| Right, but that cost is also incurred by human-written code
| that happens to have bugs.
|
| In theory experienced humans introduce less bugs. That sounds
| reasonable and believable, but anyone who's ever been paid to
| write software knows that _finding reliable humans_ is not an
| easy task unless you 're at a large established company.
| verdverm wrote:
| There was a recent study posted here that showed AI
| introduces regressions at an alarming rate, all but one
| above 50%, which indicates they spend a lot of time fixing
| their own mistakes. You've probably seen them doing this
| kind of thing, making one change that breaks another, going
| and adjusting that thing, not realizing that's making
| things worse.
| MeetingsBrowser wrote:
| The question then becomes, can LLMs generate code close to
| the same quality as professionals.
|
| In my experience, they are not even close.
| mathgeek wrote:
| We should qualify that kind of statement, as it's
| valuable to define just what percentile of "professional
| developers" the quality falls into. It will likely never
| replace p90 developers for example, but it's better than
| somewhere between there and p10. Arbitrary numbers for
| examples.
| MeetingsBrowser wrote:
| Can you quantify the quality of a p90 or p10 developer?
|
| I would frame it differently. There are developers
| successfully shipping product X. Those developer are, on
| average, as skilled as necessary to work on project X.
| else they would have moved on or the project would have
| failed.
|
| Can LLMs produce the same level of quality as project X
| developers? The only projects I know of where this is
| true are toy and hobby projects.
| mathgeek wrote:
| > Can you quantify the quality of a p90 or p10 developer?
|
| Of course not, you have switched "quality" in this
| statement to modify the developer instead of their work.
| Regarding the work, each project, as you agree with me on
| from your reply, has an average quality for its code.
| Some developers bring that down on the whole, others
| bring it up. An LLM would have a place somewhere on that
| spectrum.
| MeetingsBrowser wrote:
| This only helps if you notice the code is bad. Especially in
| overlay complex code, you have to really be paying attention to
| notice when a subtle invariant is broken, edge case missed,
| etc.
|
| Its the same reason a junior + senior engineer is about as fast
| as a senior + 100 junior engineers. The senior's review time
| becomes the bottleneck and does not scale.
|
| And even with the latest models and tooling, the _quality_ of
| the code is below what I expect from a junior. But you sure can
| get it fast.
| phillipclapham wrote:
| This is the most important point in the thread. The study
| measures code complexity but the REAL bottleneck is cognitive
| load (and drain) on the reviewer.
|
| I've been doing 10-12 hour days paired with Claude for
| months. The velocity gains are absolutely real, I am shipping
| things I would have never attempted solo before AI and
| shipping them faster then ever. BUT the cognitive cost of
| reviewing AI output is significantly higher than reviewing
| human code. It's verbose, plausible-looking, and wrong in
| ways that require sustained deep attention to catch.
|
| The study found "transient velocity increase" followed by
| "persistent complexity increase." That matches exactly. The
| speed feels incredible at first, then the review burden
| compounds and you're spending more time verifying than you
| saved generating.
|
| The fix isn't "apply traditional methods" -- it's recognizing
| that AI shifts the bottleneck from production to
| verification, and that verification under sustained cognitive
| load degrades in ways nobody's measuring yet. I think I've
| found some fixes to help me personally with this and for me
| velocity is still high, but only time will tell if this
| remains true for long.
| chrisweekly wrote:
| > _The study found "transient velocity increase" followed
| by "persistent complexity increase."_
|
| Companies facing this reality are of course typically going
| to use AI to help manage the increased complexity. But that
| leads quickly to AI becoming a crutch, without which even
| basic maintenance could pose an insurmountable challenge.
| galbar wrote:
| >The fix isn't "apply traditional methods"
|
| I would argue they are. Those traditional methods aim at
| keeping complexity low so that reading code is easier and
| requires less effort, which accelerates code review.
| tabwidth wrote:
| The part that gets me is when it passes lint, passes tests,
| and the logic is technically correct, but it quietly
| changed how something gets called. Rename a parameter. Wrap
| a return value in a Promise that wasn't there before. Add
| some intermediate type nobody asked for. None of that shows
| up as a failure anywhere. You only notice three days later
| when some other piece of code that depended on the old
| shape breaks in a way that has nothing to do with the
| original change.
| i_love_retros wrote:
| > have it write tests
|
| Just make sure it hasn't mocked so many things that nothing is
| actually being tested. Which I've witnessed.
| moregrist wrote:
| I've also seen Opus 4.5 and 4.6 churn out tons of essentially
| meaningless tests, including ones where it sets a field on a
| structure and then tests that the field was set.
|
| You have to actually care about quality with these power saws
| or you end up with poorly-fitting cabinets and might even
| lose a thumb in the process.
| teaearlgraycold wrote:
| The first thing you should do after having them write tests
| is delete half of the tests.
| AlexandrB wrote:
| > Super interesting study. One curious thing I've noticed is
| that coding agents tend to increase the code complexity of a
| project, but simultaneously massively reduce the cost of that
| code complexity.
|
| This is the same pattern I observed with IDEs. Autocomplete and
| being able to jump to a definition means spaghetti code can be
| successfully navigated so there's no "natural" barrier to
| writing spaghetti code.
| PeterStuer wrote:
| Interesting from an historical perspective. But data from 4/2025?
| Might as well have been last century.
| happycube wrote:
| I think the gist of it still applies to even Claude Code w/Opus
| 4.6.
|
| It's basically outsourcing to mediocre programmers - albeit
| very fast ones with near-infinite patience and little to no
| ego.
| Miraste wrote:
| It doesn't map well to a mediocre human programmer, I think.
| It operates in a much more jagged world between superhuman,
| and inhuman stupidity.
| andai wrote:
| A data center of geniuses on a medium dose of LSD.
| matt_heimer wrote:
| Yes, it's not surprising that warnings and complexity increased
| at a higher rate when paired with increased velocity. Increased
| velocity == increased lines of code.
|
| Does the study normalize velocity between the groups by adjusting
| the timeframes so that we could tell if complexity and warnings
| increased at a greater rate per line of code added in the AI
| group?
|
| I suspect it would, since I've had to simplify AI generated code
| on several occasions but right now the study just seems to say
| that the larger a code base grows the more complex it gets which
| is obvious.
| ex-aws-dude wrote:
| That was my thought as well, because obviously complexity
| increases when a project grows regardless of AI
| bensyverson wrote:
| Yeah, I have a more complex project I'm working on with
| Claude, but it's not that Claude is making it more complex;
| it's just that it's so complex I wouldn't attempt it without
| Claude.
| AstroBen wrote:
| "Notably, increases in codebase size are a major determinant of
| increases in static analysis warnings and code complexity, and
| absorb most variance in the two outcome variables. However,
| even with strong controls for codebase size dynamics, the
| adoption of Cursor still has a significant effect on code
| complexity, leading to a 9% baseline increase on average
| compared to projects in similar dynamics but not using Cursor."
| AstroBen wrote:
| > On average, Cursor adoption has a modestly significant positive
| impact on development velocity, particularly in terms of code
| production volume: Lines added increase by about 28.6% (Table 2).
| There is no statistically significant effect for the volume of
| commits.
|
| This doesn't equate to a faster development speed in my eyes? We
| know that AI code is incredibly verbose.
|
| More lines of code doesn't equate to faster development - even
| more so when you're comparing apples (human written) to oranges
| (AI written)
| AstroBen wrote:
| They're measuring development speed through lines of code. To
| show that's true they'd need to first show that AI and humans use
| the same number of lines to solve the same problem. That hasn't
| been my experience at all. AI is incredibly verbose.
|
| Then there's the question of if LoC is a reliable proxy for
| velocity _at all_? The common belief amongst developers is that
| it 's not.
| otabdeveloper4 wrote:
| > They're measuring development speed through lines of code.
|
| Yeah, this is the biggest facepalm.
|
| Didn't we grow out of this idiocy 40 years ago? This shit
| again? Really?
| andai wrote:
| See also
|
| -2000 lines of code
|
| https://news.ycombinator.com/item?id=26387179
|
| This is actually one thing I have found LLMs surprisingly
| useful for.
|
| I give them a code base which has one or two orders of
| magnitude of bloat, and ask them to strip it away iteratively.
| What I'm left with usually does the same thing.
|
| At this point the code base becomes small enough to navigate
| and study. Then I use it for reference and build my own
| solution.
| mellosouls wrote:
| Depends on the nature of the tool I would imagine - eg. Claude
| Code Terminal (say) would have higher entry requirements in terms
| of engineering experience (Cursor was sold as newbie-friendly) so
| I would predict higher quality code than Cursor in a similar
| survey.
|
| ofc that doesn't take into account the useful high-level and
| other advantages of IDEs that might mitigate against slop during
| review, but overall Cursor was a more natural fit for vibe-
| coders.
|
| This is said without judgement - I was a cheerleader for Cursor
| early on until it became uncompetitive in value.
| mentalgear wrote:
| > We find that the adoption of Cursor leads to a statistically
| significant, large, but transient increase in project-level
| development velocity, along with a substantial and persistent
| increase in static analysis warnings and code complexity. Further
| panel generalized-method-of-moments estimation reveals that
| increases in static analysis warnings and code complexity are
| major factors driving long-term velocity slowdown. Our study
| identifies quality assurance as a major bottleneck for early
| Cursor adopters and calls for it to be a first-class citizen in
| the design of agentic AI coding tools and AI-driven workflows.
|
| So overall seems like the pros and cons of "AI vibe coding" just
| cancel themselves out.
| mort96 wrote:
| The part you quoted doesn't support your conclusion. Per your
| quoted paragraph, the benefit of "AI vibe coding" is a large,
| but transient (i.e temporary) increase in development velocity;
| while the drawback is a persistent increase in static analysis
| warnings and code complexity.
|
| To me, this sounds like after the transient increase of
| velocity has died down, you're left with the same development
| velocity as you had when you started, but a significantly worse
| code base.
| andai wrote:
| The implication seems to be that if quality assurance is
| prioritized, the negative impact would be eliminated.
|
| This seems to assume the main cause is the accumulation of
| defects due to lack of static analysis and testing.
|
| I think a more likely cause is, the code begins to rapidly
| grow beyond the maintainers' comprehension. I don't think
| there is a technical solution for that.
| chris_money202 wrote:
| Now someone do a research study where a summary of this research
| paper is in the AGENTS.md and let's see if the overall outcomes
| are better
| dalemhurley wrote:
| I think the issue is people AI assisted code, test then commit.
|
| Traditional software dev would be build, test, refactor, commit.
|
| Even the Clean Coder recommends starting with messy code then
| tidying it up.
|
| We just need to apply traditional methods to AI assisted coding.
| duendefm wrote:
| AI is not perfect sure, one has to know how to use it. But this
| study is already flawed since models improved a lot since the
| beginning of 2026.
| Eufrat wrote:
| This is not a useful, constructive or meaningful statement.
|
| Attempting to claim the models are the future by perpetually
| arguing their limitations are because people are using the
| models wrong or that the argument has been invalidated because
| the new model fixes it might as well be part of the training
| data since Claude Opus 3.5.
| duendefm wrote:
| No no, I didn't say that at all. I'm just saying that the
| studies are irrelevant since models got a boost in their
| competence. I'm not in a fight pro or against llm's, I know
| how they work and their limitations. But the complexity of
| the problems they solve increased since opus 4.5 . If you
| can't admit that, it's your problem.
|
| Also, I'm not blaming users for their shortcomings. I'm just
| saying they are not perfect but you can get different
| outcomes according to how you use them.
| keeda wrote:
| There are actually quite a few studies out there that look at LLM
| code quality (e.g.
| https://scholar.google.com/scholar?hl=en&as_sdt=0%2C5&q=LLM+...)
| and they mostly have similar findings. This reinforces the idea
| that LLMs still require expert guidance. Note, some of these
| studies date back to 2023, which is eons ago in terms of LLM
| progress.
|
| The conclusion of this paper aligns with the emerging
| understanding that AI is simply an amplifier of your existing
| quality assurance processes: Higher discipline results in higher
| velocity, lower discipline results in lower stability (e.g.
| https://dora.dev/research/2025/) Having strong feedback and
| validation loops is more critical than ever.
|
| In this paper, for instance, they collected static analysis
| warnings using a local SonarQube server, which implies that it
| was not integrated into the projects they looked at. As such
| these warnings were not available to the agent. It's highly
| likely if these warnings were fed back into the agent it would
| fix them automatically.
|
| Another interesting thing they mention in the conclusion: the
| metrics we use for humans may not apply to agents. My go-to
| example for this is code duplication (even though this study
| finds minimal increase in duplication) -- it may actually be
| better for agents to rewrite chunks of code from scratch rather
| than use a dependency whose code is not available forcing it to
| instead rely on natural language documentation, which may or may
| not be sufficient or even accurate. What is tech debt for humans
| may actually be a boon for agents.
| woeirua wrote:
| This study's cutoff date was August 2025. I don't think this
| result is surprising given the level of coding agent ability back
| then. The whole thing just shows how out-of-date academic
| publishing is on this subject.
|
| >This yields 806 repositories with adoption dates between January
| 2024 and March 2025 that are still available on GitHub at the
| time of data analysis (August 2025).
|
| There were _very few_ people who thought that coding agents
| worked very well back then. I was not one of them, but I _do_
| think they work today.
| monkaiju wrote:
| This is the perennial excuse, and I'm sure we'll continue to
| see it. Folks will say the exact same thing when the current
| crop of slop-generators have been replaced by a newer ilk
| staticassertion wrote:
| Eh, I don't know. I mean, are we seeing _better_ models now? Of
| course. But are they truly leaps and bounds better? No, and I
| get confused by people saying that they are. They 're better
| but not like... 10x better.
|
| And when people were studying ChatGPT 3.5, everyone would go
| "Oh, but that wasn't 4!", and when people talk about Opus 4.5
| they go "4.6 is so much better!".
|
| My personal position right now is that people are extremely bad
| at evaluating model output/ changes in model capabilities.
| Model benchmarks do _not_ reflect the position that models are
| just 10x better than they were a year ago, but with how people
| discuss them you 'd think that 10x was underselling it.
| nicoburns wrote:
| I don't have personal experience, but there seems to be a
| broad consensus that Opus 4.5 was tipping point between
| "kinda bad" and "actually kinda useful".
|
| So a cutoff point of August 2025 just before that is a bit
| unfortunate (I'm sure they'll be newer studies soon).
| sentrysapper wrote:
| Evergreen excuses for tech people desperately want to work. I
| get why, it would give you agency to do other things you WANT
| to do. I tried reviewing a colleagues agent-generated code and
| it was practically unreviewable. I watched him blame himself,
| saying he just needed to adjust a parameter. He tried
| everything except admit the machine does not conceptionaly
| understand what he was asking.
___________________________________________________________________
(page generated 2026-03-16 23:01 UTC)