[HN Gopher] Speed at the cost of quality: Study of use of Cursor...
       ___________________________________________________________________
        
       Speed at the cost of quality: Study of use of Cursor AI in open
       source projects (2025)
        
       Author : wek
       Score  : 82 points
       Date   : 2026-03-16 17:07 UTC (5 hours ago)
        
 (HTM) web link (arxiv.org)
 (TXT) w3m dump (arxiv.org)
        
       | rfw300 wrote:
       | Super interesting study. One curious thing I've noticed is that
       | coding agents tend to increase the code complexity of a project,
       | but simultaneously _massively reduce_ the cost of that code
       | complexity.
       | 
       | If a module becomes unsustainably complex, I can ask Claude
       | questions about it, have it write tests and scripts that
       | empirically demonstrate the code's behavior, and worse comes to
       | worst, rip out that code entirely and replace it with something
       | better in a fraction of the time it used to take.
       | 
       | That's not to say complexity isn't bad anymore--the paper's
       | findings on diminishing returns on velocity seem well-grounded
       | and plausible. But while the newest (post-Nov. 2025) models often
       | make inadvisable design decisions, they rarely do things that are
       | outright wrong or hallucinated anymore. That makes them much more
       | useful for cleaning up old messes.
        
         | joshribakoff wrote:
         | Bad code has real world consequences. Its not limited to having
         | to rewrite it. The cost might also include sanctions, lost
         | users, attrition, and other negative consequences you don't
         | just measure in dev hours
        
           | SR2Z wrote:
           | Right, but that cost is also incurred by human-written code
           | that happens to have bugs.
           | 
           | In theory experienced humans introduce less bugs. That sounds
           | reasonable and believable, but anyone who's ever been paid to
           | write software knows that _finding reliable humans_ is not an
           | easy task unless you 're at a large established company.
        
             | verdverm wrote:
             | There was a recent study posted here that showed AI
             | introduces regressions at an alarming rate, all but one
             | above 50%, which indicates they spend a lot of time fixing
             | their own mistakes. You've probably seen them doing this
             | kind of thing, making one change that breaks another, going
             | and adjusting that thing, not realizing that's making
             | things worse.
        
             | MeetingsBrowser wrote:
             | The question then becomes, can LLMs generate code close to
             | the same quality as professionals.
             | 
             | In my experience, they are not even close.
        
               | mathgeek wrote:
               | We should qualify that kind of statement, as it's
               | valuable to define just what percentile of "professional
               | developers" the quality falls into. It will likely never
               | replace p90 developers for example, but it's better than
               | somewhere between there and p10. Arbitrary numbers for
               | examples.
        
               | MeetingsBrowser wrote:
               | Can you quantify the quality of a p90 or p10 developer?
               | 
               | I would frame it differently. There are developers
               | successfully shipping product X. Those developer are, on
               | average, as skilled as necessary to work on project X.
               | else they would have moved on or the project would have
               | failed.
               | 
               | Can LLMs produce the same level of quality as project X
               | developers? The only projects I know of where this is
               | true are toy and hobby projects.
        
               | mathgeek wrote:
               | > Can you quantify the quality of a p90 or p10 developer?
               | 
               | Of course not, you have switched "quality" in this
               | statement to modify the developer instead of their work.
               | Regarding the work, each project, as you agree with me on
               | from your reply, has an average quality for its code.
               | Some developers bring that down on the whole, others
               | bring it up. An LLM would have a place somewhere on that
               | spectrum.
        
         | MeetingsBrowser wrote:
         | This only helps if you notice the code is bad. Especially in
         | overlay complex code, you have to really be paying attention to
         | notice when a subtle invariant is broken, edge case missed,
         | etc.
         | 
         | Its the same reason a junior + senior engineer is about as fast
         | as a senior + 100 junior engineers. The senior's review time
         | becomes the bottleneck and does not scale.
         | 
         | And even with the latest models and tooling, the _quality_ of
         | the code is below what I expect from a junior. But you sure can
         | get it fast.
        
           | phillipclapham wrote:
           | This is the most important point in the thread. The study
           | measures code complexity but the REAL bottleneck is cognitive
           | load (and drain) on the reviewer.
           | 
           | I've been doing 10-12 hour days paired with Claude for
           | months. The velocity gains are absolutely real, I am shipping
           | things I would have never attempted solo before AI and
           | shipping them faster then ever. BUT the cognitive cost of
           | reviewing AI output is significantly higher than reviewing
           | human code. It's verbose, plausible-looking, and wrong in
           | ways that require sustained deep attention to catch.
           | 
           | The study found "transient velocity increase" followed by
           | "persistent complexity increase." That matches exactly. The
           | speed feels incredible at first, then the review burden
           | compounds and you're spending more time verifying than you
           | saved generating.
           | 
           | The fix isn't "apply traditional methods" -- it's recognizing
           | that AI shifts the bottleneck from production to
           | verification, and that verification under sustained cognitive
           | load degrades in ways nobody's measuring yet. I think I've
           | found some fixes to help me personally with this and for me
           | velocity is still high, but only time will tell if this
           | remains true for long.
        
             | chrisweekly wrote:
             | > _The study found "transient velocity increase" followed
             | by "persistent complexity increase."_
             | 
             | Companies facing this reality are of course typically going
             | to use AI to help manage the increased complexity. But that
             | leads quickly to AI becoming a crutch, without which even
             | basic maintenance could pose an insurmountable challenge.
        
             | galbar wrote:
             | >The fix isn't "apply traditional methods"
             | 
             | I would argue they are. Those traditional methods aim at
             | keeping complexity low so that reading code is easier and
             | requires less effort, which accelerates code review.
        
             | tabwidth wrote:
             | The part that gets me is when it passes lint, passes tests,
             | and the logic is technically correct, but it quietly
             | changed how something gets called. Rename a parameter. Wrap
             | a return value in a Promise that wasn't there before. Add
             | some intermediate type nobody asked for. None of that shows
             | up as a failure anywhere. You only notice three days later
             | when some other piece of code that depended on the old
             | shape breaks in a way that has nothing to do with the
             | original change.
        
         | i_love_retros wrote:
         | > have it write tests
         | 
         | Just make sure it hasn't mocked so many things that nothing is
         | actually being tested. Which I've witnessed.
        
           | moregrist wrote:
           | I've also seen Opus 4.5 and 4.6 churn out tons of essentially
           | meaningless tests, including ones where it sets a field on a
           | structure and then tests that the field was set.
           | 
           | You have to actually care about quality with these power saws
           | or you end up with poorly-fitting cabinets and might even
           | lose a thumb in the process.
        
             | teaearlgraycold wrote:
             | The first thing you should do after having them write tests
             | is delete half of the tests.
        
         | AlexandrB wrote:
         | > Super interesting study. One curious thing I've noticed is
         | that coding agents tend to increase the code complexity of a
         | project, but simultaneously massively reduce the cost of that
         | code complexity.
         | 
         | This is the same pattern I observed with IDEs. Autocomplete and
         | being able to jump to a definition means spaghetti code can be
         | successfully navigated so there's no "natural" barrier to
         | writing spaghetti code.
        
       | PeterStuer wrote:
       | Interesting from an historical perspective. But data from 4/2025?
       | Might as well have been last century.
        
         | happycube wrote:
         | I think the gist of it still applies to even Claude Code w/Opus
         | 4.6.
         | 
         | It's basically outsourcing to mediocre programmers - albeit
         | very fast ones with near-infinite patience and little to no
         | ego.
        
           | Miraste wrote:
           | It doesn't map well to a mediocre human programmer, I think.
           | It operates in a much more jagged world between superhuman,
           | and inhuman stupidity.
        
             | andai wrote:
             | A data center of geniuses on a medium dose of LSD.
        
       | matt_heimer wrote:
       | Yes, it's not surprising that warnings and complexity increased
       | at a higher rate when paired with increased velocity. Increased
       | velocity == increased lines of code.
       | 
       | Does the study normalize velocity between the groups by adjusting
       | the timeframes so that we could tell if complexity and warnings
       | increased at a greater rate per line of code added in the AI
       | group?
       | 
       | I suspect it would, since I've had to simplify AI generated code
       | on several occasions but right now the study just seems to say
       | that the larger a code base grows the more complex it gets which
       | is obvious.
        
         | ex-aws-dude wrote:
         | That was my thought as well, because obviously complexity
         | increases when a project grows regardless of AI
        
           | bensyverson wrote:
           | Yeah, I have a more complex project I'm working on with
           | Claude, but it's not that Claude is making it more complex;
           | it's just that it's so complex I wouldn't attempt it without
           | Claude.
        
         | AstroBen wrote:
         | "Notably, increases in codebase size are a major determinant of
         | increases in static analysis warnings and code complexity, and
         | absorb most variance in the two outcome variables. However,
         | even with strong controls for codebase size dynamics, the
         | adoption of Cursor still has a significant effect on code
         | complexity, leading to a 9% baseline increase on average
         | compared to projects in similar dynamics but not using Cursor."
        
       | AstroBen wrote:
       | > On average, Cursor adoption has a modestly significant positive
       | impact on development velocity, particularly in terms of code
       | production volume: Lines added increase by about 28.6% (Table 2).
       | There is no statistically significant effect for the volume of
       | commits.
       | 
       | This doesn't equate to a faster development speed in my eyes? We
       | know that AI code is incredibly verbose.
       | 
       | More lines of code doesn't equate to faster development - even
       | more so when you're comparing apples (human written) to oranges
       | (AI written)
        
       | AstroBen wrote:
       | They're measuring development speed through lines of code. To
       | show that's true they'd need to first show that AI and humans use
       | the same number of lines to solve the same problem. That hasn't
       | been my experience at all. AI is incredibly verbose.
       | 
       | Then there's the question of if LoC is a reliable proxy for
       | velocity _at all_? The common belief amongst developers is that
       | it 's not.
        
         | otabdeveloper4 wrote:
         | > They're measuring development speed through lines of code.
         | 
         | Yeah, this is the biggest facepalm.
         | 
         | Didn't we grow out of this idiocy 40 years ago? This shit
         | again? Really?
        
         | andai wrote:
         | See also
         | 
         | -2000 lines of code
         | 
         | https://news.ycombinator.com/item?id=26387179
         | 
         | This is actually one thing I have found LLMs surprisingly
         | useful for.
         | 
         | I give them a code base which has one or two orders of
         | magnitude of bloat, and ask them to strip it away iteratively.
         | What I'm left with usually does the same thing.
         | 
         | At this point the code base becomes small enough to navigate
         | and study. Then I use it for reference and build my own
         | solution.
        
       | mellosouls wrote:
       | Depends on the nature of the tool I would imagine - eg. Claude
       | Code Terminal (say) would have higher entry requirements in terms
       | of engineering experience (Cursor was sold as newbie-friendly) so
       | I would predict higher quality code than Cursor in a similar
       | survey.
       | 
       | ofc that doesn't take into account the useful high-level and
       | other advantages of IDEs that might mitigate against slop during
       | review, but overall Cursor was a more natural fit for vibe-
       | coders.
       | 
       | This is said without judgement - I was a cheerleader for Cursor
       | early on until it became uncompetitive in value.
        
       | mentalgear wrote:
       | > We find that the adoption of Cursor leads to a statistically
       | significant, large, but transient increase in project-level
       | development velocity, along with a substantial and persistent
       | increase in static analysis warnings and code complexity. Further
       | panel generalized-method-of-moments estimation reveals that
       | increases in static analysis warnings and code complexity are
       | major factors driving long-term velocity slowdown. Our study
       | identifies quality assurance as a major bottleneck for early
       | Cursor adopters and calls for it to be a first-class citizen in
       | the design of agentic AI coding tools and AI-driven workflows.
       | 
       | So overall seems like the pros and cons of "AI vibe coding" just
       | cancel themselves out.
        
         | mort96 wrote:
         | The part you quoted doesn't support your conclusion. Per your
         | quoted paragraph, the benefit of "AI vibe coding" is a large,
         | but transient (i.e temporary) increase in development velocity;
         | while the drawback is a persistent increase in static analysis
         | warnings and code complexity.
         | 
         | To me, this sounds like after the transient increase of
         | velocity has died down, you're left with the same development
         | velocity as you had when you started, but a significantly worse
         | code base.
        
           | andai wrote:
           | The implication seems to be that if quality assurance is
           | prioritized, the negative impact would be eliminated.
           | 
           | This seems to assume the main cause is the accumulation of
           | defects due to lack of static analysis and testing.
           | 
           | I think a more likely cause is, the code begins to rapidly
           | grow beyond the maintainers' comprehension. I don't think
           | there is a technical solution for that.
        
       | chris_money202 wrote:
       | Now someone do a research study where a summary of this research
       | paper is in the AGENTS.md and let's see if the overall outcomes
       | are better
        
       | dalemhurley wrote:
       | I think the issue is people AI assisted code, test then commit.
       | 
       | Traditional software dev would be build, test, refactor, commit.
       | 
       | Even the Clean Coder recommends starting with messy code then
       | tidying it up.
       | 
       | We just need to apply traditional methods to AI assisted coding.
        
       | duendefm wrote:
       | AI is not perfect sure, one has to know how to use it. But this
       | study is already flawed since models improved a lot since the
       | beginning of 2026.
        
         | Eufrat wrote:
         | This is not a useful, constructive or meaningful statement.
         | 
         | Attempting to claim the models are the future by perpetually
         | arguing their limitations are because people are using the
         | models wrong or that the argument has been invalidated because
         | the new model fixes it might as well be part of the training
         | data since Claude Opus 3.5.
        
           | duendefm wrote:
           | No no, I didn't say that at all. I'm just saying that the
           | studies are irrelevant since models got a boost in their
           | competence. I'm not in a fight pro or against llm's, I know
           | how they work and their limitations. But the complexity of
           | the problems they solve increased since opus 4.5 . If you
           | can't admit that, it's your problem.
           | 
           | Also, I'm not blaming users for their shortcomings. I'm just
           | saying they are not perfect but you can get different
           | outcomes according to how you use them.
        
       | keeda wrote:
       | There are actually quite a few studies out there that look at LLM
       | code quality (e.g.
       | https://scholar.google.com/scholar?hl=en&as_sdt=0%2C5&q=LLM+...)
       | and they mostly have similar findings. This reinforces the idea
       | that LLMs still require expert guidance. Note, some of these
       | studies date back to 2023, which is eons ago in terms of LLM
       | progress.
       | 
       | The conclusion of this paper aligns with the emerging
       | understanding that AI is simply an amplifier of your existing
       | quality assurance processes: Higher discipline results in higher
       | velocity, lower discipline results in lower stability (e.g.
       | https://dora.dev/research/2025/) Having strong feedback and
       | validation loops is more critical than ever.
       | 
       | In this paper, for instance, they collected static analysis
       | warnings using a local SonarQube server, which implies that it
       | was not integrated into the projects they looked at. As such
       | these warnings were not available to the agent. It's highly
       | likely if these warnings were fed back into the agent it would
       | fix them automatically.
       | 
       | Another interesting thing they mention in the conclusion: the
       | metrics we use for humans may not apply to agents. My go-to
       | example for this is code duplication (even though this study
       | finds minimal increase in duplication) -- it may actually be
       | better for agents to rewrite chunks of code from scratch rather
       | than use a dependency whose code is not available forcing it to
       | instead rely on natural language documentation, which may or may
       | not be sufficient or even accurate. What is tech debt for humans
       | may actually be a boon for agents.
        
       | woeirua wrote:
       | This study's cutoff date was August 2025. I don't think this
       | result is surprising given the level of coding agent ability back
       | then. The whole thing just shows how out-of-date academic
       | publishing is on this subject.
       | 
       | >This yields 806 repositories with adoption dates between January
       | 2024 and March 2025 that are still available on GitHub at the
       | time of data analysis (August 2025).
       | 
       | There were _very few_ people who thought that coding agents
       | worked very well back then. I was not one of them, but I _do_
       | think they work today.
        
         | monkaiju wrote:
         | This is the perennial excuse, and I'm sure we'll continue to
         | see it. Folks will say the exact same thing when the current
         | crop of slop-generators have been replaced by a newer ilk
        
         | staticassertion wrote:
         | Eh, I don't know. I mean, are we seeing _better_ models now? Of
         | course. But are they truly leaps and bounds better? No, and I
         | get confused by people saying that they are. They 're better
         | but not like... 10x better.
         | 
         | And when people were studying ChatGPT 3.5, everyone would go
         | "Oh, but that wasn't 4!", and when people talk about Opus 4.5
         | they go "4.6 is so much better!".
         | 
         | My personal position right now is that people are extremely bad
         | at evaluating model output/ changes in model capabilities.
         | Model benchmarks do _not_ reflect the position that models are
         | just 10x better than they were a year ago, but with how people
         | discuss them you 'd think that 10x was underselling it.
        
           | nicoburns wrote:
           | I don't have personal experience, but there seems to be a
           | broad consensus that Opus 4.5 was tipping point between
           | "kinda bad" and "actually kinda useful".
           | 
           | So a cutoff point of August 2025 just before that is a bit
           | unfortunate (I'm sure they'll be newer studies soon).
        
         | sentrysapper wrote:
         | Evergreen excuses for tech people desperately want to work. I
         | get why, it would give you agency to do other things you WANT
         | to do. I tried reviewing a colleagues agent-generated code and
         | it was practically unreviewable. I watched him blame himself,
         | saying he just needed to adjust a parameter. He tried
         | everything except admit the machine does not conceptionaly
         | understand what he was asking.
        
       ___________________________________________________________________
       (page generated 2026-03-16 23:01 UTC)