[HN Gopher] Google Engineers Launch "Sashiko" for Agentic AI Cod...
       ___________________________________________________________________
        
       Google Engineers Launch "Sashiko" for Agentic AI Code Review of the
       Linux Kernel
        
       Author : speckx
       Score  : 87 points
       Date   : 2026-03-18 16:17 UTC (6 hours ago)
        
 (HTM) web link (www.phoronix.com)
 (TXT) w3m dump (www.phoronix.com)
        
       | 4fterd4rk wrote:
       | oh god can we not
        
         | smlacy wrote:
         | What's your concern?
        
           | htx80nerd wrote:
           | Have you ever programmed with AI? It needs a lot of hand
           | holding for even simple things sometimes. Forgets basic
           | input, does all kinds of brain dead stuff it should know not
           | to do.
           | 
           | >"good catch - thanks for pointing that out"
        
             | asadm wrote:
             | i think it's a skill.
        
             | lame-robot-hoax wrote:
             | Can you clarify how, at all, that's relevant to the
             | article?
        
               | ablob wrote:
               | Both the curl and the SQLite project have been
               | overburdened by AI bug reports. Unless the Google
               | engineers take great care to review each potential bug
               | for validity the same fate might apply here. There have
               | been a lot of news regarding open source projects being
               | stuffed to the brim with low effort and high cost merge
               | requests or issues. You just don't see all the work that
               | is caused unless you have to deal with the fallout...
        
               | tonfa wrote:
               | This project has nothing to do with bug reports... it's
               | an opt-in tool for reviewing proposed changes that kernel
               | developers can decide to use (if they find it useful).
        
             | jamesnorden wrote:
             | Well, if it doesn't find anything it's just a waste of time
             | at best.
        
               | danielbln wrote:
               | Prevention paradox.
        
         | __tidu wrote:
         | well tbf code review is probably the most useful part of "AI
         | coding", if it catches even a single bug you missed its worth
         | it, plus false positives would waste dev time but not pollute
         | the kernel
        
       | monksy wrote:
       | I think this is a great and interesting project. However, I hope
       | that they're not doing this to submit patches to the kernel. It
       | would be much better to layer in additional tests to exploit bugs
       | and defects for verification of existance/fixes.
       | 
       | (Also tests can be focused per defect.. which prevents overload)
       | 
       | From some of the changes I'm seeing: This looks like it's doing
       | style and structure changes, which for a codebase this size is
       | going to add drag to existing development. (I'm supportive of
       | cleanups.. but done on an automated basis is a bad idea)
       | 
       | I.e.
       | https://sashiko.dev/#/message/20260318170604.10254-1-erdemhu...
        
         | rwmj wrote:
         | No, it's reviewing patches posted on LKML and offering
         | suggestions. The original patch posted corresponding to your
         | link was this, which was (presumably!) written by a human:
         | 
         | https://lkml.org/lkml/2026/3/9/1631
        
         | bjackman wrote:
         | Style and structure is not the goal here, the reason people are
         | interested in it is to find bugs.
         | 
         | Having said that, if it can save maintainers time it could be
         | useful. It's worth slowing contribution down if it lets
         | maintainers get more reviews done, since the kernel is
         | bottlenecked much more on maintainer time than on contributor
         | energy.
         | 
         | My experience with using the prototype is that it very rarely
         | comments with "opinions" it only identifies functional issues.
         | So when you get false positives it's usually of the form "the
         | model doesn't understand the code" or "the model doesn't
         | understand the context" rather than "I'm getting spammed with
         | pointless advice about C programming preferences". This may be
         | a subsystem-specific thing, as different areas of the codebase
         | have different prompts. (May also be that my coding style
         | happens to align with its "preferences").
        
       | rwmj wrote:
       | Better to link to the site itself, or one of the reviews?
       | 
       | For an example of a review (picked pretty much at random) see:
       | https://sashiko.dev/#/patchset/20260318151256.2590375-1-andr...
       | 
       | The original patch series corresponding to that is:
       | https://lkml.org/lkml/2026/3/18/1600
       | 
       | Edit: Here's a simpler and better example of a review:
       | https://sashiko.dev/#/patchset/20260318110848.2779003-1-liju...
       | 
       | I'm very glad they're not spamming the mailing list.
        
         | jeffbee wrote:
         | That is both really useful and a great example of why they
         | should have stopped writing code in C decades ago. _So many_
         | kernel bugs have arisen from people adding early returns
         | without thinking about the cleanup functions, a problem that
         | many other language platforms handle automatically on scope
         | exit.
        
           | overfeed wrote:
           | Must we do this on every thread about the Linux kernel?
        
             | RobRivera wrote:
             | The beatings will continue until morale improves
        
               | vpShane wrote:
               | yeah but Linux is love, linux is life. if you really want
               | to get the beatings going:
               | 
               | Rust > C and GNU/Linux should be Rust.
        
               | ugh123 wrote:
               | also vim > emacs
        
           | tigen wrote:
           | This ought to help with that. https://thephd.dev/c2y-the-
           | defer-technical-specification-its...
        
           | nurettin wrote:
           | > stopped writing code in C decades ago.
           | 
           | And what were they supposed to use in 2006? Free Pascal? Ada?
        
             | greenavocado wrote:
             | Someone suggested C++ and you should see the response from
             | Linus
             | 
             | https://harmful.cat-v.org/software/c++/linus
        
       | shevy-java wrote:
       | Now they want to kill the Linux kernel. :(
       | 
       | We've already seen how bug bounty projects were closed by AI
       | spam; I think it was curl? Or some other project I don't remember
       | right now.
       | 
       | I think AI tools should be required, by law, to verify that what
       | they report is actually a true bug rather than some hypothetical,
       | hallucinated context-dependent not-quite-a-real-bug bug.
        
         | tonfa wrote:
         | It's not forced upon anyone, it's a tool that patch authors or
         | reviewers can use if they want to.
        
       | ChrisArchitect wrote:
       | https://github.com/sashiko-dev/sashiko
       | (https://news.ycombinator.com/item?id=47427996)
        
       | quantium1628 wrote:
       | b2b or b2c? feels like it could go either way
        
       | withinrafael wrote:
       | Looks cool, but this site is a bit difficult for me to grok.
       | 
       | I think the table might be slightly inside-out? The Status column
       | appears to show internal pipeline states ("Pending", "In Review")
       | that really only matter to the system, while Findings are buried
       | in the column on the far right. For example, one reviewed
       | patchset with a critical and a high finding is just causally
       | hanging out below the fold. I couldn't immediately find a way to
       | filter or search for severe findings.
       | 
       | It might help to separate unreviewed patches from reviewed ones,
       | and somehow wire the findings into the visual hierarchy better.
       | Or perhaps I'm just off base and this is targeting a very
       | specific Linux kernel community workflow/mindset.
       | 
       | Just my 1c.
        
         | tonfa wrote:
         | I think it's just a dashboard, not meant to be used as is.
         | 
         | Reviewers are more likely to instead subscribe to get the
         | review inline, and then potentially incorporate that with their
         | feedback.
        
       | qainsights wrote:
       | They would have completely redesigned Google Gerrit.
        
       | kleiba wrote:
       | _> Sashiko was able to find around 53% of bugs_
       | 
       | That's cool. Another interesting metric, however, would be the
       | false positive ratio: like, I could just build a bogus system
       | that simply marks _everything_ as a bug and then claim  "my
       | system found 100% of all bugs!"
       | 
       | In practice, not just the recall of a bug finding system is
       | important but also its precision: if human reviewers get spammed
       | with piles of alleged bug reports by something like Sashiko, most
       | of which turn out not to be bugs at all, that noise binds
       | resources and could undermine trust in the usefulness of the
       | system.
        
         | i_cannot_hack wrote:
         | They mention false positives as well on github: The rate of
         | false positives is harder to measure, but based on limited
         | manual reviews it's well within 20% range and the majority of
         | it is a gray zone.
        
       | mika-el wrote:
       | the separation between who writes and who reviews is the whole
       | thing. I do same at smaller scale -- one model writes code,
       | different model reviews it. self-review misses things, same
       | reason you don't review your own PRs
        
       | throwa356262 wrote:
       | I find it interesting that this is written in Rust (not golang)
       | and co-authored with Claude (not gemini)
        
       | michaelchen58 wrote:
       | nice execution. the demo video sold me more than the text
        
       | goatyishere25 wrote:
       | cool idea. curious how you're handling the cold start problem
        
       | TacticalCoder wrote:
       | Looks like a great new tool to help ship less bugs!
       | 
       | Nitpicking on this though:
       | 
       | > "In my measurement, Sashiko was able to find 53% of bugs based
       | on a completely unfiltered set of 1000 recent upstream issues
       | based on "Fixes:" tags (using Gemini 3.1 Pro). Some might say
       | that 53% is not that impressive, but 100% of these issues were
       | missed by human reviewers."
       | 
       | That'd assume 100% of the issues that were fixed and used for
       | training were not fixed following a human review. I don't buy it:
       | it's extremely common to have a dev notice a bug in the code,
       | without a user having ever reported the bug.
       | 
       | I think the wording meant to say: "... but 100% of these issues
       | were _first_ missed by humans ".
       | 
       | My point being: the original code review by a human ain't the
       | only code review by a human. Or put it this way: it's not as if
       | we were writing code, shipping it, then never ever looking at
       | that line of code again unless a bug report were to come out.
       | It's not how development works.
        
       | simianwords wrote:
       | > Roman reports that Sashiko was able to find around 53% of bugs
       | based on an unfiltered set of 1,000 recent upstream Linux kernel
       | issues with "Fixes: " tag
       | 
       | What does this mean?
        
       ___________________________________________________________________
       (page generated 2026-03-18 23:00 UTC)