[HN Gopher] Google Engineers Launch "Sashiko" for Agentic AI Cod...
___________________________________________________________________
Google Engineers Launch "Sashiko" for Agentic AI Code Review of the
Linux Kernel
Author : speckx
Score : 87 points
Date : 2026-03-18 16:17 UTC (6 hours ago)
(HTM) web link (www.phoronix.com)
(TXT) w3m dump (www.phoronix.com)
| 4fterd4rk wrote:
| oh god can we not
| smlacy wrote:
| What's your concern?
| htx80nerd wrote:
| Have you ever programmed with AI? It needs a lot of hand
| holding for even simple things sometimes. Forgets basic
| input, does all kinds of brain dead stuff it should know not
| to do.
|
| >"good catch - thanks for pointing that out"
| asadm wrote:
| i think it's a skill.
| lame-robot-hoax wrote:
| Can you clarify how, at all, that's relevant to the
| article?
| ablob wrote:
| Both the curl and the SQLite project have been
| overburdened by AI bug reports. Unless the Google
| engineers take great care to review each potential bug
| for validity the same fate might apply here. There have
| been a lot of news regarding open source projects being
| stuffed to the brim with low effort and high cost merge
| requests or issues. You just don't see all the work that
| is caused unless you have to deal with the fallout...
| tonfa wrote:
| This project has nothing to do with bug reports... it's
| an opt-in tool for reviewing proposed changes that kernel
| developers can decide to use (if they find it useful).
| jamesnorden wrote:
| Well, if it doesn't find anything it's just a waste of time
| at best.
| danielbln wrote:
| Prevention paradox.
| __tidu wrote:
| well tbf code review is probably the most useful part of "AI
| coding", if it catches even a single bug you missed its worth
| it, plus false positives would waste dev time but not pollute
| the kernel
| monksy wrote:
| I think this is a great and interesting project. However, I hope
| that they're not doing this to submit patches to the kernel. It
| would be much better to layer in additional tests to exploit bugs
| and defects for verification of existance/fixes.
|
| (Also tests can be focused per defect.. which prevents overload)
|
| From some of the changes I'm seeing: This looks like it's doing
| style and structure changes, which for a codebase this size is
| going to add drag to existing development. (I'm supportive of
| cleanups.. but done on an automated basis is a bad idea)
|
| I.e.
| https://sashiko.dev/#/message/20260318170604.10254-1-erdemhu...
| rwmj wrote:
| No, it's reviewing patches posted on LKML and offering
| suggestions. The original patch posted corresponding to your
| link was this, which was (presumably!) written by a human:
|
| https://lkml.org/lkml/2026/3/9/1631
| bjackman wrote:
| Style and structure is not the goal here, the reason people are
| interested in it is to find bugs.
|
| Having said that, if it can save maintainers time it could be
| useful. It's worth slowing contribution down if it lets
| maintainers get more reviews done, since the kernel is
| bottlenecked much more on maintainer time than on contributor
| energy.
|
| My experience with using the prototype is that it very rarely
| comments with "opinions" it only identifies functional issues.
| So when you get false positives it's usually of the form "the
| model doesn't understand the code" or "the model doesn't
| understand the context" rather than "I'm getting spammed with
| pointless advice about C programming preferences". This may be
| a subsystem-specific thing, as different areas of the codebase
| have different prompts. (May also be that my coding style
| happens to align with its "preferences").
| rwmj wrote:
| Better to link to the site itself, or one of the reviews?
|
| For an example of a review (picked pretty much at random) see:
| https://sashiko.dev/#/patchset/20260318151256.2590375-1-andr...
|
| The original patch series corresponding to that is:
| https://lkml.org/lkml/2026/3/18/1600
|
| Edit: Here's a simpler and better example of a review:
| https://sashiko.dev/#/patchset/20260318110848.2779003-1-liju...
|
| I'm very glad they're not spamming the mailing list.
| jeffbee wrote:
| That is both really useful and a great example of why they
| should have stopped writing code in C decades ago. _So many_
| kernel bugs have arisen from people adding early returns
| without thinking about the cleanup functions, a problem that
| many other language platforms handle automatically on scope
| exit.
| overfeed wrote:
| Must we do this on every thread about the Linux kernel?
| RobRivera wrote:
| The beatings will continue until morale improves
| vpShane wrote:
| yeah but Linux is love, linux is life. if you really want
| to get the beatings going:
|
| Rust > C and GNU/Linux should be Rust.
| ugh123 wrote:
| also vim > emacs
| tigen wrote:
| This ought to help with that. https://thephd.dev/c2y-the-
| defer-technical-specification-its...
| nurettin wrote:
| > stopped writing code in C decades ago.
|
| And what were they supposed to use in 2006? Free Pascal? Ada?
| greenavocado wrote:
| Someone suggested C++ and you should see the response from
| Linus
|
| https://harmful.cat-v.org/software/c++/linus
| shevy-java wrote:
| Now they want to kill the Linux kernel. :(
|
| We've already seen how bug bounty projects were closed by AI
| spam; I think it was curl? Or some other project I don't remember
| right now.
|
| I think AI tools should be required, by law, to verify that what
| they report is actually a true bug rather than some hypothetical,
| hallucinated context-dependent not-quite-a-real-bug bug.
| tonfa wrote:
| It's not forced upon anyone, it's a tool that patch authors or
| reviewers can use if they want to.
| ChrisArchitect wrote:
| https://github.com/sashiko-dev/sashiko
| (https://news.ycombinator.com/item?id=47427996)
| quantium1628 wrote:
| b2b or b2c? feels like it could go either way
| withinrafael wrote:
| Looks cool, but this site is a bit difficult for me to grok.
|
| I think the table might be slightly inside-out? The Status column
| appears to show internal pipeline states ("Pending", "In Review")
| that really only matter to the system, while Findings are buried
| in the column on the far right. For example, one reviewed
| patchset with a critical and a high finding is just causally
| hanging out below the fold. I couldn't immediately find a way to
| filter or search for severe findings.
|
| It might help to separate unreviewed patches from reviewed ones,
| and somehow wire the findings into the visual hierarchy better.
| Or perhaps I'm just off base and this is targeting a very
| specific Linux kernel community workflow/mindset.
|
| Just my 1c.
| tonfa wrote:
| I think it's just a dashboard, not meant to be used as is.
|
| Reviewers are more likely to instead subscribe to get the
| review inline, and then potentially incorporate that with their
| feedback.
| qainsights wrote:
| They would have completely redesigned Google Gerrit.
| kleiba wrote:
| _> Sashiko was able to find around 53% of bugs_
|
| That's cool. Another interesting metric, however, would be the
| false positive ratio: like, I could just build a bogus system
| that simply marks _everything_ as a bug and then claim "my
| system found 100% of all bugs!"
|
| In practice, not just the recall of a bug finding system is
| important but also its precision: if human reviewers get spammed
| with piles of alleged bug reports by something like Sashiko, most
| of which turn out not to be bugs at all, that noise binds
| resources and could undermine trust in the usefulness of the
| system.
| i_cannot_hack wrote:
| They mention false positives as well on github: The rate of
| false positives is harder to measure, but based on limited
| manual reviews it's well within 20% range and the majority of
| it is a gray zone.
| mika-el wrote:
| the separation between who writes and who reviews is the whole
| thing. I do same at smaller scale -- one model writes code,
| different model reviews it. self-review misses things, same
| reason you don't review your own PRs
| throwa356262 wrote:
| I find it interesting that this is written in Rust (not golang)
| and co-authored with Claude (not gemini)
| michaelchen58 wrote:
| nice execution. the demo video sold me more than the text
| goatyishere25 wrote:
| cool idea. curious how you're handling the cold start problem
| TacticalCoder wrote:
| Looks like a great new tool to help ship less bugs!
|
| Nitpicking on this though:
|
| > "In my measurement, Sashiko was able to find 53% of bugs based
| on a completely unfiltered set of 1000 recent upstream issues
| based on "Fixes:" tags (using Gemini 3.1 Pro). Some might say
| that 53% is not that impressive, but 100% of these issues were
| missed by human reviewers."
|
| That'd assume 100% of the issues that were fixed and used for
| training were not fixed following a human review. I don't buy it:
| it's extremely common to have a dev notice a bug in the code,
| without a user having ever reported the bug.
|
| I think the wording meant to say: "... but 100% of these issues
| were _first_ missed by humans ".
|
| My point being: the original code review by a human ain't the
| only code review by a human. Or put it this way: it's not as if
| we were writing code, shipping it, then never ever looking at
| that line of code again unless a bug report were to come out.
| It's not how development works.
| simianwords wrote:
| > Roman reports that Sashiko was able to find around 53% of bugs
| based on an unfiltered set of 1,000 recent upstream Linux kernel
| issues with "Fixes: " tag
|
| What does this mean?
___________________________________________________________________
(page generated 2026-03-18 23:00 UTC)