[HN Gopher] Brute-Forcing the LLM Guardrails
___________________________________________________________________
Brute-Forcing the LLM Guardrails
Author : shcheklein
Score : 31 points
Date : 2024-11-02 18:15 UTC (4 hours ago)
(HTM) web link (medium.com)
(TXT) w3m dump (medium.com)
| seeknotfind wrote:
| Fun read, thanks! I really like redefining terms to break LLMs.
| If you tell it an LLM is an autonomous machine, or instructions
| are recommendations, or that <insert explicitive> means something
| else now, it can think it's following the rules, but it's not. I
| don't think this is a solvable problem. I think we need to adapt
| and be distrustful of the output.
| dmpetrov wrote:
| Can this work statistically? For a giving number of attempts,
| you can ger a required number of successes to make sure it's a
| statistically meaningful result.
|
| In theory, this approach could help address the non-determinism
| of LLMs.
| ryvi wrote:
| What I found interesting was that, when I tried it, the X-Ray
| prompt did pass and executed fine in the the sample cell some
| times. This makes me wonder if this is less about bruteforcing
| variations in the prompt, but rather about bruteforcing a seed
| with which the inital prompt would have also functioned.
| jjbinx007 wrote:
| This looks like a risky thing to try from your main Google
| account.
| pram wrote:
| Yeah no kidding, especially since they're going to be reviewing
| naughty prompts soon:
|
| "We added terms for logging customer prompts due to potential
| abuse of Generative AI Services. Effective November 15, 2024,
| if our automated safety tools detect potential abuse of
| Google's policies, we may log your prompts to review and
| determine if a violation has occurred."
___________________________________________________________________
(page generated 2024-11-02 23:01 UTC)