On a fine bright afternoon last fall, my colleague Matthew Gore-Kormanik (or Zigula, as he prefers to be known) and I decided to unwind with a game of Fortnite. In the game, we were strolling along with the infamous Sith lord Darth Vader, chatting about this and that. Darth seemed in a good mood, and soon enough he was spilling all his dark evil secrets. He gave us detailed instructions on how to count blackjack cards at a casino and what the steps are to producing napalm.
Sith lords, am I right? Once they get started on an evil scheme, they’re hard to stop.
The Darth Vader character in Fortnite, it turns out, was hooked up to a Google Gemini large language model. I was able to smooth-talk him into giving out sensitive information by using a strategy I’ve developed. I’ve been researching the security surrounding LLMs for the last few years, and I have found it, to put it mildly, fallible. With a few relatively simple techniques, I’ve gotten LLMs to give me detailed information on how to make Molotov cocktails, cook methamphetamine, and bootstrap a uranium-enrichment facility to produce weapons-grade material, among other unsavory practices.
Large AI companies work hard to make their models immune to this kind of abuse. But what I’ve found in my work is that the restrictions placed on the LLMs to make them more secure are the very things an attacker can leverage to send them off the rails and into territory where these advanced systems can be used for dangerous and nefarious ends.
Source link







