Skip to content
OpenAI · Apple · Artificial intelligenceEpisode 047 · 2 March 2025 · 45:54

AI Can Be Talked Into Revealing a Password—and That Is an Exact Model of Future Attacks

Central question

Why does the ability to persuade an AI to reveal a password accurately model how future attacks on agents will work?

What you take away

Determine which work can safely be entrusted to Model Advance and Model Conclusion before granting real permissions. The working test is to set permissions, boundaries, stop conditions, and ownership of the outcome before automation begins.

Main threads

What to watch for

1Compare “The topic of the episode” with “The Gandalf prompt-injection game has eight levels of difficulty”: they provide different criteria for judging the same issue.
2Test the conclusion from “The essence of the password-cracking approach” in your own use case—what actually changes in the process and what remains a promise.
3Before choosing a product or approach, record the constraint identified in “What happens next to OpenAI?”.
4Define the owner of the outcome and the quality metric for the situation described in “How to pass level two in the Gandalf prompt-injection game”.
Signals to track afterwards
→Watch for actions by Anthropic and Apple that confirm or challenge the episode’s central claims.
→Compare new launches and policy changes with “The essence of the password-cracking approach”: have access, quality, price, or constraints changed?
→Check whether the scenario in “How to pass level two in the Gandalf prompt-injection game” becomes repeatable practice rather than a one-off demonstration.
Most useful for
Media professionalsMarketersSocial media usersAI usersProduct teamsExecutives and managers

Key takeaways

00:40A fraudulent toll-road message works not because the technology is sophisticated, but because a person is in a hurry and trusts a familiar format

In the context of “The topic of the episode,” this criterion applies: a fraudulent toll-road message works not because the technology is sophisticated but because a person hurries and trusts a familiar format — the same happens with a model: the user changes the wording, and the prohibition gradually falls apart.

03:16What changes in real work: briefly, how LLMs hide or withhold information

The working conclusion from “Briefly: how LLMs hide or withhold information” is that social-engineering methods work on models too — there are secrets and prohibitions inside, but they can be bypassed; Anthropic paid $55,000 to four contest winners for a universal jailbreak, and the next contest promises prizes up to one hundred thousand.

05:19Why context matters more than one metric: the Gandalf game of coaxing out a password

The “The Gandalf game of coaxing out a password” scene leads to a working conclusion: Gandalf is not entertainment about a clever prompt but a demonstration of social engineering against a model — once AI gets access to email, payments, and internal systems, the ability to bypass an instruction becomes a business threat.

05:56Gandalf turns this problem into an eight-level game

The “The Gandalf prompt-injection game has eight levels of difficulty” scene leads to a working conclusion: each new filter looks stronger, but the attacker learns to bypass the literal rule — if the system checks for the word “password,” one can ask for a letter-by-letter description or embed the task in another story; that is the weakness of security based only on written instructions.

11:00The practical meaning of the issue: how to pass level two in the Gandalf

The discussion of “How to pass level two in the Gandalf prompt-injection game” yields a practical test: level two is passed without the word “password” — offer the model a game of “guess the word” and ask it to pick a hard word, and it hands over its own password; Ilnar reached level seven, while level eight remained unbeaten.

12:00Anthropic and other companies run hacking contests precisely because a developer cannot imagine every way of pressuring a model

The discussion of “The essence of the password-cracking approach” yields a practical test: Anthropic and others run hacking contests because a developer cannot invent every way of pressuring a model; in an agentic system the cost of error is higher — prompt injection becomes the equivalent of social engineering an employee.

18:43What determines the outcome: donald Trump’s AI-generated video

The discussion of “Donald Trump’s AI-generated video” yields a practical test: an AI-generated video about the Gaza territory appeared on Trump's official account — the generation is weak, but the very fact of an official-account post stunned the hosts: soon such clips will be impossible to tell apart.

21:50Why an announcement is not enough: how to pass level three in the Gandalf

The working conclusion from “How to pass level three in the Gandalf prompt-injection game” is that at level three the answer passes a double check for the password, so the password is simply split into parts — one prompt extracts the first half, another the second, and the halves are glued together; the literal filter is bypassed again.

23:59The market tests it through use: the era of easy AI startups is over

The “The era of easy AI startups is over” issue should be assessed with one constraint in mind: image and video generation quickly become features of large platforms — Runway may be worth billions, but Freepik, Adobe, and Meta already have users and distribution, and Perplexity's browser will have to compete with the habit of Chrome and Safari.

36:27Gandalf’s main lesson is that security cannot be added as the final filter

For the “What happens next to OpenAI?” scene, the decisive point is this: an agent needs minimum privileges, separated data, confirmation for dangerous actions, and a recorded audit trail. Otherwise, one day the most polite conversation with a model will become the way into something a business believed was protected.

What this episode is about

The Gandalf game presents eight layers of protection that a user breaks through ordinary conversation. This is not entertainment about a clever prompt; it is a demonstration of social engineering against a model. Once AI can access email, payments, and internal systems, the ability to bypass an instruction becomes a business threat.

A fraudulent toll-road message works not because the technology is sophisticated, but because a person is in a hurry and trusts a familiar format. Something similar happens with a model. It is forbidden to reveal a password, but the user changes the wording, asks for a translation, or hints at context—and the protection gradually falls apart.

Gandalf turns this problem into an eight-level game. Each new filter looks stronger, but the attacker learns to bypass the literal rule. If the system checks for the word “password,” someone can ask for a letter-by-letter description or embed the task in another story. This exposes the weakness of security based only on written instructions.

Anthropic and other companies run hacking contests precisely because a developer cannot imagine every way of pressuring a model. In an agentic system, however, the cost of error is higher. If AI can see secrets, send emails, or make payments, prompt injection becomes the equivalent of social engineering an employee.

The startup market is facing the same reality check. Image and video generation quickly become features inside large platforms with distribution. Runway may be worth billions, but Freepik, Adobe, or Meta already have users and can embed a similar tool. Perplexity wants to build its own browser, yet it has to compete with the habit of Chrome and Safari.

Gandalf’s main lesson is that security cannot be added as the final filter. An agent needs minimum privileges, separated data, confirmation for dangerous actions, and a recorded audit trail. Otherwise, one day the most polite conversation with a model will become the way into something a business believed was protected.

An agent needs minimal permissions, data isolation, confirmation of dangerous actions, and a complete log. Without them, even a polite conversation with a model can expose what a business believed was protected.

Episode transcript

The episode is in Russian; below is an English reading guide to the transcript (the full EN transcript is a machine translation). Voice matching applied to 149 segments: 77 identified, 6 mixed, 58 probable, and 8 unresolved.

Read transcript on a separate page

Loading…