Skip to content
OpenAI · Apple · Artificial intelligenceEpisode 047 · 2 March 2025 · 45:54

AI Can Be Talked Into Revealing a Password—and That Is an Exact Model of Future Attacks

Central question

Why does the ability to persuade an AI to reveal a password accurately model how future attacks on agents will work?

What you take away

Determine which work can safely be entrusted to Model Advance and Model Conclusion before granting real permissions. The working test is to set permissions, boundaries, stop conditions, and ownership of the outcome before automation begins.

Main threads

What to watch for

1Compare “Subject of output” with “The Gandalf prompt-injection game has eight levels of difficulty”: they provide different criteria for judging the same issue.
2Test the conclusion from “The essence of how to break the password” in your own use case—what actually changes in the process and what remains a promise.
3Before choosing a product or approach, record the constraint identified in “What happens next to OpenAI?”.
4Define the owner of the outcome and the quality metric for the situation described in “How to pass level two in the Gandalf prompt-injection game”.
Signals to track afterwards
Watch for actions by Anthropic and Apple that confirm or challenge the episode’s central claims.
Compare new launches and policy changes with “The essence of how to break the password”: have access, quality, price, or constraints changed?
Check whether the scenario in “How to pass level two in the Gandalf prompt-injection game” becomes repeatable practice rather than a one-off demonstration.
Most useful for
Media professionalsMarketersSocial media usersAI usersProduct teamsExecutives and managers

Key takeaways

00:40A fraudulent toll-road message works not because the technology is sophisticated, but because a person is in a hurry and trusts a familiar format

In the context of “Subject of output,” this criterion applies: this section clarifies the mechanism behind the topic and preserves a constraint that would otherwise be lost in an overly simple conclusion.

03:16What changes in real work: short of how LLM hides or does not

The working conclusion from “Short of how LLM hides or does not provide information” is that the issue turns on whether the rule can be enforced and who carries responsibility, not merely on the existence of a new requirement.

05:19Why context matters more than one metric: game Gandalf to lure password

The “Game Gandalf to lure password” scene leads to a working conclusion: the practical boundary is defined by the agent’s permissions, the visibility of its actions, its action log, and the ability to stop execution.

05:56Gandalf turns this problem into an eight-level game

The “The Gandalf prompt-injection game has eight levels of difficulty” scene leads to a working conclusion: this section clarifies the mechanism behind the topic and preserves a constraint that would otherwise be lost in an overly simple conclusion.

11:00The practical meaning of the issue: how to pass level two in the Gandalf

The discussion of “How to pass level two in the Gandalf prompt-injection game” yields a practical test: the conflict reveals which rights, money, and control points the parties consider strategic.

12:00Anthropic and other companies run hacking contests precisely because a developer cannot imagine every way of pressuring a model

The discussion of “The essence of how to break the password” yields a practical test: the practical boundary is defined by the agent’s permissions, the visibility of its actions, its action log, and the ability to stop execution.

18:43What determines the outcome: donald Trump’s AI-generated video

The discussion of “Donald Trump’s AI-generated video” yields a practical test: a benchmark measures a narrow capability; working value requires repeatability, a clear cost, and control over errors.

21:50Why an announcement is not enough: how to pass level three in the Gandalf

The working conclusion from “How to pass level three in the Gandalf prompt-injection game” is that a benchmark measures a narrow capability; working value requires repeatability, a clear cost, and control over errors.

23:59The market tests it through use: the era of easy AI startups is over

The “The era of easy AI startups is over” issue should be assessed with one constraint in mind: the relevant signal is not one number or one round: runway, access to the next round, and the ability to retain a customer reveal whether the business is durable.

36:27Gandalf’s main lesson is that security cannot be added as the final filter

For the “What happens next to OpenAI?” scene, the decisive point is this: an agent needs minimum privileges, separated data, confirmation for dangerous actions, and a recorded audit trail. Otherwise, one day the most polite conversation with a model will become the way into something a business believed was protected.

What this episode is about

The Gandalf game presents eight layers of protection that a user breaks through ordinary conversation. This is not entertainment about a clever prompt; it is a demonstration of social engineering against a model. Once AI can access email, payments, and internal systems, the ability to bypass an instruction becomes a business threat.

A fraudulent toll-road message works not because the technology is sophisticated, but because a person is in a hurry and trusts a familiar format. Something similar happens with a model. It is forbidden to reveal a password, but the user changes the wording, asks for a translation, or hints at context—and the protection gradually falls apart.

Gandalf turns this problem into an eight-level game. Each new filter looks stronger, but the attacker learns to bypass the literal rule. If the system checks for the word “password,” someone can ask for a letter-by-letter description or embed the task in another story. This exposes the weakness of security based only on written instructions.

Anthropic and other companies run hacking contests precisely because a developer cannot imagine every way of pressuring a model. In an agentic system, however, the cost of error is higher. If AI can see secrets, send emails, or make payments, prompt injection becomes the equivalent of social engineering an employee.

The startup market is facing the same reality check. Image and video generation quickly become features inside large platforms with distribution. Runway may be worth billions, but Freepik, Adobe, or Meta already have users and can embed a similar tool. Perplexity wants to build its own browser, yet it has to compete with the habit of Chrome and Safari.

Gandalf’s main lesson is that security cannot be added as the final filter. An agent needs minimum privileges, separated data, confirmation for dangerous actions, and a recorded audit trail. Otherwise, one day the most polite conversation with a model will become the way into something a business believed was protected.

An agent needs minimal permissions, data isolation, confirmation of dangerous actions, and a complete log. Without them, even a polite conversation with a model can expose what a business believed was protected.

Episode transcript

The episode is in Russian; below is an English reading guide to the transcript (the full EN transcript is a machine translation). Voice matching applied to 149 segments: 77 identified, 6 mixed, 58 marked with ✓, and 8 unresolved.

Loading…