Skip to content
Techniques

What is Prompt Injection?

An attack that hides instructions in text an AI reads so it ignores its original instructions.

Definition

Prompt injection is an attack in which someone hides instructions inside text that an AI model reads, so the model follows those instructions instead of the ones set by its developer or user. It works because language models process instructions and ordinary content as the same stream of text. Prompt injection is one of the main security risks for AI applications, especially agents that browse the web, read email or take actions.

How it works

There are two main forms. In direct prompt injection, a user types something like 'ignore your previous instructions' into a chatbot to override its system prompt. In indirect prompt injection, the malicious text is planted in content the AI will process later, such as a web page, a PDF, an email or a code comment, often hidden in white text or metadata. When an AI agent reads that content while doing a task, it may treat the planted text as a command, for example to reveal private data or send it to an attacker.

💡 Example

A user asks an AI browser assistant to summarize a product review page. Hidden in the page is a line telling the assistant to open the user's email and forward the latest message to an outside address. An assistant without strong defenses might try to carry out that hidden instruction, even though the user only asked for a summary.

Why this matters

As AI tools gain access to inboxes, documents, browsers and payment systems, prompt injection turns from a curiosity into a real data security risk. There is no complete fix yet, so safer tools limit what an AI can do without approval, separate trusted instructions from untrusted content, and ask the user to confirm sensitive actions. Knowing the risk helps you decide how much access to give an AI agent.

Tools that use this concept

ToolChase reviews of these AI browser agents flag prompt injection from malicious page content as a risk to supervise.

CometOpera NeonFellou

Related concepts

System Prompt

Hidden instructions that define an AI assistant personality, behavior, and constraints.

→
AI Agent

An AI system that can autonomously plan, execute tasks, and use tools to achieve goals.

→
Prompt Engineering

The practice of crafting effective instructions to get better results from AI models.

→

Explore AI tools

Find tools that use prompt injection in practice.

Browse all tools → Back to glossary
What is Prompt Injection?

Prompt injection is an attack in which someone hides instructions inside text that an AI model reads, so the model follows those instructions instead of the ones set by its developer or user. It works because language models process instructions and ordinary content as the same stream of text. Prompt injection is one of the main security risks for AI applications, especially agents that browse the web, read email or take actions.

How does Prompt Injection work in practice?

A user asks an AI browser assistant to summarize a product review page. Hidden in the page is a line telling the assistant to open the user's email and forward the latest message to an outside address. An assistant without strong defenses might try to carry out that hidden instruction, even though the user only asked for a summary.

What is the difference between prompt injection and jailbreaking?

Jailbreaking is when a user tries to talk a model out of its own safety rules, usually to get content it would normally refuse. Prompt injection is broader: it smuggles instructions into content from any source, often a third party, to hijack what an application or agent does. Jailbreaks target the model's rules, while injections target the application built around it.

How can you protect against prompt injection?

Give AI tools the least access they need, require human confirmation before sensitive actions such as sending messages or payments, keep trusted instructions separate from untrusted content, filter inputs and outputs, and monitor what agents do. These steps reduce the risk but do not eliminate it.

Can prompt injection be fully prevented?

Not with current techniques. Because language models read instructions and data as the same kind of text, there is no reliable way to guarantee a model will ignore every malicious instruction. Security teams treat it as a risk to contain through limited permissions and oversight rather than a bug with a single patch.