What is prompt injection?

In short

Prompt injection is an attack where instructions hidden in text a language model reads steer it away from its real task. For coding agents, those instructions can come from a web page, an issue, a document or tool output, and can make the agent take actions you never asked for.

Direct and indirect injection

For coding agents the real risk is the indirect kind, because agents read outside content all the time.

  • Direct: the user tells the model to ignore its rules.
  • Indirect: the instruction is buried in content the model processes, such as a README, a code comment, an issue, a web page or an MCP tool’s response.

Why it’s hard to solve

The model sees its instructions and the data it reads in the same stream of text and can’t always tell them apart. That’s why there’s no known method that fully prevents prompt injection. OWASP’s Top 10 for large language model applications ranks prompt injection first.

What can go wrong with agents

The more tools and access an agent has, the bigger the damage a successful injection can do.

  • Leaking secrets such as .env files or tokens.
  • Installing a malicious package or running a destructive command.
  • Slipping a hard-to-spot backdoor into the code.
  • Using an MCP tool to act on a service the agent has access to.

Defenses

No single measure is enough; the aim is to limit the damage even if an attack gets through.

  • Least privilege: give the agent only the tools and access the job needs.
  • Always require approval for destructive actions and anything that reaches outside.
  • Handle untrusted repos and content in an isolated environment, such as a container or VM.
  • Install MCP servers only from sources you trust.
  • Review the agent’s changes before merging.

Safeguards in AgentVera

In AgentVera, agent access to your servers and databases is off by default, and risky actions come to you.

  • Auto approval only approves safe requests such as reads and tests. Commands with redirects (>), $() or backticks are never auto-approved, and commands like git push and rm always ask you (Pro and Team).
  • SSH agent access is off for every server by default. When on, agents never see the connection details or password, and every request goes through your rules and is logged (Pro and Team).
  • Database access is off by default too. When on, it starts read-only and every query an agent runs is logged (Pro and Team).
  • Secret values for MCP servers are kept in your operating system’s keychain.

FAQ

Can prompt injection be fully prevented?

Not with any method known today. You reduce the risk by limiting permissions, requiring approval for critical actions and isolating the agent.

Can my coding agent be attacked while reading a web page?

Yes. Hidden text on a page can look like an instruction to the agent. That’s why an agent that reads web content shouldn’t be able to run destructive commands without approval.

Are MCP servers vulnerable to prompt injection?

Yes, because tool descriptions and tool results go into the model’s context. Use servers you trust and be careful with tools that pull in outside content.

What’s the difference between prompt injection and a jailbreak?

A jailbreak is a user trying to get around the model’s safety rules. Prompt injection is usually a third party hiding instructions in content the model processes to turn the application to their own ends.

Related terms

All terms

See these ideas at work.

Download AgentVera for free, add a project folder and run Claude Code, Codex and other agents side by side.