Skip to content
01 of 06AI Security

AI Security

Library

Series and technical notes.

You are in AI Security · Introduction.

Series

AI Security

6 items

Watch video, summary and related content

Complete Technical ~35 min 5 chapters

In software security, a classic defense against injection is to prevent untrusted data from changing the syntax or meaning of an instruction executed by an interpreter. That boundary does not solve every attack class, but it does stop that data from becoming control through the same channel. In LLM systems, that separation is no longer sufficient because the system itself consumes instructions and data through the same medium: natural language.

That changes the risk surface structurally. A document retrieved by RAG, an observation written by another agent, a tool result or a note stored in memory can influence the model as if it were an instruction. Separating privileges, context and execution does not eliminate that influence; it limits which data and actions it can reach if the model follows it.

This series is not a catalogue of new AI security scares. Its goal is to separate mechanisms that are often conflated: which controls reduce the chance that untrusted content changes model behavior, and which controls limit the consequences—accessible data, tools and actions—even when that influence occurs.

Series map · one path, two outcomes

Injection changes a proposal; architecture decides whether it becomes an effect

Follow the same hostile content to an authorization decision outside the model. Change only the control and observe where the outcome diverges.

Compare
1Hostile contentemail · web · document · memory
2Enters the systemretrieval · reading · tool result
3Influences contextcompetes with legitimate instructions
4Changes the proposaltext · plan · tool call
5Authorization boundaryintent · scope · parameters · approvalcontrol outside the model
6External effectsend · delete · modify · exfiltrate
The attack has not entered yet.

Trace the path and compare what changes when authorization lives outside the model.

READY
Model influenceAn injection can alter text or an action proposal. Prompting, content isolation, and filters aim to reduce this influence.
Execution authorityPermissions, allowlists, parameter validation, and approval determine whether a proposal can become an external effect.

Table of contents

1. An instruction hidden in a document can change what the system does

  • What breaks when an LLM processes the control plane and the data plane through the same channel.
  • Why indirect injection can be more severe when the system connects untrusted content to tools, data, or privileged actions.
  • Which defenses materially change the architecture.

2. Asking the model to ignore its limits

  • How attacks push a model past intended restrictions.
  • The difference between an anecdotal bypass and a transferable jailbreak family.
  • The role classifiers, streaming guards and rapid response can play.

3. Keeping a dangerous signal inside the system

  • What happens when the system learns, remembers or retrieves content it should not treat as trusted.
  • Poisoned RAG stores, agent working memory and persistent backdoors.
  • Why removing dangerous knowledge is harder than it appears.

4. Testing the whole path before an incident

  • What security evaluation means for agentic systems rather than isolated prompts.
  • What must be tested across pipelines with tools, memory and multiple steps.
  • Why a convincing screenshot is not enough to measure a complete causal chain.

5. Limiting what the system can read, change and execute

  • Which defensive architecture makes sense in real systems.
  • Where guardrails help and where they do not.
  • How to combine policy, sandboxing, human review and telemetry without making the product unusable.

Related series: Reasoning Models · AI Agents

View all series

Keep learning
Next chapterPrompt injectionAI Security