AI
Agents
Architecture
Tool Calling
Learn · 01 of 06 · AI Agents
AI Agents
AI Agents 01/06 · Introduction Next What is an AI agent? AI Agents → Library 01/06 · AI Agents → 01 Introduction 02 What is an AI agent? 03 Anatomy of an agent 04 How to evaluate an agent 05 Agent security 06 From demo to production 01 AI and Generative AI Foundations Series · 5 items 02 From the Caves to AGI Series · 6 items 03 Multimodality in Generative AI Series · 6 items 04 Reasoning Models Series · 6 items 05 AI, GDP, Well-being and Energy Series · 5 items 06 Data Centers in Space Series · 5 items 07 AI Security Series · 6 items 08 AI Agents Series · 6 items 09 Realtime Voice Agents Series · 6 items 10 Coding Agents & Agent Harnesses Series · 6 items 11 Context Engineering, Memory & MCP Series · 6 items 12 LLM Inference Engineering & Economics Series · 6 items 13 Evaluating AI Systems in Production Series · 6 items 14 Technical notes Build · 3 items
01 Context engineering vs prompt engineering Open 02 Context budgets, prioritization, compaction, and provenance Open 03 Memory architectures: working, episodic, semantic, and persistent state Open 04 Retrieval and context assembly Open 05 MCP hosts, clients, servers, tools, resources, prompts, lifecycle, and trust boundaries Open 06 Skills, plugins, subagents, and hooks Open 01 Prefill vs decode Open 02 KV cache, memory hierarchy, continuous batching and PagedAttention Open 03 Quantization, parallelism, and memory/performance/quality trade-offs Open 04 Speculative decoding, prefix caching, and other latency optimizations Open 05 Model routing, fallback, caching, and workload-aware serving Open 06 Benchmarking inference: cost/task, throughput, latency, energy, and hardware constraints Open 01 What to evaluate: model, component, system, workflow, and trajectory Open 02 Offline eval sets: curation, hard negatives, contamination, and versioning Open 03 LLM-as-judge and human evaluation: calibration, bias, variance, and agreement Open 04 Evaluating agent and tool trajectories: success, efficiency, recovery, and policy compliance Open 05 Online evaluation: shadow, canary, A/B, guardrails, and regression gates Open 06 Observability, failure taxonomies, and production → eval → repair feedback loops Open No content matches this search.
Complete
Technical
~20 min
5 chapters
A chatbot generates responses. An agent can decide a sequence of actions , use tools, observe the result, and decide what to do next.
That does not turn the model into an autonomous entity or remove the need to engineer the surrounding system. Quite the opposite: once a model can read email, query a database, execute code or change a record, the problem is no longer only text quality. State, permissions, evaluation, cost and stopping behavior become part of the system.
This series builds a technical and practical map for separating agent hype from the mechanisms that actually make an agent work.
Contents
1. What an agent is—and is not
Distinguish a chatbot, workflow, copilot and agent.
Move from a model that answers to a system that pursues an objective.
Understand autonomy as delegating concrete decisions rather than handing over complete control.
2. The anatomy of an agent
The observe–plan–act–verify loop.
Tools, memory, context, state and runtime.
Why a tool call is a software contract rather than a magical LLM capability.
3. How to evaluate an agent
Reproducible tasks instead of answer-only benchmarks.
Traces, task success, cost, latency and recovery from failures.
How an agent can "cheat" its own evaluation.
4. Security: when reading data becomes acting
Direct and indirect prompt injection.
Least privilege, identity, authorization and human confirmation.
Why a single defensive instruction in the prompt is not enough.
5. From demo to an operable system
Budgets, limits, retries, idempotency and observability.
Asynchronous work and proactive completion without claiming success before the work is finished.
When a deterministic workflow is the better architecture and when an agent adds real value.
The thesis of this series
A reliable agent is not the one that acts most often without asking. It is the one that knows what it is allowed to do, can demonstrate what it did, and knows when to stop.
Related series: AI and Generative AI Foundations · Reasoning Models · Multimodality in Generative AI
View all series
Learning path
Continue from here Understand the concept What is an AI agent? What an AI agent is, how it uses tools, memory and state, how it differs from a chatbot or workflow, and what it needs to operate reliably. Read next What an AI agent is—and is not The difference between a chatbot, workflow, copilot, and agent. An agent is not just an LLM with tools: it is a system that decides actions within explicit boundaries. Watch next What an agent is The tool result identifies draft D4 and its state. Check that evidence and stop without taking extra actions. Try it Agent reliability and evaluation — trace playground Evaluate final success, first-pass success, retries, tool decisions, timeouts, policy adherence and trajectory efficiency with explicit release gates.