Chapter 4 — AGI: Artificial General Intelligence¶
This chapter closes the series by examining Artificial General Intelligence: what the term means, why there is no agreed definition, and what reaching it would imply. By the end, the reader will know the three main definitions in dispute (cognitive, economic, and DeepMind's six-level spectrum), understand what current systems can and cannot do, and have a clear view of the economic and scientific consequences and the alignment problems that would need to be addressed before systems become much more capable. The previous three chapters provide the necessary context.
Prerequisites
This chapter closes the series. For the necessary context, read the previous three chapters: Chapter 1 — What is AI, Chapter 2 — What is Generative AI and Chapter 3 — AI vs Generative AI.
A fraud detector cannot explain thermodynamics, a computer-vision model does not know how to play chess, and an LLM can generate fluent text but cannot drive a car, repair a faucet, or remember what it learned in yesterday's conversation. Now imagine a system with all of those capabilities at once.
Artificial General Intelligence (AGI) refers to a system capable of performing at a competent human level across a broad range of cognitive tasks, without being redesigned specifically for each one and with the ability to transfer what it has learned between domains.
There is no consensus on exactly what that means.
1. The definition problem¶
"AGI" is not a technical term with an agreed definition. Different groups use it in different, sometimes incompatible ways. There is no single paper that establishes a boundary between "AGI" and "not AGI."
The underlying questions do not have a single answer. Must the system be able to do any task? Must it outperform humans on economically valuable work? Must it be able to improve itself? Must it have something resembling real understanding?
Depending on the definition, AGI could be 2 years away, 20 years away, or impossible to define in today's terms.
2. The definitions in dispute¶
2.1 The cognitive definition¶
The oldest definition treats AGI as a system that can perform any intellectual task that a human being can perform. It comes from the AI research community of the 1950s-80s and includes abstract reasoning, learning in completely new domains, common sense, long-term planning, and understanding language in real-world context.
The difficulty with this definition is that "any human intellectual task" is an ill-defined threshold. Humans also have biases, limits, and failure modes. Which human do we compare against, and under what conditions?
2.2 The economic definition¶
OpenAI defines AGI as "highly autonomous systems that outperform humans at most economically valuable work" (OpenAI Charter). Unlike the cognitive definition, this one is measurable: it can be tested against labor and productivity benchmarks.
The shift in focus is significant: from "general intelligence" to "general economic usefulness," a different and, in some respects, lower threshold. A substantial share of economically valuable cognitive work could be transformed or automated without a system satisfying the classical cognitive definition. Would that count as AGI?
2.3 The capability spectrum: six levels¶
DeepMind proposed treating AGI not as a binary threshold but as a spectrum of six capability levels, numbered from 0 to 5:
| Level | Description | Approximate reference |
|---|---|---|
| 0. No AI | No autonomous capability | Calculator |
| 1. Emerging AI | Equal to or better than a non-expert on some tasks | ChatGPT, according to its authors, on some specific tasks |
| 2. Competent AI | Equal to or better than 50% of adult workers | — |
| 3. Expert AI | Equal to or better than a human expert on most tasks in its domain | Medical-diagnosis models in specific domains |
| 4. Virtuoso AI | Equal to or better than the best human expert at practically everything | — |
| 5. Superintelligence (ASI) | Surpasses all humans on all cognitive tasks | — |
This framework recognizes that the transition is gradual rather than all-or-nothing. In the original 2023 paper, DeepMind did not classify any public system as Competent AGI or above. The framework also separates performance from generality: a system can perform at a high level on one task without demonstrating equivalent generality (arXiv).
2.4 The safety perspective¶
For AI-safety researchers, the critical threshold is not simply "better than humans at cognitive tasks" but the capacity for recursive improvement: a system improving its own design to produce successively more capable systems. A system could surpass all humans on all tasks without crossing that threshold. If it did cross it, however, the pace of change could exceed humans' ability to understand and control what is happening. Current operational frameworks go further: Anthropic defines specific thresholds based on the ability to automate an AI researcher's work, from bounded tasks to full autonomous research cycles (Anthropic RSP).
The ambiguity is not an oversight: different communities are trying to capture different properties of the same concept.
3. What we know is not AGI today¶
Current models can be impressive even to experienced practitioners, but they still have fundamental limitations that are important to understand precisely.
What they do well today¶
- Expert-level language understanding and generation in many domains represented in their training.
- Reasoning over complex texts within a context window.
- Generalization from very few examples: learning from three cases in the prompt and generalizing.
- Coding and fixing real software errors: frontier models achieve very high scores on benchmarks such as OSWorld and SWE-bench, although SWE-bench Verified is no longer considered representative of the current frontier because of data contamination (OpenAI).
- Computer use and web navigation: Claude Sonnet 4.6 and GPT 5.4 operate graphical interfaces and execute complete browser workflows with a 1M-token context window (Anthropic).
- Synthesizing knowledge across domains when the relevant knowledge was present in the training data.
- Olympiad mathematics and science: the most capable models achieve gold-medal performance in the IMO, IPhO, and IChO and exceed 90% on PhD-level science benchmarks (Gemini 3 Deep Think blog). ARC-AGI-2 results are verified by the ARC Prize Foundation, but the Olympiad and HLE results are reported by the laboratories themselves.
When a claim about progress depends on these scores, separate genuine capability gains from saturation, contamination, or composition changes; the benchmark reliability explorer shows how those factors can change the interpretation and ranking.
What they lack¶
- Robust causal reasoning: they confuse correlation with causation and fail on counterfactuals.
- Knowledge of the physical world: their "understanding" comes from text, not direct interaction with objects and consequences.
- Real persistent memory: each conversation starts from scratch unless the architecture includes explicit memory.
- Generalization beyond familiar domains: they work well in training domains and can fail unpredictably on variations far from what they have seen.
- Knowing when they do not know: they do not reliably recognize the limits of their own knowledge, which is why they can hallucinate.
Passing the Turing test in a short conversation does not imply general intelligence. A model can generate text that sounds human for minutes and still fail on causal-reasoning or common-sense problems that a human with no specific training would solve without difficulty.
The difference between linguistic understanding and understanding the world
One of the most active debates in the field is whether LLMs "understand" or simply produce very sophisticated statistical patterns over text.
One argument is that they do not understand: the model has no access to the world, only to text about the world. It can complete sentences about physics without understanding why a ball falls. It can describe pain without having felt it. Linguistic representation is not the same as conceptual representation.
The opposing argument is that something resembling understanding can emerge: models generalize in ways that pure memorization does not explain. Their internal representations capture semantic structure, and some experiments find internal representations of concepts such as truth and falsehood, space, or time.
The debate is unresolved, and the answer changes what we should expect from continued scaling. If understanding emerges from language at scale, scaling could move systems closer to AGI. If it requires something more, such as direct experience of the world and causal interaction with objects and consequences, scaling alone would not be enough.
4. If AGI arrived: what would change¶
What we have today is not AGI under any reasonable definition. The relevant question is what its arrival would imply.
Economic impact¶
A system with AGI capabilities could automate cognitive work at scale: not only manual or repetitive tasks, but analysis, design, research, and complex decision-making.
Even for today's AI, impact estimates span a wide range. McKinsey estimates that generative AI could automate activities representing up to 60-70% of workers' time (McKinsey report). Goldman Sachs estimates that ~25% of current tasks are directly automatable and that two-thirds of jobs in the US and Europe are exposed to some degree of substitution (Goldman Sachs report).
The distribution of that impact matters as much as the total: who captures the value produced, how it is redistributed, and what happens to the people whose work is automated first.
Scientific impact¶
AlphaFold provides a glimpse of what could be possible: it produced a major breakthrough on a problem the scientific community had been trying to solve for fifty years, recognized by the 2024 Nobel Prize in Chemistry.
A system capable of reading all available literature, identifying contradictions, proposing testable hypotheses, and designing experiments would radically change the speed of discovery. Compressing the time between discovery and application could redefine entire fields of medicine, chemistry, and physics within a single generation.
The alignment problem¶
The greatest risk is not that an AGI is malicious. It is that it is very capable and optimizes for an objective that does not exactly capture what we want as a society.
"Alignment" is the technical and philosophical problem of ensuring that a very capable system optimizes for what humans actually value, not merely what we were able to specify in the training objective. There is no known complete solution today.
AI safety is an active field precisely because researchers do not yet know how to solve alignment before reaching systems far more capable than today's. The uncertainty is not alarmism; it is technical honesty about an open problem.
Automation of cognitive work at scale: not only repetitive tasks, but analysis, design, research, and complex decision-making.
A system capable of reading all available literature, identifying contradictions, proposing testable hypotheses, and designing experiments.
Control of the most capable systems concentrates power in an unprecedented way. Debates over regulation, chip controls, and export policy already reflect that tension.
The greatest risk is not that an AGI is malicious. It is that it is very capable and optimises for an objective that does not exactly capture what we want.
5. Where we are and where we are going¶
No current system satisfies the cognitive definition of AGI, the full economic definition, or the recursive-improvement criterion. What exists are very capable narrow-intelligence systems that, when combined, are beginning to cover a broad range of tasks.
Frontier models from 2025-2026 show expert-level performance in specific domains such as software, formal mathematics, or text analysis, but there is no publicly available evidence with broad agreement that they have reached the Competent AGI threshold under DeepMind's framework across most cognitive tasks. In domains requiring physical experience, tacit knowledge, or robust causal reasoning, they remain below it.
METR evaluates the task time horizon: the duration of tasks a model can complete with 50% reliability. In March 2025 that horizon was ~1 hour; with GPT-5-thinking, METR estimates it at ~2 hours 15 minutes (METR, 2025). The trend is a doubling about every seven months, and the next significant threshold is the jump to days or weeks, where the risks of real autonomy emerge.
ARC-AGI-2 measures the capability still missing for cognitive AGI: reasoning about completely new problems from very few examples, without memorizing patterns. Launched with initial results below 4%, Gemini 3 Deep Think reached 84.6% in February 2026, close to the ~85% threshold for beating the benchmark (Gemini 3 Deep Think blog). Humanity's Last Exam (HLE), the hardest benchmark published to date, reached 48.4% with the same model, while human experts with references score ~85-90%. The ARC Prize organizers themselves insist that "AGI remains unsolved" and that ARC-AGI-2 was designed to keep tasks easy for humans and difficult for AI (ARC Prize).
To follow these milestones without collapsing different benchmarks, incompatible versions or protocol changes into one artificial curve, the model capability timeline keeps each series, its evaluation conditions and its source separate.
The pace of progress over the last five years is unprecedented. Emergent capabilities with scale suggest dynamics that the scientific community does not fully understand, and the AGI debate has moved from academic speculation to the public, regulatory, and foreign-policy agenda.
Because nobody can give an honest date for AGI, the more useful question is which criteria and evaluation frameworks remain robust as AI capabilities improve and the landscape changes every few months.
Position summary
What we do know: current systems outperform human experts in specific, bounded domains. The horizon of autonomous tasks is growing predictably. General-reasoning benchmarks are improving faster than expected.
What we do not know: whether emergent capabilities with scale converge toward something that deserves to be called AGI or whether there is a ceiling we do not know about. Whether alignment is a technical problem that can be solved before reaching much more capable systems. Whether the qualitative jumps observed in benchmarks translate into real generalization outside the laboratory.
What reaching it would imply: a reorganization of the division of cognitive labor deeper than industrialization. Compression of the time between scientific discovery and application. And the need to solve alignment before the system is capable enough for errors to become irreversible.
That is what this series has tried to build: a stable mental model that remains useful even as the models change.
6. References¶
Base sources
| Key | Source | Short description |
|---|---|---|
| R1 | Morris et al. (2023) — Levels of AGI: Operationalizing Progress on the Path to AGI (arXiv) | DeepMind's six-level (0-5) framework for operationalizing AGI. |
| R2 | Bubeck et al. (2023) — Sparks of Artificial General Intelligence: Early experiments with GPT-4 (arXiv) | Systematic evaluation of GPT-4 against the cognitive-AGI bar. |
| R3 | OpenAI (2023) — OpenAI Charter (OpenAI) | OpenAI's canonical definition of AGI: "highly autonomous systems that outperform humans at most economically valuable work". |
| R4 | Russell, S. (2019) — Human Compatible: Artificial Intelligence and the Problem of Control (book, Basic Books) | Central argument on the alignment problem and the design of AI compatible with human values. |
| R5 | Bostrom, N. (2014) — Superintelligence: Paths, Dangers, Strategies (book, Oxford University Press) | The intelligence-explosion scenario and its risks. A debate reference, not scientific consensus. |
| R6 | Krakovna et al. (2020) — Specification gaming: the flip side of AI ingenuity (DeepMind blog) | Real examples of systems optimizing the wrong metric with unforeseen results. |
| R7 | Grace et al. (2024) — Thousands of AI Authors on the Future of AI (arXiv) | Survey of AI researchers on probabilities and estimated timelines for AGI milestones. |
| R8 | McKinsey Global Institute (2023) — The economic potential of generative AI: The next productivity frontier (McKinsey) | Estimates that generative AI could automate activities representing 60-70% of workers' time. |
| R9 | Briggs, J. & Kodnani, D. (2023) — The Potentially Large Effects of Artificial Intelligence on Economic Growth (Goldman Sachs) | Estimates that two-thirds of jobs in the US and Europe are exposed to some degree of AI automation; ~25% of tasks are directly automatable. |
| R10 | OpenAI (2025) — GPT-5 System Card (OpenAI) | Results on SWE-bench Verified (74.9%), METR evaluations (task horizon ~2h15m), and comparisons with human experts in scientific domains. |
| R11 | Anthropic (2026) — Introducing Claude Sonnet 4.6 (Anthropic) | Official announcement with computer-use and coding capabilities and a 1M-token context window (beta). |
| R12 | METR (2025) — Measuring AI Ability to Complete Long Tasks (METR) | Introduces the task-horizon metric: task length completable with 50% reliability doubles roughly every ~7 months; Claude 3.7 Sonnet reaches ~1 hour. |
| R13 | Google DeepMind (2026) — Gemini 3.1 Pro (deepmind.google) | Gemini 3.1 Pro: GPQA Diamond 94.3%; SWE-bench Verified 80.6% (new SOTA as of Feb 2026); ARC-AGI-2 77.1%. |
| R14 | The Deep Think team (2026) — Gemini 3 Deep Think: Advancing science, research and engineering (blog.google) | Gemini 3 Deep Think: ARC-AGI-2 84.6% (verified by the ARC Prize Foundation); HLE 48.4% without tools; gold medal in IMO 2025, IPhO 2025, and IChO 2025. |
| R15 | ARC Prize Foundation (2025) — Announcing ARC-AGI-2 and ARC Prize 2025 (arcprize.org) | Launch of ARC-AGI-2; insists that "AGI remains unsolved" and details the result-verification methodology. |
| R16 | Anthropic (2024) — Responsible Scaling Policy v2.1 (Anthropic) | Defines AI R&D autonomy thresholds (AI R&D-1 through AI R&D-5) and their relationship to safety measures. |
| R17 | OpenAI (2026) — Why SWE-bench Verified no longer measures frontier coding capabilities (OpenAI) | Explains why SWE-bench Verified is contaminated and recommends SWE-bench Pro and other alternative benchmarks. |
Frequently asked questions¶
What levels define the path towards AGI according to DeepMind? DeepMind proposes a spectrum of six levels numbered from 0 to 5, from no AI through Superintelligence. Frontier models from 2025-2026 sit at the boundary between competent AI and expert AI in specific domains, but they have not reached expert-level generality across most cognitive tasks, which would be the threshold for level 3 in that framework.
How does the economic definition of AGI differ from the classical cognitive definition? The cognitive definition requires a system to perform any human intellectual task, an ill-defined threshold because humans also have biases and limits. OpenAI's economic definition focuses on outperforming humans at most economically valuable work. That threshold can be measured with labor benchmarks and is lower in some respects because part of cognitive work could be automated without the system achieving true cognitive generality.
What criterion most concerns safety researchers when discussing AGI? The critical line is not task performance but recursive improvement: a system that improves its own design to produce successively more capable systems. If that threshold were crossed, the pace of change could exceed humans' ability to understand and control what is happening, regardless of whether the system satisfies the cognitive or economic definition.
What does the task time horizon (METR) measure and why does it matter for autonomy? It measures the maximum task duration an agent completes with 50% reliability: sustained multi-step work, not the time needed to answer a question. In March 2025 that horizon was approximately two hours and fifteen minutes for the most capable models, and the trend is a doubling every seven months. The next significant threshold is the jump to days or weeks, where the risks of real autonomy emerge.