Skip to content
02 of 05AI and Generative AI Foundations

Chapter 1 — What is AI?

Library

Series and technical notes.

You are in AI and Generative AI Foundations · What is AI?.

Series

AI and Generative AI Foundations

5 items

Watch video, summary and related content

Estimated reading11 min

This chapter presents a framework for understanding any modern artificial-intelligence system, from a spam filter to a large language model. By the end, the reader will be able to distinguish the four main technological families (AI, ML, DL and GenAI), understand where a system's learning signal comes from, and know what changes internally in each type of algorithm during training. No prior technical experience is required; familiarity with the terminology will make the following chapters easier to follow. The chapter closes with MLOps: the engineering practices that turn a trained model into a reliable product.

Artificial intelligence is not a “mind” or an autonomous entity. It is a family of systems built to optimize a task against a measurable objective, often using data and some form of learning.

AI systems may classify, predict, make decisions or, more recently, generate content.

To separate products, models and marketing claims, we'll use a simple framework that works across modern AI systems.

Any “AI system” can be understood by answering:

  • What type of AI application is it? Which family or technology does it use?
  • How does it learn? Where does the learning signal come from?
  • What changes during training? Which parts of the model are adjusted?

We'll take those questions in order.


1. The General Framework: AI, ML, DL and GenAI

A useful way to see the hierarchy is:

  • Artificial Intelligence (AI): the broadest concept. It refers to machines or software that imitate capabilities associated with human intelligence in order to reason, solve problems or make decisions. In the human-body analogy, AI would be the “brain”.

  • Machine Learning (ML): a branch of AI that lets systems learn from data and improve with experience, instead of depending on explicit rules. In the analogy, ML would be the “training” of that brain.

  • Deep Learning (DL): a specialized type of ML that uses neural networks with many layers to handle complex data such as images, audio or language. In the analogy, those layers resemble the interconnected neurons of the brain.

  • Generative Artificial Intelligence (GenAI): a part of DL focused on generating content (text, image, audio, code).

AI / ML / DL describe the system's technological family. In this series, “generative” is used as a practical label for systems whose primary output is new content, although technically it also refers to a family of generative models that model distributions and generate samples.

AI
The umbrella: deciding and planning
Broad framework: techniques that let a system reason, plan or decide. It can use rules, search, optimization or (sometimes) learning. If the focus is making decisions/plans without “training a model”, it is usually AI.
Typical input
•Goal (what you want to achieve)
•Constraints (time, cost, rules)
•Environment state (map, resources, conditions)
Typical output
•Decision (what to do)
•Plan / route (ordered steps)
•Action / policy (how to act in each case)
Capabilities
-Learns from data -Neural network -Generates content
  • 🤖 Robot vacuum: efficient route
    Plans where to go to cover the house while avoiding obstacles.
  • ♟️ Sudoku/chess: best move
    Explores options and chooses the one that maximizes a score.
  • 🗓️ Shop shifts
    Assigns schedules while satisfying rules and constraints.
  • 🚚 Delivery: multi-stop route
    Optimizes visit order to minimize time or distance.
  • 📋 Rule-based support (FAQ)
    If it detects X, it recommends Y by following a defined flow.
  • 🚦 Coordinated traffic lights
    Adjusts cycles with control logic to reduce congestion.

We now know how to identify which technological family a system uses. The second question is: where does its learning signal come from?

2. How do these systems learn?

These are not model types. They describe how the learning signal is constructed—in other words, where the supervision comes from.

Supervised / unsupervised / self-supervised / reinforcement learning (RL) describe where the learning signal comes from.

Where the teacher comes from
Quick rule: change the teacher and you change the way the system learns.
Core idea

It learns by comparing its answer with the correct one.

When it fits

When you can obtain reliable labels and measure the error.

What it returns

A prediction: class, value or probability.

Example

Spam, churn, default risk.

That answers where the learning signal comes from. The third question completes the framework: what exactly changes inside the model during training? The answer depends on the algorithm family.

3. What changes during training

Knowing where the learning signal comes from (supervised / self-supervised / RL) is not enough.

The next question is what changes inside the model and how those changes improve performance.

3.1 The universal learning loop

  1. Train on the current data and make predictions.
  2. Measure the error (or how well it separates / groups).
  3. Adjust something internal to reduce that error.
  4. Repeat many times.

Learning = changing internal parameters to make fewer mistakes on data similar to the training data.

3.2 Why a model can learn today and work tomorrow

Models are trained on a sample of the world (the data available today) and are expected to capture general patterns that continue to hold in the future.

  • If future data is similar, the model generalizes well.
  • If it changes substantially (data drift), performance can fall, so the system needs monitoring and sometimes retraining.

What gets adjusted matters because different algorithm families learn in different ways. AI systems are not static products: they need ongoing monitoring and maintenance.


3.3 What is adjusted depending on the type of algorithm

Think of each algorithm as a system with its own kind of parameters. Training means updating those parameters so that its predictions improve.

1. Adjust rules / decisions: Decision Trees, Random Forest, XGBoost

What changes internally:

  • The questions it asks (which variable to inspect).
  • The thresholds for those questions (e.g. “more than X?”).
  • The structure of the tree (which branches exist and how deep it goes).

It is like building a questionnaire: “if A happens, ask B; otherwise ask C”.

Loan-approval example:

  • First candidate rule: “monthly income above X?” This separates applications with greater repayment capacity.
  • Then: “debt ratio below Y?” This refines the separation.
  • Training = trying many questions/thresholds and keeping those that best distinguish approvable applications from those that should be rejected.
Decision Trees
Quick guide

Ask questions first; decide at the end

Training chooses splits. Prediction follows a path down to a final leaf.

Train: choose useful splits. Predict: reach a leaf.

2. Adjust probabilities learned by counting: Naive Bayes

What changes internally:

  • Tables of frequencies/probabilities: which signals appear more often in each class.
  • It treats signals as almost independent given the class, so the evidence can be combined with a simple calculation.

It is like keeping a count: “when it is spam, how often do I see ‘free’? How often do I see ‘urgent’?”

Spam example:

  • If “free” appears very often in spam and rarely in non-spam, that pushes the prediction toward spam.
  • Training = updating those counts with many examples and turning them into probabilities.
Classification with Naive Bayes
Quick guide

The prediction comes from combining simple clues

Count which words appear most often, then combine those signals to decide.

First it learns from short examples. Then it compares which signals carry more weight.

3. Adjust groups by feature similarity: Clustering, k-means

What changes internally:

  • The position of the group “centers” (each center is called a prototype).

Note: what “similar” means depends on the distance you use and on how you scale the variables.

It is like placing magnets on a map: each data point goes to the nearest magnet, then you move the magnets to the center of each group.

Example:

  • Group customers by behavior (frequency, spending, channels) without prior labels, using only the raw data.
  • Training = moving the centers so that points are as close as possible to their group (more similar customers, closer together).
Clustering Algorithms
Quick guide

The groups reposition themselves until they fit the customers

First assign each customer to the nearest group. Then reposition each group and repeat until almost nothing changes.

Each color groups similar customers. Trying more or fewer groups changes the result.

4. Adjust numerical weights (Neural networks)

What changes internally:

  • The weights (and biases) in the connections are numbers that determine how much influence each input signal has when signals are combined.
  • In deep networks, there are millions of weights distributed across layers.

Each neuron calculates a weighted sum and then applies an activation function, which lets the system learn non-linear concepts.

Technical deep dive (optional)

Two typical roles of the activation function:

  • In internal layers: it adds non-linearity (capacity).
  • At the output: it turns a score into something interpretable (e.g. a probability with sigmoid/softmax).

Why the activation function matters:

  • Without activation, several consecutive layers would be equivalent to a single linear transformation, so the model would be too rigid.
  • The activation introduces non-linearity, which lets the model capture relationships such as “if A and B happen, but not C…”, curves, soft thresholds, etc.
  • It also affects training: the type of activation influences how easy or difficult it is to adjust weights in deep layers.

Four core pieces make those updates possible:

  • Activation function: lets systems learn non-linear concepts.
  • Loss function: measures how wrong the model's output is.
  • Backpropagation: propagates the error signal backward through the network to determine how the weights should change.
  • Optimizer: decides how much to move each weight at each step (small, repeated steps).

Spam example:

  • Signals: “free”, “urgent”, “many links”…
  • The network combines signals with weights, passes through activations and produces a score/probability.
  • If it fails, it adjusts weights/biases so that next time “free” carries more or less weight, etc.
Step 1 · 1 neuron

A single neuron learns linear relationships

It adjusts only one weight (slope) and one bias (offset). That is enough when the data follows a straight line.

Temperature: °C → °F · real points · — fitted curve

What it adjusts 1 weight + 1 bias → 2 parameters
Limit It can only learn straight lines

Epoch — one complete pass through all the training data. Loss — how far the predictions are from the true values; the lower it is, the better the fit.


3.4 What type of data each family is useful for

Not every family is equally suitable for every problem. The type of data is often the first decision filter:

Family Data where it works well Where it fails or is not the first choice
Trees (Decision Tree, Random Forest, XGBoost) Structured tabular data: numbers, categories, mixed variables. A favorite for business data and Kaggle competitions with tables. Images, audio, raw text without preprocessing.
Naive Bayes Text (bag of words, token frequencies), categorical data with few correlations between variables. Very fast with little data. Continuous data with strong correlations; complex relationships between variables.
K-means (clustering) Continuous numerical data where Euclidean distance makes sense: coordinates, scaled behavioral metrics. Text, high-dimensional data without prior reduction, purely categorical variables.
Neural networks Images, audio, text, time series, video. They perform especially well when data volumes are large and the pattern is complex. Small tabular datasets: trees often win with lower computational cost.

These four families illustrate the spectrum of adjustment mechanisms, not the whole map. There are dozens more: SVMs, logistic/linear regression, Gaussian mixture models, Bayesian networks, time-series models (ARIMA, Prophet), ensemble methods, etc. Choosing an algorithm always starts by understanding the type of data and the objective of the problem.

Together, these three axes (technological family, learning type, adjustment mechanism) let us describe any modern AI system. They also expose a deeper distinction: where the system's logic comes from.

4. Classical software vs AI

Everything above changes where a solution's logic comes from. The distinction is not yet about how developers write code; it is about whether the rules are written explicitly or learned from data.

In classical software:

  • Input data + human-written rules → output

Example: converting Fahrenheit to Celsius with a fixed formula.

  • The programmer explicitly writes the rule: C = (F - 32) x 5/9
  • If the same Fahrenheit value comes in, the same Celsius value always comes out.

In AI:

  • Input data + output data → learned rules
  • The “algorithm”, the mathematical formula, emerges from training.

Example: using many (Fahrenheit, Celsius) pairs so that the system learns the conversion.

  • You no longer write the exact formula by hand.
  • The model adjusts parameters and learns an approximate rule that then generalizes to new values.

This is the basic principle of so-called Software 2.0: the logic is no longer written, it is learned.

This does not yet change how software itself is built, but how solutions are built using AI. LLMs will later change how software itself is developed.

This shift emerged over decades of advances, failures and technical leaps that explain where we are today and where we are going.

5. Major milestones

There is no need to memorize the whole chronology. The important thing is to see what changed in each wave.

Date Major milestone What changes
1950 Turing (paper) Establishes the conceptual framework for “intelligence in machines”.
1955–1956 Dartmouth Conference (proposal) AI becomes a formal research field.
1958–1959 Perceptron and early demonstrations of machine learning (paper · Rosenblatt, paper · Samuel) Demonstrates learning from data as an alternative to relying only on rules.
1980s Expert systems (XCON case) First business wave of rule-based AI.
1986 Backpropagation (paper) Training multilayer neural networks becomes viable.
1997 Deep Blue (IBM) A specialized AI defeats the world chess champion and makes the power of narrow AI visible.
2012 AlexNet + ImageNet (paper, ILSVRC) The modern era of deep learning scaled with data and GPUs begins.
2017 Transformer — “Attention is all you need” (paper) Introduces the base architecture used by modern language models.
2020–2022 GPT-3, AlphaFold and ChatGPT (article · GPT-3, CASP14, Nature paper · AlphaFold, announcement · ChatGPT) Foundation models scale, AI delivers direct scientific impact and generative AI reaches mass adoption.

If you want an even simpler mental picture, read it like this:

rules -> statistical learning -> deep learning -> foundation models -> useful AI at scale

Those advances produce capable and efficient systems. But a capable model is not yet a product. For an AI system to work reliably in the real world, a complete engineering cycle is required.

6. MLOps: the complete cycle for AI to work in the real world

MLOps is the engineering discipline that keeps AI systems reliable in the real world: not only today, but also three months from now when the data, the market or user behavior has changed. (Google Cloud) The key idea:

A trained model ≠ AI in production. In production you need a complete cycle: data → training → deployment → monitoring → improvement.

The clearest way to understand MLOps is as a chain of 8 steps. If one is missing, you normally have a demo, not a product.

  1. Data (capture): collect signals from the real world (events, transactions, documents, logs).
  2. Data (prepare): clean and turn data into useful variables (what the model “understands”).
  3. Train: the model learns patterns from historical examples.
  4. Evaluate: check that it works “well enough” before touching production.
  5. Version: record which model it is and which data/version it was trained with (traceability).
  6. Deploy: put it to work (API or batch), ideally gradually.
  7. Monitor: check whether the world changes (data), whether the system is healthy (latency/errors) and whether performance falls.
  8. Feedback and improvement: when the ground truth arrives (real labels), correct, retrain or roll back.
MLOps at a glance Training is not enough: you need the full lifecycle.
First you prepare and train. Then you deploy, monitor, and iterate again.
1 Data 2 Prepare 3 Train 4 Evaluate 5 Version 6 Deploy 7 Monitor 8 Feedback lifecycle

Step

Description

Example (fraud) ...
What comes out of this step ...
If it is not implemented correctly ...

Next reading

The next chapter goes deeper into the most disruptive type of system of the last decade: Chapter 2 — What is Generative AI? →

7. References

Core sources
Key Source Short description
R1 OECD — Explanatory Memorandum on the Updated OECD Definition of an AI System (OECD) Clarifies the modern definition of an AI system.
R2 ISO/IEC 22989:2022 — Artificial intelligence — Concepts and terminology (ISO) Core vocabulary and concepts in the field.
R3 Tom M. Mitchell — Machine Learning (CMU School of Computer Science) Formally defines learning with E/T/P.
R4 Y. LeCun, Y. Bengio, G. Hinton (2015) — Deep Learning (Nature) Short overview of the deep-learning revolution.
R5 I. Goodfellow, Y. Bengio, A. Courville — Deep Learning (Deep Learning Book) Technical foundation for modern neural networks.
R6 R. S. Sutton, A. G. Barto — Reinforcement Learning: An Introduction (Incomplete Ideas) Classical reference for reinforcement learning.
R7 R. Wirth, J. Hipp — CRISP-DM: Towards a Standard Process Model for Data Mining (cs.unibo.it) Standard process for data and ML projects.
R8 D. Sculley et al. (2015) — Hidden Technical Debt in Machine Learning Systems (NeurIPS Papers) Explains why a model alone is not enough for a real system.

Milestone references

Milestone sources
Milestone Source Short description
H1 Alan Turing — Computing Machinery and Intelligence (paper) Conceptual framework for intelligence in machines.
H2 Dartmouth Summer Research Project on Artificial Intelligence (1955) (proposal) Foundational act of the field of AI.
H3 Frank Rosenblatt — The Perceptron: A Probabilistic Model for Information Storage and Organization in the Brain (paper) First major step toward learning from data.
H4 Arthur Samuel — Some Studies in Machine Learning Using the Game of Checkers (paper) First practical demonstration that a machine can learn to play better than its programmer.
H5 XCON / expert systems in production at Digital Equipment Corporation (XCON case) Industrial rise of rule-based AI.
H6 Rumelhart, Hinton, Williams — Learning representations by back-propagating errors (paper) Makes it viable to train multilayer neural networks.
H7 IBM — Deep Blue vs Kasparov (1997) (IBM) Demonstrates the power of specialized AI in a concrete domain.
H8 AlexNet — ImageNet Classification with Deep Convolutional Neural Networks (2012) (paper) Triggers the modern era of visual deep learning.
H9 ImageNet / ILSVRC (ILSVRC) Benchmark that accelerates progress in computer vision.
H10 AlphaGo (2016) — Mastering the game of Go with deep neural networks and tree search (Nature paper) Combines deep learning and search: first system to surpass the best humans at Go.
H11 AlphaGo vs Lee Sedol (2016) (DeepMind page) The public milestone that made that technical leap visible globally.
H12 Attention Is All You Need — Transformer (2017) (paper) Base architecture of today's language and image models.
H13 BERT (2019) — Pre-training of Deep Bidirectional Transformers for Language Understanding (paper) Consolidates bidirectional pretraining in NLP.
H14 GPT-3 (2020) — Language Models are Few-Shot Learners (OpenAI article) Scales language foundation models to an unprecedented level.
H15 CASP14 / AlphaFold (2020–2021) (CASP14, Nature paper, DeepMind blog) Solves the protein-folding problem: the first direct scientific impact of AI at that scale.
H16 ChatGPT (2022) (OpenAI announcement) Popularizes generative AI at scale with the general public.
H17 University of Reading — Turing Test 2014 (article) Popular reference in the debate around the Turing test.
H18 AlphaGo — The Movie (2017, dir. Greg Kohs) (documentary · YouTube) Documentary following the preparation for the match against Lee Sedol from the inside. Recommended for understanding the human and technical impact of the milestone.
H19 AlphaFold: The making of a scientific breakthrough — DeepMind (2021) (video · YouTube) DeepMind documentary video about the process and impact of AlphaFold. Recommended before reading the paper.

Frequently asked questions

What is the difference between Deep Learning and Machine Learning? Deep Learning is a specialized type of Machine Learning that uses neural networks with many layers to handle complex data such as images, audio or language. ML learns patterns in order to generalize; DL uses layered neural networks for advanced perception tasks where simpler algorithms often fall short.

How can a system learn without anyone telling it the correct answer? Through two different routes. In unsupervised learning, the system looks for structure in unlabeled data, such as grouping customers by behavior. In self-supervised learning, the data itself creates the signal: for example, a language model predicts the next token using the preceding text as the “answer”, without any human having annotated it.

Why is logic learned rather than written in AI? In traditional software the programmer defines the rules explicitly. In AI, input and output data cause the logic to emerge from training: the programmer designs the learning process, but the concrete logic that solves the problem is discovered by the model itself. You do not write the formula; you learn it.

What physically changes inside a neural network when the model learns? Millions of numerical weights distributed across the connections between layers are adjusted. Backpropagation propagates the error signal through the network, and the optimizer decides how much to move each weight at each step to reduce the accumulated error.