๐Ÿ“– Glossary

Every bit of jargon from the lessons, in plain English.

Activation function
A function applied to a neuron's weighted sum (e.g. ReLU, sigmoid) that adds non-linearity, letting stacked layers learn complex patterns.
Agent
An AI system that reasons and takes actions, often by calling tools in a loop, to complete a goal.
Attention
A mechanism that lets each token weigh and pull information from other tokens. The heart of the transformer.
Backpropagation
The algorithm that computes every weight's gradient by passing error signals backwards through a network using the chain rule.
Batch / mini-batch
A small group of training examples processed together in one optimisation step.
Bias (in data)
Systematic skew in data or outcomes that can make a model unfair or unreliable for some groups.
Classification
Predicting a category, like spam vs. not spam.
CNN
Convolutional neural network. Uses sliding filters to detect local patterns, mainly in images.
Context window
How many tokens an LLM can consider at once, both prompt and response.
Diffusion model
A generative model that learns to turn noise into data by reversing a gradual noising process.
Embedding
A vector of numbers representing something (a word, image, user) so that similar things are close together.
Epoch
One full pass through the training dataset.
Fine-tuning
Further training a pre-trained model on a smaller, task-specific dataset.
Gradient
The slope of the loss with respect to each parameter; it points in the direction of steepest increase.
Gradient descent
An optimisation method that repeatedly steps parameters opposite to the gradient to reduce loss.
Hallucination
When a model confidently produces false or invented information.
Inference
Using a trained model to make predictions, with its parameters frozen.
Learning rate
The step size used in gradient descent.
LLM
Large language model. A big transformer trained on text to predict the next token.
LoRA
Low-Rank Adaptation. Fine-tunes by training small low-rank matrices while freezing the base weights.
Loss function
A number measuring how wrong a model's predictions are. Training minimises it.
Overfitting
When a model memorises training data, including its noise, and performs poorly on new data.
Parameters / weights
The learned numbers inside a model that determine its behaviour.
Precision
Of the items predicted positive, the fraction that truly are positive: TP / (TP + FP).
Prompt
The input text you give an LLM to steer its output.
Quantisation
Storing model weights with fewer bits (e.g. 4-bit) to save memory and speed up inference.
RAG
Retrieval-augmented generation. Fetches relevant documents and includes them in the prompt so answers are grounded.
Recall
Of the truly positive items, the fraction the model caught: TP / (TP + FN).
Regression
Predicting a continuous number, like a price or temperature.
Regularisation
Techniques (L2 penalty, dropout, etc.) that discourage overly complex models to reduce overfitting.
Reinforcement learning
Learning to act by trial and error to maximise reward from an environment.
RLHF
Reinforcement learning from human feedback. Uses human preference data to fine-tune model behaviour.
Temperature
A sampling setting controlling randomness. Lower is more deterministic, higher is more varied.
Token
A chunk of text (a word, part of a word or punctuation) that language models read and write.
Transformer
A neural network architecture built on attention, used by virtually all modern LLMs.
Validation set
Data held out from training to tune settings and check for overfitting.

Want more? Google's Machine Learning Glossary covers hundreds of terms.