๐ Glossary
Every bit of jargon from the lessons, in plain English.
- Activation function
- A function applied to a neuron's weighted sum (e.g. ReLU, sigmoid) that adds non-linearity, letting stacked layers learn complex patterns.
- Agent
- An AI system that reasons and takes actions, often by calling tools in a loop, to complete a goal.
- Attention
- A mechanism that lets each token weigh and pull information from other tokens. The heart of the transformer.
- Backpropagation
- The algorithm that computes every weight's gradient by passing error signals backwards through a network using the chain rule.
- Batch / mini-batch
- A small group of training examples processed together in one optimisation step.
- Bias (in data)
- Systematic skew in data or outcomes that can make a model unfair or unreliable for some groups.
- Classification
- Predicting a category, like spam vs. not spam.
- CNN
- Convolutional neural network. Uses sliding filters to detect local patterns, mainly in images.
- Context window
- How many tokens an LLM can consider at once, both prompt and response.
- Diffusion model
- A generative model that learns to turn noise into data by reversing a gradual noising process.
- Embedding
- A vector of numbers representing something (a word, image, user) so that similar things are close together.
- Epoch
- One full pass through the training dataset.
- Fine-tuning
- Further training a pre-trained model on a smaller, task-specific dataset.
- Gradient
- The slope of the loss with respect to each parameter; it points in the direction of steepest increase.
- Gradient descent
- An optimisation method that repeatedly steps parameters opposite to the gradient to reduce loss.
- Hallucination
- When a model confidently produces false or invented information.
- Inference
- Using a trained model to make predictions, with its parameters frozen.
- Learning rate
- The step size used in gradient descent.
- LLM
- Large language model. A big transformer trained on text to predict the next token.
- LoRA
- Low-Rank Adaptation. Fine-tunes by training small low-rank matrices while freezing the base weights.
- Loss function
- A number measuring how wrong a model's predictions are. Training minimises it.
- Overfitting
- When a model memorises training data, including its noise, and performs poorly on new data.
- Parameters / weights
- The learned numbers inside a model that determine its behaviour.
- Precision
- Of the items predicted positive, the fraction that truly are positive: TP / (TP + FP).
- Prompt
- The input text you give an LLM to steer its output.
- Quantisation
- Storing model weights with fewer bits (e.g. 4-bit) to save memory and speed up inference.
- RAG
- Retrieval-augmented generation. Fetches relevant documents and includes them in the prompt so answers are grounded.
- Recall
- Of the truly positive items, the fraction the model caught: TP / (TP + FN).
- Regression
- Predicting a continuous number, like a price or temperature.
- Regularisation
- Techniques (L2 penalty, dropout, etc.) that discourage overly complex models to reduce overfitting.
- Reinforcement learning
- Learning to act by trial and error to maximise reward from an environment.
- RLHF
- Reinforcement learning from human feedback. Uses human preference data to fine-tune model behaviour.
- Temperature
- A sampling setting controlling randomness. Lower is more deterministic, higher is more varied.
- Token
- A chunk of text (a word, part of a word or punctuation) that language models read and write.
- Transformer
- A neural network architecture built on attention, used by virtually all modern LLMs.
- Validation set
- Data held out from training to tune settings and check for overfitting.
Want more? Google's Machine Learning Glossary covers hundreds of terms.