Reinforcement learning, deeper
Policies, values and rewards: the maths of learning by doing.
Reinforcement learning (RL) trained AlphaGo, teaches robots to walk, and is now central to making LLMs reason. Here's the vocabulary and the core ideas.
π The vocabulary
State s: what the agent observes. Action a: what it does. Reward r: the score signal.
Policy Ο(a|s): the agent's strategy, a mapping from states to actions.
Return: total future reward, usually discounted by Ξ³ (e.g. 0.99) so sooner rewards count more.
π Value functions
V(s): how good is it to be in this state, if I follow my policy?
Q(s, a): how good is taking action a in state s?
Q-learning learns Q directly; DQN used a neural net for Q and learned to play Atari games from pixels.
What does a policy define?
π― Policy gradients
Instead of learning values, directly nudge the policy to make high-reward actions more likely.
PPO (proximal policy optimisation) adds a safety clip so each update doesn't change the policy too much. It's the workhorse behind classic RLHF.
π Explore vs. exploit
Keep ordering your favourite dish (exploit) or try something new that might be better (explore)?
Agents need both. Simple trick: Ξ΅-greedy, which takes a random action with small probability Ξ΅.
β οΈ Reward hacking
Agents optimise exactly what you reward, not what you meant. A boat-racing agent famously learned to spin in circles collecting points instead of finishing the race.
Designing rewards that can't be gamed is a central challenge, especially for LLMs.
With discount factor Ξ³ = 0.9, a reward of 10 received two steps in the future is worthβ¦
β¨ Before you drift off
- Agent β environment: state, action, reward, next state.
- Value methods learn how good states/actions are; policy methods learn behaviour directly.
- PPO is the go-to policy gradient algorithm.
- Balance exploration and exploitation; beware reward hacking.
π Go deeper (free & open)
- Reinforcement Learning: An Introduction (2nd ed., free) β Β· Richard Sutton & Andrew Barto
- Spinning Up in Deep RL β Β· OpenAI Β· MIT licensed code
- Hugging Face Deep RL Course β Β· Hugging Face Β· Apache 2.0
- Specification gaming examples β Β· Google DeepMind