Reinforcement learning, deeper

Policies, values and rewards: the maths of learning by doing.

⏱ 7 min read

Reinforcement learning (RL) trained AlphaGo, teaches robots to walk, and is now central to making LLMs reason. Here's the vocabulary and the core ideas.

πŸ€– Agentpolicy Ο€(a|s)🌍 Environmentgame, robot, worldaction aβ‚œnew state sβ‚œβ‚Šβ‚ + reward rβ‚œβ‚Šβ‚ ⭐goal: maximisetotal future reward
The agent observes a state, takes an action, and the environment returns a reward and a new state. Repeat.

πŸ“– The vocabulary

State s: what the agent observes. Action a: what it does. Reward r: the score signal.

Policy Ο€(a|s): the agent's strategy, a mapping from states to actions.

Return: total future reward, usually discounted by Ξ³ (e.g. 0.99) so sooner rewards count more.

πŸ’Ž Value functions

V(s): how good is it to be in this state, if I follow my policy?

Q(s, a): how good is taking action a in state s?

Q-learning learns Q directly; DQN used a neural net for Q and learned to play Atari games from pixels.

🧩 Quick quiz

What does a policy define?

🎯 Policy gradients

Instead of learning values, directly nudge the policy to make high-reward actions more likely.

PPO (proximal policy optimisation) adds a safety clip so each update doesn't change the policy too much. It's the workhorse behind classic RLHF.

πŸ” Explore vs. exploit

Keep ordering your favourite dish (exploit) or try something new that might be better (explore)?

Agents need both. Simple trick: Ξ΅-greedy, which takes a random action with small probability Ξ΅.

⚠️ Reward hacking

Agents optimise exactly what you reward, not what you meant. A boat-racing agent famously learned to spin in circles collecting points instead of finishing the race.

Designing rewards that can't be gamed is a central challenge, especially for LLMs.

🧩 Quick quiz

With discount factor Ξ³ = 0.9, a reward of 10 received two steps in the future is worth…

✨ Before you drift off

  • Agent ↔ environment: state, action, reward, next state.
  • Value methods learn how good states/actions are; policy methods learn behaviour directly.
  • PPO is the go-to policy gradient algorithm.
  • Balance exploration and exploitation; beware reward hacking.

πŸ“š Go deeper (free & open)