Seeing and remembering: CNNs & RNNs

Special network shapes for images and sequences, and why they mattered.

โฑ 7 min read

A plain network treats every input pixel as unrelated. But pixels next to each other belong together, and words in a sentence come in order. Two architectures were built to respect that structure.

imagefilter-101-101-101slides overevery patchfeature mapbright = edge found
A 3ร—3 filter slides across the image, producing a feature map that lights up wherever its pattern (here, a vertical edge) appears.

๐Ÿ–ผ๏ธ Convolutional Neural Networks (CNNs)

A small filter (say 3ร—3) slides across the image, checking each patch for a pattern like an edge.

The same filter is reused everywhere, so a cat ear is recognised wherever it appears and the model needs far fewer weights.

๐Ÿ—๏ธ Feature hierarchy

Stack conv layers with pooling (downsizing) in between.

Early layers find edges โ†’ then textures โ†’ parts (eyes, wheels) โ†’ whole objects.

AlexNet's win in the 2012 ImageNet competition kick-started the deep learning boom.

๐Ÿงฉ Quick quiz

Why do CNNs reuse the same filter across the whole image?

๐Ÿ” Recurrent Neural Networks (RNNs)

Read a sequence one step at a time, carrying a hidden state, a running memory of what came before.

Used for text, speech, and time series before 2017.

๐Ÿง  LSTMs: better memory

Plain RNNs forget long-ago inputs (vanishing gradients again).

LSTMs add gates that learn what to remember, forget and output, so they hold on to context much longer.

โฉ Why transformers took over

RNNs read word by word, so they can't easily be parallelised and still struggle with long text.

Transformers (Advanced level!) look at every word at once with attention. Today they dominate language and are increasingly used for images too (Vision Transformers).

๐Ÿงฉ Quick quiz

What does an RNN's hidden state do?

โœจ Before you drift off

  • CNNs slide shared filters over images to detect local patterns.
  • Layers build from edges to objects.
  • RNNs/LSTMs process sequences with a running memory.
  • Transformers replaced RNNs for most language tasks.

๐Ÿ“š Go deeper (free & open)