Skip to content
AI Atlas

TIMELINE

Timeline of AI

From the Turing test to diffusion models: the events that actually changed the direction.

  • Breakthrough
  • Foundation
  • Setback
1950

In “Computing Machinery and Intelligence”, Turing replaced “can machines think?” with an operational question: if a conversation cannot reveal whether the other party is human, on what grounds do we deny it thought? That reframing shaped seventy years of debate.

1956

Dartmouth: the term “AI” is coined

McCarthy, Minsky, Shannon and others named the field in a summer workshop, optimistically expecting a handful of bright people to solve most of it in one season. It was the founding of a discipline — and of its first overestimation.

1957

The Perceptron

Deep Learning

Rosenblatt built the first learning machine that adjusted its weights from data, and did it in hardware. It proved that learning from examples was physically realisable — the first cornerstone of the neural-network line.

1969

“Perceptrons” exposes the single-layer limit

Deep Learning

Minsky and Papert proved a single-layer perceptron cannot represent XOR or any linearly inseparable function. The result was correct, but widely misread as “neural networks are a dead end”, starving the field of funding for over a decade.

1980

Expert systems go commercial

Machine Learning

Rule-base-plus-inference-engine systems entered industry (such as DEC’s XCON). They proved AI could pay for itself, and exposed the ceiling of the purely symbolic route: rules had to be written by hand and could not be learned from data.

1986

Backpropagation popularised for multilayer networks

Deep Learning

Rumelhart, Hinton and Williams showed that propagating error backwards through the chain rule could actually train networks with hidden layers. XOR fell, and the neural-network route reopened.

1989

LeNet and the convolutional network

Computer Vision

LeCun applied convolution and weight sharing to handwritten-digit recognition and deployed it in a real postal system. The core structure of every vision model today was already conceived more than three decades ago.

1997

LSTM, and Deep Blue

Reinforcement Learning

Two things happened: Hochreiter and Schmidhuber introduced LSTM, whose gates alleviate vanishing gradients in recurrent networks; and IBM’s Deep Blue beat world chess champion Kasparov by brute-force search — a victory for search, not for learning.

2006

“Deep learning” enters the lexicon

Deep Learning

Hinton proposed deep belief networks and layer-wise pre-training, showing that deep networks could be trained. The specific trick was later superseded by better initialisation and normalisation, but depth itself returned to centre stage.

2009

The ImageNet dataset

Computer Vision

Fei-Fei Li’s team released ImageNet with fourteen million labelled images, plus a competition. In hindsight the scale of the data, not any single algorithm, was the detonator.

2012

AlexNet wins

Computer Vision

AlexNet cut top-5 error on ImageNet from 26% to 15.3%, while the runner-up still used hand-crafted features. GPUs, big data and deep networks converged in that moment, and deep learning became mainstream.

2014

GANs, and attention takes shape

Generative AI

Goodfellow introduced generative adversarial networks, pitting two networks against each other to approach a data distribution. The same year, Seq2Seq and attention appeared, solving the loss of information when long sequences are squeezed into a single vector.

2015

ResNet takes networks to 152 layers

Deep Learning

Residual connections let gradients bypass layers directly, so training hundred-layer networks no longer degraded. It pushed ImageNet top-5 error to 3.57% — below the human rate of roughly 5%. Depth stopped being the variable that needed explaining.

2016

AlphaGo defeats Lee Sedol

Reinforcement Learning

Go’s state space dwarfs chess, defeating brute-force search. AlphaGo evaluated positions with deep networks, decided with Monte Carlo tree search, and improved by playing itself. The famous move in game four showed it had learned something other than human play.

2017

Transformer: attention is all you need

NLP & Large Language Models

“Attention Is All You Need” replaced recurrence with pure attention, making training parallelisable and dependencies explicit. Nearly every large model today rests on this architecture — the single most influential AI paper of the century so far.

2018

BERT and GPT diverge

NLP & Large Language Models

BERT used bidirectional masking for understanding tasks, GPT autoregressive modelling for generation. Their divergence was really two answers to one question: should the model see the whole sentence before answering, or emit one token at a time?

2020

AlphaFold 2 and GPT-3

Generative AI

AlphaFold 2 reached experimental accuracy in protein structure prediction, widely described as solving a fifty-year problem. That same year GPT-3, with 175 billion parameters, demonstrated in-context learning: new tasks solved from a few prompt examples, with no parameter updates.

2021

CLIP: unifying vision through language

Computer Vision

CLIP aligned image and text representations by contrastive learning over four hundred million pairs. The result was zero-shot classification: categorising with no training examples at all, using only the textual names of the classes. Multimodality became the dominant paradigm.

2022

Diffusion models reach the public

Generative AI

Stable Diffusion moved the diffusion process into latent space and open-sourced it, putting image generation within reach of any consumer GPU. Generative AI turned from research demo into mass-market tool — and brought its controversies with it.

RLHF makes language models usable

AI Engineering, Safety & Ethics

ChatGPT’s key ingredient was not merely a bigger model but training a reward model on human preferences and then steering output with reinforcement learning. The same base model went from “can continue text” to “can be helpful”. It was the first time alignment produced a difference the public could feel.

2023

The open-model wave, and tool use

NLP & Large Language Models

The LLaMA family and many open-weight models brought usable language models to local machines, while quantisation and LoRA made fine-tuning possible on consumer hardware. Models also began calling external tools and retrieving documents — turning from text generators into components of a larger system.

2024

Video, 3D, and compute at inference time

Generative AI

Generation extended from images to video and 3D. Along another axis, models began to “think longer” at inference time, trading extra sampling and search for more reliable answers. A second route to improvement appeared beside pure scaling.