In plain English
The transformer is the blueprint most of today's AI models are built on. Its key trick is paying "attention" to how every word in a passage relates to every other word, which helps it understand context.
In practice
You rarely work with transformers directly, but the design explains practical limits. Attention is why models have a context window, and why very long inputs cost more and run slower.
Under the hood
Introduced in the 2017 paper "Attention Is All You Need", the transformer uses multi-head self-attention and feed-forward layers instead of recurrence, so it trains efficiently in parallel. The "T" in GPT stands for transformer. Self-attention cost grows quadratically with sequence length.
Example
"Almost every major language model today is a transformer."