The transformer architecture
technology
The neural network design published by Vaswani and seven co-authors in 2017 and presented at NeurIPS in Long Beach that December, which replaced recurrence and convolution with attention alone. Its consequence is as much economic as scientific: because a transformer processes a whole sequence in parallel rather than one step at a time, training scales with available accelerators instead of with sequence length, and that single property is what made it worth spending vastly more compute on a language model than had ever been spent before. The architecture was introduced for machine translation and then proved indifferent to what it was fed — the same design underlies contemporary language models, most protein and molecular models, and much of image and audio generation. It is the clearest case in the field of an engineering choice about parallelism becoming, by way of the hardware it fits, a change in what was attemptable.
A technology is not a place. Drawing it on a map would assert something about the world that no stored fact supports.
Read in · 1
Evidence · 1
Timeline
No dated observations are stored for this object. Atlas shows what was observed and when — it does not infer a history.
Connections · 2
- The AI acceleratordepends_on
- Deep learninginstance_of
Assembled narrative · 1
An assembled narrative is available for this object.
Observed changes · 0
No public Signals are attached to this object. Signals show what changed and when it was observed — never a direction or a rank.
Actions
/atlas?object=TECH_TRANSFORMER