The transformer architecture
technology
The neural network design published by Vaswani and seven co-authors in 2017 and presented at NeurIPS in Long Beach that December, which replaced recurrence and convolution with attention alone. Its consequence is as much economic as scientific: because a transformer processes a whole sequence in parallel rather than one step at a time, training scales with available accelerators instead of with sequence length, and that single property is what made it worth spending vastly more compute on a language model than had ever been spent before. The architecture was introduced for machine translation and then proved indifferent to what it was fed — the same design underlies contemporary language models, most protein and molecular models, and much of image and audio generation. It is the clearest case in the field of an engineering choice about parallelism becoming, by way of the hardware it fits, a change in what was attemptable.
A technology is not a place. Drawing it on a map would assert something about the world that no stored fact supports.
Read in · 1
Evidence · 1
Timeline
No dated observations are stored for this object. Atlas shows what was observed and when — it does not infer a history.
Connections · 2
- The AI acceleratordepends_on
- Deep learninginstance_of
Assembled narrative · 1
Assembled from 29 blocks · 1 evidence · 43 related
- Story
- The neural network design published by Vaswani and seven co-authors in 2017 and presented at NeurIPS in Long Beach that December, which replaced recurrence and convolution with attention alone. Its consequence is as much economic as scientific: because a transformer processes a whole sequence in parallel rather than one step at a time, training scales with available accelerators instead of with sequence length, and that single property is what made it worth spending vastly more compute on a language model than had ever been spent before. The architecture was introduced for machine translation and then proved indifferent to what it was fed — the same design underlies contemporary language models, most protein and molecular models, and much of image and audio generation. It is the clearest case in the field of an engineering choice about parallelism becoming, by way of the hardware it fits, a change in what was attemptable.
- Knowledge
- Deep learning
- The transformer architecture
- The AI accelerator
- Connections
- The AI accelerator
- Deep learning
- Machine translation
- Machine-generated code
- Growth in frontier training compute
- The transformer is presented at NeurIPS
- GLUE is saturated and SuperGLUE is built to replace it
- Deep learning
- The transformer architecture
- Training compute
- Taiwan
- A deep convolutional network wins the ImageNet challenge
- The AI accelerator
- ImageNet
- The transformer architecture
- AlphaFold
- Machine-learned weather forecasting
- Machine learning in medical imaging
- A deep convolutional network wins the ImageNet challenge
- AlphaGo defeats Lee Sedol in Seoul
- Evidence
- The transformer replaces recurrence and convolution with attention alone, is more parallelisable, and was shown superior in quality on two machine translation tasks while requiring significantly less time to train. V55 verification basis: the proceedings were not fetched; search results carried the author list, the venue, the Long Beach dates of 4–9 December 2017, the page range and the substance of the abstract and attributed them to the NeurIPS proceedings. The paper's arXiv identifier, its 2017 preprint date and its reported BLEU scores were explicitly NOT confirmed by retrieval and are not asserted anywhere in this pack.
Observed changes · 0
No public Signals are attached to this object. Signals show what changed and when it was observed — never a direction or a rank.
Actions
/atlas?object=TECH_TRANSFORMER&experience=TECH_TRANSFORMER