Skip to content
RichseenAtlasAtlasSign in

The transformer architecture

technology

The neural network design published by Vaswani and seven co-authors in 2017 and presented at NeurIPS in Long Beach that December, which replaced recurrence and convolution with attention alone. Its consequence is as much economic as scientific: because a transformer processes a whole sequence in parallel rather than one step at a time, training scales with available accelerators instead of with sequence length, and that single property is what made it worth spending vastly more compute on a language model than had ever been spent before. The architecture was introduced for machine translation and then proved indifferent to what it was fed — the same design underlies contemporary language models, most protein and molecular models, and much of image and audio generation. It is the clearest case in the field of an engineering choice about parallelism becoming, by way of the hardware it fits, a change in what was attemptable.

A technology is not a place. Drawing it on a map would assert something about the world that no stored fact supports.

Read in · 1

Evidence · 1
Timeline

No dated observations are stored for this object. Atlas shows what was observed and when — it does not infer a history.

Connections · 2
Assembled narrative · 1

Assembled from 29 blocks · 1 evidence · 43 related

  1. Story
  2. The neural network design published by Vaswani and seven co-authors in 2017 and presented at NeurIPS in Long Beach that December, which replaced recurrence and convolution with attention alone. Its consequence is as much economic as scientific: because a transformer processes a whole sequence in parallel rather than one step at a time, training scales with available accelerators instead of with sequence length, and that single property is what made it worth spending vastly more compute on a language model than had ever been spent before. The architecture was introduced for machine translation and then proved indifferent to what it was fed — the same design underlies contemporary language models, most protein and molecular models, and much of image and audio generation. It is the clearest case in the field of an engineering choice about parallelism becoming, by way of the hardware it fits, a change in what was attemptable.
  3. Knowledge
  4. Deep learning
  5. The transformer architecture
  6. The AI accelerator
  7. Connections
  8. The AI accelerator
  9. Deep learning
  10. Machine translation
  11. Machine-generated code
  12. Growth in frontier training compute
  13. The transformer is presented at NeurIPS
  14. GLUE is saturated and SuperGLUE is built to replace it
  15. Deep learning
  16. The transformer architecture
  17. Training compute
  18. Taiwan
  19. A deep convolutional network wins the ImageNet challenge
  20. The AI accelerator
  21. ImageNet
  22. The transformer architecture
  23. AlphaFold
  24. Machine-learned weather forecasting
  25. Machine learning in medical imaging
  26. A deep convolutional network wins the ImageNet challenge
  27. AlphaGo defeats Lee Sedol in Seoul
  28. Evidence
  29. The transformer replaces recurrence and convolution with attention alone, is more parallelisable, and was shown superior in quality on two machine translation tasks while requiring significantly less time to train. V55 verification basis: the proceedings were not fetched; search results carried the author list, the venue, the Long Beach dates of 4–9 December 2017, the page range and the substance of the abstract and attributed them to the NeurIPS proceedings. The paper's arXiv identifier, its 2017 preprint date and its reported BLEU scores were explicitly NOT confirmed by retrieval and are not asserted anywhere in this pack.
Close the narrative
Observed changes · 0

No public Signals are attached to this object. Signals show what changed and when it was observed — never a direction or a rank.

Actions

Read the assembled narrativeContinue in StudioOpen TwinTwin does not start a decision from this kind of object.ShareSaveSaved objects are part of the authenticated projection, which is declared and not yet built.

/atlas?object=TECH_TRANSFORMER&experience=TECH_TRANSFORMER