Skip to content
RichseenAtlasAtlasSign in

The transformer is presented at NeurIPS

node

Vaswani and seven co-authors presented "Attention Is All You Need" at the 31st Conference on Neural Information Processing Systems in Long Beach, proposing a sequence architecture built on attention alone, without recurrence or convolution, and demonstrating it on two machine translation tasks with better quality and substantially less training time. The removal of sequential dependency was the consequential part: it made training throughput a function of how many accelerators could be pointed at the problem, and thereby made very large training runs an engineering question rather than a scheduling impossibility.

A node is not a place. Drawing it on a map would assert something about the world that no stored fact supports.

Read in · 1

Evidence · 1
Timeline

No dated observations are stored for this object. Atlas shows what was observed and when — it does not infer a history.

Connections · 2
Assembled narrative · 1

Assembled from 23 blocks · 1 evidence · 26 related

  1. Story
  2. Vaswani and seven co-authors presented "Attention Is All You Need" at the 31st Conference on Neural Information Processing Systems in Long Beach, proposing a sequence architecture built on attention alone, without recurrence or convolution, and demonstrating it on two machine translation tasks with better quality and substantially less training time. The removal of sequential dependency was the consequential part: it made training throughput a function of how many accelerators could be pointed at the problem, and thereby made very large training runs an engineering question rather than a scheduling impossibility.
  3. Knowledge
  4. The transformer architecture
  5. Machine translation
  6. The transformer is presented at NeurIPS
  7. Connections
  8. The transformer architecture
  9. Machine translation
  10. The AI accelerator
  11. Deep learning
  12. Machine translation
  13. Machine-generated code
  14. Growth in frontier training compute
  15. The transformer is presented at NeurIPS
  16. GLUE is saturated and SuperGLUE is built to replace it
  17. Labour-market effects of machine learning
  18. The transformer architecture
  19. The transformer is presented at NeurIPS
  20. A human-parity claim in translation fails under document-level evaluation
  21. The second International AI Safety Report reports mixed labour findings
  22. Evidence
  23. The transformer replaces recurrence and convolution with attention alone, is more parallelisable, and was shown superior in quality on two machine translation tasks while requiring significantly less time to train. V55 verification basis: the proceedings were not fetched; search results carried the author list, the venue, the Long Beach dates of 4–9 December 2017, the page range and the substance of the abstract and attributed them to the NeurIPS proceedings. The paper's arXiv identifier, its 2017 preprint date and its reported BLEU scores were explicitly NOT confirmed by retrieval and are not asserted anywhere in this pack.
Close the narrative
Observed changes · 0

No public Signals are attached to this object. Signals show what changed and when it was observed — never a direction or a rank.

Actions

Read the assembled narrativeContinue in StudioOpen TwinTwin does not start a decision from this kind of object.ShareSaveSaved objects are part of the authenticated projection, which is declared and not yet built.

/atlas?object=EVT_TRANSFORMER_PUBLISHED_2017&experience=EVT_TRANSFORMER_PUBLISHED_2017