The transformer is presented at NeurIPS
node
Vaswani and seven co-authors presented "Attention Is All You Need" at the 31st Conference on Neural Information Processing Systems in Long Beach, proposing a sequence architecture built on attention alone, without recurrence or convolution, and demonstrating it on two machine translation tasks with better quality and substantially less training time. The removal of sequential dependency was the consequential part: it made training throughput a function of how many accelerators could be pointed at the problem, and thereby made very large training runs an engineering question rather than a scheduling impossibility.
A node is not a place. Drawing it on a map would assert something about the world that no stored fact supports.
Read in · 1
Evidence · 1
Timeline
No dated observations are stored for this object. Atlas shows what was observed and when — it does not infer a history.
Connections · 2
- The transformer architectureconcerns
- Machine translationconcerns
Assembled narrative · 1
Assembled from 23 blocks · 1 evidence · 26 related
- Story
- Vaswani and seven co-authors presented "Attention Is All You Need" at the 31st Conference on Neural Information Processing Systems in Long Beach, proposing a sequence architecture built on attention alone, without recurrence or convolution, and demonstrating it on two machine translation tasks with better quality and substantially less training time. The removal of sequential dependency was the consequential part: it made training throughput a function of how many accelerators could be pointed at the problem, and thereby made very large training runs an engineering question rather than a scheduling impossibility.
- Knowledge
- The transformer architecture
- Machine translation
- The transformer is presented at NeurIPS
- Connections
- The transformer architecture
- Machine translation
- The AI accelerator
- Deep learning
- Machine translation
- Machine-generated code
- Growth in frontier training compute
- The transformer is presented at NeurIPS
- GLUE is saturated and SuperGLUE is built to replace it
- Labour-market effects of machine learning
- The transformer architecture
- The transformer is presented at NeurIPS
- A human-parity claim in translation fails under document-level evaluation
- The second International AI Safety Report reports mixed labour findings
- Evidence
- The transformer replaces recurrence and convolution with attention alone, is more parallelisable, and was shown superior in quality on two machine translation tasks while requiring significantly less time to train. V55 verification basis: the proceedings were not fetched; search results carried the author list, the venue, the Long Beach dates of 4–9 December 2017, the page range and the substance of the abstract and attributed them to the NeurIPS proceedings. The paper's arXiv identifier, its 2017 preprint date and its reported BLEU scores were explicitly NOT confirmed by retrieval and are not asserted anywhere in this pack.
Observed changes · 0
No public Signals are attached to this object. Signals show what changed and when it was observed — never a direction or a rank.
Actions
/atlas?object=EVT_TRANSFORMER_PUBLISHED_2017&experience=EVT_TRANSFORMER_PUBLISHED_2017