Skip to content
RichseenAtlasAtlasSign in

GLUE is saturated and SuperGLUE is built to replace it

node

The GLUE language-understanding suite was released in 2018 and within about a year pretrained transformer models had passed its non-expert human baseline. SuperGLUE was constructed in 2019 explicitly as a harder successor, opening with a gap of roughly eighteen points between the best available model and the human baseline — and closed similarly quickly. The two-year cycle established here has repeated at every subsequent level of difficulty and is the origin of the pattern the Journey later examines: an instrument built to discriminate, answered, and replaced faster than the interpretation of its scores could be established.

A node is not a place. Drawing it on a map would assert something about the world that no stored fact supports.

Read in · 1

Evidence · 1
Timeline

No dated observations are stored for this object. Atlas shows what was observed and when — it does not infer a history.

Connections · 2
Assembled narrative · 1

Assembled from 24 blocks · 1 evidence · 37 related

  1. Story
  2. The GLUE language-understanding suite was released in 2018 and within about a year pretrained transformer models had passed its non-expert human baseline. SuperGLUE was constructed in 2019 explicitly as a harder successor, opening with a gap of roughly eighteen points between the best available model and the human baseline — and closed similarly quickly. The two-year cycle established here has repeated at every subsequent level of difficulty and is the origin of the pattern the Journey later examines: an instrument built to discriminate, answered, and replaced faster than the interpretation of its scores could be established.
  3. Knowledge
  4. The transformer architecture
  5. Benchmark saturation
  6. GLUE is saturated and SuperGLUE is built to replace it
  7. Connections
  8. Benchmark saturation
  9. The transformer architecture
  10. The difficulty of evaluating machine learning systems
  11. Measured adoption of AI by firms
  12. Fei-Fei Li
  13. GLUE is saturated and SuperGLUE is built to replace it
  14. An audit finds the most-cited preference leaderboard is optimised rather than merely measured
  15. The AI Index records both fast benchmark movement and doubts about benchmarks
  16. The AI accelerator
  17. Deep learning
  18. Machine translation
  19. Machine-generated code
  20. Growth in frontier training compute
  21. The transformer is presented at NeurIPS
  22. GLUE is saturated and SuperGLUE is built to replace it
  23. Evidence
  24. A systematic review of 445 large-language-model benchmarks, conducted by about forty researchers led from the Oxford Internet Institute, found that roughly half claim to measure abstract phenomena such as reasoning or harmlessness without defining them or specifying how they are measured, that 27% rely on convenience sampling, and that only 16% use uncertainty estimates or statistical tests when comparing model performance; the paper offers eight recommendations for benchmark construction. V55 verification basis: the paper was not fetched, but search results carried the 445-benchmark count, the three statistics, the arXiv identifier 2511.04703, the NeurIPS 2025 venue and the OII lead and attributed them to this work and to the Institute's own announcement.
Close the narrative
Observed changes · 0

No public Signals are attached to this object. Signals show what changed and when it was observed — never a direction or a rank.

Actions

Read the assembled narrativeContinue in StudioOpen TwinTwin does not start a decision from this kind of object.ShareSaveSaved objects are part of the authenticated projection, which is declared and not yet built.

/atlas?object=EVT_GLUE_SUPERGLUE_2018_2019&experience=EVT_GLUE_SUPERGLUE_2018_2019