Skip to content
RichseenAtlasAtlasSign in

CASP — Critical Assessment of Structure Prediction

node

The blind, biennial community experiment that has assessed protein structure prediction since 1994, and the reason the AlphaFold claim is the best-evidenced capability claim in machine learning. Its design is the point: experimental structures that have been solved but not published are held back by their determining laboratories, predictors submit models against sequence alone with no access to the answer, and independent assessors score the submissions afterwards. A predictor cannot train on the test set, cannot tune against the leaderboard, and cannot choose which targets to be measured on. That is exactly the set of protections that public machine-learning benchmarks lack, and CASP14 in late 2020 is where AlphaFold2 produced accuracy that assessors described as competitive with experiment. Not a place: CASP is an experiment and a community, run through a coordinating centre rather than a building.

A node is not a place. Drawing it on a map would assert something about the world that no stored fact supports.

Read in · 1

Evidence · 1
Timeline

No dated observations are stored for this object. Atlas shows what was observed and when — it does not infer a history.

Connections · 2
Assembled narrative · 1

Assembled from 32 blocks · 1 evidence · 35 related

  1. Story
  2. The blind, biennial community experiment that has assessed protein structure prediction since 1994, and the reason the AlphaFold claim is the best-evidenced capability claim in machine learning. Its design is the point: experimental structures that have been solved but not published are held back by their determining laboratories, predictors submit models against sequence alone with no access to the answer, and independent assessors score the submissions afterwards. A predictor cannot train on the test set, cannot tune against the leaderboard, and cannot choose which targets to be measured on. That is exactly the set of protections that public machine-learning benchmarks lack, and CASP14 in late 2020 is where AlphaFold2 produced accuracy that assessors described as competitive with experiment. Not a place: CASP is an experiment and a community, run through a coordinating centre rather than a building.
  3. Knowledge
  4. CASP — Critical Assessment of Structure Prediction
  5. AlphaFold
  6. The difficulty of evaluating machine learning systems
  7. Connections
  8. AlphaFold
  9. The difficulty of evaluating machine learning systems
  10. AlphaFold2 is assessed blind at CASP14
  11. CASP — Critical Assessment of Structure Prediction
  12. Google DeepMind
  13. John Jumper
  14. Deep learning
  15. AlphaFold Protein Structure Database
  16. AlphaFold2 is assessed blind at CASP14
  17. AlphaFold2 is published and the database opens
  18. The AlphaFold database expands to over 200 million structures
  19. The Nobel Prize in Chemistry recognises protein structure prediction
  20. CASP — Critical Assessment of Structure Prediction
  21. Construct validity in AI evaluation
  22. Benchmark saturation
  23. METR
  24. European Union Artificial Intelligence Act
  25. A human-parity claim in translation fails under document-level evaluation
  26. A randomised trial finds experienced developers slower with AI assistance
  27. An audit finds the most-cited preference leaderboard is optimised rather than merely measured
  28. A systematic review of 445 benchmarks reports pervasive validity failures
  29. The Commission proposes deferring the high-risk obligations
  30. The AI Index records both fast benchmark movement and doubts about benchmarks
  31. Evidence
  32. AlphaFold2 predicts protein structure from sequence at accuracy competitive with experimental determination for most targets, as assessed blind at CASP14 in 2020, where predictions are reported to have achieved a median domain GDT_TS of 92.4 including on free-modelling targets. V55 verification basis: the paper was NOT fetched — nature.com is blocked to this session — but search results carried the title, journal, volume 596 and pages 583–589 and attributed them to this paper, and carried the CASP14 accuracy characterisation. The 92.4 median GDT_TS figure is attributed in retrieved summaries to the CASP14 assessment literature rather than to this paper, and a curator should confirm which document reports it before it is quoted as the paper's own statistic.
Close the narrative
Observed changes · 0

No public Signals are attached to this object. Signals show what changed and when it was observed — never a direction or a rank.

Actions

Read the assembled narrativeContinue in StudioOpen TwinTwin does not start a decision from this kind of object.ShareSaveSaved objects are part of the authenticated projection, which is declared and not yet built.

/atlas?object=ORG_CASP&experience=ORG_CASP