CASP — Critical Assessment of Structure Prediction
node
The blind, biennial community experiment that has assessed protein structure prediction since 1994, and the reason the AlphaFold claim is the best-evidenced capability claim in machine learning. Its design is the point: experimental structures that have been solved but not published are held back by their determining laboratories, predictors submit models against sequence alone with no access to the answer, and independent assessors score the submissions afterwards. A predictor cannot train on the test set, cannot tune against the leaderboard, and cannot choose which targets to be measured on. That is exactly the set of protections that public machine-learning benchmarks lack, and CASP14 in late 2020 is where AlphaFold2 produced accuracy that assessors described as competitive with experiment. Not a place: CASP is an experiment and a community, run through a coordinating centre rather than a building.
A node is not a place. Drawing it on a map would assert something about the world that no stored fact supports.
Read in · 1
Timeline
No dated observations are stored for this object. Atlas shows what was observed and when — it does not infer a history.
Connections · 2
- AlphaFoldassesses
- The difficulty of evaluating machine learning systemscontradicts
Assembled narrative · 1
Assembled from 32 blocks · 1 evidence · 35 related
- Story
- The blind, biennial community experiment that has assessed protein structure prediction since 1994, and the reason the AlphaFold claim is the best-evidenced capability claim in machine learning. Its design is the point: experimental structures that have been solved but not published are held back by their determining laboratories, predictors submit models against sequence alone with no access to the answer, and independent assessors score the submissions afterwards. A predictor cannot train on the test set, cannot tune against the leaderboard, and cannot choose which targets to be measured on. That is exactly the set of protections that public machine-learning benchmarks lack, and CASP14 in late 2020 is where AlphaFold2 produced accuracy that assessors described as competitive with experiment. Not a place: CASP is an experiment and a community, run through a coordinating centre rather than a building.
- Knowledge
- CASP — Critical Assessment of Structure Prediction
- AlphaFold
- The difficulty of evaluating machine learning systems
- Connections
- AlphaFold
- The difficulty of evaluating machine learning systems
- AlphaFold2 is assessed blind at CASP14
- CASP — Critical Assessment of Structure Prediction
- Google DeepMind
- John Jumper
- Deep learning
- AlphaFold Protein Structure Database
- AlphaFold2 is assessed blind at CASP14
- AlphaFold2 is published and the database opens
- The AlphaFold database expands to over 200 million structures
- The Nobel Prize in Chemistry recognises protein structure prediction
- CASP — Critical Assessment of Structure Prediction
- Construct validity in AI evaluation
- Benchmark saturation
- METR
- European Union Artificial Intelligence Act
- A human-parity claim in translation fails under document-level evaluation
- A randomised trial finds experienced developers slower with AI assistance
- An audit finds the most-cited preference leaderboard is optimised rather than merely measured
- A systematic review of 445 benchmarks reports pervasive validity failures
- The Commission proposes deferring the high-risk obligations
- The AI Index records both fast benchmark movement and doubts about benchmarks
- Evidence
- AlphaFold2 predicts protein structure from sequence at accuracy competitive with experimental determination for most targets, as assessed blind at CASP14 in 2020, where predictions are reported to have achieved a median domain GDT_TS of 92.4 including on free-modelling targets. V55 verification basis: the paper was NOT fetched — nature.com is blocked to this session — but search results carried the title, journal, volume 596 and pages 583–589 and attributed them to this paper, and carried the CASP14 accuracy characterisation. The 92.4 median GDT_TS figure is attributed in retrieved summaries to the CASP14 assessment literature rather than to this paper, and a curator should confirm which document reports it before it is quoted as the paper's own statistic.
Observed changes · 0
No public Signals are attached to this object. Signals show what changed and when it was observed — never a direction or a rank.
Actions
/atlas?object=ORG_CASP&experience=ORG_CASP