Skip to content
RichseenAtlasFutureSign in

What AI Changes

The strongest evidence in this Journey was produced by a protocol that structural biologists have been running since 1994, and its whole method is to make sure that whoever is being tested cannot choose the questions, cannot see the answers, and does not do the scoring.

Anyone who follows this subject arrives braced for a system doing something impressive, and meets none of that; the opening screen is occupied instead by the arrangement that decides whether an impressive result counts for anything — a sealed answer, targets set by people with no stake in the outcome, scoring done by a third party. Procedural, cool, slightly austere, in the unhurried register of a method section. The yardstick comes first and the thing to be measured comes later.

What machine learning has demonstrably changed is narrow, specific and unusually well evidenced; what is claimed for it is broad, general and measured largely by instruments the claimants also built — and those two facts are the same fact, because a claim about these systems gets corrected only where somebody outside can pick up the object and examine it.

Why Richseen chose this

Machine learning is now arguing about its own measuring instruments in public, and the corpus caught that argument from several directions at once: an audit of the most-cited preference leaderboard, a systematic review of hundreds of benchmarks, a national statistical agency changing a question and watching its own number jump, and a statutory regime deferred because the standards for assessing these systems were not ready. What is actually in dispute is not whether the systems are impressive — the protein-structure case settles that — but whether anything currently in circulation is strong enough to carry the weight being put on it. That question has an answer, the answer varies enormously from claim to claim, and almost nobody sorts the claims before deciding.

The Richseen lens

The capability is real and the measurement of it is largely self-administered — both halves hold at once, which is why a claim here is worth what its design is worth rather than what its result looks like: who chose the test, who held the answers, who did the scoring, and who would have had to answer for being wrong.

  1. held answersWhether the party being measured could choose the targets, see the answers, or score itself. The one design that removes all three came from structural biology, not from this field.
  2. the melting rulerWhat a rising score establishes once everyone is optimising against the instrument that produces it — and how invisible that transition is from the score alone.
  3. accountabilityWhich deployments are backed by an institution that would have to answer publicly for being wrong, and how much narrower those results are than the announcements that preceded them.
  4. who can checkWhether a specialist community exists that is positioned to inspect the object of a claim. Crystals have one. An economy does not.
  5. the missing instrumentWhat happens to a rule, a survey or a projection when the means of measuring the thing it governs does not yet exist.

The chapters

Roughly fifty years, closed exactly

The reader sees what a claim looks like when it satisfies the design in full — and notices that the best-evidenced result in the field is also the most tightly scoped one, with its own limits published alongside it.

AlphaFold · AlphaFold Protein Structure Database · CASP — Critical Assessment of Structure Prediction · Google DeepMind · John Jumper · Demis Hassabis

Why 2012 and not the 1980s

The story stops being about ideas about intelligence and becomes a story about supply — a dataset somebody assembled, processors somebody could buy, and an architecture that converted both into training throughput.

Deep learning · ImageNet · The transformer architecture · The AI accelerator · Growth in frontier training compute · Training compute · Data-centre electricity demand · Northern Virginia data-centre corridor · Epoch AI · Geoffrey Hinton · Fei-Fei Li

When the ruler is also the target

A published score stops being a reading and becomes a claim with a provenance — compatible with real capability, with contamination, and with tuning against the instrument, in a way the number cannot distinguish.

Benchmark saturation · Construct validity in AI evaluation · The difficulty of evaluating machine learning systems · Machine translation · Machine-generated code

Somebody has to answer for it

The reader acquires a second-best test to use when a blind assessment is unavailable — an institution putting its own accountability behind the output — and sees that the results which pass it are consistently smaller and more specific than the announcements were.

Machine-learned weather forecasting · European Centre for Medium-Range Weather Forecasts · Machine learning in medical imaging · Measured adoption of AI by firms

Four measurements, four populations, no contradiction

The labour question stops looking like a disagreement between good and bad research and starts looking like four instruments pointed at four different things — including one trial in which the participants were confidently wrong about their own productivity.

Labour-market effects of machine learning · METR · Machine-generated code · Measured adoption of AI by firms

Where the correction can reach

The Journey's mechanism becomes explicit: a claim is corrected when there is an object a specialist can pick up, and survives when the object is a profession, an economy or a decade.

GNoME and the materials-discovery claim · Machine learning in medical imaging · Geoffrey Hinton · Labour-market effects of machine learning

Fact, rate, statute, unknown

The reader leaves with the material sorted into four columns of very different strength, and with the deferred statutory deadline read as what it is — the measurement problem in this Journey, restated as law.

European Union Artificial Intelligence Act · AI standards bodies · Data-centre electricity demand · Growth in frontier training compute · Labour-market effects of machine learning · Benchmark saturation

What to look at

31 records in this Journey.

Showing 8 of 19 featured records. Atlas does not choose which of the rest matter.

Connections you would not expect

  • Two claims about structures, decided in opposite directions by the same kind of community. In one, chemists and biologists arranged in advance to hold the answers, and the prediction survived. In the other, a structural claim was announced first and crystallographers sampled it afterwards, and it did not. The discipline is nearly the same; what differs is whether the checking was designed in before the result or attempted after it.

    CASP — Critical Assessment of Structure Prediction · GNoME and the materials-discovery claim

  • A fortnightly business survey and a machine-learning leaderboard look like unrelated artefacts, and they fail in exactly the same way. Both produced a sharp rise that had nothing to do with the world changing: one because the wording of the question was rewritten, the other because systems were tuned toward the prompts the instrument happened to contain. Neither movement is visible in the number itself, and both are routinely quoted as evidence of acceleration.

    Measured adoption of AI by firms · Benchmark saturation

  • The learned forecast model is measured against the physics-based system it improves on — and the agency has to keep that system running, because it produces the reanalysis the learned model is trained from. The clearest case of a machine-learning system outperforming its predecessor is also a case of it being unable to exist without it, which is a relationship no benchmark reports.

    Machine-learned weather forecasting · European Centre for Medium-Range Weather Forecasts

  • The compute trend is usually read as a property of the field — a curve in a database, drawn from figures somebody had to reconstruct because they are not published. At the far end of that curve it stops being a curve: it is an electricity demand series, and then it is a county with transmission capacity. The same quantity is a mathematical abstraction at one end and a planning dispute at the other.

    Training compute · Data-centre electricity demand · Northern Virginia data-centre corridor

  • The most specific occupational displacement prediction in this corpus was made about radiology from an image-classification result, and it is the one occupation where a randomised trial eventually said what the technology actually did to the work. It removed a large share of the reading load and left the accountability where it was. The claim was about employment; the measurable change was about workload — and that substitution is what most current labour claims are still doing.

    Machine learning in medical imaging · Labour-market effects of machine learning

What is not settled

  • Is the entry-level employment signal the leading edge of a large displacement, or a modest reallocation inside normal churn?

    Four capable teams have looked and reached four answers, because they measured different populations over different windows: one occupation under a controlled rollout, one age band in an administrative payroll panel, an entire labour market over a short period, and an economy in a model. The payroll figure itself differs between versions of the analysis and the corpus records both. None of the four establishes causation from machine learning as such, and the most authoritative available summary describes the picture as mixed and unresolved, which is where this Journey leaves it.

  • Can an evaluation regime be built that a well-resourced developer cannot optimise against?

    The design that works is documented and has been running in another discipline for decades, but it depends on withheld answers, targets chosen by others and independent scoring — conditions that are expensive and that no general-purpose benchmark has reproduced. The audit of the largest preference leaderboard and the review of hundreds of benchmarks both describe failures of process and construction rather than of arithmetic, which is what makes them hard to patch. Nothing in the corpus reports a general solution.

  • Does the growth in training compute meet a physical limit before it meets a limit of usefulness?

    The trend is a reconstruction, not a disclosure, and it measures inputs rather than capability — a mapping the corpus says has never been stable. Against it sit constraints that do not scale the same way: grid connections, capital, and leading-edge fabrication concentrated in a small number of facilities. The energy series has its own projection attached, and a projection is not an observation. Which limit arrives first is genuinely unknown and this Journey does not guess.

  • How much of this Journey rests on sources that were read at second hand?

    More than is comfortable, and the corpus says so in every block. The headline accuracy figure for the protein result reached the pack through assessment literature rather than the paper; the adoption figures for the structure database came through institutional communications carried in secondary reporting; the benchmark history came from summaries of a paper that was not retrieved; the size of the crystallographers' sample was not established; the decade-old radiology remark is paraphrased from later journalism; and the official citation for the amending act was not obtained. A Journey whose subject is evidential strength has to state its own, and this is it.

What to carry out of this

The finding is an asymmetry. Wherever an outside party could hold the answers in advance, randomise the assignment, or pick up the sample afterwards, the claim came back smaller than it went in — and survived: a fold predicted blind, a forecast an agency puts its name to, half a reading workload removed with a human still accountable. Wherever no such party exists — a profession, an economy, a decade — the claim is larger, older and still uncontradicted, because nothing has been in a position to contradict it. That gap has a price and somebody is paying it. The design that closes it is expensive in money and coordination, which is why it has been running for decades in structural biology and nowhere in general-purpose evaluation; the cost of its absence falls on whoever happens to be the unmeasured object. The experienced developers who worked more slowly and were confident they had worked faster. The entry-level workers whose situation four capable teams measured four different ways without converging. The county at the far end of the compute curve, where an abstraction resolves into transmission capacity. Everyone covered by an obligation that could not begin because the means of assessing it were not built. None of this is a verdict on the technology. It is an account of who is currently carrying the difference between what can be shown and what is being said.

This is unfinished in the world, not only in the telling.

Other ways to look at this

where they are — not shown, because the editorial brief calls it incidental to this subject.

Where this leads

Continue

  • Think this through in StudioFor a given claim about what a system can now do, who would have to hold the answers in advance for the claim to be assessable — and what would arranging that cost?

Ask another question

Ask the world another question. Dynamic keeps this Journey as its context and never changes what is written above.

Open Dynamic Atlas →