← Signals
Recordedyearscience

A human-parity claim in translation fails under document-level evaluation

Läubli, Sennrich and Volk re-examined a claim that machine translation had reached parity with professional human translation on the Chinese–English news task, and found the claim depended on the evaluation unit. Raters comparing isolated sentences saw little difference; the same raters comparing whole documents preferred human translation clearly, because the failures that separate the two are failures of cohesion, reference and consistency across sentences — invisible by construction to a sentence-level protocol. The paper is a template for the whole Journey: the system did not change between the two evaluations, and the finding reversed.

Historical — it happened, and the record is settled.

Subjects

Evidence

Verified — Every field is supported by a cited authoritative source.

Recorded on Machine translation

  1. The transformer is presented at NeurIPS
  2. A human-parity claim in translation fails under document-level evaluation
  3. An audit finds the most-cited preference leaderboard is optimised rather than merely measured
  4. A systematic review of 445 benchmarks reports pervasive validity failures
  5. A randomised trial finds experienced developers slower with AI assistance
  6. The Commission proposes deferring the high-risk obligations
  7. The second International AI Safety Report reports mixed labour findings
  8. The AI Index records both fast benchmark movement and doubts about benchmarks

Related

Related because they share a subject in the Atlas — never because the text looks similar.