A human-parity claim in translation fails under document-level evaluation
Läubli, Sennrich and Volk re-examined a claim that machine translation had reached parity with professional human translation on the Chinese–English news task, and found the claim depended on the evaluation unit. Raters comparing isolated sentences saw little difference; the same raters comparing whole documents preferred human translation clearly, because the failures that separate the two are failures of cohesion, reference and consistency across sentences — invisible by construction to a sentence-level protocol. The paper is a template for the whole Journey: the system did not change between the two evaluations, and the finding reversed.
Historical — it happened, and the record is settled.
Subjects
- Machine translation · technology · not located
- The difficulty of evaluating machine learning systems · phenomenon · not located
- Construct validity in AI evaluation · concept · not located
Evidence
Verified — Every field is supported by a cited authoritative source.
Has Machine Translation Achieved Human Parity? A Case for Document-level Evaluation
Testing a human-parity claim on the WMT Chinese–English news task with alternative evaluation protocols, human raters assessing adequacy and fluency showed a stronger preference for human over machine translation when evaluating whole documents than when evaluating isolated sentences, indicating that errors decisive for quality are frequently invisible at sentence level. V55 verification basis: the paper was not fetched; search results carried the title, the ACL Anthology identifier D18-1512, the EMNLP 2018 Brussels venue and the substance of the finding and attributed them to this paper.
Proceedings of EMNLP 2018 (Association for Computational Linguistics), Brussels · 2018 · Verified
Läubli, S., Sennrich, R. and Volk, M., EMNLP 2018; ACL Anthology D18-1512, https://aclanthology.org/D18-1512/
Recorded on Machine translation
- The transformer is presented at NeurIPS
- A human-parity claim in translation fails under document-level evaluation
- An audit finds the most-cited preference leaderboard is optimised rather than merely measured
- A systematic review of 445 benchmarks reports pervasive validity failures
- A randomised trial finds experienced developers slower with AI assistance
- The Commission proposes deferring the high-risk obligations
- The second International AI Safety Report reports mixed labour findings
- The AI Index records both fast benchmark movement and doubts about benchmarks
Related
Related because they share a subject in the Atlas — never because the text looks similar.
- The second International AI Safety Report reports mixed labour findings · 2026-02
- The AI Index records both fast benchmark movement and doubts about benchmarks · 2026
- The Commission proposes deferring the high-risk obligations · 2025-11-19
- A randomised trial finds experienced developers slower with AI assistance · 2025-07-10
- An audit finds the most-cited preference leaderboard is optimised rather than merely measured · 2025-04
- A systematic review of 445 benchmarks reports pervasive validity failures · 2025-11
- The transformer is presented at NeurIPS · 2017-12-04 to 2017-12-09