A human-parity claim in translation fails under document-level evaluation
node
Läubli, Sennrich and Volk re-examined a claim that machine translation had reached parity with professional human translation on the Chinese–English news task, and found the claim depended on the evaluation unit. Raters comparing isolated sentences saw little difference; the same raters comparing whole documents preferred human translation clearly, because the failures that separate the two are failures of cohesion, reference and consistency across sentences — invisible by construction to a sentence-level protocol. The paper is a template for the whole Journey: the system did not change between the two evaluations, and the finding reversed.
A node is not a place. Drawing it on a map would assert something about the world that no stored fact supports.
Read in · 1
Timeline
No dated observations are stored for this object. Atlas shows what was observed and when — it does not infer a history.
Connections · 3
Assembled narrative · 1
Assembled from 33 blocks · 1 evidence · 28 related
- Story
- Läubli, Sennrich and Volk re-examined a claim that machine translation had reached parity with professional human translation on the Chinese–English news task, and found the claim depended on the evaluation unit. Raters comparing isolated sentences saw little difference; the same raters comparing whole documents preferred human translation clearly, because the failures that separate the two are failures of cohesion, reference and consistency across sentences — invisible by construction to a sentence-level protocol. The paper is a template for the whole Journey: the system did not change between the two evaluations, and the finding reversed.
- Knowledge
- Construct validity in AI evaluation
- Machine translation
- The difficulty of evaluating machine learning systems
- A human-parity claim in translation fails under document-level evaluation
- Connections
- Machine translation
- The difficulty of evaluating machine learning systems
- Construct validity in AI evaluation
- Labour-market effects of machine learning
- The transformer architecture
- The transformer is presented at NeurIPS
- A human-parity claim in translation fails under document-level evaluation
- The second International AI Safety Report reports mixed labour findings
- CASP — Critical Assessment of Structure Prediction
- Construct validity in AI evaluation
- Benchmark saturation
- METR
- European Union Artificial Intelligence Act
- A human-parity claim in translation fails under document-level evaluation
- A randomised trial finds experienced developers slower with AI assistance
- An audit finds the most-cited preference leaderboard is optimised rather than merely measured
- A systematic review of 445 benchmarks reports pervasive validity failures
- The Commission proposes deferring the high-risk obligations
- The AI Index records both fast benchmark movement and doubts about benchmarks
- The difficulty of evaluating machine learning systems
- A human-parity claim in translation fails under document-level evaluation
- An audit finds the most-cited preference leaderboard is optimised rather than merely measured
- A systematic review of 445 benchmarks reports pervasive validity failures
- Evidence
- Testing a human-parity claim on the WMT Chinese–English news task with alternative evaluation protocols, human raters assessing adequacy and fluency showed a stronger preference for human over machine translation when evaluating whole documents than when evaluating isolated sentences, indicating that errors decisive for quality are frequently invisible at sentence level. V55 verification basis: the paper was not fetched; search results carried the title, the ACL Anthology identifier D18-1512, the EMNLP 2018 Brussels venue and the substance of the finding and attributed them to this paper.
Observed changes · 0
No public Signals are attached to this object. Signals show what changed and when it was observed — never a direction or a rank.
Actions
/atlas?object=EVT_LAUBLI_PARITY_2018&experience=EVT_LAUBLI_PARITY_2018