A human-parity claim in translation fails under document-level evaluation
node
Läubli, Sennrich and Volk re-examined a claim that machine translation had reached parity with professional human translation on the Chinese–English news task, and found the claim depended on the evaluation unit. Raters comparing isolated sentences saw little difference; the same raters comparing whole documents preferred human translation clearly, because the failures that separate the two are failures of cohesion, reference and consistency across sentences — invisible by construction to a sentence-level protocol. The paper is a template for the whole Journey: the system did not change between the two evaluations, and the finding reversed.
A node is not a place. Drawing it on a map would assert something about the world that no stored fact supports.
Read in · 1
Timeline
No dated observations are stored for this object. Atlas shows what was observed and when — it does not infer a history.
Connections · 3
Assembled narrative · 1
An assembled narrative is available for this object.
Observed changes · 0
No public Signals are attached to this object. Signals show what changed and when it was observed — never a direction or a rank.
Actions
/atlas?object=EVT_LAUBLI_PARITY_2018