Has Machine Translation Achieved Human Parity? A Case for Document-level Evaluation
evidence
Testing a human-parity claim on the WMT Chinese–English news task with alternative evaluation protocols, human raters assessing adequacy and fluency showed a stronger preference for human over machine translation when evaluating whole documents than when evaluating isolated sentences, indicating that errors decisive for quality are frequently invisible at sentence level. V55 verification basis: the paper was not fetched; search results carried the title, the ACL Anthology identifier D18-1512, the EMNLP 2018 Brussels venue and the substance of the finding and attributed them to this paper.
A evidence is not a place. Drawing it on a map would assert something about the world that no stored fact supports.
Evidence · 0
This object cites no evidence.
Timeline
No dated observations are stored for this object. Atlas shows what was observed and when — it does not infer a history.
Connections · 0
Assembled narrative · 0
This object does not clear the publishing floor: an assembled narrative needs a description and at least one cited piece of evidence.
Observed changes · 0
No public Signals are attached to this object. Signals show what changed and when it was observed — never a direction or a rank.
Actions
/atlas?object=EVIDENCE_EV_LAUBLI_2018