2019Unpublished venueOpen access

Agreement is overrated: A plea for correlation to assess human evaluation reliability

Jacopo Amidei, Paul Piwek, Alistair Willis

Open full text 27 citations

Abstract

Inter-Annotator Agreement (IAA) is used as a means of assessing the quality of NLG evaluation data, in particular, its reliability.According to existing scales of IAA interpretationsee, for example, Lommel et al. (2014), Liu et al. (2016), Sedoc et al. (2018) and Amidei et al. (2018a) -most data collected for NLG evaluation fail the reliability test.We confirmed this trend by analysing papers published over the last 10 years in NLG-specific conferences (in total 135 papers that included some sort of human evaluation study).Following Sampson and Babarczy (2008), Lommel et al. (2014), Joshi et al. (2016) and Amidei et al. ( 2018b), such phenomena can be explained in terms of irreducible human language variability.Using three case studies, we show the limits of considering IAA as the only criterion for checking evaluation reliability.Given human language variability, we propose that for human evaluation of NLG, correlation coefficients and agreement coefficients should be used together to obtain a better assessment of the evaluation data reliability.This is illustrated using the three case studies.

Open-access reader

About this research paper

What this paper is about

Inter-Annotator Agreement (IAA) is used as a means of assessing the quality of NLG evaluation data, in particular, its reliability.According to existing scales of IAA interpretationsee, for example, Lommel et al. (2014), Liu et al. (2016), Sedoc et al. (2018) and Amidei et al. (2018a) -most data collected for NLG evaluation fail the reliability test.We confirmed this trend by analysing papers published over the last 10 years in NLG-specific conferences (in total 135 papers that included some sort of human evaluation study).Following Sampson and Babarczy (2008), Lommel et al. (2014), Joshi et al. (2016) and Amidei et al. ( 2018b), such phenomena can be explained in terms of irreducible human language variability.Using three case studies, we show the limits of considering IAA as the only criterion for checking evaluation reliability.Given human language variability, we propose that for human evaluation of NLG, correlation coefficients and agreement coefficients should be used together to obtain a better assessment of the evaluation data reliability.This is illustrated using the three case studies.

Why it matters

OpenAlex reports 27 citations for this work. Citation counts describe recorded attention and do not establish research quality.

Key contribution

A contribution statement is not available in the OpenAlex record.

Method / approach

Method details are not available in the OpenAlex metadata.

Main findings

Findings are not separately available in the OpenAlex metadata.

Limitations

Limitations are not available in the OpenAlex metadata.

Applications

Application details are not available in the OpenAlex metadata.

Available abstract

Inter-Annotator Agreement (IAA) is used as a means of assessing the quality of NLG evaluation data, in particular, its reliability.According to existing scales of IAA interpretationsee, for example, Lommel et al. (2014), Liu et al. (2016), Sedoc et al. (2018) and Amidei et al. (2018a) -most data collected for NLG evaluation fail the reliability test.We confirmed this trend by analysing papers published over the last 10 years in NLG-specific conferences (in total 135 papers that included some sort of human evaluation study).Following Sampson and Babarczy (2008), Lommel et al. (2014), Joshi et al. (2016) and Amidei et al. ( 2018b), such phenomena can be explained in terms of irreducible human language variability.Using three case studies, we show the limits of considering IAA as the only criterion for checking evaluation reliability.Given human language variability, we propose that for human evaluation of NLG, correlation coefficients and agreement coefficients should be used together to obtain a better assessment of the evaluation data reliability.This is illustrated using the three case studies.

Key concepts: Reliability (semiconductor), Plea, Computer science, Interpretation (philosophy), sort, Correlation, Agreement, Statistics

Related papers

Back to paper searchBrowse research topicsOriginal source
Agreement is overrated: A plea for correlation to assess human evaluation reliability — Research Paper | ScholarLens