2017Unpublished venueOpen access

Manipulating the alpha level cannot cure significance testing – comments on "Redefine statistical significance"

David Trafimow, Valentin Amrhein, Corson N. Areshenkoff, Carlos Barrera-Causil, Eric J. Beh, Yusuf Bilgiç, Roser Bono, M. T. Bradley, William M. Briggs, Héctor A. Cepeda-Freyre, Sergio E. Chaigneau, Daniel R. Ciocca, Juan Carlos Correa, Denis Cousineau, Michiel R. de Boer, Subhra Sankar Dhar, Igor Dolgov, Juana Gómez‐Benito, Marián Grendár, James W. Grice, Martin E. Guerrero-Gimenez, Andrés Gutiérrez, Tania B. Huedo–Medina, Klaus Jaffé, Armina Janyan, Ali Karimnezhad, Fränzi Korner‐Nievergelt, Koji Kosugi, Martin Lachmair, Rubén Daniel Ledesma, Roberto Limongi, Marco Tullio Liuzza, Rosaria Lombardo, Michael J. Marks, Gunther Meinlschmidt, Ladislas Nalborczyk, Hung T. Nguyen, Raydonal Ospina, J Perezgonzalez, Roland Pfister, Juan José Rahona, David Alberto Rodríguez Medina, Xavier Romão, Susana Ruiz Fernández, Isabel Suárez, Marion Tegethoff, Mauricio Tejo, Rens van de Schoot, Ivan Vankov, Santiago Velasco-Forero, Tonghui Wang, Yuki Yamada, Felipe Carlos Martín Zoppino, Fernando Marmolejo‐Ramos

Open full text 11 citations

Abstract

We argue that depending on p-values to reject null hypotheses, including a recent call for changing the canonical alpha level for statistical significance from .05 to .005, is deleterious for the finding of new discoveries and the progress of science. Given that blanket and variable criterion levels both are problematic, it is sensible to dispense with significance testing altogether. There are alternatives that address study design and determining sample sizes much more directly than significance testing does; but none of the statistical tools should replace significance testing as the new magic method giving clear-cut mechanical answers. Inference should not be based on single studies at all, but on cumulative evidence from multiple independent studies. When evaluating the strength of the evidence, we should consider, for example, auxiliary assumptions, the strength of the experimental design, or implications for applications. To boil all this down to a binary decision based on a p-value threshold of .05, .01, .005, or anything else, is not acceptable.

Open-access reader

About this research paper

What this paper is about

We argue that depending on p-values to reject null hypotheses, including a recent call for changing the canonical alpha level for statistical significance from .05 to .005, is deleterious for the finding of new discoveries and the progress of science. Given that blanket and variable criterion levels both are problematic, it is sensible to dispense with significance testing altogether. There are alternatives that address study design and determining sample sizes much more directly than significance testing does; but none of the statistical tools should replace significance testing as the new magic method giving clear-cut mechanical answers. Inference should not be based on single studies at all, but on cumulative evidence from multiple independent studies. When evaluating the strength of the evidence, we should consider, for example, auxiliary assumptions, the strength of the experimental design, or implications for applications. To boil all this down to a binary decision based on a p-value threshold of .05, .01, .005, or anything else, is not acceptable.

Why it matters

OpenAlex reports 11 citations for this work. Citation counts describe recorded attention and do not establish research quality.

Key contribution

A contribution statement is not available in the OpenAlex record.

Method / approach

Method details are not available in the OpenAlex metadata.

Main findings

Findings are not separately available in the OpenAlex metadata.

Limitations

Limitations are not available in the OpenAlex metadata.

Applications

Application details are not available in the OpenAlex metadata.

Available abstract

We argue that depending on p-values to reject null hypotheses, including a recent call for changing the canonical alpha level for statistical significance from .05 to .005, is deleterious for the finding of new discoveries and the progress of science. Given that blanket and variable criterion levels both are problematic, it is sensible to dispense with significance testing altogether. There are alternatives that address study design and determining sample sizes much more directly than significance testing does; but none of the statistical tools should replace significance testing as the new magic method giving clear-cut mechanical answers. Inference should not be based on single studies at all, but on cumulative evidence from multiple independent studies. When evaluating the strength of the evidence, we should consider, for example, auxiliary assumptions, the strength of the experimental design, or implications for applications. To boil all this down to a binary decision based on a p-value threshold of .05, .01, .005, or anything else, is not acceptable.

Key concepts: Null hypothesis, Statistical significance, Significance testing, Statistical hypothesis testing, p-value, Statistical inference, Econometrics, Multiple comparisons problem

Related papers

Back to paper searchBrowse research topicsOriginal source
Manipulating the alpha level cannot cure significance testing – comments on "Redefine statistical significance" — Research Paper | ScholarLens