2005AAP Grand RoundsRequires access

Weighing the Evidence: Power Calculations

James A. Taylor

Open publisher page 0 citations

Abstract

When comparing the efficacy of 2 treatments in a clinical trial, 4 outcomes are possible: 1) the study detected a “true” difference; 2) the study found a difference, but there is no “true” difference; 3) the study found no difference, and there is none; and 4) the study demonstrated no difference, but there is a “true” difference. Calculation of the “power” of a study can help determine which of these outcomes is most likely. The power of a study is influenced principally by the number of study participants and the size of the difference to be detected.Power calculations are used to determine the likelihood that, if a predetermined clinically meaningful difference is present, it will be detected. The most contentious part of a power calculation is deciding what constitutes a clinically meaningful difference. A power of 80% or 90% to detect this difference generally is assumed to be sufficient to validate that there is no clinically meaningful difference between the 2 drugs tested.For the study comparing racemic albuterol and levalbuterol, the authors chose a 10% improvement in FEV1 measurements as a clinically meaningful difference. Their study showed no statistically significant difference between levalbuterol and racemic albuterol. The finding could be “true” (outcome #3 above) or wrong (outcome #4). The authors state that, given the sample size, they had a 90% chance of detecting a 10% difference in change in FEV1 among children treated with the different medications, if the difference is actually present. Thus, there is a 10% chance that a difference in efficacy at least this large was not detected due to chance. In addition, there is more than a 10% chance that smaller differences in efficacy could be present but not detected.How should readers use power calculations? First, in a study that demonstrates no difference between 2 treatments, check to see whether the authors include a power calculation. Lack of a power calculation represents an important weakness. If a power calculation is included, consider whether the authors’ designation of a “clinically meaningful” difference makes sense. If the difference seems too large, then it is possible that an important difference in efficacy may have been missed. In this situation, a larger sample size is needed to answer the research question definitively.

About this research paper

What this paper is about

When comparing the efficacy of 2 treatments in a clinical trial, 4 outcomes are possible: 1) the study detected a “true” difference; 2) the study found a difference, but there is no “true” difference; 3) the study found no difference, and there is none; and 4) the study demonstrated no difference, but there is a “true” difference. Calculation of the “power” of a study can help determine which of these outcomes is most likely. The power of a study is influenced principally by the number of study participants and the size of the difference to be detected.Power calculations are used to determine the likelihood that, if a predetermined clinically meaningful difference is present, it will be detected. The most contentious part of a power calculation is deciding what constitutes a clinically meaningful difference. A power of 80% or 90% to detect this difference generally is assumed to be sufficient to validate that there is no clinically meaningful difference between the 2 drugs tested.For the study comparing racemic albuterol and levalbuterol, the authors chose a 10% improvement in FEV1 measurements as a clinically meaningful difference. Their study showed no statistically significant difference between levalbuterol and racemic albuterol. The finding could be “true” (outcome #3 above) or wrong (outcome #4). The authors state that, given the sample size, they had a 90% chance of detecting a 10% difference in change in FEV1 among children treated with the different medications, if the difference is actually present. Thus, there is a 10% chance that a difference in efficacy at least this large was not detected due to chance. In addition, there is more than a 10% chance that smaller differences in efficacy could be present but not detected.How should readers use power calculations? First, in a study that demonstrates no difference between 2 treatments, check to see whether the authors include a power calculation. Lack of a power calculation represents an important weakness. If a power calculation is included, consider whether the authors’ designation of a “clinically meaningful” difference makes sense. If the difference seems too large, then it is possible that an important difference in efficacy may have been missed. In this situation, a larger sample size is needed to answer the research question definitively.

Why it matters

A significance statement is not available in the OpenAlex record.

Key contribution

A contribution statement is not available in the OpenAlex record.

Method / approach

Method details are not available in the OpenAlex metadata.

Main findings

Findings are not separately available in the OpenAlex metadata.

Limitations

Limitations are not available in the OpenAlex metadata.

Applications

Application details are not available in the OpenAlex metadata.

Available abstract

When comparing the efficacy of 2 treatments in a clinical trial, 4 outcomes are possible: 1) the study detected a “true” difference; 2) the study found a difference, but there is no “true” difference; 3) the study found no difference, and there is none; and 4) the study demonstrated no difference, but there is a “true” difference. Calculation of the “power” of a study can help determine which of these outcomes is most likely. The power of a study is influenced principally by the number of study participants and the size of the difference to be detected.Power calculations are used to determine the likelihood that, if a predetermined clinically meaningful difference is present, it will be detected. The most contentious part of a power calculation is deciding what constitutes a clinically meaningful difference. A power of 80% or 90% to detect this difference generally is assumed to be sufficient to validate that there is no clinically meaningful difference between the 2 drugs tested.For the study comparing racemic albuterol and levalbuterol, the authors chose a 10% improvement in FEV1 measurements as a clinically meaningful difference. Their study showed no statistically significant difference between levalbuterol and racemic albuterol. The finding could be “true” (outcome #3 above) or wrong (outcome #4). The authors state that, given the sample size, they had a 90% chance of detecting a 10% difference in change in FEV1 among children treated with the different medications, if the difference is actually present. Thus, there is a 10% chance that a difference in efficacy at least this large was not detected due to chance. In addition, there is more than a 10% chance that smaller differences in efficacy could be present but not detected.How should readers use power calculations? First, in a study that demonstrates no difference between 2 treatments, check to see whether the authors include a power calculation. Lack of a power calculation represents an important weakness. If a power calculation is included, consider whether the authors’ designation of a “clinically meaningful” difference makes sense. If the difference seems too large, then it is possible that an important difference in efficacy may have been missed. In this situation, a larger sample size is needed to answer the research question definitively.

Key concepts: Significant difference, Medicine, Mean difference, Sample size determination, Potential difference, Outcome (game theory), Minimal clinically important difference, Confidence interval

Related papers

Back to paper searchBrowse research topicsOriginal source
Weighing the Evidence: Power Calculations — Research Paper | ScholarLens