Weighing the Evidence: Power Calculations
James A. Taylor
Abstract
James A. Taylor
Abstract
When comparing the efficacy of 2 treatments in a clinical trial, 4 outcomes are possible: 1) the study detected a “true” difference; 2) the study found a difference, but there is no “true” difference; 3) the study found no difference, and there is none; and 4) the study demonstrated no difference, but there is a “true” difference. Calculation of the “power” of a study can help determine which of these outcomes is most likely. The power of a study is influenced principally by the number of study participants and the size of the difference to be detected.Power calculations are used to determine the likelihood that, if a predetermined clinically meaningful difference is present, it will be detected. The most contentious part of a power calculation is deciding what constitutes a clinically meaningful difference. A power of 80% or 90% to detect this difference generally is assumed to be sufficient to validate that there is no clinically meaningful difference between the 2 drugs tested.For the study comparing racemic albuterol and levalbuterol, the authors chose a 10% improvement in FEV1 measurements as a clinically meaningful difference. Their study showed no statistically significant difference between levalbuterol and racemic albuterol. The finding could be “true” (outcome #3 above) or wrong (outcome #4). The authors state that, given the sample size, they had a 90% chance of detecting a 10% difference in change in FEV1 among children treated with the different medications, if the difference is actually present. Thus, there is a 10% chance that a difference in efficacy at least this large was not detected due to chance. In addition, there is more than a 10% chance that smaller differences in efficacy could be present but not detected.How should readers use power calculations? First, in a study that demonstrates no difference between 2 treatments, check to see whether the authors include a power calculation. Lack of a power calculation represents an important weakness. If a power calculation is included, consider whether the authors’ designation of a “clinically meaningful” difference makes sense. If the difference seems too large, then it is possible that an important difference in efficacy may have been missed. In this situation, a larger sample size is needed to answer the research question definitively.
A significance statement is not available in the OpenAlex record.
A contribution statement is not available in the OpenAlex record.
Method details are not available in the OpenAlex metadata.
Findings are not separately available in the OpenAlex metadata.
Limitations are not available in the OpenAlex metadata.
Application details are not available in the OpenAlex metadata.
When comparing the efficacy of 2 treatments in a clinical trial, 4 outcomes are possible: 1) the study detected a “true” difference; 2) the study found a difference, but there is no “true” difference; 3) the study found no difference, and there is none; and 4) the study demonstrated no difference, but there is a “true” difference. Calculation of the “power” of a study can help determine which of these outcomes is most likely. The power of a study is influenced principally by the number of study participants and the size of the difference to be detected.Power calculations are used to determine the likelihood that, if a predetermined clinically meaningful difference is present, it will be detected. The most contentious part of a power calculation is deciding what constitutes a clinically meaningful difference. A power of 80% or 90% to detect this difference generally is assumed to be sufficient to validate that there is no clinically meaningful difference between the 2 drugs tested.For the study comparing racemic albuterol and levalbuterol, the authors chose a 10% improvement in FEV1 measurements as a clinically meaningful difference. Their study showed no statistically significant difference between levalbuterol and racemic albuterol. The finding could be “true” (outcome #3 above) or wrong (outcome #4). The authors state that, given the sample size, they had a 90% chance of detecting a 10% difference in change in FEV1 among children treated with the different medications, if the difference is actually present. Thus, there is a 10% chance that a difference in efficacy at least this large was not detected due to chance. In addition, there is more than a 10% chance that smaller differences in efficacy could be present but not detected.How should readers use power calculations? First, in a study that demonstrates no difference between 2 treatments, check to see whether the authors include a power calculation. Lack of a power calculation represents an important weakness. If a power calculation is included, consider whether the authors’ designation of a “clinically meaningful” difference makes sense. If the difference seems too large, then it is possible that an important difference in efficacy may have been missed. In this situation, a larger sample size is needed to answer the research question definitively.
Key concepts: Significant difference, Medicine, Mean difference, Sample size determination, Potential difference, Outcome (game theory), Minimal clinically important difference, Confidence interval