The ongoing tyranny of statistical significance testing in biomedical research
- 1.7k Downloads
Since its introduction into the biomedical literature, statistical significance testing (abbreviated as SST) caused much debate. The aim of this perspective article is to review frequent fallacies and misuses of SST in the biomedical field and to review a potential way out of the fallacies and misuses associated with SSTs. Two frequentist schools of statistical inference merged to form SST as it is practised nowadays: the Fisher and the Neyman-Pearson school. The P-value is both reported quantitatively and checked against the α-level to produce a qualitative dichotomous measure (significant/nonsignificant). However, a P-value mixes the estimated effect size with its estimated precision. Obviously, it is not possible to measure these two things with one single number. For the valid interpretation of SSTs, a variety of presumptions and requirements have to be met. We point here to four of them: study size, correct statistical model, correct causal model, and absence of bias and confounding. It has been stated that the P-value is perhaps the most misunderstood statistical concept in clinical research. As in the social sciences, the tyranny of SST is still highly prevalent in the biomedical literature even after decades of warnings against SST. The ubiquitous misuse and tyranny of SST threatens scientific discoveries and may even impede scientific progress. In the worst case, misuse of significance testing may even harm patients who eventually are incorrectly treated because of improper handling of P-values. For a proper interpretation of study results, both estimated effect size and estimated precision are necessary ingredients.
KeywordsStatistics P-value Confidence intervals
None of the authors reports any conflict of interest.
- 2.Hogben LT. Statistical theory: an examination of the contemporary crisis in statistical theory from a behaviourist viewpoint. London: George Allen & Unwin; 1957.Google Scholar
- 3.Morrison DE, Henkel RE. The significance test controversy: a reader. Chicago: Aldine Pub; 1970.Google Scholar
- 5.Greenland S, Rothman KJ. Fundamentals of epidemiologic data analysis. In: Rothman KJ, Greenland S, Lash TL, editors. Modern epidemiology. 3rd ed. Philadelphia: Wolters Kluwer, Lippincott Williams & Wilkins; 2008. p. 213–37.Google Scholar
- 7.Miettinen OS. Theoretical epidemiology. Albany: Delmar Publishers Inc.; 1985.Google Scholar
- 12.Fisher RA. Statistical methods and scientific inference. Edingburgh: Oliver & Boyd; 1956.Google Scholar
- 15.Neyman J, Pearson ES. On the use and interpretation of certain test criteria for purposes of statistical inference. Part I. Biometrika. 1928;20A:175–240.Google Scholar
- 18.Sobin LH, Wittekind Ch. TNM classification of malignant tumours. 6th ed. New York: Wiley-Liss, Inc.; 2002.Google Scholar
- 23.Fisher RA. The design of experiments. Edinburgh: Oliver & Boyd; 1935.Google Scholar
- 31.Loftus GR. On the tyranny of hypothesis testing in the social sciences. Contemp Psychol. 1991;36(2):102–5.Google Scholar