We are running an experiment testing two triage pathways. There are 30 patients in each group. We will run the study a hundred times to see how often we find a statistically significant difference with an alpha of 0.05.

Out of 100 studies, how often did each test call it significant anyway?

t-test parametric: compares means - of 100
Mann-Whitney U non-parametric: compares ranks - of 100

On normal data they agree almost perfectly. On skewed data the two tests come apart: a handful of boarded patients yank the mean around, so the t-test misses a difference that is really there, while the rank-based test finds it most of the time.

← Back On to the questions →