# Method: Sample Size & Power Calculator (method version 1.0.0)

## Calculation

1. **t-tests (two independent groups, or paired).** Power is computed exactly from the noncentral t distribution. For n per group in two independent groups, df = 2n − 2 and the noncentrality is δ = d √(n/2). For n pairs, df = n − 1 and δ = d √n. With the critical value t* = t(1 − α/2, df), two-sided power = 1 − F(t*; df, δ) + F(−t*; df, δ), where F is the noncentral t CDF.
2. **Required n for a t-test** is the smallest n (at least 2) whose power reaches the target. The search starts from the normal-approximation estimate 2((z₁₋α/₂ + z₁₋β)/d)² per group (paired: ((z₁₋α/₂ + z₁₋β)/d)²) and moves up or down one at a time. Above 10 million per group, where power changes by less than the calculation's error from one subject to the next, Guenther's closed form is used instead: the estimate plus z₁₋α/₂² / 4 per group (paired: plus z₁₋α/₂² / 2), rounded up; achieved power for such samples uses the normal distribution.
3. **Two proportions.** n per group = [z₁₋α/₂ √(2 p̄ q̄) + z₁₋β √(p₁q₁ + p₂q₂)]² / (p₁ − p₂)², rounded up, with p̄ = (p₁ + p₂)/2 and q = 1 − p, without the continuity correction (Fleiss, Levin & Paik 2003). Achieved power solves the same equation for z₁₋β.
4. **Numerics.** The noncentral t CDF integrates Φ(t s − δ) against the density of s = √(V/df), V ~ χ²(df), with Simpson's rule over the mean of s ± 12 standard deviations. Student t quantiles come from the regularized incomplete beta function by bisection. Against scipy 1.17.1, CDF values agree to about 1e-7 and quantiles to within about 5e-10 (relative).

## Outside this record

- **ANOVA.** The page's ANOVA option applies the two-group normal formula to Cohen's d instead of the F test with Cohen's f. The page labels it an approximation, and it is not validated here.

## Conventions

- Tests are two-sided.
- Effect size is Cohen's d for t-tests (its sign is ignored): for two groups the mean difference over the pooled SD, and for pairs dz, the mean difference over the SD of the differences. For the proportions test it is the difference in proportions.
- n is per group for two independent groups and the number of pairs for the paired design.
- A t-test needs at least 2 observations per group; a smaller n returns no power and a message.
- An effect size so small that the sample size overflows returns no sample size and a message.

## Worked examples

Six cases in `cases.json`, each with the expected answer and where it comes from: exact sample sizes for two groups at d = 0.5 and d = 2.0, a paired design at d = 0.5, achieved power with 5 per group at d = 2.0, two proportions worked by hand, and a very small effect (d = 0.1, edge case).

## References

- Faul F, Erdfelder E, Lang AG, Buchner A. G*Power 3: a flexible statistical power analysis program for the social, behavioral, and biomedical sciences. Behavior Research Methods. 2007;39(2):175-191. https://doi.org/10.3758/bf03193146
- Guenther WC. Sample size formulas for normal theory t tests. The American Statistician. 1981;35(4):243-244. https://doi.org/10.1080/00031305.1981.10479363
- Cohen J. Statistical Power Analysis for the Behavioral Sciences. 2nd ed. Lawrence Erlbaum Associates; 1988. ISBN 978-0-8058-0283-2.
- Fleiss JL, Levin B, Paik MC. Statistical Methods for Rates and Proportions. 3rd ed. Wiley; 2003. ISBN 0-471-52629-0.
- Virtanen P, Gommers R, Oliphant TE, et al. SciPy 1.0: fundamental algorithms for scientific computing in Python. Nature Methods. 2020;17(3):261-272. https://doi.org/10.1038/s41592-019-0686-2 (cross-check values, scipy 1.17.1)

## Changes from the page before this record

- t-test sample sizes and power are now exact. The page used the normal approximation, which returns too few subjects. At 80% power and α 0.05, two groups with d = 0.2, 0.5, 0.8, 1.0, 1.5 and 2.0 gave 393, 63, 25, 16, 7 and 4 per group against the exact 394, 64, 26, 17, 9 and 6. Paired designs gave 197, 32, 13, 8, 4 and 2 against 199, 34, 15, 10, 6 and 5. Achieved power and the power curve use the same exact calculation.
- The page's method note said the approximation matched exact answers within 1–2 subjects. At d = 2.0 it was 2 per group short for two groups and 3 short for pairs; the note now describes the exact method.
- The ANOVA option is labelled an approximation.
- A t-test with n = 1 returns a message instead of stopping the page.
- Proportions results are unchanged.
