Research Methodology
How to calculate sample size for your thesis (with worked examples)
The three formulas most PG theses need, the assumptions examiners check, and the mistakes that get synopses sent back.
Resora Research Team · 15 Sept 2026 · 3 min read
"How many patients do I need?" is usually the first statistical question a resident asks — and the one ethics committees and university boards scrutinise most closely. A sample size isn't a number you pick; it's a number you justify, with a formula, assumptions from published literature, and an allowance for real-world losses.
What every sample size calculation needs
Whatever your design, you'll need to state:
- The primary outcome — one outcome drives the calculation, not all of them.
- An expected value from prior studies — a prevalence, a mean and standard deviation, or proportions in two groups. Cite the study you took it from.
- The confidence level — almost always 95% (Z = 1.96).
- Precision or power — how narrow your estimate should be (descriptive studies), or the probability of detecting a real difference (comparative studies, usually 80%).
- An allowance for non-response or drop-out — typically 10–20%.
1. Estimating a single proportion (prevalence studies)
For cross-sectional studies estimating how common something is:
n = Z² × p × (1 − p) / d²
where p is the expected prevalence and d the absolute precision.
Example. A previous study found anaemia in 30% of antenatal women. You want your estimate within ±5 percentage points:
n = 1.96² × 0.30 × 0.70 / 0.05² = 0.8067 / 0.0025 ≈ 323
Allowing for 10% non-response: 323 / 0.9 ≈ 359 participants.
Absolute vs relative precision. "5% precision" can mean ±5 percentage points (absolute) or ±5% of 30%, i.e. ±1.5 points (relative). The second gives a sample roughly eleven times larger. Always say which one you used.
2. Comparing two means
For comparing a continuous outcome (blood pressure, HbA1c, pain score) between two groups:
n per group = 2 × (Zα/2 + Zβ)² × σ² / Δ²
With 95% confidence and 80% power, (1.96 + 0.84)² = 7.84.
Example. Systolic BP has a standard deviation of 10 mmHg, and a 5 mmHg difference is clinically meaningful:
n = 2 × 7.84 × 10² / 5² = 1568 / 25 ≈ 63 per group
3. Comparing two proportions
For a binary outcome (complication yes/no) between two groups:
n per group = (Zα/2 + Zβ)² × [p₁(1 − p₁) + p₂(1 − p₂)] / (p₁ − p₂)²
Example. Complications occur in 40% with standard care and you expect 20% with the new technique:
n = 7.84 × (0.24 + 0.16) / 0.20² = 3.136 / 0.04 ≈ 79 per group
Software that applies a continuity correction will give a somewhat larger number — that's expected, and more conservative.
Tools that help
You don't need to calculate by hand. G*Power, OpenEpi and nMaster (from CMC Vellore) all handle these designs and more. Save a screenshot of the inputs and output for your records — committees sometimes ask for it.
Mistakes that get synopses sent back
- No source for the assumptions. An expected prevalence or SD must come from a cited study, ideally in a similar population.
- Using a "convenient" number. "All patients admitted during the study period" is a sampling frame, not a sample size justification — though it can be acceptable if you show it exceeds the calculated minimum.
- Ignoring the design. Cluster sampling needs a design effect (often 1.5–2); matched designs need different formulas.
- Powering for every objective. Calculate for the primary objective and say so.
- Forgetting feasibility. If your department sees 150 eligible patients a year and you need 359, revisit the design, duration or precision before submitting.
A template paragraph
"Based on a prevalence of 30% reported by [Author, Year], with 95% confidence and an absolute precision of 5%, the minimum sample size calculated using the formula n = Z²p(1−p)/d² was 323. Allowing for 10% non-response, 359 participants will be enrolled."
If your design doesn't fit these formulas — diagnostic accuracy, correlation, non-inferiority, survival — the principle is the same but the formula changes. That's where a statistician saves you a round of corrections.