SamplingShala
Evaluation methods

Sample size for program evaluation

A program evaluation's sample size depends first on the question. Estimating a rate ("what share of children read at level?") needs about 385 respondents for ±5 points at 95 percent confidence. Detecting an impact ("did the program move reading by 0.3 standard deviations?") needs about 175 per arm at 80 percent power. Both then grow with clustering (design effect) and expected attrition, and a real evaluation sizes every stakeholder group, students, teachers, parents, officials, by its own logic.

The most common sizing mistake in the sector is answering an impact question with a descriptive survey's sample size, or vice versa. The two formulas measure different things and can differ by an order of magnitude once small effects are involved. Here is the map.

Estimating a rate: the descriptive survey

n₀ = z² × p(1 - p) / e² About 385 for ±5 points at 95 percent confidence. Apply the finite population correction for small populations, and multiply by DEFF if fieldwork is clustered. See the calculator page for the full table.

Detecting an impact: powered comparisons

n per arm = 2 × (z₀ + z₁)² / MDES² z₀ = 1.96 (two-sided 5 percent), z₁ = 0.84 (80 percent power), MDES in standard deviations.
Minimum detectable effectPer armTotal (two arms)
0.5 SD (large)about 63about 126
0.3 SD (typical education program)about 175about 350
0.2 SD (modest but meaningful)about 393about 786

Choosing the MDES is the honest heart of the design: it is the smallest effect worth detecting, not the effect you hope for. Halving it quadruples the sample, the same square law as the margin of error.

The adjustments that make it real

  • Design effect. Sampling through schools multiplies the sample by DEFF = 1 + (m - 1) × ICC, often 3 to 5 in practice. This is usually the largest single adjustment; see how many schools to survey.
  • Attrition. If you expect to lose 15 percent of the panel by endline, recruit n / 0.85 at baseline. Longitudinal designs compound this across waves: retention (1 - a) per wave becomes (1 - a)waves - 1 overall.
  • Panel correlation. Baseline-endline designs that track the same children get a discount: the correlation between a child's two scores removes shared noise, so the sample shrinks relative to two independent cross-sections.
  • Matching inflation. Quasi-experimental designs pay a premium (often 1.1x to 1.5x) because matched comparisons are less efficient than randomised ones.
  • Finite population. When the population is small, the correction n = n₀ / (1 + (n₀ - 1)/N) reduces the sample; a design that exceeds the population is telling you to run a census.

Several stakeholder groups, one plan

An education evaluation rarely samples only students. A defensible plan sizes each group by its own logic and states which one binds:

GroupTypical methodTypical size logic
StudentsStatistical, clusteredPowers the main claim; usually binds
TeachersStatistical or quotaOwn formula, smaller n
Head teachersCensus of sampled schoolsOne per school visited
Parents / SMCQuota or purposive FGDs6 to 10 per FGD, to saturation
OfficialsPurposive KIIsSaturation across strata
Qualitative components are sized by saturation, the point where new interviews stop yielding new themes, not by confidence formulas. Pretending otherwise, in either direction, is the second most common sizing mistake. A mixed-methods plan states both logics side by side.

Frequently asked questions

What sample size does an impact evaluation need?

About 175 per arm for a 0.3 SD effect at 80 percent power and 5 percent two-sided significance, before clustering and attrition. Multiply by DEFF for school-based sampling and divide by expected retention. For a 0.2 SD effect, plan for roughly 393 per arm.

Why does a panel design need fewer people than two cross-sections?

Because each child serves as their own comparison. The baseline-endline correlation removes the noise the two measurements share, so detecting change takes fewer children, provided the panel survives attrition, which you plan for by over-recruiting at baseline.

Which stakeholder group determines my total budget?

Usually the group carrying the strongest claim, students powering the impact estimate, because clustering and power inflate it far beyond the others. Sampling Shala's Build mode computes every group and names the binding constraint explicitly.

Size all eight designs, honestly

Descriptive, baseline-endline, longitudinal, RCT, quasi-experimental, LQAS, mixed methods and qualitative, with every assumption stated and the justification written for you.

Open Build mode