Sample size for program evaluation
A program evaluation's sample size depends first on the question. Estimating a rate ("what share of children read at level?") needs about 385 respondents for ±5 points at 95 percent confidence. Detecting an impact ("did the program move reading by 0.3 standard deviations?") needs about 175 per arm at 80 percent power. Both then grow with clustering (design effect) and expected attrition, and a real evaluation sizes every stakeholder group, students, teachers, parents, officials, by its own logic.
The most common sizing mistake in the sector is answering an impact question with a descriptive survey's sample size, or vice versa. The two formulas measure different things and can differ by an order of magnitude once small effects are involved. Here is the map.
Estimating a rate: the descriptive survey
Detecting an impact: powered comparisons
| Minimum detectable effect | Per arm | Total (two arms) |
|---|---|---|
| 0.5 SD (large) | about 63 | about 126 |
| 0.3 SD (typical education program) | about 175 | about 350 |
| 0.2 SD (modest but meaningful) | about 393 | about 786 |
Choosing the MDES is the honest heart of the design: it is the smallest effect worth detecting, not the effect you hope for. Halving it quadruples the sample, the same square law as the margin of error.
The adjustments that make it real
- Design effect. Sampling through schools multiplies the sample by DEFF = 1 + (m - 1) × ICC, often 3 to 5 in practice. This is usually the largest single adjustment; see how many schools to survey.
- Attrition. If you expect to lose 15 percent of the panel by endline, recruit n / 0.85 at baseline. Longitudinal designs compound this across waves: retention (1 - a) per wave becomes (1 - a)waves - 1 overall.
- Panel correlation. Baseline-endline designs that track the same children get a discount: the correlation between a child's two scores removes shared noise, so the sample shrinks relative to two independent cross-sections.
- Matching inflation. Quasi-experimental designs pay a premium (often 1.1x to 1.5x) because matched comparisons are less efficient than randomised ones.
- Finite population. When the population is small, the correction n = n₀ / (1 + (n₀ - 1)/N) reduces the sample; a design that exceeds the population is telling you to run a census.
Several stakeholder groups, one plan
An education evaluation rarely samples only students. A defensible plan sizes each group by its own logic and states which one binds:
| Group | Typical method | Typical size logic |
|---|---|---|
| Students | Statistical, clustered | Powers the main claim; usually binds |
| Teachers | Statistical or quota | Own formula, smaller n |
| Head teachers | Census of sampled schools | One per school visited |
| Parents / SMC | Quota or purposive FGDs | 6 to 10 per FGD, to saturation |
| Officials | Purposive KIIs | Saturation across strata |
Frequently asked questions
What sample size does an impact evaluation need?
About 175 per arm for a 0.3 SD effect at 80 percent power and 5 percent two-sided significance, before clustering and attrition. Multiply by DEFF for school-based sampling and divide by expected retention. For a 0.2 SD effect, plan for roughly 393 per arm.
Why does a panel design need fewer people than two cross-sections?
Because each child serves as their own comparison. The baseline-endline correlation removes the noise the two measurements share, so detecting change takes fewer children, provided the panel survives attrition, which you plan for by over-recruiting at baseline.
Which stakeholder group determines my total budget?
Usually the group carrying the strongest claim, students powering the impact estimate, because clustering and power inflate it far beyond the others. Sampling Shala's Build mode computes every group and names the binding constraint explicitly.
Size all eight designs, honestly
Descriptive, baseline-endline, longitudinal, RCT, quasi-experimental, LQAS, mixed methods and qualitative, with every assumption stated and the justification written for you.
Open Build mode