How the estimate works
Take the before-to-after change among participants, subtract the before-to-after change among the comparison group. Shared shocks such as a drought or price rise cancel out, because both groups experienced them.
For programme managers, CSR teams and evaluation commissioners
A quasi experimental impact evaluation design estimates what a programme changed when you cannot randomly assign who receives it. This guide compares matching, difference in differences and randomised trials, then covers participatory methods such as outcome harvesting and most significant change, so programme and CSR teams can pick a design that fits their data, budget and timeline.

It is an evaluation that builds a comparison group without random assignment, then estimates impact as the difference between participants and that group. Credibility rests on how similar the groups are before the programme. Common versions are propensity score matching, difference in differences and regression discontinuity, each with its own assumptions.
Impact is the gap between what happened to participants and what would have happened to them without the programme. You never observe the second part. Every design on this page is a way to approximate it, and each approximation fails in a known way.
Matching estimates each household's probability of joining the programme from characteristics measured before it started, such as land holding, household size or distance to the block office, and compares participants with non-participants who had similar probabilities. It only adjusts for what you measured. If joining depended on motivation nobody recorded, matched groups can still differ, which is why matching works best with rich baseline data and a clear enrolment rule.
Difference in differences compares the change over time in a participant group with the change in a comparison group. It needs data from before and after for both, and the assumption that both groups were on similar trends before the programme.
Take the before-to-after change among participants, subtract the before-to-after change among the comparison group. Shared shocks such as a drought or price rise cancel out, because both groups experienced them.
The method assumes outcomes would have moved in parallel without the programme. With two or more pre-programme rounds, or administrative data from earlier years, you can check this instead of assuming it.
The same villages or households at baseline and endline, identical questions and recall periods, and stable IDs so records can be linked. Changing the questionnaire between rounds quietly breaks the comparison.
A cluster randomised trial assigns whole villages, schools or health facilities, rather than individuals, to receive the programme or not. It gives the most credible estimate when rollout can genuinely be randomised, for example when a programme expands in phases and not every cluster can start at once.
People inside one village share markets, schools and weather, so their outcomes move together. That intra-cluster correlation means 40 villages with 20 households each carry less information than 800 households drawn independently. Power calculations must include the number of clusters, households per cluster and an assumed correlation, and adding clusters usually helps more than adding households.
A cluster randomised trial India survey firm is usually responsible for listing, baseline, endline and data quality, not for the randomisation itself, which the evaluator should run and document. Field teams must stay blind to the design where possible and use exactly the same instruments and procedures in treatment and control clusters, or the difference found may reflect the survey rather than the programme.
Randomising who receives a benefit raises ethical questions. Phased rollout, where control clusters receive the programme later, is a common answer. Research involving human participants in India may need review by an ethics committee, and informed consent is recorded before every interview.
| Method | Best when | Main data need | Main weakness |
|---|---|---|---|
| Cluster randomised trial | Rollout can be randomised or phased | Baseline and endline in treatment and control clusters | Needs planning before launch |
| Difference in differences | Programme already started, baseline exists | Before and after data for both groups | Parallel trends may not hold |
| Propensity score matching | Rich data on who enrolled and why | Pre-programme characteristics | Ignores unobserved differences |
| Regression discontinuity | Eligibility set by a score or cut-off | Many units near the cut-off | Estimate applies near the cut-off only |
| Outcome harvesting | Outcomes are unknown in advance | Documented changes and verification | Does not measure average effect size |
| Most significant change | Understanding what change people value | Collected stories and selection panels | Selected stories are not representative |
Many evaluations combine one quantitative design with one participatory method.
The OECD DAC Network on Development Evaluation defines six criteria, revised in 2019 to add coherence. They frame the questions an evaluation asks, while the designs above answer the impact question within them.
When outcomes are hard to predict or the question is why change happened, participatory and qualitative methods add what counterfactual designs miss. They are often run alongside a survey rather than instead of one.

The most significant change technique, set out by Rick Davies and Jess Dart, collects stories of change from participants and has stakeholders select the most significant ones, discussing why. The discussion itself is part of the evidence.
The community score card, documented by the World Bank as a social accountability tool, brings service users and providers together to score a service such as a health centre or school, compare scores, and agree an action plan. It combines a participatory meeting with a simple scoring of indicators the community chooses.
Preparation and mobilisation come first, then an input tracking matrix comparing entitlements with what reached the facility. Users score the service in groups, providers score themselves separately, and an interface meeting compares both and ends with agreed actions. Repeating the round shows whether actions happened.
Scores, attendance and action points are easy to lose on paper. Recording them on an offline Android app keeps each round's scores by facility and date, with photos of the meeting where consent is given, so the next round compares like with like.
Design support and fieldwork for baseline, endline and impact studies.
Learn moreIndependent evaluation of government schemes and CSR programmes.
Learn moreMonitoring and evaluation data collection for non-profits and CSR teams.
Learn moreStraight answers to the questions people ask most about this topic.
A logframe is a compact table of objectives, indicators, means of verification and assumptions, useful for monitoring and donor reporting. A theory of change is a fuller explanation of how and why activities should lead to outcomes, including the assumptions between each step. Many programmes use the theory of change to design the evaluation and the logframe to track delivery.
M&E covers monitoring and evaluation: tracking delivery and periodically assessing results. MEAL adds accountability and learning, meaning feedback and complaint channels for the people served, and routines for using evidence to change the programme. In practice a MEAL framework includes everything in an M&E plan plus these two explicit functions.
It can, if you have reliable data on characteristics that were fixed before the programme, such as land records or census-linked household data. Matching on characteristics the programme itself could have changed, like current income, biases the estimate. Without either baseline data or stable pre-programme variables, matching is weak evidence.
It depends on the effect size you want to detect, outcome variability, intra-cluster correlation and households per cluster. Trials with very few clusters per arm give unstable estimates. A power calculation using realistic correlation values, ideally from earlier surveys in similar districts, should set the number before fieldwork is budgeted.
Qualitative methods explain how and why change happened and can establish contribution, but they do not measure an average effect size across a population. For questions about how much a programme changed an outcome, pair them with a quantitative design. For questions about what changed that nobody anticipated, they are often the better tool.

Send your programme description, enrolment rules and any baseline data. We will suggest a feasible design and field plan.