Skip to main content
Bharat Surveyभारत सर्वे
Guide · Evaluation methods

For programme managers, CSR teams and evaluation commissioners

A quasi experimental impact evaluation design estimates what a programme changed when you cannot randomly assign who receives it. This guide compares matching, difference in differences and randomised trials, then covers participatory methods such as outcome harvesting and most significant change, so programme and CSR teams can pick a design that fits their data, budget and timeline.

  • When each design is credible
  • Data each one needs
  • Participatory and qualitative options
  • A method selection table
Constituency 172 · Survey round 2Updated 12s ago · sample
3,000
Target homes
2,184
Visits
1,962
Completed
1,871
Accepted
Daily accepted surveysLast 12 days
Quality
Accepted1871
In review64
Outside area19
Weak signal8
Top issues · by panchayat
Water62%
Roads48%
Power31%
Drains24%
Location evidence
Survey #BS-0412
  • In areaInside approx. boundary
  • Accuracy±12 m
  • Time09:42 · fresh fix
  • Mock GPSNot detected
  • Interview length11 min
Added to accepted count
Quasi experimental impact evaluation design, and the methods around it
Counterfactual thinkingBaseline and endlineComparison groupsParallel trendsCluster samplingOECD DAC criteriaParticipatory evidenceMixed methods
Quick answer

It is an evaluation that builds a comparison group without random assignment, then estimates impact as the difference between participants and that group. Credibility rests on how similar the groups are before the programme. Common versions are propensity score matching, difference in differences and regression discontinuity, each with its own assumptions.

The counterfactual problem

Impact is the gap between what happened to participants and what would have happened to them without the programme. You never observe the second part. Every design on this page is a way to approximate it, and each approximation fails in a known way.

Propensity score matching

Matching estimates each household's probability of joining the programme from characteristics measured before it started, such as land holding, household size or distance to the block office, and compares participants with non-participants who had similar probabilities. It only adjusts for what you measured. If joining depended on motivation nobody recorded, matched groups can still differ, which is why matching works best with rich baseline data and a clear enrolment rule.

  • Needs pre-programme characteristics
  • Needs overlap between groups
  • Cannot fix unobserved differences
  • Pairs well with difference in differences

Before choosing a design

  • How were participants selected?
  • Is there baseline data?
  • Can a rollout be phased?
  • Which outcomes, measured how?
  • When is the decision due?
Experimental and quasi-experimental

Difference in differences compares the change over time in a participant group with the change in a comparison group. It needs data from before and after for both, and the assumption that both groups were on similar trends before the programme.

Design

How the estimate works

Take the before-to-after change among participants, subtract the before-to-after change among the comparison group. Shared shocks such as a drought or price rise cancel out, because both groups experienced them.

Constituency 172 · Survey round 2Updated 12s ago · sample
3,000
Target homes
2,184
Visits
1,962
Completed
1,871
Accepted
Daily accepted surveysLast 12 days
Assumption

Parallel trends

The method assumes outcomes would have moved in parallel without the programme. With two or more pre-programme rounds, or administrative data from earlier years, you can check this instead of assuming it.

  • Plot pre-period trends
  • Pick comparison blocks with similar history
Field

What it needs from a survey

The same villages or households at baseline and endline, identical questions and recall periods, and stable IDs so records can be linked. Changing the questionnaire between rounds quietly breaks the comparison.

Survey 3/6 · AmenitiesOffline
How is drinking water in your area?
  • Comes regularly
  • Comes sometimes
  • Rarely / never
  • Don't know
✓ Section saved on phone · autosave
BackNext
Randomised designs

A cluster randomised trial assigns whole villages, schools or health facilities, rather than individuals, to receive the programme or not. It gives the most credible estimate when rollout can genuinely be randomised, for example when a programme expands in phases and not every cluster can start at once.

Why clusters change sample size

People inside one village share markets, schools and weather, so their outcomes move together. That intra-cluster correlation means 40 villages with 20 households each carry less information than 800 households drawn independently. Power calculations must include the number of clusters, households per cluster and an assumed correlation, and adding clusters usually helps more than adding households.

The survey firm's job in a trial

A cluster randomised trial India survey firm is usually responsible for listing, baseline, endline and data quality, not for the randomisation itself, which the evaluator should run and document. Field teams must stay blind to the design where possible and use exactly the same instruments and procedures in treatment and control clusters, or the difference found may reflect the survey rather than the programme.

  • Same questionnaire in every arm
  • Same interviewer training and schedule
  • Tracking of attrition by arm
  • Records with GPS and time for audit

Ethics and approvals

Randomising who receives a benefit raises ethical questions. Phased rollout, where control clusters receive the programme later, is a common answer. Research involving human participants in India may need review by an ethics committee, and informed consent is recorded before every interview.

Decision aid

Impact evaluation methods compared by use case, data needs and limits.
MethodBest whenMain data needMain weakness
Cluster randomised trialRollout can be randomised or phasedBaseline and endline in treatment and control clustersNeeds planning before launch
Difference in differencesProgramme already started, baseline existsBefore and after data for both groupsParallel trends may not hold
Propensity score matchingRich data on who enrolled and whyPre-programme characteristicsIgnores unobserved differences
Regression discontinuityEligibility set by a score or cut-offMany units near the cut-offEstimate applies near the cut-off only
Outcome harvestingOutcomes are unknown in advanceDocumented changes and verificationDoes not measure average effect size
Most significant changeUnderstanding what change people valueCollected stories and selection panelsSelected stories are not representative

Many evaluations combine one quantitative design with one participatory method.

Our view on method choice

Evaluation criteria

The OECD DAC Network on Development Evaluation defines six criteria, revised in 2019 to add coherence. They frame the questions an evaluation asks, while the designs above answer the impact question within them.

01

Relevance

Is the intervention doing the right things for the people and institutions it serves, and does it still fit as circumstances change?
02

Coherence

How well does it fit with other interventions in the same place or sector, including government schemes?
03

Effectiveness

Is it achieving its objectives and results, including differences across groups such as women or remote villages?
04

Efficiency

Are resources converted into results economically and in a timely way?
05

Impact

What higher-level effects, positive or negative, intended or not, has it generated? This is where quasi-experimental designs sit.
06

Sustainability

Will the benefits continue once funding or implementation support ends?
Participatory methods

When outcomes are hard to predict or the question is why change happened, participatory and qualitative methods add what counterfactual designs miss. They are often run alongside a survey rather than instead of one.

Working backwards from observed change

Outcome harvesting, developed by Ricardo Wilson-Grau and colleagues, collects evidence of what changed, such as a new rule adopted by a panchayat, and then works back to how the programme contributed. Harvested outcomes are verified with independent informants. It suits advocacy and governance work where outcomes cannot be listed in advance.
  • Describe each outcome precisely
  • Establish contribution, not attribution
  • Verify with people outside the programme
Working backwards from observed change
Participatory method

The most significant change technique, set out by Rick Davies and Jess Dart, collects stories of change from participants and has stakeholders select the most significant ones, discussing why. The discussion itself is part of the evidence.

01

1. Define domains of change

Agree two to four broad areas, such as livelihoods or women's participation, and a reporting period.
02

2. Collect stories

Field staff ask participants what the most significant change was in each domain and why it matters to them.
03

3. Select at each level

Panels at village, district and programme level choose the most significant stories and record their reasons.
04

4. Feed back the choices

Share which stories were selected and why with the people who told them and the staff who collected them.
05

5. Verify and analyse

Check selected stories in the field, then look for patterns across all collected stories, not only the winners.
Social accountability

The community score card, documented by the World Bank as a social accountability tool, brings service users and providers together to score a service such as a health centre or school, compare scores, and agree an action plan. It combines a participatory meeting with a simple scoring of indicators the community chooses.

Steps in a score card round

Preparation and mobilisation come first, then an input tracking matrix comparing entitlements with what reached the facility. Users score the service in groups, providers score themselves separately, and an interface meeting compares both and ends with agreed actions. Repeating the round shows whether actions happened.

  1. Choose the service and facility
  2. Track inputs against entitlements
  3. Community groups score indicators
  4. Providers self-assess
  5. Interface meeting and action plan
  6. Repeat after an agreed period

Where a survey app helps

Scores, attendance and action points are easy to lose on paper. Recording them on an offline Android app keeps each round's scores by facility and date, with photos of the meeting where consent is given, so the next round compares like with like.

Not the right fit if

  • Attributing impact without any comparison group or baseline
  • Randomising a programme that has already reached everyone
  • Clinical trials or laboratory-based outcome measurement
  • Designs promised to deliver a significant result

A good fit if

  • Baseline, midline and endline household surveys
  • Treatment and control fieldwork with identical procedures
  • Offline data collection with GPS and time stamps
  • Qualitative interviews and score card rounds alongside surveys
FAQ

Straight answers to the questions people ask most about this topic.

Logframe vs theory of change: which do I need?

A logframe is a compact table of objectives, indicators, means of verification and assumptions, useful for monitoring and donor reporting. A theory of change is a fuller explanation of how and why activities should lead to outcomes, including the assumptions between each step. Many programmes use the theory of change to design the evaluation and the logframe to track delivery.

MEAL framework vs M&E: what is the difference?

M&E covers monitoring and evaluation: tracking delivery and periodically assessing results. MEAL adds accountability and learning, meaning feedback and complaint channels for the people served, and routines for using evidence to change the programme. In practice a MEAL framework includes everything in an M&E plan plus these two explicit functions.

Can propensity score matching be done without a baseline?

It can, if you have reliable data on characteristics that were fixed before the programme, such as land records or census-linked household data. Matching on characteristics the programme itself could have changed, like current income, biases the estimate. Without either baseline data or stable pre-programme variables, matching is weak evidence.

How many clusters does a cluster randomised trial need?

It depends on the effect size you want to detect, outcome variability, intra-cluster correlation and households per cluster. Trials with very few clusters per arm give unstable estimates. A power calculation using realistic correlation values, ideally from earlier surveys in similar districts, should set the number before fieldwork is budgeted.

Is qualitative evidence enough to show impact?

Qualitative methods explain how and why change happened and can establish contribution, but they do not measure an average effect size across a population. For questions about how much a programme changed an outcome, pair them with a quantitative design. For questions about what changed that nobody anticipated, they are often the better tool.

शाम के समय कस्बे का हवाई दृश्य
India Nationwide, Hindi-first

Send your programme description, enrolment rules and any baseline data. We will suggest a feasible design and field plan.