Skip to main content
Evidence For Social Impact Lab

Data Quality & Management

From raw submissions to verified, reproducible datasets.

The problem we solve

Raw submissions are not a dataset. Duplicates, inconsistent responses and unexplained outliers need to be found, resolved and documented while the field team is still reachable.

Scope — what is included

  • High-frequency checks run during data collection
  • Duplicate identification and reconciliation
  • Independent verification and back-checks
  • Data cleaning with a documented audit trail
  • Progress and quality dashboards

Methodology

  • Consistency checks
  • Back-checks
  • Statistical diagnostics
  • Audit trails
  • Reproducible pipelines

Deliverables

  • Quality assurance reports per round
  • Cleaned dataset with variable documentation
  • Reproducible cleaning code
  • Monitoring dashboard
  • Documentation of every cleaning decision

What we need from you

Research decisions stay with your team. To deliver this work we need you to:

  • Approve cleaning rules and how flagged cases should be resolved
  • Confirm the variable list and format required for analysis
  • Provide secure storage for any identifiable data

Illustrative workflow

A synthetic example, included to show how the work runs. It is not a past engagement and does not describe a real client.

A four-week collection round produces daily submissions that must be verified before the field team demobilises.

  1. 1Run high-frequency checks each night and circulate flags by morning
  2. 2Back-check a sampled share of interviews and compare responses
  3. 3Resolve duplicates and inconsistencies while the team is still in the field
  4. 4Release a cleaned dataset with code and a decision log