A modular data pipeline for processing and analyzing humanitarian perception survey data. This project simulates a humanitarian feedback workflow, transforming raw field responses into structured insights that can support program improvement, accountability, and evidence-based decision-making.
Humanitarian organizations collect large volumes of survey responses to assess aid effectiveness across areas such as food assistance, WASH, and cash transfers. However, field-level data can face challenges including:
- Low-Quality Submissions: Rapid or incomplete responses can reduce data reliability.
- Data Noise: Duplicate records and input errors can affect analysis.
- Limited Visibility: Disconnected data can make it difficult to identify differences in beneficiary satisfaction across regions and sectors.
I developed an automated pipeline that:
- Synthetic Data Generation: Uses Python to simulate humanitarian survey data across regions, operational sectors, and data-quality scenarios.
- Automated Data Cleaning: Uses R to deduplicate records, identify low-quality submissions, and validate data structure.
- Statistical Analysis: Generates performance metrics to compare beneficiary satisfaction across agencies, regions, and sectors.
- Visual Analysis: Produces visualizations that highlight satisfaction patterns and performance gaps for reporting and decision-making.
- Languages: Python (Data Generation), R (Cleaning, Statistics, Visualization)
- Data Processing:
tidyverse,dplyr,ggplot2 - Workflow: Modular script orchestration (
main_pipeline.R) - Environment: Designed for reproducible data analysis workflows
Analysis of the simulated data highlights differences in beneficiary satisfaction across selected humanitarian hubs in Ethiopia, including Somali, Afar, and Gambella regions. By identifying agency-sector combinations with lower satisfaction levels, the pipeline helps prioritize areas for further investigation and targeted program improvement.
A potential extension is to integrate the structured feedback data with Retrieval-Augmented Generation (RAG) systems and qualitative reports. This could help humanitarian teams combine quantitative survey results with narrative feedback to identify emerging service gaps and support more informed decisions.
Developed as a simulated humanitarian Monitoring, Evaluation, Accountability, and Learning (MEAL) data analytics project.