Data quality · Statistical programming · Data integration
Data cleaning and quality-monitoring system for a BCEAO financial-services supply survey
Designing a workflow for anomaly detection, review, correction, reintegration and consolidation in institutional survey data.
- Role
- R pipeline development · Collaborative Stata workflow
- Period
- April–May 2026
Context
Institutional survey quality control
As part of a subcontracting assignment delivered by a team in April–May 2026, I contributed to cleaning and quality control for a BCEAO survey on the supply of financial services.
The work connected several survey databases, consistency checks, atypical-value review, correction workflows and consolidation into an analysis-ready dataset.
Problem
From separate databases to a controlled workflow
How can multiple institutional survey databases be transformed into a structured process that identifies anomalies, organizes human review, reintegrates corrections and produces a consolidated dataset?
My role
An individually owned R pipeline within a team assignment
I independently developed the R pipeline used to structure, reorganize and prepare the data before exporting monitoring tables to Excel for review.
I also prepared review tables, integrated returned corrections, and merged and consolidated the databases. Stata processing, do-file improvements, consistency checks and the broader cleaning workflow were collaborative contributions.
Workflow
A repeatable review and reintegration cycle
The operational layer linked programmed controls with human review and correction.
- Survey databases
- Harmonization
- Validation rules
- Anomaly review
- Excel correction
- Reintegration
- Consolidation
Distinctive contribution
Connecting statistical processing and operational review
My main contribution was the R-based layer that moved observations from statistical processing into review tables and then reintegrated corrected values. The work therefore went beyond one-off cleaning.
Outcome
A structured quality-monitoring process
The assignment produced a structured cleaning and quality-monitoring process combining programmatic controls with operational review. No survey figures, anomaly counts or information about responding institutions are published.
Collaboration
Individual and collective contributions kept distinct
Team project. The R pipeline described here is my individual contribution; Stata processing and the overall cleaning and validation workflow were collaborative.
Technologies
Tools and methods
R · Stata · Excel · Data Quality · Data Validation · Data Integration
Confidentiality
Operational information remains private
Raw databases, observed values, respondent names, contact details, detailed anomalies, operational files and original scripts are not published.