Data quality · Statistical programming · Data integration

Data cleaning and quality-monitoring system for a BCEAO financial-services supply survey

Designing a workflow for anomaly detection, review, correction, reintegration and consolidation in institutional survey data.

Role
R pipeline development · Collaborative Stata workflow
Period
April–May 2026

Context

Institutional survey quality control

As part of a subcontracting assignment delivered by a team in April–May 2026, I contributed to cleaning and quality control for a BCEAO survey on the supply of financial services.

The work connected several survey databases, consistency checks, atypical-value review, correction workflows and consolidation into an analysis-ready dataset.

Problem

From separate databases to a controlled workflow

How can multiple institutional survey databases be transformed into a structured process that identifies anomalies, organizes human review, reintegrates corrections and produces a consolidated dataset?

My role

An individually owned R pipeline within a team assignment

I independently developed the R pipeline used to structure, reorganize and prepare the data before exporting monitoring tables to Excel for review.

I also prepared review tables, integrated returned corrections, and merged and consolidated the databases. Stata processing, do-file improvements, consistency checks and the broader cleaning workflow were collaborative contributions.

Workflow

A repeatable review and reintegration cycle

The operational layer linked programmed controls with human review and correction.

  1. Survey databases
  2. Harmonization
  3. Validation rules
  4. Anomaly review
  5. Excel correction
  6. Reintegration
  7. Consolidation

Distinctive contribution

Connecting statistical processing and operational review

My main contribution was the R-based layer that moved observations from statistical processing into review tables and then reintegrated corrected values. The work therefore went beyond one-off cleaning.

Outcome

A structured quality-monitoring process

The assignment produced a structured cleaning and quality-monitoring process combining programmatic controls with operational review. No survey figures, anomaly counts or information about responding institutions are published.

Collaboration

Individual and collective contributions kept distinct

Team project. The R pipeline described here is my individual contribution; Stata processing and the overall cleaning and validation workflow were collaborative.

Technologies

Tools and methods

R · Stata · Excel · Data Quality · Data Validation · Data Integration

Confidentiality

Operational information remains private

Raw databases, observed values, respondent names, contact details, detailed anomalies, operational files and original scripts are not published.