Ml Engineering

Customer Churn Prediction Platform

From customer data to deployable churn-risk predictions.

A production-style machine learning platform that ingests telecom customer data, engineers model-ready features, trains and tunes churn models, tracks experiments with MLflow, persists model artifacts, and serves churn predictions through FastAPI and Docker.

Problem

Context

Customer churn prediction requires more than training a classifier: data ingestion, feature preparation, reproducible modeling, artifact management, serving, testing, and deployment workflows must work together.

Challenge

Build a modular machine learning system that converts telecom customer data into reproducible churn predictions while separating experimentation from the model actually served in production-style inference.

Objective

Create an end-to-end churn prediction platform covering ingestion, feature engineering, baseline modeling, hyperparameter tuning, experiment tracking, persisted artifacts, API inference, containerization, and automated quality checks.

Architecture

A modular pipeline ingests and validates telecom customer data, engineers model-ready features, trains a Logistic Regression baseline and separately tunes XGBoost, tracks experiments in MLflow, persists artifacts, and serves the baseline model through FastAPI inside Docker.

  1. 01

    Telecom customer dataset

  2. 02

    Validated ingestion

  3. 03

    Engineered model features

  4. 04

    Baseline and tuned model training

  5. 05

    MLflow experiment tracking

  6. 06

    Persisted model artifacts

  7. 07

    FastAPI churn inference

Data Ingestion

Loads and validates the telecom churn dataset before downstream processing.

  • Python
  • Pandas

Feature Engineering

Transforms customer records into model-ready features, including categorical encoding and target preparation.

  • Python
  • Pandas
  • scikit-learn

Model Training

Trains a reproducible Logistic Regression baseline with preprocessing and persisted evaluation artifacts.

  • scikit-learn
  • Joblib

Model Tuning

Tunes an XGBoost classifier with GridSearchCV as a separate development model artifact.

  • XGBoost
  • scikit-learn

Experiment Tracking

Records machine learning experiments and artifacts for reproducibility and comparison.

  • MLflow

Inference API

Loads the persisted baseline model and exposes churn predictions through an HTTP API.

  • FastAPI
  • Uvicorn

Containerized Runtime

Packages and runs the inference service reproducibly using container tooling.

  • Docker
  • Docker Compose

Quality Automation

Validates pipeline behavior through automated tests and continuous integration.

  • pytest
  • GitHub Actions

Technology

Language

  • Python 3.11

Data

  • Pandas

Ml Ai

  • scikit-learn
  • XGBoost
  • MLflow
  • Joblib

Framework

  • FastAPI
  • Uvicorn

Devops

  • Docker
  • Docker Compose
  • GitHub Actions

Testing

  • pytest

Engineering Contribution

Built a modular ML data pipeline

Implemented ingestion and feature-engineering stages that transform telecom customer data into reproducible model-ready inputs.

Implemented reproducible baseline training

Built a scikit-learn training pipeline with median imputation and Logistic Regression, persisted the trained model, and recorded training metadata.

Added separate XGBoost model tuning

Implemented GridSearchCV-based XGBoost tuning as a separate development workflow rather than conflating experimental tuning with the currently served model.

Added experiment tracking and artifact persistence

Integrated MLflow for experiment tracking and persisted model and reporting artifacts for repeatable development workflows.

Delivered containerized model inference

Exposed churn predictions through FastAPI and packaged the inference service with Docker and Docker Compose.

Added automated testing and CI

Implemented automated tests and GitHub Actions checks to validate the ML platform continuously.

Engineering Evidence

Code

Validated Telecom Data Ingestion

The ingestion workflow processes the 7,043-row, 21-column telecom churn dataset into a validated processed dataset.

Code

Model Feature Pipeline

The feature pipeline performs categorical encoding, converts Churn to a binary target, excludes customerID, and produces model-ready feature data.

Code

Logistic Regression Baseline

The baseline uses a scikit-learn Pipeline with median SimpleImputer preprocessing and LogisticRegression configured with random_state=42, max_iter=1000, and liblinear.

Code

XGBoost Hyperparameter Tuning

The project tunes an XGBoost classifier using GridSearchCV and persists the tuned model separately from the serving baseline.

Documentation

MLflow Experiment Tracking

Training workflows log machine learning experiment information and artifacts through MLflow.

Validation

Persisted Model and Report Artifacts

The project persists trained models and machine learning reports so development results can be reproduced and inspected.

Code

FastAPI Prediction Service

The API exposes churn inference and returns a predicted class, a Yes/No churn label, and churn probability.

Documentation

Explicit Serving Model Boundary

The current API loads the persisted Logistic Regression baseline for inference; the tuned XGBoost model is produced and persisted separately as a development artifact.

Ci Cd

Dockerized Inference Runtime

The FastAPI service is packaged for containerized execution with Docker and Docker Compose.

Test

Automated ML Platform Tests

The audited repository includes automated tests across ingestion, feature engineering, training, and serving workflows.

Ci Cd

GitHub Actions Continuous Integration

GitHub Actions runs automated project quality checks as part of the repository workflow.

Documentation

Current API Input Boundary

The current prediction API expects engineered model features rather than accepting raw customer records directly.

What It Proves

Machine learning engineering

Demonstrates modular ingestion, feature engineering, training, tuning, artifact persistence, and model-serving workflows.

MLOps foundations

Demonstrates experiment tracking, containerized inference, automated testing, and continuous integration around a machine learning system.

Production-aware model serving

Demonstrates explicit separation between experimental model tuning and the artifact currently used by the inference API.