Data Engineering

Urban Mobility Data Lakehouse

From raw mobility data to analytics-ready intelligence.

A reproducible data engineering platform that transforms raw NYC taxi data through Raw, Bronze, Silver, and Gold layers and loads curated analytical outputs into DuckDB.

Problem

Context

Raw urban mobility data requires structured ingestion, transformation, validation, and analytical modeling before it can support reliable analysis.

Challenge

Build a reproducible local-first pipeline that preserves raw inputs while progressively transforming them into validated, analytics-ready datasets.

Objective

Create a modular data lakehouse workflow that produces curated Gold marts and makes them available through an analytical DuckDB warehouse.

Architecture

A configuration-driven medallion pipeline moves NYC taxi data through Raw, Bronze, Silver, and Gold layers before loading curated Gold outputs into DuckDB for analytical querying.

  1. 01

    Raw source data

  2. 02

    Bronze structured data

  3. 03

    Silver cleaned and validated data

  4. 04

    Gold analytical marts

  5. 05

    DuckDB analytics warehouse

Raw Layer

Preserves source mobility data as the starting point for reproducible processing.

  • Python

Bronze Layer

Creates the first structured lakehouse representation while retaining source-oriented data.

  • Python
  • Pandas
  • PyArrow
  • Parquet

Silver Layer

Applies cleaning, standardization, transformation, and validation to produce analysis-ready records.

  • Python
  • Pandas
  • PySpark
  • Parquet

Gold Layer

Produces curated analytical marts for trip performance, borough demand, and payment-type revenue.

  • Python
  • Parquet

Analytics Warehouse

Loads Gold-layer outputs into DuckDB for analytical SQL queries and downstream exploration.

  • DuckDB
  • SQL

Technology

Language

  • Python 3.11
  • SQL

Data

  • Pandas
  • PySpark
  • PyArrow
  • Parquet
  • DuckDB
  • YAML

Testing

  • Pytest

Devops

  • Make

Engineering Contribution

Designed a medallion-style ETL workflow

Implemented a Raw → Bronze → Silver → Gold pipeline that progressively transforms source mobility data into curated analytical datasets.

Built analytics-ready Gold marts

Implemented daily trip, borough-hour demand, and payment-type revenue marts for downstream analysis.

Integrated an analytical DuckDB warehouse

Loaded Gold-layer outputs into a DuckDB analytics warehouse with documented SQL query workflows.

Added automated validation and testing

Implemented automated tests covering pipeline components, transformations, validation behavior, and warehouse behavior.

Made the workflow reproducible

Provided configuration-driven execution and Make-based commands for installation, bootstrap, testing, downloading, and pipeline execution.

Engineering Evidence

Architecture

Raw → Bronze → Silver → Gold Architecture

Repository documentation describes the medallion-style progression from raw mobility data through curated Gold outputs.

Code

Curated Gold Analytical Marts

The project implements daily_trip_summary, borough_hour_demand, and payment_type_revenue analytical marts.

Code

DuckDB Analytics Warehouse

Gold outputs are loaded into a DuckDB analytics warehouse with documented analytical SQL workflows.

Test

Automated Pipeline Tests

The audited repository includes automated tests for pipeline components, transformations, validation, and warehouse behavior.

Documentation

Reproducible Execution Workflow

The repository documents Make-based installation, bootstrap, test, download, and full pipeline execution commands.

What It Proves

Data pipeline engineering

Demonstrates design of a modular multi-stage pipeline that turns raw source data into curated analytical outputs.

Analytical data modeling

Demonstrates creation of purpose-built Gold marts for operational and revenue-oriented analysis.

Data quality and reproducibility

Demonstrates automated validation, testing, configuration-driven execution, and repeatable project workflows.