Data Engineering
Urban Mobility Data Lakehouse
From raw mobility data to analytics-ready intelligence.
A reproducible data engineering platform that transforms raw NYC taxi data through Raw, Bronze, Silver, and Gold layers and loads curated analytical outputs into DuckDB.
Problem
Context
Raw urban mobility data requires structured ingestion, transformation, validation, and analytical modeling before it can support reliable analysis.
Challenge
Build a reproducible local-first pipeline that preserves raw inputs while progressively transforming them into validated, analytics-ready datasets.
Objective
Create a modular data lakehouse workflow that produces curated Gold marts and makes them available through an analytical DuckDB warehouse.
Architecture
A configuration-driven medallion pipeline moves NYC taxi data through Raw, Bronze, Silver, and Gold layers before loading curated Gold outputs into DuckDB for analytical querying.
- 01
Raw source data
- 02
Bronze structured data
- 03
Silver cleaned and validated data
- 04
Gold analytical marts
- 05
DuckDB analytics warehouse
Raw Layer
Preserves source mobility data as the starting point for reproducible processing.
- Python
Bronze Layer
Creates the first structured lakehouse representation while retaining source-oriented data.
- Python
- Pandas
- PyArrow
- Parquet
Silver Layer
Applies cleaning, standardization, transformation, and validation to produce analysis-ready records.
- Python
- Pandas
- PySpark
- Parquet
Gold Layer
Produces curated analytical marts for trip performance, borough demand, and payment-type revenue.
- Python
- Parquet
Analytics Warehouse
Loads Gold-layer outputs into DuckDB for analytical SQL queries and downstream exploration.
- DuckDB
- SQL
Technology
Language
- Python 3.11
- SQL
Data
- Pandas
- PySpark
- PyArrow
- Parquet
- DuckDB
- YAML
Testing
- Pytest
Devops
- Make
Engineering Contribution
Designed a medallion-style ETL workflow
Implemented a Raw → Bronze → Silver → Gold pipeline that progressively transforms source mobility data into curated analytical datasets.
Built analytics-ready Gold marts
Implemented daily trip, borough-hour demand, and payment-type revenue marts for downstream analysis.
Integrated an analytical DuckDB warehouse
Loaded Gold-layer outputs into a DuckDB analytics warehouse with documented SQL query workflows.
Added automated validation and testing
Implemented automated tests covering pipeline components, transformations, validation behavior, and warehouse behavior.
Made the workflow reproducible
Provided configuration-driven execution and Make-based commands for installation, bootstrap, testing, downloading, and pipeline execution.
Engineering Evidence
Architecture
Raw → Bronze → Silver → Gold Architecture
Repository documentation describes the medallion-style progression from raw mobility data through curated Gold outputs.
Code
Curated Gold Analytical Marts
The project implements daily_trip_summary, borough_hour_demand, and payment_type_revenue analytical marts.
Code
DuckDB Analytics Warehouse
Gold outputs are loaded into a DuckDB analytics warehouse with documented analytical SQL workflows.
Test
Automated Pipeline Tests
The audited repository includes automated tests for pipeline components, transformations, validation, and warehouse behavior.
Documentation
Reproducible Execution Workflow
The repository documents Make-based installation, bootstrap, test, download, and full pipeline execution commands.
What It Proves
Data pipeline engineering
Demonstrates design of a modular multi-stage pipeline that turns raw source data into curated analytical outputs.
Analytical data modeling
Demonstrates creation of purpose-built Gold marts for operational and revenue-oriented analysis.
Data quality and reproducibility
Demonstrates automated validation, testing, configuration-driven execution, and repeatable project workflows.