project Β· 2021
Enterprise data lake on GCP
Architected an enterprise data lake with a metadata-driven PySpark framework, schema versioning, and reusable data-quality patterns.
The problem
A growing organisation needed a dependable central platform for data ingestion, transformation, and analytics, with conventions that would support more teams over time.
Approach
- Designed a layered lake architecture on managed cloud data services.
- Built a metadata-driven PySpark framework to generate consistent ingestion and transformation workflows.
- Established schema-versioning, validation, review, and knowledge-sharing practices for the engineering team.
Outcome
The project created a scalable foundation for new data products and a repeatable engineering model for building them.