Data Engineer with a background in software engineering and 3+ years in the tech industry spanning fintech, Web3,
and AI automation. Self-taught across the modern data stack with hands on experience building end to end pipelines,
distributed processing systems, and AI-powered workflows.
Currently pursuing a CS degree while actively building production grade data engineering projects.
Category Tools
- processing - Apache Spark, PySpark, Pandas
- Cloud - Microsoft Azure, Azure Data Lake Storage
- Platform - DataBricks, Delta lake
- Database - Postgresql, MySql
- Languages - Python, Sql
- Automation - n8n
- BI/ Analysis - Power BI, Excel, DataBricks Dashboard
- End to end ELT pipeline on Databricks using the real world Berka dataset (1M+ transactions, 8 CSV files).
- Implements Bronze/Silver/Gold Medallion architecture with a full star schema data model (5 dimensions + 2 fact tables).
- Includes real data cleaning challenges, encoded Czech values, malformed nulls, and date decoding.
Exploratory analysis on 9,355 job postings using PostgreSQL.
- Top paying job titles
- Salary by experience level
- Country-based salary comparison
- Remote vs in-person trends
Large scale analysis on 1.6 million job postings across 4 tables.
- Most in-demand skills
- Top paying skills
- Top hiring companies
- Top 5 Data Engineer skills
End-to-end ELT pipeline on Azure using the Bronze/Silver/Gold Medallion architecture. Built with PySpark, Delta Live Tables, and Databricks Workflows. 📁 View Project
- Python basics
- Pandas
- SQL fundamentals
- Advanced SQL (Subqueries, CTEs, Window Functions)
- PySpark
- Databricks
- Cloud (AWS/GCP)
Email: aladejoseph656@gmail.com
Built in public as part of a focused data engineering transition. Follow the journey branch by branch.