Economics student transitioning into Data Science & Analytics.
I combine statistical rigor, econometrics, and machine learning to solve real-world problems β with a strong passion for the intersection of economics, sports analytics, and data engineering.
- π Currently Upskilling: Python, Advanced SQL, pandas, scikit-learn, and Data Engineering best practices.
- π― Goal: Building robust end-to-end data pipelines and predictive models with honest, production-ready validation.
β½π Football Performance Persistence
Sports Econometrics & Time-Series Machine Learning
- A sports econometrics study examining whether a football team's performance in La Liga persists year-over-year or reverts to the mean.
- Features rigorous statistical tests (Pearson/Spearman) alongside a machine learning pipeline utilizing temporal validation (
TimeSeriesSplit) to predict match outcomes, evaluated honestly against a baseline model.
Tech Stack: Python Β· pandas Β· scikit-learn Β· SciPy Β· Seaborn
Production-Ready Data Cleaning Engine
- An end-to-end data sanitization pipeline handling exact duplicates, text inconsistencies, mixed date formats, and extreme outliers.
- Implements automated quality controls (
assertstatements) to programmatically guarantee dataset integrity before export.
Tech Stack: Python Β· pandas Β· NumPy
π’ Titanic EDA Analysis
Exploratory Data Analysis & Statistical Hypothesis Testing
- In-depth exploration of survival drivers on the Titanic, employing statistical hypothesis testing (Chi-Square, t-tests) to validate empirical findings.
- Includes dynamic multi-variable visualizations in Plotly and an integrated SQL querying workflow.
Tech Stack: Python Β· pandas Β· SQL Β· Seaborn Β· Plotly Β· SciPy
β‘ Portfolio in active development β stay tuned for new projects! β‘