Repository containing portfolio of data science projects completed by me for self learning and hobby purposes. Presented in the form of Jupyter notebooks and Python files.
- Install dependencies using requirements.txt.
- Run notebooks as usual by using a jupyter notebook server, Vscode etc.
-
- Retail Sales Efficiency Analytics | Demo: Every retail manager schedules staff by footfall instinct. This project proves or disproves that instinct with numbers from 421,570 weekly sales records across 45 stores and 99 departments spanning 3 years of real transaction data.
-
- Retail Sales Efficiency Analytics | Demo: Every retail manager schedules staff by footfall instinct. This project proves or disproves that instinct with numbers from 421,570 weekly sales records across 45 stores and 99 departments spanning 3 years of real transaction data.
- Bank Marketing — Revenue Intelligence Pipeline | Demo: Small business owners make pricing and staffing decisions based on gut feel about their "busy seasons" — but rarely quantify how much revenue they leave on the table by under-resourcing during peak demand cycles.
- Global Company Financial Intelligence Pipeline | Demo: An automated end-to-end pipeline that unifies 5 global financial datasets into a single "Investment Intelligence" framework.
- E-Commerce Shipment Delay Prediction | Demo: 59.7% of 10,999 shipments arrived late across a 2-year e-commerce logistics dataset. The analysis surfaces one structural pattern: discount depth and operational friction before dispatch predict failure more reliably than shipment mode or product weight.
- Football Player Performance Analytics | Demo: Among 1039 players analysed, only 3 (3.7%) qualify as Elite tier, confirming the Pareto principle in football talent. Average performance score across 81 qualified players: 0.142.
- Spotify Global Streams Analytics | Demo: 730 tracks. 2.04B daily streams. The analysis surfaces three distinct predictive patterns in streaming velocity. Collaborative tracks register 18–24% higher daily share percentages than solo releases. Top 5 momentum tracks isolated through priority scoring show breakout potential within 14–30 day windows. Rank volatility exhibits 67% correlation with algorithmic playlist reshuffles, indicating external driver dominance over track permanence. Isolation Forest detected 12–15% of catalog exhibiting anomalous velocity spike patterns inconsistent with longevity trend
- Heart Disease Analytics | Demo : Disease prevalence in this cohort is 44.4%. Asymptomatic chest pain (CP=0) carries the highest subgroup disease rate. ML model enables automated triage with strong AUC performance.
- Screen Time & Mental Health Analytics | Demo : In a cohort of 999 subjects, 16.2% meet BDI depression criteria. Average leisure screen time is 3.96 hours/day; high-screen users show markedly higher risk.
- EEG Eye State Analysis | Demo : Eyes-closed state accounts for 31.7% of 1000 EEG readings. Occipital and frontal channels are the strongest discriminators. Gradient Boosting achieves best AUC.
- Cruise Reservation Analytics | demo: 77,040 reservations. 2018–01-05 to 2022–02–06. 20 features. The analysis identifies Transatlántico as the highest-value route at $7,922 per reservation.
- Airline Passenger No-Show Prediction : 10,000 passenger records. 24 raw features. 11 engineered features. The analysis quantifies a 4.7% no‑show rate and $16,321.43 revenue at risk.
-
- Facebook Ad Analytics & AI Assistant: Facebook Ad Analytics AI Assistant is a Python-based tool designed to enhance and automate the analysis of Facebook advertising campaign data. Leveraging AI-driven insights, this assistant streamlines data processing, uncovers actionable trends, and supports marketing teams in optimizing their ad strategies for better ROI.
- Google Analytics Dashboard & Predictor: Developed an end-to-end Google Analytics model using Jupyter Notebook and Python, focusing on data extraction, preprocessing, analysis, and visualization. The project demonstrates a full workflow from raw analytics data to actionable insights, leveraging machine learning techniques for predictive analysis and reporting.
- Real Estate Insights: End-to-End ML Deployment: Developed a machine learning model to predict house prices based on key features such as location, size, and amenities. Leveraged Python and Jupyter Notebook for data analysis, preprocessing, and model development, and containerized the application using Docker for seamless deployment.
- Video Game Sale Predictor: Streamlit Deployment: Developed a data science pipeline to analyze and predict video game sales based on historical data and key features such as platform, genre, and region. Leveraged Jupyter Notebook and Python for data preprocessing, feature engineering, and visualization, delivering actionable insights through machine learning models.
-
-
- Consumer Behaviour Analysis And Model Building: Designed a data science solution to analyze consumer behavior and build predictive models for actionable insights. Leveraged Jupyter Notebook for exploratory analysis and Python for implementing machine learning models to identify key trends and predict consumer actions.
- MSFT Stock Price Prediction: Developed a machine learning model to predict the future stock prices of Microsoft Corporation (MSFT) using historical data
-
- DocQuery: Document Oracle is an intelligent document analysis system that enables natural conversation with PDF documents. Using RAG architecture and advanced embedding techniques, it allows users to ask questions about their documents and receive context-aware responses.
-
- 👋I'm contuniously developing skills along with building real world projects. Please explore above projects and share your feedback, I'd like to get your suggestion🙏!