Skip to content
View Atharva-512's full-sized avatar
🏠
Working from home
🏠
Working from home

Block or report Atharva-512

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Atharva-512/README.md
Typing SVG

LinkedIn Gmail GitHub



01  🧰 Technical Toolbox

Languages  
Cloud & Data Engineering  
Databases & Warehousing  
BI, Analytics & DevOps  

02  🚀 Featured Projects

🍽️ Restaurant POS ELT Pipeline & Analytics Warehouse

Problem → Restaurant operators reconciled sales, ops, and KPI reports manually across disconnected POS exports. Built → A production-grade, metadata-driven ELT pipeline staging raw POS reports into an analytics-ready DuckDB warehouse. Impact → One automated source of truth for sales performance, ops efficiency, and KPI reporting — replacing manual reconciliation entirely. Architecture → Bronze → Silver → Gold staging · Star schema — 3 fact tables, 6 dimension tables, 14 reporting views · Containerized, CI-automated deployment · 3 interactive Power BI dashboards.

Python SQL Pandas DuckDB Apache Parquet Power BI Docker GitHub Actions

Repo


☁️ End-to-End E-Commerce Data Pipeline on Azure (Olist)

Problem → E-commerce order data lived across HTTP APIs and SQL Server with no unified, low-latency reporting layer. Built → Parameterized Azure Data Factory pipelines ingesting 100K+ order records into a Databricks-processed, Synapse-served warehouse. Impact → Reliable, low-latency business reporting at scale with fault-tolerant multi-source ingestion. Architecture → Medallion architecture on Databricks (PySpark) · Partitioned Synapse SQL views · Fault-tolerant batch ingestion into ADLS Gen2.

Azure Data Factory ADLS Gen2 Azure Databricks PySpark Synapse Analytics MongoDB Power BI

Repo


🛍️ Retail Sales Analytics Platform

Problem → Retail transaction data held revenue and segmentation insight that wasn't surfaced anywhere. Built → SQL and Python-driven analysis of 2,000+ transactions feeding interactive Power BI dashboards. Impact → Revenue trends, customer segments, and discount impact surfaced across 6 categories, 5 regions, 3 channels to support growth decisions. Architecture → SQL-based ETL workflows · EDA & statistical analysis for analytics-ready datasets · KPI-tracking executive dashboards.

SQL Python Pandas NumPy Matplotlib Seaborn Power BI MySQL Jupyter

Repo


03  🎯 Current Focus

🎯 Building

  • Scalable cloud-native data platforms on Azure
  • Distributed processing with Databricks + PySpark
  • Dimensional data modeling & modern ELT
  • Executive-grade BI dashboards

📈 Snapshot

  • 3 production-style ELT/ETL pipelines shipped
  • 100K+ records processed in Azure pipeline
  • 14 reporting views across a star-schema warehouse
  • National Finalist — Smart India Hackathon

04  🧭 Engineering Focus Matrix

Area Focus
☁️ Cloud Data Engineering Azure Data Factory · Azure Databricks · ADLS Gen2
🏛️ Data Warehousing Azure Synapse Analytics · DuckDB · Star Schema Modeling
Distributed Processing PySpark · Staged Batch Pipelines
🧮 Analytics Engineering SQL Modeling · Data Quality · Incremental Processing
📊 Business Intelligence Power BI · Tableau · DAX · KPI Reporting
🔁 Pipeline Automation Docker · GitHub Actions · CI/CD

05  📊 GitHub Analytics





⏳ Stats occasionally take a few seconds to load or need a page refresh — that's GitHub's image cache, not a broken widget. If a card ever shows blank, refresh the page once.


06  🎓 Certifications

  • 🎓 Big Data Engineering Bootcamp with GCP & Azure Cloud (74 Hours) — Udemy
  • 🎓 Ultimate Job-Ready AI-Powered Data Analytics Course — CodeWithHarry

💌 Let's Build Something Data-Driven

Open to conversations on data engineering, analytics platforms, and cloud architecture.

LinkedIn Email




Pinned Loading

  1. azure-ecommerce-data-pipeline azure-ecommerce-data-pipeline Public

    End-to-End Azure Data Engineering Pipeline using Azure Data Factory, Databricks, ADLS Gen2, PySpark, and Synapse Analytics for E-Commerce Analytics.

    Python 1

  2. real-time-financial-monitoring real-time-financial-monitoring Public

    Real-Time Financial Transaction Monitoring System using Kafka, Spark Structured Streaming and MongoDB

    Python 1

  3. retail-supply-chain-pipeline retail-supply-chain-pipeline Public

    End-to-End Retail Supply Chain Data Pipeline using Apache Airflow, PySpark and SQL

    Python 1

  4. customer-segmentation-analytics customer-segmentation-analytics Public

    Customer Segmentation & Marketing Analytics System using Python, SQL, Tableau and Machine Learning.

    Jupyter Notebook

  5. retail-sales-analytics retail-sales-analytics Public

    End-to-End Retail Sales Analytics & Business Intelligence Platform using SQL, Python, Power BI and Excel.

    Jupyter Notebook 1

  6. smart-incident-management-system smart-incident-management-system Public

    Real-Time Incident Management & Escalation System built using Java, Spring Boot, MySQL, Gmail API and n8n workflow automation.

    Java