Skip to content

Latest commit

 

History

1 Commit

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Customer Segmentation Analysis

Overview

This project applies machine learning techniques to classify customers into different segmentation groups based on demographic and behavioural data. The workflow includes data preprocessing, feature engineering, model selection, hyperparameter tuning, and performance evaluation.

The project was developed as part of a university machine learning/data science assignment and focuses on building a complete end-to-end classification pipeline.


Project Structure

customer-segmentation-analysis/
├── README.md
├── requirements.txt
├── segmentation_pipeline.ipynb
├── Segmentation.csv
└── images/
    └── segmentation_pipeline.png

Dataset

The dataset contains customer information used to predict segmentation classes.

Main steps performed on the dataset:

  • Data inspection and cleaning
  • Missing value analysis
  • Feature engineering
  • Encoding categorical variables
  • Feature scaling
  • Handling class imbalance
  • Train/test split preparation

Technologies & Libraries

  • Python
  • pandas
  • NumPy
  • scikit-learn
  • XGBoost
  • imbalanced-learn
  • mlxtend
  • matplotlib
  • missingno

Dependencies are listed in requirements.txt.


Machine Learning Workflow

1. Data Preprocessing

  • Missing value inspection
  • Feature transformations
  • Profession mapping
  • New feature creation using co-occurrence information
  • Scaling and encoding pipelines

2. Model Selection

Different machine learning models were evaluated to compare classification performance.

3. Model Refinement

The selected model (XGBoost) was further improved through:

  • Hyperparameter tuning
  • Cross-validation
  • Performance evaluation

4. Evaluation

The notebook includes:

  • Accuracy and F1-score evaluation
  • Classification reports
  • Confusion matrices
  • Learning curves
  • Validation curves

Results

The project demonstrates the complete process of developing a supervised machine learning pipeline for customer segmentation tasks.

Key outcomes:

  • Built a reproducible ML workflow
  • Compared multiple classification approaches
  • Improved model performance through tuning
  • Generated visual analyses and evaluation metrics

Pipeline Overview

Add your pipeline image here:

![Pipeline Overview](image/segmentation_pipeline.png)

How to Run

1. Clone the repository

git clone https://github.com/your-username/customer-segmentation-analysis.git
cd customer-segmentation-analysis

2. Install dependencies

pip install -r requirements.txt

3. Launch Jupyter Notebook

jupyter notebook

Open:

segmentation_pipeline.ipynb

Repository Notes

  • This repository is intended for educational and portfolio purposes.
  • The project was developed as part of a university assignment.
  • Some preprocessing and modelling choices were made for experimentation and learning purposes.

About

Machine learning pipeline for customer segmentation using Python, XGBoost, and scikit-learn.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages