Ở trường đại học, sinh viên thường chỉ nhận ra mình sắp rớt môn khi môn học đã gần kết thúc (buổi 8 - 9 đối với môn 9 tuần, hoặc tuần 13 - 15 đối với học kỳ 15 tuần truyền thống). Lúc này thì điểm quá trình đã chốt, số buổi vắng đã vượt khung 20% quy chế và gần như không còn cơ hội cứu vãn.
Dự án này giải quyết bài toán: Dựa vào dữ liệu ở mốc 50% thời lượng môn học (sau 4 - 5 buổi đối với môn 9 tuần, hoặc tuần 6 - 8 đối với kỳ 15 tuần, khi chưa hề có điểm thi cuối kỳ), làm sao phát hiện sớm sinh viên đang có nguy cơ rớt môn để kịp thời hỗ trợ?
- Nếu chờ có điểm thi rồi nhân hệ số cộng lại thì đó chỉ là phép tính số học, không giúp ích gì cho việc can thiệp sớm.
- Ở giai đoạn giữa môn, hành vi học tập của sinh viên mang tính phi tuyến tính:
- Có bạn điểm danh đầy đủ nhưng điểm kiểm tra lại tụt dần qua các buổi.
- Có bạn vắng 2 buổi nhưng bài tập thực hành vẫn đạt điểm tốt.
- Có bạn đang gánh quá nhiều tín chỉ kết hợp nợ môn ở kỳ trước, dẫn đến quá tải.
- Machine Learning giúp kết hợp các tín hiệu hành vi (chuyên cần, xu hướng điểm, tải học vụ, giờ tự học) để ước lượng xác suất rủi ro trước khi quá muộn.
Thử nghiệm được thực hiện trên 1,200 hồ sơ sinh viên, chia theo tỷ lệ 80% huấn luyện và 20% kiểm thử (Stratified Split để giữ nguyên tỷ lệ nhãn).
| Mô hình | Accuracy | Precision | Recall | F1-Score | ROC-AUC |
|---|---|---|---|---|---|
| Logistic Regression (Baseline) | 89.17% | 84.48% | 92.45% | 88.29% | 0.9807 |
| LightGBM | 89.17% | 85.71% | 90.57% | 88.07% | 0.9714 |
| Random Forest | 88.75% | 84.96% | 90.57% | 87.67% | 0.9673 |
Trong bài toán hỗ trợ sinh viên:
- False Negative (bỏ sót sinh viên sắp rớt môn): Hậu quả rất lớn vì sinh viên đó sẽ rớt môn thật do không được nhắc nhở.
- False Positive (nhắc nhở sinh viên chăm): Hậu quả không đáng kể, chỉ là một email hỏi thăm từ cố vấn học tập.
Do đó, chúng ta thực hiện quét ngưỡng xác suất (Threshold Tuning):
- Ở ngưỡng tối ưu
$\tau = 0.75$ : Mô hình đạt Precision 95.88% mà vẫn giữ được Recall 87.74% và F1-Score 91.63%. - Việc này giúp bộ phận cố vấn học tập lọc ra đúng nhóm sinh viên thực sự cần can thiệp mà không tạo ra cảnh báo giả tràn lan.
Một mô hình đưa vào giáo dục không thể chỉ trả về con số chung chung "bạn có 80% nguy cơ rớt môn". Chúng ta sử dụng SHAP (SHapley Additive exPlanations) để chỉ ra nguyên nhân cụ thể:
- Điểm Quiz 2 và Điểm phạt vắng học (
engagement_penalty): Là hai chỉ báo mạnh nhất. Điểm Quiz 2 thấp (màu xanh) đẩy mạnh giá trị SHAP dương (tăng nguy cơ rớt môn). - Xu hướng điểm (
score_trend): Sinh viên có điểm bài kiểm tra sau thấp hơn bài kiểm tra trước có nguy cơ rớt môn cao hơn rõ rệt so với sinh viên có điểm ổn định. - Giải thích theo từng cá nhân (Waterfall Plot):
student-academic-risk-predictor/
├── data/
│ ├── raw/ # Dữ liệu 1,200 sinh viên
│ └── processed/ # Dữ liệu sau khi trích xuất đặc trưng
├── src/
│ ├── data_loader.py # Khởi tạo và nạp dữ liệu
│ ├── features.py # Kỹ thuật tạo đặc trưng (Feature Engineering)
│ ├── train.py # Huấn luyện, đánh giá và tối ưu ngưỡng
│ ├── evaluate.py # Sinh các biểu đồ đánh giá mô hình
│ └── explain.py # Tính toán SHAP values toàn cục và cục bộ
├── notebooks/
│ └── student_risk_analysis.ipynb # Notebook phân tích dữ liệu chi tiết
├── app/
│ ├── main.py # FastAPI service phục vụ endpoint /api/predict
│ └── static/
│ └── index.html # Giao diện chẩn đoán nguy cơ học vụ
├── reports/
│ └── figures/ # Lưu trữ toàn bộ biểu đồ thực nghiệm
├── requirements.txt # Danh sách thư viện
└── README.md # Tài liệu dự án
.\.venv\Scripts\Activate.ps1# Huấn luyện mô hình và lưu metadata
python src/train.py
# Sinh các biểu đồ đánh giá
python src/evaluate.py
# Sinh các biểu đồ giải thích SHAP
python src/explain.pypython -m uvicorn app.main:app --port 8000Truy cập trình duyệt: http://localhost:8000 để thử nghiệm nạp các hồ sơ sinh viên mẫu.
In higher education, struggling students often realize they are failing only when the course is nearly finished (sessions 8 - 9 in 9-week modular courses, or weeks 13 - 15 in traditional 15-week semesters). By that point, formative grades are locked, attendance deficits exceed institutional thresholds (e.g., 20% absence limits), and recovery is virtually impossible.
This project addresses the problem: Using academic and behavioral data collected at the 50% milestone (after sessions 4 - 5 in 9-week courses, or weeks 6 - 8 in 15-week semesters, before any final exams take place), how can we identify students at risk of course failure early enough to provide targeted academic support?
- Final grade calculation is simple arithmetic, but it only happens after final exams are taken.
- At mid-course, student performance signals are non-linear:
- Some students attend all lectures but show declining test scores.
- Some students miss 2 sessions but maintain strong practical lab marks.
- Some carry an overloaded credit burden combined with past course failures.
- Machine Learning combines these early signals (attendance patterns, score trajectory, credit load, study hours) to estimate risk probabilities while intervention is still possible.
The models were evaluated on 1,200 student records using an 80/20 stratified train/test split.
| Model | Accuracy | Precision | Recall | F1-Score | ROC-AUC |
|---|---|---|---|---|---|
| Logistic Regression (Baseline) | 89.17% | 84.48% | 92.45% | 88.29% | 0.9807 |
| LightGBM | 89.17% | 85.71% | 90.57% | 88.07% | 0.9714 |
| Random Forest | 88.75% | 84.96% | 90.57% | 87.67% | 0.9673 |
In educational intervention:
- False Negative (missing an at-risk student): Serious consequence, as the student receives no support and may fail the course.
- False Positive (flagging a passing student): Low consequence, resulting only in a supportive check-in message from an academic advisor.
By conducting probability threshold tuning:
- At optimal threshold
$\tau = 0.75$ : The model reaches Precision 95.88% while maintaining Recall 87.74% and F1-Score 91.63%. - This enables advisors to focus resources on students who genuinely need support without causing alarm fatigue.
Educational models require transparency. We use SHAP (SHapley Additive exPlanations) to explain both global feature importance and individual student diagnoses:
- Key Drivers: Second quiz score, attendance deficit penalty, and negative score trend are the strongest indicators of academic difficulty.
- Individual Waterfall Explanations: For each flagged student, the system isolates the specific factors elevating their risk score to guide constructive advisory conversations.
# 1. Activate environment
.\.venv\Scripts\Activate.ps1
# 2. Run training and report generation
python src/train.py
python src/evaluate.py
python src/explain.py
# 3. Start web interface
python -m uvicorn app.main:app --port 8000Open http://localhost:8000 in your browser.




