Student mental health has a direct bearing on academic performance and overall well-being, yet problems such as stress, anxiety and depression usually go unnoticed until they have already affected a student\'s grades or attendance. Counsellors and academic staff typically rely on manual observation, self-reported questionnaires, or referrals to identify students who may be struggling, an approach that is reactive rather than preventive and does not scale well across large student populations. This paper presents MindTrack, a machine learning based framework for predicting student mental health risk from a combination of academic and psychological indicators, namely age, gender, study hours, sleep hours, attendance, CGPA, stress level, anxiety level and depression level. Five supervised classifiers — Support Vector Machine (SVM), Random Forest, Decision Tree, K-Nearest Neighbors (KNN) and Naive Bayes — were trained on a dataset of 12,000 student records and were independently validated on a separate set of 399 records so that the reported performance reflects behaviour on unseen data rather than the training data itself. Among the five models, the Support Vector Machine produced the most balanced and reliable results, reaching a validation accuracy of 80.45% and an F1-score of 0.8203. A Random Forest based feature importance analysis further showed that stress level, anxiety level and depression level were the strongest predictors of risk, together accounting for more than 71% of the total feature contribution. All five trained models were deployed inside a Flask based web dashboard that walks a user through an Overview, Academic Info, Health Indicators and Results screen, returning an instant High Risk / Low Risk classification along with a side-by-side comparison of every model and a feature importance chart. The results indicate that combining routinely available academic records with short psychological self-assessments can support earlier and more consistent identification of at-risk students than manual counselling alone.
Introduction
MindTrack is a machine-learning-based system designed to help educational institutions identify students who may be at high risk of mental health difficulties such as stress, anxiety, and depression. The system is intended to support counsellors and educators by providing an early, consistent, data-driven warning rather than replacing professional judgement.
Problem
Student mental health problems can be influenced by academic pressure, examination workload, poor sleep, personal circumstances, and changes in study habits. Traditionally, institutions identify these problems through faculty observation, student complaints, questionnaires, or counselling referrals. This approach is:
Reactive rather than preventive.
Time-consuming and difficult to scale.
Dependent on individual staff judgement.
Unable to continuously combine academic and psychological information.
Likely to identify problems only after they affect attendance, grades, or behaviour.
Proposed Solution
MindTrack uses supervised machine learning to predict whether a student belongs to a High Risk or Low Risk mental-health category. It combines academic information with psychological self-assessment data and provides the prediction through an interactive Flask web dashboard.
The system uses nine main attributes:
Age
Gender
Study hours
Sleep hours
Attendance
CGPA
Stress level
Anxiety level
Depression level
Dataset and Methodology
The project uses:
12,000 student records for training.
399 independent records for validation.
3,425 training records are High Risk and 8,575 are Low Risk.
The data is standardized using a scaler fitted only on the training data to avoid data leakage.
Five machine-learning algorithms are compared:
Support Vector Machine (SVM) with an RBF kernel.
Random Forest with 200 decision trees.
Decision Tree.
K-Nearest Neighbors (KNN).
Naive Bayes.
The models are evaluated using measures such as accuracy, F1-score, sensitivity, and specificity, rather than relying only on accuracy because the dataset is imbalanced.
Important Features
The analysis indicates that stress, anxiety, and depression levels are particularly important indicators of mental-health risk. Sleep and academic factors such as attendance, CGPA, and study hours are also considered because psychological and academic factors can interact with each other.
A Random Forest feature-importance panel is included so that users can see which factors contribute most to the prediction. This makes the system more transparent than simply providing a High Risk/Low Risk result.
Advantages of MindTrack
Compared with traditional manual screening, MindTrack provides:
Early identification of potentially at-risk students.
Automated and consistent predictions.
Comparison of multiple ML models.
Real-time results through a web dashboard.
Feature-importance explanations to support human decision-making.
Scalability for large student populations.
Conclusion
This paper presented MindTrack, a machine learning based system for predicting student mental health risk from nine academic and psychological attributes. Five supervised classifiers were trained on 12,000 student records and validated independently on 399 separate records, with the Support Vector Machine emerging as the best-performing model at 80.45% validation accuracy and an F1-score of 0.8203. A Random Forest based feature importance analysis identified stress, anxiety and depression scores as the dominant predictors, together contributing more than 71% of the total feature importance, ahead of sleep hours, CGPA, attendance, age, study hours and gender. All five models were integrated into a Flask-based interactive dashboard that walks a user through an Overview, Academic Info, Health Indicators and Results workflow, returning a real-time risk classification along with a transparent comparison of every model and the underlying feature importance.
The system demonstrates that combining routinely available academic records with a short psychological self-assessment can support faster and more consistent identification of at-risk students than manual, referral-based observation alone, without requiring any specialised statistical knowledge from the person using the dashboard.
References
[1] Gururaj G, Varghese M, Benegal V, et al. National Mental Health Survey of India: Prevalence, Pattern and Outcomes. Bengaluru: National Institute of Mental Health and Neuro Sciences (NIMHANS); 2016:129.
[2] Patel V, Ramasundarahettige C, Vijayakumar L, et al. Suicide mortality in India: a nationally representative survey. Lancet. 2012;379(9834):2343–2351.
[3] Mareiniss DP. Decreasing GME training stress to foster residents\' professionalism. Acad Med. 2004;79(9):825–831.
[4] Shanafelt TD, Bradley KA, Wipf JE, Back AL. Burnout and self-reported patient care in an internal medicine residency program. Ann Intern Med. 2002;136(5):358–367.
[5] Ibrahim AK, Kelly SJ, Adams CE, Glazebrook C. A systematic review of studies of depression prevalence in university students. J Psychiatr Res. 2013;47(3):391–400.
[6] Islam MR, Rahman MM, Mondal MRH. Deep learning for health informatics. IEEE J Biomed Health Inform. 2017;21(5):1246–1259.
[7] Fatani H, Alharthi R, Alghamdi A, et al. A comprehensive review on machine learning in healthcare industry: classification, restrictions, opportunities and challenges. MDPI Sensors. 2023;23:4178.
[8] IHME (Institute for Health Metrics and Evaluation). Global Burden of Disease Study 2019 (GBD 2019) Results. Seattle: IHME; 2021.
[9] Shetty A, Rao R, Sweekritha KC, Zuha F, Rai S. Machine Learning Based Student Mental Health Predictor. 2025 IEEE International Conference on Distributed Computing, VLSI, Electrical Circuits and Robotics (DISCOVER).
[10] Pedregosa F, Varoquaux G, Gramfort A, et al. Scikit-learn: Machine Learning in Python. Journal of Machine Learning Research. 2011;12:2825–2830.
[11] Grinberg M. Flask Web Development: Developing Web Applications with Python. O\'Reilly Media; 2018.
[12] McKinney W. Data Structures for Statistical Computing in Python (Pandas). Proceedings of the 9th Python in Science Conference. 2010:51–56.
[13] Harris CR, Millman KJ, van der Walt SJ, et al. Array programming with NumPy. Nature. 2020;585:357–362.
[14] Cortes C, Vapnik V. Support-vector networks. Machine Learning. 1995;20(3):273–297.
[15] Breiman L. Random Forests. Machine Learning. 2001;45(1):5–32.
[16] Alani M. Guide to OSI and TCP/IP Models. Springer International Publishing; 2014.