Ijraset Journal For Research in Applied Science and Engineering Technology
Authors: Perseus Bhavnagri, Shraddha Ayare Shirke
DOI Link: https://doi.org/10.22214/ijraset.2026.84783
Certificate: View Certificate
Universities are being offered artificial intelligence that flags students likely to fail, and the case for buying it rests on accuracy figures drawn from papers that split one pooled pile of student records at random. This study asks how much of that accuracy a university would actually see. Using the Open University Learning Analytics Dataset, 32,593 student registrations across seven modules and twenty-two presentations with 10,655,280 recorded interactions, a gradient-boosted model was trained to predict failure or withdrawal from what was known on day 30 of a course. Under a random split it reaches an area under the ROC curve of 0.8450. Three sources of inflation were then separated on identical test rows. Losing familiarity with the course costs almost nothing, 0.0085 of area under the curve, consistent in direction across 17 of 22 presentations but small by every effect size used (Wilcoxon p = 0.00219, Cliff\'s delta 0.223, bootstrap interval spanning zero); training on 2013 and predicting 2014 costs 0.0037 and is not significant (p = 0.12720). What does matter is arithmetic rather than transfer. Pooling students across presentations that fail at rates from 0.274 to 0.658 adds 0.0267 that no single course ever sees, and a model told only which presentation a student is enrolled on, nothing about the student, already reaches 0.5777. Larger still, 15.7 percent of students had already unregistered before day 30; removing them costs 0.0575. The model is better at recognising students who have already left than students who will fail while trying (0.8799 against 0.7948). At an alert budget of one student in five it misses 58.6 percent of those at risk, and misses the least deprived more often than the most deprived, a pooled gap of 0.118 that shrinks to 0.086 and reverses in one module once deprivation is compared within courses. Applied together, an unseen course and a population restricted to students still enrolled bring the figure to 0.7520, a fall of 0.0931 of which 62.2 percent is departed students, 28.7 percent is pooling and 9.1 percent is the unfamiliar course. The paper concludes that reported accuracy in this literature is largely a property of how the test set was assembled, and recommends within-course reporting, the exclusion of students already gone, and a stated alert budget.
The text examines a key problem in early-warning systems for student dropout and failure: high predictive accuracy on a benchmark dataset may not translate into equally good performance when a university deploys the model on new courses or future cohorts.
The study uses the Open University Learning Analytics Dataset (OULAD), containing 32,593 student registrations, 10.6 million learning-platform interactions, and seven modules, to investigate why reported accuracy can be misleading.
The central issue is the difference between:
A random split can make a model look better because the training and test sets contain students from the same courses and cohorts.
The researchers test the same model and students under four evaluation designs:
They also investigate whether performance is inflated because of:
Using gradient-boosted trees, the conventional random split produces:
Logistic regression achieves:
However, the study emphasizes that 0.8450 should not automatically be interpreted as the performance a university would obtain in practice.
The strongest evidence comes from a simple experiment where the model is given only the student's module/presentation, with no clicks, demographics, or assessment information.
That model achieves an ROC-AUC of 0.5777.
This means that about 22.5% of the discrimination above random chance achieved by the full model can be obtained simply from knowing which course the student is taking. Since courses have very different failure rates—roughly 27.4% to 65.8% across presentations—a pooled random test can benefit from differences between courses.
Therefore, the study supports its hypothesis that part of the impressive headline performance comes from distinguishing students in high-risk courses from students in low-risk courses, rather than purely identifying individual students who are struggling.
The paper argues that this is not necessarily a modelling failure. A model may genuinely perform well while the evaluation protocol answers the wrong practical question.
For example, a university wants to know:
“Who among the students currently taking this course needs help?”
A random split instead partly answers:
“How well can the model distinguish students from records drawn from a mixture of courses with different historical failure rates?”
Those questions can produce substantially different results.
Previous studies have shown that:
However, the authors argue that previous work has not sufficiently quantified how much performance changes when evaluation becomes more deployment-realistic.
The study aims to provide a deployment-realistic benchmark rather than another high benchmark score. Its central message is:
An impressive predictive score on a randomly split educational dataset does not necessarily represent the performance a university should expect after deployment.
The study specifically seeks to separate the effects of course differences, temporal drift, student withdrawal, and demographic disparities, allowing universities to understand whether a performance gap is caused by the model itself or by the way it was evaluated.
An at-risk model built on the Open University Learning Analytics Dataset and evaluated the way this literature evaluates such models reaches an area under the ROC curve of 0.8450 on day 30 of a course. Taking that figure apart shows that the part of it a university would actually experience is smaller, and that the shortfall is not where it was expected to be. Model transfer, the failure mode the portability literature warns about, turns out to be cheap: an unfamiliar course costs 0.0085 of area under the curve and a year of drift costs 0.0037, neither of which a support office would notice. The expensive terms are structural. Pooling students from courses that fail at rates between 27.4 and 65.8 percent contributes discrimination that no single course contains, to the point that a model told only which presentation a student is on reaches 0.5777 while knowing nothing about the student. And 15.7 percent of the students in the file had already unregistered before the prediction was made; removing them costs 0.0575, the largest single term, because the model is markedly better at recognising a student who has already gone than one who is still present and failing. Put together, a model scoring students on a module it has never seen, among only those still enrolled on the day the prediction is made, averages 0.7520. The 0.0931 that separates that from the published-style figure is 62.2 percent departed students, 28.7 percent pooled courses and 9.1 percent unfamiliarity. At an alert budget of one student in five, the model misses 58.6 percent of those who fail or withdraw. It misses the least deprived more often than the most deprived, and raises six times as many false alarms about the most deprived, both consequences of the same difference in base rates; a quarter of that pooled gap dissolves when deprivation is compared inside courses rather than across them, and on one module it reverses. None of this means early-warning systems should not be built. It means the number on the brochure is a property of an evaluation protocol, and that a university reading it is entitled to ask which students were in the test set, how many courses were stirred together to produce it, and how many people the institution can actually telephone.
[1] V. Kaushal and R. Mall, \"AI-driven early warning systems for student success: discovering static feature dominance in temporal prediction models,\" arXiv:2512.12493, 2025. doi: 10.48550/arXiv.2512.12493. [2] M. El Jihaoui, O. E. K. Abra, and K. Mansouri, \"Towards a fair and transparent early warning system for identifying students at risk of academic failure,\" Frontiers in Education, vol. 11, art. 1878686, 2026. doi: 10.3389/feduc.2026.1878686. [3] J. Kuzilek, M. Hlosta, and Z. Zdrahal, \"Open University Learning Analytics dataset,\" Scientific Data, vol. 4, art. 170171, 2017. doi: 10.1038/sdata.2017.171. [4] E. Howard, \"ouladFormat R package: preparing the Open University Learning Analytics Dataset for analysis,\" arXiv:2501.08366, 2025. doi: 10.48550/arXiv.2501.08366. [5] M. Hlosta, D. Herrmannova, L. Vachova, J. Kuzilek, Z. Zdrahal, and A. Wolff, \"Modelling student online behaviour in a virtual learning environment,\" arXiv:1811.06369, 2018. doi: 10.48550/arXiv.1811.06369. [6] R. da Silva, J. Eicher, and G. Longo, \"A unified survival benchmark for temporal dropout risk prediction in learning analytics,\" arXiv:2604.08870, 2026. doi: 10.48550/arXiv.2604.08870. [7] C. Labba and A. Boyer, \"Towards an online incremental approach to predict students performance,\" arXiv:2407.10256, 2024. doi: 10.48550/arXiv.2407.10256. [8] L. Aulck, N. Velagapudi, J. Blumenstock, and J. West, \"Predicting student dropout in higher education,\" arXiv:1606.06364, 2016. doi: 10.48550/arXiv.1606.06364. [9] J. Gardner and C. Brooks, \"Student success prediction in MOOCs,\" User Modeling and User-Adapted Interaction, vol. 28, no. 2, pp. 127-203, 2018. doi: 10.1007/s11257-018-9203-z. [10] D. Gasevic, S. Dawson, T. Rogers, and D. Gasevic, \"Learning analytics should not promote one size fits all: the effects of instructional conditions in predicting academic success,\" The Internet and Higher Education, vol. 28, pp. 68-84, 2016. doi: 10.1016/j.iheduc.2015.10.002. [11] V. Swamy, M. Marras, and T. Kaser, \"Meta transfer learning for early success prediction in MOOCs,\" Proc. 9th ACM Conf. on Learning @ Scale, pp. 121-132, 2022. doi: 10.1145/3491140.3528273. [12] J. Futoma, M. Simons, T. Panch, F. Doshi-Velez, and L. A. Celi, \"The myth of generalisability in clinical research and machine learning in health care,\" The Lancet Digital Health, vol. 2, no. 9, pp. e489-e492, 2020. doi: 10.1016/S2589-7500(20)30186-2. [13] M. Roberts, D. Driggs, M. Thorpe, J. Gilbey, M. Yeung, S. Ursprung, et al., \"Common pitfalls and recommendations for using machine learning to detect and prognosticate for COVID-19 using chest radiographs and CT scans,\" Nature Machine Intelligence, vol. 3, no. 3, pp. 199-217, 2021. doi: 10.1038/s42256-021-00307-0. [14] R. F. Kizilcec and H. Lee, \"Algorithmic fairness in education,\" arXiv:2007.05443, 2020. doi: 10.48550/arXiv.2007.05443. [15] R. Yu, H. Lee, and R. F. Kizilcec, \"Should college dropout prediction models include protected attributes?,\" Proc. 8th ACM Conf. on Learning @ Scale, pp. 91-100, 2021. doi: 10.1145/3430895.3460139. [16] S. Riazy, K. Simbeck, and V. Schreck, \"Fairness in learning analytics: student at-risk prediction in virtual learning environments,\" Proc. 12th Int. Conf. on Computer Supported Education, pp. 15-25, 2020. doi: 10.5220/0009324100150025. [17] H. Anahideh, N. Nezami, P. Haghighat, and D. Gandara, \"Auditing the imputation effect on fairness of predictive analytics in higher education,\" arXiv:2109.07908, 2021. doi: 10.48550/arXiv.2109.07908. [18] K. A. Bird, B. L. Castleman, and Y. Song, \"Are algorithms biased in education? Exploring racial bias in predicting community college student success,\" Annenberg Institute at Brown University, EdWorkingPaper 23-717; published version in the Journal of Policy Analysis and Management, vol. 44, pp. 379-402, 2024, 2023. doi: 10.26300/yd7z-6e20. [19] S. M. Jayaprakash, E. W. Moody, E. J. M. Lauria, J. R. Regan, and J. D. Baron, \"Early alert of academically at-risk students: an open source analytics initiative,\" Journal of Learning Analytics, vol. 1, no. 1, pp. 6-47, 2014. doi: 10.18608/jla.2014.11.3. [20] R. S. Baker, N. Bosch, S. Hutt, A. F. Zambrano, and A. J. Bowers, \"On fixing the right problems in predictive analytics: AUC is not the problem,\" arXiv:2404.06989, 2024. doi: 10.48550/arXiv.2404.06989. [21] O. Zawacki-Richter, V. I. Marin, M. Bond, and F. Gouverneur, \"Systematic review of research on artificial intelligence applications in higher education - where are the educators?,\" International Journal of Educational Technology in Higher Education, vol. 16, art. 39, 2019. doi: 10.1186/s41239-019-0171-0. [22] D. Lakens, \"Calculating and reporting effect sizes to facilitate cumulative science: a practical primer for t-tests and ANOVAs,\" Frontiers in Psychology, vol. 4, art. 863, 2013. doi: 10.3389/fpsyg.2013.00863. [23] R. A. Kievit, W. E. Frankenhuis, L. J. Waldorp, and D. Borsboom, \"Simpson\'s paradox in psychological science: a practical guide,\" Frontiers in Psychology, vol. 4, art. 513, 2013. doi: 10.3389/fpsyg.2013.00513.
Copyright © 2026 Perseus Bhavnagri, Shraddha Ayare Shirke. This is an open access article distributed under the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.
Paper Id : IJRASET84783
Publish Date : 2026-09-01
ISSN : 2321-9653
Publisher Name : IJRASET
DOI Link : Click Here
Submit Paper Online
