Authors :
Augustin Pambi Tadiamba; Pierre Katalay Kafunda; David Mayoya Kutangila; Joseph Shanga Minga; Bebe Bebe Ngwaba; Joseph Bimbala Ngwaba; Grace-Roven Gracia Tshimanga; Jean-Jacques Matondo Katshitshi; Roger Kwenge Mboma
Volume/Issue :
Volume 11 - 2026, Issue 7 - July
Google Scholar :
https://tinyurl.com/ysyvkeue
Scribd :
https://tinyurl.com/5fz3rw85
DOI :
https://doi.org/10.38124/ijisrt/26jul246
Note : A published paper may take 4-5 working days from the publication date to appear in PlumX Metrics, Semantic Scholar, and ResearchGate.
Abstract :
Educational data can help identify, before the end of a semester, academic trajectories that deserve timely
attention. This study develops an academic risk profiling and tailored student guidance using machine learning techniques
for universities in Kinshasa. A quantitative, experimental, and predictive design was applied to a harmonized dataset of
9,000 student-semester observations and fourteen predictors after removing a redundant composite engagement index. Data
preparation, modelling, and visualisation were performed in a reproducible Python/Jupyter Notebook workflow using
pandas, NumPy, scikit-learn, and Matplotlib. A stratified 80/20 training-test split and five-fold stratified cross-validation
were used to compare multinomial logistic regression, decision tree, random forest, and multilayer perceptron models. On
the independent test set, the multilayer perceptron achieved the highest accuracy (0.720) and macro-AUC-ROC (0.945),
while logistic regression achieved the highest balanced accuracy (0.714). Random forest was retained because it achieved
the highest macro F1 score (0.686), the prespecified criterion for protecting attention to minority classes, together with
macro precision of 0.713 and macro One-vs-Rest AUC-ROC of 0.941. Continuous assessment average, midterm average,
and the previous validated-credit rate were the most informative signals. Recommendations are derived from the predicted
class, while uncertain cases are flagged for human review. The proposed model is therefore a decision-support tool rather
than an automated decision-maker.
Keywords :
Machine Learning; Educational Data; Academic Risk Profiling; Random Forest; Tailored Student Guidance.
References :
- P. Long and G. Siemens, “Penetrating the fog: Analytics in learning and education,” EDUCAUSE Review, vol. 46, no. 5, pp. 30–40, 2011.
- G. Siemens and R. S. J. d. Baker, “Learning analytics and educational data mining: Towards communication and collaboration,” in Proceedings of the 2nd International Conference on Learning Analytics and Knowledge, 2012, pp. 252–254, doi: 10.1145/2330601.2330661.
- C. Romero and S. Ventura, “Educational data mining and learning analytics: An updated survey,” WIREs Data Mining and Knowledge Discovery, vol. 10, no. 3, e1355, 2020, doi: 10.1002/widm.1355.
- R. S. J. d. Baker and P. S. Inventado, “Educational data mining and learning analytics,” in Learning Analytics: From Research to Practice, J. A. Larusson and B. White, Eds. Springer, 2014, pp. 61–75, doi: 10.1007/978-1-4614-3305-7_4.
- S. B. Kotsiantis, C. J. Pierrakeas, and P. E. Pintelas, “Predicting students’ performance in distance learning using machine learning techniques,” Applied Artificial Intelligence, vol. 18, no. 5, pp. 411–426, 2004, doi: 10.1080/08839510490442058.
- P. Cortez and A. M. G. Silva, “Using data mining to predict secondary school student performance,” in Proceedings of the 5th Annual Future Business Technology Conference, 2008, pp. 5–12.
- S. M. Jayaprakash, E. W. Moody, E. J. M. Lauría, J. R. Regan, and J. D. Baron, “Early alert of academically at-risk students: An open source analytics initiative,” Journal of Learning Analytics, vol. 1, no. 1, pp. 6–47, 2014.
- D. Ifenthaler and J. Y.-K. Yau, “Using learning analytics to support study success in higher education: A systematic review,” Educational Technology Research and Development, vol. 68, pp. 1961–1990, 2020, doi: 10.1007/s11423-020-09788-z.
- D. Gašević, S. Dawson, and G. Siemens, “Let’s not forget: Learning analytics are about learning,” TechTrends, vol. 59, no. 1, pp. 64–71, 2015, doi: 10.1007/s11528-014-0822-x.
- G. Siemens, “Learning analytics: The emergence of a discipline,” American Behavioral Scientist, vol. 57, no. 10, pp. 1380–1400, 2013, doi: 10.1177/0002764213498851.
- S. Slade and P. Prinsloo, “Learning analytics: Ethical issues and dilemmas,” American Behavioral Scientist, vol. 57, no. 10, pp. 1510–1529, 2013, doi: 10.1177/0002764213479366.
- N. Sclater, “Developing a code of practice for learning analytics,” Journal of Learning Analytics, vol. 3, no. 1, pp. 16–42, 2016, doi: 10.18608/jla.2016.31.3.
- UNESCO, Recommendation on the Ethics of Artificial Intelligence. Paris, France: UNESCO, 2021.
- S. Kaufman, S. Rosset, C. Perlich, and O. Stitelman, “Leakage in data mining: Formulation, detection, and avoidance,” ACM Transactions on Knowledge Discovery from Data, vol. 6, no. 4, Art. no. 15, 2012, doi: 10.1145/2382577.2382579.
- M. Kuhn and K. Johnson, Feature Engineering and Selection: A Practical Approach for Predictive Models. Boca Raton, FL: CRC Press, 2019.
- L. Breiman, “Random forests,” Machine Learning, vol. 45, no. 1, pp. 5–32, 2001.
- T. Hastie, R. Tibshirani, and J. Friedman, The Elements of Statistical Learning: Data Mining, Inference, and Prediction, 2nd ed. New York, NY: Springer, 2009.
- T. Fawcett, “An introduction to ROC analysis,” Pattern Recognition Letters, vol. 27, no. 8, pp. 861–874, 2006, doi: 10.1016/j.patrec.2005.10.010.
- T. Kluyver et al., “Jupyter Notebooks – a publishing format for reproducible computational workflows,” in Positioning and Power in Academic Publishing: Players, Agents and Agendas, F. Loizides and B. Schmidt, Eds. IOS Press, 2016, pp. 87–90, doi: 10.3233/978-1-61499-649-1-87.
- W. McKinney, “Data structures for statistical computing in Python,” in Proceedings of the 9th Python in Science Conference, 2010, pp. 56–61, doi: 10.25080/Majora-92bf1922-00a.
- C. R. Harris, K. J. Millman, S. J. van der Walt, et al., “Array programming with NumPy,” Nature, vol. 585, pp. 357–362, 2020, doi: 10.1038/s41586-020-2649-2.
- F. Pedregosa, G. Varoquaux, A. Gramfort, et al., “Scikit-learn: Machine learning in Python,” Journal of Machine Learning Research, vol. 12, pp. 2825–2830, 2011.
- J. D. Hunter, “Matplotlib: A 2D graphics environment,” Computing in Science & Engineering, vol. 9, no. 3, pp. 90–95, 2007, doi: 10.1109/MCSE.2007.55.
Educational data can help identify, before the end of a semester, academic trajectories that deserve timely
attention. This study develops an academic risk profiling and tailored student guidance using machine learning techniques
for universities in Kinshasa. A quantitative, experimental, and predictive design was applied to a harmonized dataset of
9,000 student-semester observations and fourteen predictors after removing a redundant composite engagement index. Data
preparation, modelling, and visualisation were performed in a reproducible Python/Jupyter Notebook workflow using
pandas, NumPy, scikit-learn, and Matplotlib. A stratified 80/20 training-test split and five-fold stratified cross-validation
were used to compare multinomial logistic regression, decision tree, random forest, and multilayer perceptron models. On
the independent test set, the multilayer perceptron achieved the highest accuracy (0.720) and macro-AUC-ROC (0.945),
while logistic regression achieved the highest balanced accuracy (0.714). Random forest was retained because it achieved
the highest macro F1 score (0.686), the prespecified criterion for protecting attention to minority classes, together with
macro precision of 0.713 and macro One-vs-Rest AUC-ROC of 0.941. Continuous assessment average, midterm average,
and the previous validated-credit rate were the most informative signals. Recommendations are derived from the predicted
class, while uncertain cases are flagged for human review. The proposed model is therefore a decision-support tool rather
than an automated decision-maker.
Keywords :
Machine Learning; Educational Data; Academic Risk Profiling; Random Forest; Tailored Student Guidance.