⚠ Official Notice: www.ijisrt.com is the official website of the International Journal of Innovative Science and Research Technology (IJISRT) Journal for research paper submission and publication. Please beware of fake or duplicate websites using the IJISRT name.



Academic Risk Profiling and Tailored Student Guidance Through Machine Learning in Kinshasa Universities


Authors : Augustin Pambi Tadiamba; Pierre Katalay Kafunda; David Mayoya Kutangila; Joseph Shanga Minga; Bebe Bebe Ngwaba; Joseph Bimbala Ngwaba; Grace-Roven Gracia Tshimanga; Jean-Jacques Matondo Katshitshi; Roger Kwenge Mboma

Volume/Issue : Volume 11 - 2026, Issue 7 - July


Google Scholar : https://tinyurl.com/ysyvkeue

Scribd : https://tinyurl.com/5fz3rw85

DOI : https://doi.org/10.38124/ijisrt/26jul246

Note : A published paper may take 4-5 working days from the publication date to appear in PlumX Metrics, Semantic Scholar, and ResearchGate.


Abstract : Educational data can help identify, before the end of a semester, academic trajectories that deserve timely attention. This study develops an academic risk profiling and tailored student guidance using machine learning techniques for universities in Kinshasa. A quantitative, experimental, and predictive design was applied to a harmonized dataset of 9,000 student-semester observations and fourteen predictors after removing a redundant composite engagement index. Data preparation, modelling, and visualisation were performed in a reproducible Python/Jupyter Notebook workflow using pandas, NumPy, scikit-learn, and Matplotlib. A stratified 80/20 training-test split and five-fold stratified cross-validation were used to compare multinomial logistic regression, decision tree, random forest, and multilayer perceptron models. On the independent test set, the multilayer perceptron achieved the highest accuracy (0.720) and macro-AUC-ROC (0.945), while logistic regression achieved the highest balanced accuracy (0.714). Random forest was retained because it achieved the highest macro F1 score (0.686), the prespecified criterion for protecting attention to minority classes, together with macro precision of 0.713 and macro One-vs-Rest AUC-ROC of 0.941. Continuous assessment average, midterm average, and the previous validated-credit rate were the most informative signals. Recommendations are derived from the predicted class, while uncertain cases are flagged for human review. The proposed model is therefore a decision-support tool rather than an automated decision-maker.

Keywords : Machine Learning; Educational Data; Academic Risk Profiling; Random Forest; Tailored Student Guidance.

References :

  1. P. Long and G. Siemens, “Penetrating the fog: Analytics in learning and education,” EDUCAUSE Review, vol. 46, no. 5, pp. 30–40, 2011.
  2. G. Siemens and R. S. J. d. Baker, “Learning analytics and educational data mining: Towards communication and collaboration,” in Proceedings of the 2nd International Conference on Learning Analytics and Knowledge, 2012, pp. 252–254, doi: 10.1145/2330601.2330661.
  3. C. Romero and S. Ventura, “Educational data mining and learning analytics: An updated survey,” WIREs Data Mining and Knowledge Discovery, vol. 10, no. 3, e1355, 2020, doi: 10.1002/widm.1355.
  4. R. S. J. d. Baker and P. S. Inventado, “Educational data mining and learning analytics,” in Learning Analytics: From Research to Practice, J. A. Larusson and B. White, Eds. Springer, 2014, pp. 61–75, doi: 10.1007/978-1-4614-3305-7_4.
  5. S. B. Kotsiantis, C. J. Pierrakeas, and P. E. Pintelas, “Predicting students’ performance in distance learning using machine learning techniques,” Applied Artificial Intelligence, vol. 18, no. 5, pp. 411–426, 2004, doi: 10.1080/08839510490442058.
  6. P. Cortez and A. M. G. Silva, “Using data mining to predict secondary school student performance,” in Proceedings of the 5th Annual Future Business Technology Conference, 2008, pp. 5–12.
  7. S. M. Jayaprakash, E. W. Moody, E. J. M. Lauría, J. R. Regan, and J. D. Baron, “Early alert of academically at-risk students: An open source analytics initiative,” Journal of Learning Analytics, vol. 1, no. 1, pp. 6–47, 2014.
  8. D. Ifenthaler and J. Y.-K. Yau, “Using learning analytics to support study success in higher education: A systematic review,” Educational Technology Research and Development, vol. 68, pp. 1961–1990, 2020, doi: 10.1007/s11423-020-09788-z.
  9. D. Gašević, S. Dawson, and G. Siemens, “Let’s not forget: Learning analytics are about learning,” TechTrends, vol. 59, no. 1, pp. 64–71, 2015, doi: 10.1007/s11528-014-0822-x.
  10. G. Siemens, “Learning analytics: The emergence of a discipline,” American Behavioral Scientist, vol. 57, no. 10, pp. 1380–1400, 2013, doi: 10.1177/0002764213498851.
  11. S. Slade and P. Prinsloo, “Learning analytics: Ethical issues and dilemmas,” American Behavioral Scientist, vol. 57, no. 10, pp. 1510–1529, 2013, doi: 10.1177/0002764213479366.
  12. N. Sclater, “Developing a code of practice for learning analytics,” Journal of Learning Analytics, vol. 3, no. 1, pp. 16–42, 2016, doi: 10.18608/jla.2016.31.3.
  13. UNESCO, Recommendation on the Ethics of Artificial Intelligence. Paris, France: UNESCO, 2021.
  14. S. Kaufman, S. Rosset, C. Perlich, and O. Stitelman, “Leakage in data mining: Formulation, detection, and avoidance,” ACM Transactions on Knowledge Discovery from Data, vol. 6, no. 4, Art. no. 15, 2012, doi: 10.1145/2382577.2382579.
  15. M. Kuhn and K. Johnson, Feature Engineering and Selection: A Practical Approach for Predictive Models. Boca Raton, FL: CRC Press, 2019.
  16. L. Breiman, “Random forests,” Machine Learning, vol. 45, no. 1, pp. 5–32, 2001.
  17. T. Hastie, R. Tibshirani, and J. Friedman, The Elements of Statistical Learning: Data Mining, Inference, and Prediction, 2nd ed. New York, NY: Springer, 2009.
  18. T. Fawcett, “An introduction to ROC analysis,” Pattern Recognition Letters, vol. 27, no. 8, pp. 861–874, 2006, doi: 10.1016/j.patrec.2005.10.010.
  19. T. Kluyver et al., “Jupyter Notebooks – a publishing format for reproducible computational workflows,” in Positioning and Power in Academic Publishing: Players, Agents and Agendas, F. Loizides and B. Schmidt, Eds. IOS Press, 2016, pp. 87–90, doi: 10.3233/978-1-61499-649-1-87.
  20. W. McKinney, “Data structures for statistical computing in Python,” in Proceedings of the 9th Python in Science Conference, 2010, pp. 56–61, doi: 10.25080/Majora-92bf1922-00a.
  21. C. R. Harris, K. J. Millman, S. J. van der Walt, et al., “Array programming with NumPy,” Nature, vol. 585, pp. 357–362, 2020, doi: 10.1038/s41586-020-2649-2.
  22. F. Pedregosa, G. Varoquaux, A. Gramfort, et al., “Scikit-learn: Machine learning in Python,” Journal of Machine Learning Research, vol. 12, pp. 2825–2830, 2011.
  23. J. D. Hunter, “Matplotlib: A 2D graphics environment,” Computing in Science & Engineering, vol. 9, no. 3, pp. 90–95, 2007, doi: 10.1109/MCSE.2007.55.

Educational data can help identify, before the end of a semester, academic trajectories that deserve timely attention. This study develops an academic risk profiling and tailored student guidance using machine learning techniques for universities in Kinshasa. A quantitative, experimental, and predictive design was applied to a harmonized dataset of 9,000 student-semester observations and fourteen predictors after removing a redundant composite engagement index. Data preparation, modelling, and visualisation were performed in a reproducible Python/Jupyter Notebook workflow using pandas, NumPy, scikit-learn, and Matplotlib. A stratified 80/20 training-test split and five-fold stratified cross-validation were used to compare multinomial logistic regression, decision tree, random forest, and multilayer perceptron models. On the independent test set, the multilayer perceptron achieved the highest accuracy (0.720) and macro-AUC-ROC (0.945), while logistic regression achieved the highest balanced accuracy (0.714). Random forest was retained because it achieved the highest macro F1 score (0.686), the prespecified criterion for protecting attention to minority classes, together with macro precision of 0.713 and macro One-vs-Rest AUC-ROC of 0.941. Continuous assessment average, midterm average, and the previous validated-credit rate were the most informative signals. Recommendations are derived from the predicted class, while uncertain cases are flagged for human review. The proposed model is therefore a decision-support tool rather than an automated decision-maker.

Keywords : Machine Learning; Educational Data; Academic Risk Profiling; Random Forest; Tailored Student Guidance.

Paper Submission Last Date
31 - July - 2026

SUBMIT YOUR PAPER CALL FOR PAPERS
Video Explanation for Published paper

Never miss an update from Papermashup

Get notified about the latest tutorials and downloads.

Subscribe by Email

Get alerts directly into your inbox after each post and stay updated.
Subscribe
OR

Subscribe by RSS

Add our RSS to your feedreader to get regular updates from us.
Subscribe