Authors :
Patrick Ufomba Nwogu; Chigozirim Ajaegbu; Dr. Faruk Umar Ambursa; Dr. Femi Adeluyi
Volume/Issue :
Volume 11 - 2026, Issue 9 - September
Google Scholar :
https://tinyurl.com/3rc4ahkp
DOI :
https://doi.org/10.38124/ijisrt/26sep377
Note : A published paper may take 4-5
working days from the publication date to appear in PlumX Metrics, Semantic Scholar, and
ResearchGate.
Abstract :
Dyslexia is a specific learning difficulty associated with persistent problems in accurate and fluent word
recognition, decoding and spelling. Early identification can support timely educational intervention; however, computational
screening models must address high-dimensional behavioural data, class imbalance and sensitivity to model configuration.
This study investigates the optimization of dyslexia prediction through hyperparameter-tuned ensemble machine learning
using the Dyt-desktop behavioural dataset derived from the Rello et al. dyslexia-screening study.
Keywords :
Dyslexia Prediction; Machine Learning; Hyperparameter Optimization; Random Forest; Xgboost; Extra Trees; RFRFE; Ensemble Learning; Stacking; Behavioural Data; Educational Data Mining.
References :
- Rello, L., Baeza-Yates, R., Ali, A., Bigham, J. P., & Serra, M. (2020). Predicting risk of dyslexia with an online gamified test. PLOS ONE, 15(12), e0241687.
- Breiman, L. (2001). Random forests. Machine Learning, 45, 5–32.
- Chen, T., & Guestrin, C. (2016). XGBoost: A scalable tree boosting system. Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 785–794.
- Geurts, P., Ernst, D., & Wehenkel, L. (2006). Extremely randomized trees. Machine Learning, 63, 3–42.
- Wolpert, D. H. (1992). Stacked generalization. Neural Networks, 5(2), 241–259.
- Bergstra, J., & Bengio, Y. (2012). Random search for hyper-parameter optimization. Journal of Machine Learning Research, 13, 281–305.
- Saito, T., & Rehmsmeier, M. (2015). The precision-recall plot is more informative than the ROC plot when evaluating binary classifiers on imbalanced datasets. PLOS ONE, 10(3), e0118432.
- Chicco, D., & Jurman, G. (2020). The advantages of the Matthews correlation coefficient (MCC) over F1 score and accuracy in binary classification evaluation. BM C Genomics, 21, 6.
- Pedregosa, F., Varoquaux, G., Gramfort, A., et al. (2011). Scikit-learn: Machine learning in Python. Journal of Machine Learning Research, 12, 2825–2830.
- Lundberg, S. M., & Lee, S.-I. (2017). A unified approach to interpreting model predictions. Advances in Neural Information Processing Systems, 30.
Dyslexia is a specific learning difficulty associated with persistent problems in accurate and fluent word
recognition, decoding and spelling. Early identification can support timely educational intervention; however, computational
screening models must address high-dimensional behavioural data, class imbalance and sensitivity to model configuration.
This study investigates the optimization of dyslexia prediction through hyperparameter-tuned ensemble machine learning
using the Dyt-desktop behavioural dataset derived from the Rello et al. dyslexia-screening study.
Keywords :
Dyslexia Prediction; Machine Learning; Hyperparameter Optimization; Random Forest; Xgboost; Extra Trees; RFRFE; Ensemble Learning; Stacking; Behavioural Data; Educational Data Mining.