QA711 : Predicting Students’ Academic Performance Using Data Science and Machine Learning Approaches
Thesis > Central Library of Shahrood University > Computer Engineering > MSc > 2026
Authors:
Abstarct: This study aimed to design and validate a localized frxamework for analyzing educational data and predicting students’ academic performance (pass/fail) using data science and machine learning algorithms. The dataset consisted of 708 student records, reduced to 10 main features after preprocessing. The class distribution exhibited severe imbalance (95.76% pass, 4.24% fail), which was addressed by applying the Adaptive Synthetic Sampling (ADASYN) technique on the training set. Six models—Random Forest, Support Vector Machine, K-Nearest Neighbors, Artificial Neural Network, XGBoost, and an Ensemble model (soft voting)—were trained with strong regularization to prevent overfitting. Results showed that the Ensemble model achieved the best overall balance with 95.31% accuracy, 97.58% F1-score, and 99.02% recall, while SVM demonstrated the highest discriminative power with an AUC-ROC of 0.8034. Permutation Importance analysis revealed that weekly study hours (importance 0.016432), past exam scores (0.010329), and attendance rate (0.007042) were the top three predictors. The findings confirmed the research hypotheses and demonstrated that combining behavioral and academic features significantly improves prediction accuracy. This frxamework can serve as a foundation for early warning systems and targeted educational interventions in the Iraqi education system.
Keywords:
#Academic performance prediction #Machine learning #Class imbalance #Permutation importance #Ensemble model #ADASYN #Iraqi education system Keeping place: Central Library of Shahrood University
Visitor:
Visitor: