The conjugacy decision problem is fundamental in combinatorial group theory and has applications in algebraic cryptography. Previous machine-learning studies demonstrated that conjugacy can be classified in selected non-free groups, but a statistical prediction is not an algebraic proof. This paper introduces a reproducible certificate-guided learning pipeline for random two-generator one-relator presentations satisfying the metric condition $C'(1/6)$. Positive instances are constructed with an explicit conjugator and accepted only after Dehn reduction verifies $g^{-1}ug =_G v$. Negative instances are retained only when the images of $u$ and $v$ in the abelianization are unequal, which provides an exact obstruction to conjugacy. A symmetry-preserving feature map combines pairwise word-length statistics, exponent-sum information, and aggregated labelled-walk features of orders one to three. Logistic regression and a random forest are evaluated using a presentation-disjoint split. The principal experiment contains $4{,}200$ certified pairs from six presentations: $2{,}800$ pairs from four presentations are used for training, while $1{,}400$ pairs from two previously unseen presentations are used for testing. Logistic regression obtains an accuracy of $0.9964$ and an F1 score of $0.9964$, while the random forest obtains $1.0000$ for accuracy, F1 score, and ROC–AUC. An independent experiment using a second random seed reproduces this performance pattern. These high scores characterize the certified pilot distribution and do not constitute a solution to unrestricted conjugacy. The principal contribution is a transparent computational framework that separates statistical prediction from algebraic proof, prevents the use of unverifiable class labels, and identifies the construction of difficult certified negative instances as the central problem for subsequent research.
Michael Nsikan JohnKnowledgeTrend Preprints · 2026
Accreditation; Quality assurance; Machine learning; Higher education; Early warning; Nigeria
Programme accreditation by the National Universities Commission (NUC) is the main external quality assurance mechanism in Nigerian universities, but it is periodic and retrospective: weaknesses are usually discovered during the visit rather than before it. This study aimed to develop and evaluate a machine-learning (ML) framework that predicts accreditation outcomes from routinely collected pre-visit indicators so that institutions can remediate deficiencies early. Because programme-level accreditation records are not publicly available, we built a reproducible synthetic panel of 4,321 accreditation visits to 1,515 programmes in 149 simulated universities (2014–2024), with outcomes (full, interim, denied) generated by NUC-style scoring rules. Fifteen indicators covering staffing, curriculum, facilities, library, funding, research and employer rating, together with ownership, discipline and prior status, were used as predictors. Logistic regression (LR), support vector machine, multilayer perceptron, random forest and extreme gradient boosting (XGBoost) were trained on 2014–2021 visits with university-grouped cross-validation and tested on 2022–2024 visits. LR performed best on the temporal test set (accuracy 0.773; macro-F1 0.708, 95% CI 0.669–0.744; area under the curve [AUC] 0.908), exceeding a prior-status rule (macro-F1 0.537). For the binary task of identifying programmes at risk of not receiving full accreditation, LR achieved an AUC of 0.900 and good calibration; the 20% of programmes with the highest predicted risk included 52 of the 53 denied programmes. Proportion of PhD-holding staff and laboratory provision were the most influential predictors. ML-based risk scoring could support continuous, pre-emptive quality monitoring, but validation on real NUC data is required before operational use.
Osemengbe Oyaimare Uddin, Susan Konyeha, Glory Nosa EdegbeAfrican Journal of Mathematics, Statistics and Computer Science · 2026
Missing values in rainfall datasets reduce the reliability of statistical inference and undermine decision-making in agriculture, climate studies, and environmental management. This study evaluated the performance of five imputation techniques; Expectation-Maximization (EM), Multiple Imputation (MI), Regression Imputation (RI), Bootstrap Expectation-Maximization (BEM), and Random Forest (RF) using the 2019 Nigerian rainfall dataset obtained from the National Bureau of Statistics. The methods were compared using Raw Bias, Mean Squared Error (MSE), Root Mean Squared Error (RMSE), and Variance to assess estimation accuracy, predictive performance, and stability. The results revealed that MI produced the lowest bias (-0.003), making it the most suitable method for minimizing systematic estimation error. In contrast, RF achieved the highest predictive accuracy, recording the lowest MSE (0.9530) and RMSE (0.9762). Although BEM exhibited the lowest variance (0.9844), indicating greater stability, it was associated with relatively high bias, limiting its overall effectiveness. The findings demonstrate that no single method is universally optimal; rather, the choice of imputation technique should be guided by the primary analytical objective. RF is recommended for applications requiring high predictive accuracy, whereas MI is preferable when unbiased parameter estimation is essential. The study provides empirical evidence to support the adoption of robust imputation techniques by agencies such as the Nigerian Meteorological Agency (NiMet), thereby improving the quality of national climate databases and strengthening evidence-based agricultural and environmental decision-making.
Taiwo Wale Ayanniyi, Efosa Michael OgbeideKtrend - International Journal of Mathematics and Statistics (IJMS) · 2026