Prediction Intervals and Group Variable Importance for Classification Models in University Enrollment Yield
Pub. online: 15 July 2026
Type: Statistical Data Science
Open Access
Received
2 June 2026
2 June 2026
Accepted
1 July 2026
1 July 2026
Published
15 July 2026
15 July 2026
Abstract
Accurate forecasting of student yield is critical for academic institutions because enrollment projections directly affect resource allocation, course offerings, housing capacity, and budget decisions. Existing enrollment models primarily focus on predicting individual matriculation probabilities, whereas institutional planning depends on accurate prediction of the aggregate enrollment count together with its uncertainty quantification. Here we develop statistical methods for aggregate enrollment prediction and interval estimation under logistic, regularized regression, tree-based, and ensemble classification models. For unpenalized logistic regression, we derive an asymptotic prediction interval and a parametric-simulation interval that propagates coefficient uncertainty. For regularized and tree-based learners, we develop a model-agnostic bootstrap interval that captures selection, estimation, and predictive uncertainty in a unified framework. To interpret predictive structure, we introduce partial and marginal AUC measures to quantify the unique and standalone contributions of feature groups. Simulation studies show that the proposed intervals attain close-to-nominal coverage under well-calibrated and bootstrap-stable learners. Applied to University of Connecticut in-state freshman cohorts from 2020–2025, the proposed intervals contain the observed 2025 enrollment count while identifying the feature groups that contribute most strongly to the enrollment prediction.
Supplementary material
Supplementary MaterialThe online supplementary material consists of two files. The file supplement.pdf contains introduction of MoE (Section S.1), the theoretical derivations referenced in the main text (Section S.2), additional simulation results including the MoE DGP robustness study (Table S.1 and Figure S.1 in Section S.3), and expanded variable-importance figures (Figure S.2 in Section S.4). The file code.zip contains the complete R reproduction pipeline, including data generation, model fitting for all seven candidate methods, construction of the asymptotic, parametric-simulation, and bootstrap prediction intervals, and computation of the mAUC and pAUC feature-group importance measures. The archive also includes a synthetic dataset that preserves the structure of the applicant records used in the empirical analysis while protecting individual-level information.
References
Ab Ghani NL, Che Cob Z, Mohd Drus S, Sulaiman H (2019). Student enrolment prediction model in higher education institution: A data mining approach. In: Othman MA, Abd Aziz MZA, Md Saat MS, Misran MH (Eds.), Proceedings of the 3rd International Symposium of Information and Internet Technology (SYMINTECH 2018), volume 565 of Lecture Notes in Electrical Engineering, 43–52. Springer International Publishing.
Abdar M, Pourpanah F, Hussain S, Rezazadegan D, Liu L, …, Nahavandi S (2021). A review of uncertainty quantification in deep learning: Techniques, applications and challenges. Information Fusion, 76: 243–297. https://doi.org/10.1016/j.inffus.2021.05.008
Aulck L, Nambi D, West J (2020). Increasing enrollment by optimizing scholarship allocations using machine learning and genetic algorithms. In: Rafferty AN, Whitehill J, Cavalli-Sforza V, Romero C (Eds.), Proceedings of the 13th International Conference on Educational Data Mining (EDM 2020), 29–38. ERIC.
Breiman L (2001). Random forests. Machine Learning, 45(1): 5–32. https://doi.org/10.1023/A:1010933404324
Brinkman PT, McIntyre C (1997). Methods and techniques of enrollment forecasting. New Directions for Institutional Research, 93: 67–80. https://doi.org/10.1002/ir.9305
Chatfield C (1993). Calculating interval forecasts. Journal of Business & Economic Statistics, 11(2): 121–135. https://doi.org/10.1080/07350015.1993.10509938
Efron B, Tibshirani R (1997). Improvements on cross-validation: The 632+ bootstrap method. Journal of the American Statistical Association, 92(438): 548–560. https://doi.org/10.2307/2965703
Esquivel JA, Esquivel JA (2021). A machine learning based DSS in predicting undergraduate freshmen enrolment in a Philippine university. International Journal of Computer Trends and Technology, 69: 50–54. https://doi.org/10.14445/22312803/IJCTT-V69I5P107
Gawlikowski J, Tassi CRN, Ali M, Lee J, Humt M, …, Zhu XX (2023). A survey of uncertainty in deep neural networks. Artificial Intelligence Review, 56(Suppl 1): 1513–1589. https://doi.org/10.1007/s10462-023-10562-9
Gneiting T, Katzfuss M (2014). Probabilistic forecasting. Annual Review of Statistics and Its Application, 1: 125–151. https://doi.org/10.1146/annurev-statistics-062713-085831
Goenner CF, Pauls K (2006). A predictive model of inquiry to enrollment. Research in Higher Education, 47: 935–956. https://doi.org/10.1007/s11162-006-9021-8
Hoenack SA, Pierro DJ (1990). An econometric model of a public university’s income and enrollments. Journal of Economic Behavior & Organization, 14(3): 403–423. https://doi.org/10.1016/0167-2681(90)90067-N
Jacobs RA, Jordan MI, Nowlan SJ, Hinton GE (1991). Adaptive mixtures of local experts. Neural Computation, 3(1): 79–87. https://doi.org/10.1162/neco.1991.3.1.79
Janitza S, Hornung R (2018). On the overestimation of random forest’s out-of-bag error. PLoS ONE, 13(8): e0201904. https://doi.org/10.1371/journal.pone.0201904
Langston R, Loreto D (2017). Seamless integration of predictive analytics and CRM within an undergraduate admissions recruitment and marketing plan. Strategic Enrollment Management Quarterly, 4(4): 161–172. https://doi.org/10.1002/sem3.20095
Langston R, Wyant R, Scheid J (2016). Strategic enrollment management for chief enrollment officers: Practical use of statistical and mathematical data in forecasting first year and transfer college enrollment. Strategic Enrollment Management Quarterly, 4(2): 74–89. https://doi.org/10.1002/sem3.20085
Lei J, G’Sell M, Rinaldo A, Tibshirani RJ, Wasserman L (2018). Distribution-free predictive inference for regression. Journal of the American Statistical Association, 113(523): 1094–1111. https://doi.org/10.1080/01621459.2017.1307116
Liu S, Lu D, Painter SL, Griffiths NA, Pierce EM (2023). Uncertainty quantification of machine learning models to improve streamflow prediction under changing climate and environmental conditions. Frontiers in Water, 5: 1150126. https://doi.org/10.3389/frwa.2023.1150126
Lopez LJL, Elsharief S, Jorf DA, Darwish F, Ma C, Shamout FE (2025). Uncertainty quantification for machine learning in healthcare: A survey. In: Xu XO, Choi E, Singhal P, Gerych W, Tang S, Agrawal M, Subbaswamy A, Sizikova E, Dunn J, Daneshjou R, Sarker T, McDermott M, Chen I (Eds.), Proceedings of the Sixth Conference on Health, Inference, and Learning, 862–907.
Mnich K, Kitlas Golińska A, Polewko-Klim A, Rudnicki WR (2020). Bootstrap bias corrected cross validation applied to Super Learning. ArXiv preprint: arXiv:2003.08342.
Pencina MJ, D’Agostino RB, Vasan RS (2008). Evaluating the added predictive ability of a new marker: From area under the ROC curve to reclassification and beyond. Statistics in Medicine, 27(2): 157–172. https://doi.org/10.1002/sim.2929
Pepe MS, Janes H, Longton G, Leisenring W, Newcomb P (2004). Limitations of the odds ratio in gauging the performance of a diagnostic, prognostic, or screening marker. American Journal of Epidemiology, 159(9): 882–890. https://doi.org/10.1093/aje/kwh101
Van der Vaart AW Dudoit S, Van der Laan MJ (2006). Oracle inequalities for multi-fold cross validation. Statistics & Decisions, 24(3): 351–371. https://doi.org/10.1524/stnd.2006.24.3.351
Wright MN, Ziegler A (2017). Ranger: A fast implementation of random forests for high dimensional data in C++ and R. Journal of Statistical Software, 77(1): 1–17. https://doi.org/10.18637/jss.v077.i01
Zou H, Hastie T (2005). Regularization and variable selection via the elastic net. Journal of the Royal Statistical Society, Series B, Statistical Methodology, 67(2): 301–320. https://doi.org/10.1111/j.1467-9868.2005.00503.x