Journal of Data Science logo


Login Register

  1. Home
  2. To appear
  3. Prediction Intervals and Group Variable ...

Journal of Data Science

Submit your article Information
  • Article info
  • Related articles
  • More
    Article info Related articles

Prediction Intervals and Group Variable Importance for Classification Models in University Enrollment Yield
Xiaohui Yin   Yingfa Xie   Jun Yan     All authors (10)

Authors

 
Placeholder
https://doi.org/10.6339/26-JDS1241
Pub. online: 15 July 2026      Type: Statistical Data Science      Open accessOpen Access

Received
2 June 2026
Accepted
1 July 2026
Published
15 July 2026

Abstract

Accurate forecasting of student yield is critical for academic institutions because enrollment projections directly affect resource allocation, course offerings, housing capacity, and budget decisions. Existing enrollment models primarily focus on predicting individual matriculation probabilities, whereas institutional planning depends on accurate prediction of the aggregate enrollment count together with its uncertainty quantification. Here we develop statistical methods for aggregate enrollment prediction and interval estimation under logistic, regularized regression, tree-based, and ensemble classification models. For unpenalized logistic regression, we derive an asymptotic prediction interval and a parametric-simulation interval that propagates coefficient uncertainty. For regularized and tree-based learners, we develop a model-agnostic bootstrap interval that captures selection, estimation, and predictive uncertainty in a unified framework. To interpret predictive structure, we introduce partial and marginal AUC measures to quantify the unique and standalone contributions of feature groups. Simulation studies show that the proposed intervals attain close-to-nominal coverage under well-calibrated and bootstrap-stable learners. Applied to University of Connecticut in-state freshman cohorts from 2020–2025, the proposed intervals contain the observed 2025 enrollment count while identifying the feature groups that contribute most strongly to the enrollment prediction.

Supplementary material

 Supplementary Material
The online supplementary material consists of two files. The file supplement.pdf contains introduction of MoE (Section S.1), the theoretical derivations referenced in the main text (Section S.2), additional simulation results including the MoE DGP robustness study (Table S.1 and Figure S.1 in Section S.3), and expanded variable-importance figures (Figure S.2 in Section S.4). The file code.zip contains the complete R reproduction pipeline, including data generation, model fitting for all seven candidate methods, construction of the asymptotic, parametric-simulation, and bootstrap prediction intervals, and computation of the mAUC and pAUC feature-group importance measures. The archive also includes a synthetic dataset that preserves the structure of the applicant records used in the empirical analysis while protecting individual-level information.

References

 
Ab Ghani NL, Che Cob Z, Mohd Drus S, Sulaiman H (2019). Student enrolment prediction model in higher education institution: A data mining approach. In: Othman MA, Abd Aziz MZA, Md Saat MS, Misran MH (Eds.), Proceedings of the 3rd International Symposium of Information and Internet Technology (SYMINTECH 2018), volume 565 of Lecture Notes in Electrical Engineering, 43–52. Springer International Publishing.
 
Abdar M, Pourpanah F, Hussain S, Rezazadegan D, Liu L, …, Nahavandi S (2021). A review of uncertainty quantification in deep learning: Techniques, applications and challenges. Information Fusion, 76: 243–297. https://doi.org/10.1016/j.inffus.2021.05.008
 
Abelt J, Browning D, Dyer C, Haines M, Ross J, …, Gerber M (2015). Predicting likelihood of enrollment among applicants to the UVa undergraduate program. In: 2015 Systems and Information Engineering Design Symposium, 194–199. IEEE.
 
Aulck L, Nambi D, West J (2020). Increasing enrollment by optimizing scholarship allocations using machine learning and genetic algorithms. In: Rafferty AN, Whitehill J, Cavalli-Sforza V, Romero C (Eds.), Proceedings of the 13th International Conference on Educational Data Mining (EDM 2020), 29–38. ERIC.
 
Breiman L (2001). Random forests. Machine Learning, 45(1): 5–32. https://doi.org/10.1023/A:1010933404324
 
Brinkman PT, McIntyre C (1997). Methods and techniques of enrollment forecasting. New Directions for Institutional Research, 93: 67–80. https://doi.org/10.1002/ir.9305
 
Chatfield C (1993). Calculating interval forecasts. Journal of Business & Economic Statistics, 11(2): 121–135. https://doi.org/10.1080/07350015.1993.10509938
 
Chen T, Guestrin C (2016). XGBoost: A scalable tree boosting system. In: Krishnapuram B, Shah M, Smola AJ, Aggarwal CC, Shen D, Rastogi R (Eds.), Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 785–794.
 
Cirelli J, Konkol AM, Aqlan F, Nwokeji JC (2018). Predictive analytics models for student admission and enrollment. In: Proceedings of the International Conference on Industrial Engineering and Operations Management, 1395–1403.
 
Efron B, Tibshirani R (1997). Improvements on cross-validation: The 632+ bootstrap method. Journal of the American Statistical Association, 92(438): 548–560. https://doi.org/10.2307/2965703
 
Esquivel JA, Esquivel JA (2021). A machine learning based DSS in predicting undergraduate freshmen enrolment in a Philippine university. International Journal of Computer Trends and Technology, 69: 50–54. https://doi.org/10.14445/22312803/IJCTT-V69I5P107
 
Fisher A, Rudin C, Dominici F (2019). All models are wrong, but many are useful: Learning a variable’s importance by studying an entire class of prediction models simultaneously. Journal of Machine Learning Research, 20(177): 1–81.
 
Gawlikowski J, Tassi CRN, Ali M, Lee J, Humt M, …, Zhu XX (2023). A survey of uncertainty in deep neural networks. Artificial Intelligence Review, 56(Suppl 1): 1513–1589. https://doi.org/10.1007/s10462-023-10562-9
 
Gneiting T, Katzfuss M (2014). Probabilistic forecasting. Annual Review of Statistics and Its Application, 1: 125–151. https://doi.org/10.1146/annurev-statistics-062713-085831
 
Goenner CF, Pauls K (2006). A predictive model of inquiry to enrollment. Research in Higher Education, 47: 935–956. https://doi.org/10.1007/s11162-006-9021-8
 
Harrington CF, Schibik T (2013). An econometric approach to optimizing student enrollment. Journal of Higher Education Theory and Practice, 13(1): 45–55.
 
Hayes JB, Price RA, York RP (2013). A simple model for estimating enrollment yield from a list of freshman prospects. Academy of Educational Leadership Journal, 17(2): 61–68.
 
Hoenack SA, Pierro DJ (1990). An econometric model of a public university’s income and enrollments. Journal of Economic Behavior & Organization, 14(3): 403–423. https://doi.org/10.1016/0167-2681(90)90067-N
 
Jacobs RA, Jordan MI, Nowlan SJ, Hinton GE (1991). Adaptive mixtures of local experts. Neural Computation, 3(1): 79–87. https://doi.org/10.1162/neco.1991.3.1.79
 
Jamison J (2017). Applying machine learning to predict Davidson College’s admissions yield. In: Caspersen ME, Edwards SH, Barnes T, Garcia DD (Eds.), Proceedings of the 2017 ACM SIGCSE Technical Symposium on Computer Science Education, 765–766.
 
Janitza S, Hornung R (2018). On the overestimation of random forest’s out-of-bag error. PLoS ONE, 13(8): e0201904. https://doi.org/10.1371/journal.pone.0201904
 
Kye A (2023). Comparative analysis of classification performance for US college enrollment predictive modeling using four machine learning algorithms (logistic regression, decision tree, support vector machine, artificial neural network), Ph.D. thesis, Loyola University Chicago.
 
Langston R, Loreto D (2017). Seamless integration of predictive analytics and CRM within an undergraduate admissions recruitment and marketing plan. Strategic Enrollment Management Quarterly, 4(4): 161–172. https://doi.org/10.1002/sem3.20095
 
Langston R, Wyant R, Scheid J (2016). Strategic enrollment management for chief enrollment officers: Practical use of statistical and mathematical data in forecasting first year and transfer college enrollment. Strategic Enrollment Management Quarterly, 4(2): 74–89. https://doi.org/10.1002/sem3.20085
 
Lei J, G’Sell M, Rinaldo A, Tibshirani RJ, Wasserman L (2018). Distribution-free predictive inference for regression. Journal of the American Statistical Association, 113(523): 1094–1111. https://doi.org/10.1080/01621459.2017.1307116
 
Liu S, Lu D, Painter SL, Griffiths NA, Pierce EM (2023). Uncertainty quantification of machine learning models to improve streamflow prediction under changing climate and environmental conditions. Frontiers in Water, 5: 1150126. https://doi.org/10.3389/frwa.2023.1150126
 
Lopez LJL, Elsharief S, Jorf DA, Darwish F, Ma C, Shamout FE (2025). Uncertainty quantification for machine learning in healthcare: A survey. In: Xu XO, Choi E, Singhal P, Gerych W, Tang S, Agrawal M, Subbaswamy A, Sizikova E, Dunn J, Daneshjou R, Sarker T, McDermott M, Chen I (Eds.), Proceedings of the Sixth Conference on Health, Inference, and Learning, 862–907.
 
Mnich K, Kitlas Golińska A, Polewko-Klim A, Rudnicki WR (2020). Bootstrap bias corrected cross validation applied to Super Learning. ArXiv preprint: arXiv:2003.08342.
 
Pencina MJ, D’Agostino RB, Vasan RS (2008). Evaluating the added predictive ability of a new marker: From area under the ROC curve to reclassification and beyond. Statistics in Medicine, 27(2): 157–172. https://doi.org/10.1002/sim.2929
 
Pepe MS, Janes H, Longton G, Leisenring W, Newcomb P (2004). Limitations of the odds ratio in gauging the performance of a diagnostic, prognostic, or screening marker. American Journal of Epidemiology, 159(9): 882–890. https://doi.org/10.1093/aje/kwh101
 
Platt J (1999). Probabilistic outputs for support vector machines and comparisons to regularized likelihood methods. In: Smola A.J., Bartlett P.L., Schölkopf B., Schuurmans D. (Eds.), Advances in Large Margin Classifiers, 61–74.
 
Prokhorenkova L, Gusev G, Vorobev A, Dorogush AV, Gulin A (2018). Catboost: Unbiased boosting with categorical features. In: Bengio S, Wallach HM, Larochelle H, Grauman K, Cesa-Bianchi N, Garnett R (Eds.), Advances in Neural Information Processing Systems, volume 31.
 
Tang Z, Westling T (2024). Consistency of the bootstrap for asymptotically linear estimators based on machine learning.
 
Trusheim D, Rylee C (2011). Predictive modeling: Linking enrollment and budgeting. Planning for Higher Education, 40(1): 12–21.
 
Van der Laan MJ, Polley EC, Hubbard AE (2007). Super learner. Statistical Applications in Genetics and Molecular Biology, 6(1): 1–23.
 
Van der Vaart AW Dudoit S, Van der Laan MJ (2006). Oracle inequalities for multi-fold cross validation. Statistics & Decisions, 24(3): 351–371. https://doi.org/10.1524/stnd.2006.24.3.351
 
Wright MN, Ziegler A (2017). Ranger: A fast implementation of random forests for high dimensional data in C++ and R. Journal of Statistical Software, 77(1): 1–17. https://doi.org/10.18637/jss.v077.i01
 
Wyner AJ, Olson M, Bleich J, Mease D (2017). Explaining the success of AdaBoost and random forests as interpolating classifiers. Journal of Machine Learning Research, 18(48): 1–33.
 
Zadrozny B, Elkan C (2002). Transforming classifier scores into accurate multiclass probability estimates. In: Grossman RL, Han J, Kumar V, Motwani R, Agrawal R (Eds.), Proceedings of the 8th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 694–699.
 
Zou H, Hastie T (2005). Regularization and variable selection via the elastic net. Journal of the Royal Statistical Society, Series B, Statistical Methodology, 67(2): 301–320. https://doi.org/10.1111/j.1467-9868.2005.00503.x

Related articles PDF XML
Related articles PDF XML

Copyright
2026 The Author(s). Published by the School of Statistics and the Center for Applied Statistics, Renmin University of China.
by logo by logo
Open access article under the CC BY license.

Keywords
classification enrollment prediction prediction interval uncertainty quantification variable importance

Metrics
since February 2021
59

Article info
views

28

PDF
downloads

Export citation

Copy and paste formatted citation
Placeholder

Download citation in file


Share


RSS

Journal of data science

  • Online ISSN: 1683-8602
  • Print ISSN: 1680-743X

About

  • About journal
  • Renmin University of China homepage
  • Academic Journal Management
    and Development Center homepage

For contributors

  • Submit
  • OA Policy
  • Become a Peer-reviewer

Contact us

  • JDS@ruc.edu.cn
  • Contact person: Jing Zhou
  • Phone: +86-10-62511318
  • No. 59 Zhongguancun Street, Haidian District Beijing, 100872, P.R. China
Powered by PubliMill  •  Privacy policy