Pub. online:15 Jul 2026Type:Statistical Data ScienceOpen Access
Journal:Journal of Data Science
Volume 24, Issue 3 (2026): Special Issue: 2025 GASP Conference, pp. 523–543
Abstract
Accurate forecasting of student yield is critical for academic institutions because enrollment projections directly affect resource allocation, course offerings, housing capacity, and budget decisions. Existing enrollment models primarily focus on predicting individual matriculation probabilities, whereas institutional planning depends on accurate prediction of the aggregate enrollment count together with its uncertainty quantification. Here we develop statistical methods for aggregate enrollment prediction and interval estimation under logistic, regularized regression, tree-based, and ensemble classification models. For unpenalized logistic regression, we derive an asymptotic prediction interval and a parametric-simulation interval that propagates coefficient uncertainty. For regularized and tree-based learners, we develop a model-agnostic bootstrap interval that captures selection, estimation, and predictive uncertainty in a unified framework. To interpret predictive structure, we introduce partial and marginal AUC measures to quantify the unique and standalone contributions of feature groups. Simulation studies show that the proposed intervals attain close-to-nominal coverage under well-calibrated and bootstrap-stable learners. Applied to University of Connecticut in-state freshman cohorts from 2020–2025, the proposed intervals contain the observed 2025 enrollment count while identifying the feature groups that contribute most strongly to the enrollment prediction.
Historical data or real-world data are often available in clinical trials, genetics, health care, psychology, environmental health, engineering, economics, and business. The power priors have emerged as a useful class of informative priors for a variety of situations in which historical data are available. In this paper, an overview of the development of the power priors is provided. Various variations of the power priors are derived under a binomial regression model and a normal linear regression model. The development of software on the power priors is also briefly reviewed. Throughout this paper, the data from the Kociba study and the National Toxicology Program study as well as the data from the Alzheimer’s Disease Neuroimaging Initiative (ADNI) study are used to demonstrate the derivations of the power priors and their variations. Detailed analyses of the data from these studies are carried out to further demonstrate the usefulness of the power priors and their variations in these real applications. Finally, the directions of future research on the power priors are discussed.
The complexity of energy infrastructure at large institutions increasingly calls for data-driven monitoring of energy usage. This article presents a hybrid monitoring algorithm for detecting consumption surges using statistical hypothesis testing, leveraging the posterior distribution and its information about uncertainty to introduce randomness in the parameter estimates, while retaining the frequentist testing framework. This hybrid approach is designed to be asymptotically equivalent to the Neyman-Pearson test. We show via extensive simulation studies that the hybrid approach enjoys control over type-1 error rate even with finite sample sizes whereas the naive plug-in method tends to exceed the specified level, resulting in overpowered tests. The proposed method is applied to the natural gas usage data at the University of Connecticut.