Pub. online:21 May 2026Type:Statistical Data ScienceOpen Access
Journal:Journal of Data Science
Volume 24, Issue 3 (2026): Special Issue: 2025 GASP Conference, pp. 482–503
Abstract
Statistical survey metadata contains essential contextual information that underpins the accurate interpretation, discovery, and reuse of statistical data. However, traditional metadata formats are not optimized for consumption by large language models (LLMs), which increasingly function as interfaces for data exploration, question-answering, and decision support. This work introduces a knowledge graph-based approach to modeling survey metadata using semantic web standards and linked data principles, specifically designed to make metadata machine-understandable and LLM-compatible. The core metadata entities, including surveys, datasets, variables, concepts, populations, and provenance, are modeled as rich interlinked nodes that allow reasoning, contextual enrichment, and structured prompting. The graph integrates established ontologies such as the Resource Description Framework (RDF) to promote interoperability and alignment with global standards. We demonstrate how this structure allows LLMs to surface relevant metadata, ground their outputs in authoritative sources, and generate semantically precise responses. This approach enhances transparency, facilitates metadata reuse, and supports the development of artificial intelligence (AI) applications powered by statistical products.
Pub. online:9 Jun 2026Type:Statistical Data ScienceOpen Access
Journal:Journal of Data Science
Volume 24, Issue 3 (2026): Special Issue: 2025 GASP Conference, pp. 504–522
Abstract
Privacy-preserving machine learning methods seek to train useful models that do not disclose information about the data on which they were trained. Such methods are vital when organizations train neural networks on sensitive individual-level data and seek to release the models publicly. Their goal poses a trade-off between predictive performance (utility) and privacy protection. That trade-off makes privacy-preserving machine learning methods difficult to apply in practice, usually requiring extensive iteration and hyperparameter tuning. Yet, practitioners often have little guidance for navigating competing statistical, computational, and privacy demands. We present an implementation algorithm for the Stochastic Weight Averaging–Gaussian Pseudo Posterior Mechanism (SWAG-PPM), a Bayesian differentially private deep learning method. The implementation algorithm focuses on the joint tuning of two key hyperparameters whose interaction governs model convergence and the privacy–utility trade-off. We introduce novel diagnostic tools to evaluate convergence and guide hyperparameter adjustments. Using a transformer model for occupational injury classification, we demonstrate that diagnostic-guided tuning with SWAG-PPM can achieve strong privacy protection and utility. While our case study uses a specific dataset and model architecture, all methodological steps can apply to other settings where privacy risk is heterogeneously distributed.
Pub. online:15 Jul 2026Type:Statistical Data ScienceOpen Access
Journal:Journal of Data Science
Volume 24, Issue 3 (2026): Special Issue: 2025 GASP Conference, pp. 523–543
Abstract
Accurate forecasting of student yield is critical for academic institutions because enrollment projections directly affect resource allocation, course offerings, housing capacity, and budget decisions. Existing enrollment models primarily focus on predicting individual matriculation probabilities, whereas institutional planning depends on accurate prediction of the aggregate enrollment count together with its uncertainty quantification. Here we develop statistical methods for aggregate enrollment prediction and interval estimation under logistic, regularized regression, tree-based, and ensemble classification models. For unpenalized logistic regression, we derive an asymptotic prediction interval and a parametric-simulation interval that propagates coefficient uncertainty. For regularized and tree-based learners, we develop a model-agnostic bootstrap interval that captures selection, estimation, and predictive uncertainty in a unified framework. To interpret predictive structure, we introduce partial and marginal AUC measures to quantify the unique and standalone contributions of feature groups. Simulation studies show that the proposed intervals attain close-to-nominal coverage under well-calibrated and bootstrap-stable learners. Applied to University of Connecticut in-state freshman cohorts from 2020–2025, the proposed intervals contain the observed 2025 enrollment count while identifying the feature groups that contribute most strongly to the enrollment prediction.
Pub. online:8 Jun 2026Type:Statistical Data ScienceOpen Access
Journal:Journal of Data Science
Volume 24, Issue 3 (2026): Special Issue: 2025 GASP Conference, pp. 544–563
Abstract
Survey researchers are increasingly adopting hybrid sampling designs to address the limitations of traditional probability sampling, especially when studying rare or hard-to-reach populations. Challenges such as high screening costs, low statistical efficiency, and operational constraints make purely probability-based approaches impractical in many contexts. This article uses public data from the National Health and Nutrition Examination Survey to demonstrate how one can make population estimates from a hybrid sampling strategy that combines data from a stratified, multistage probability sample with data from a non-probability sample within the same primary sampling units as the probability sample. We outline a framework and discuss methods for analyzing data from a hybrid sample such as this, where covariates and survey outcomes are observed in both the probability and non-probability samples. We present a case study to illustrate the framework. We provide the case study R code in the supplementary material.
Pub. online:21 Jul 2026Type:Statistical Data ScienceOpen Access
Journal:Journal of Data Science
Volume 24, Issue 3 (2026): Special Issue: 2025 GASP Conference, pp. 564–583
Abstract
Parameter estimation and inference from complex survey samples typically focuses on global model parameters whose estimators have asymptotic properties, such as from fixed effects regression models. The central challenge is to both mitigate bias induced from potentially unbalanced samples and to incorporate adjustments for differences in effective sample size to get correct variance and interval estimates. We present a motivating example of Bayesian inference for a multi-level or mixed effects model in which estimates of both the local parameters (e.g. group level random effects) and the global parameters need to be adjusted for the complex sampling design. We evaluate the limitations of the survey-weighted pseudo-posterior and an existing automated post-processing method to improve the uncertainty quantification. We propose modifications to the automated process and demonstrate their improvements for multi-level models via a simulation study and a motivating example from the National Survey on Drug Use and Health. Reproduction examples are included in the supplementary material and the updated R package is available via github: https://github.com/RyanHornby/csSampling
Pub. online:8 Jun 2026Type:Data Science In ActionOpen Access
Journal:Journal of Data Science
Volume 24, Issue 3 (2026): Special Issue: 2025 GASP Conference, pp. 584–599
Abstract
The United States Department of Agriculture’s (USDA’s) National Agricultural Statistics Service (NASS) conducted a pilot study in 2024 to obtain data collected onboard farm machinery and explore their uses for statistical purposes. NASS has recognized high value in these machine-logged data (MLD) systems as they can potentially augment, or even replace, traditional survey efforts while providing additional benefits of reducing respondent burden and improving crop-related estimates. This pilot study ultimately addressed four topics: 1) understanding the obstacles in obtaining MLD from farmers; 2) creating geographic workflows to manage inherent geospatial MLD; 3) developing the linkages to NASS’s tabular list frame information; and 4) assessing the use of MLD to replace survey data for time-sensitive estimates. To study each topic, field-level information was gathered from the MLD systems of dozens of producers over hundreds of fields across the central United States (US) for the 2023 growing season. Results showed that 90% of the fields could be linked to a producer on the NASS list frame. Of those producers, the consistency of MLD versus traditional survey reporting was highly variable for those who were selected for a survey in 2023. Comparisons showed median MLD values were larger than historical NASS survey values. Approximately 48% of survey comparisons showed a difference of 25% or less between MLD and historical NASS survey values. MLD shows promise for use in official statistics; however, further analyses with additional producers’ data and enhancements to MLD collection processes are needed before supplementing traditional survey methods.
Pub. online:10 Jul 2026Type:Data Science In ActionOpen Access
Journal:Journal of Data Science
Volume 24, Issue 3 (2026): Special Issue: 2025 GASP Conference, pp. 600–615
Abstract
High quality record linkages are critical for enriching survey data with alternative data sources. To enhance the American Community Survey (ACS) (US Census Bureau, 2025), we evaluated several options to match commercially available property data to housing unit records in the ACS and the Census Master Address File (MAF) (US Census Bureau, 2022) data, with the ultimate goal of supplementing data collected in the survey with the commercial property data. Techniques to match address data come in two flavors: spatial matching and address matching. Spatial matching is done by overlaying commercial boundary shape files on the lat-long coordinates on the MAF to associate Census housing unit records to commercial property parcel records. This method is useful because it does not require matching of text fields, but performs poorly when parcels include many housing units (e.g., large apartment buildings). Address matching, or entity resolution at the address level, links records across the two sources based on the content of various address fields, and offers the possibility of disambiguating multiple matches and matching objects that cannot be successfully assigned a unique match through spatial matching. Several approaches are available for address matching, including rule-based, deterministic, fuzzy, and probabilistic methods. In our article we summarize literature comparing various combinations of spatial and address matching techniques to illustrate the trade-off between linkage rates and linkage quality. We also consider hybrid solutions that leverage spatial matching as well as several types of address matching to maximize high quality linkages for our research and discuss future directions such as incorporating probabilistic matching as an added step.
Pub. online:17 Aug 2026Type:Data Science In ActionOpen Access
Journal:Journal of Data Science
Volume 24, Issue 3 (2026): Special Issue: 2025 GASP Conference, pp. 616–629
Abstract
In the evolving field of survey research, leveraging machine learning to predict response behavior has transformative potential for the efficiency of survey operations. Integrating multiple data sources may improve response prediction by providing more nuanced insights for household outreach. This study presents a model-driven approach to enhancing respondent cooperation in the Medical Expenditure Panel Survey (MEPS) by combining features from disparate data sources. MEPS is a longitudinal household survey with 5 rounds of interviewing over 2.5 years. Its sample is derived prior National Health Interview Survey (NHIS) participants. MEPS Round 1 response rates are critical for sustaining representativeness throughout each panel. In this study, we constructed a multimodal machine learning model to predict (1) the likelihood of a positive response for an upcoming contact attempt and (2) the likelihood that new panel households complete a Round 1 interview. The model integrates tract-level data from the American Community Survey (ACS), outcomes from the Advance Call Records (ACR) made prior to MEPS Round 1, and paradata from the early contact period. We also explored the relative contributions of these sources to model performance. Our model aims to help manage field labor by identifying complex cases needing specialized support. It can also assist in determining the optimal mode for the next contact to increase the chance of a completed interview. Beyond improving MEPS operations, this study offers a roadmap for incorporating additional data sources to support fieldwork.