Density deconvolution is a fundamental problem in measurement error settings, arising when the goal is to recover the distribution of unobserved latent variables based on their error-contaminated proxies. An important application arises in nutritional epidemiology, where the objective is to estimate the long-term average intake of nutritional components using 24-hour dietary recall data. This task involves complex multivariate models that account for conditionally heteroscedastic measurement errors, zero-inflated proxies for episodically consumed dietary components, diverse marginal distribution shapes across components, etc. While flexible Bayesian hierarchical methods, coupled with Markov chain Monte Carlo techniques, have been increasingly successful in recent years to enable deconvolution under such intricate but realistic scenarios, the development of rapid and user-friendly software implementations has lagged behind. The $\mathsf{R}$ package BayesDecon has been developed to address this challenge, providing a fast and accessible implementation of flexible Bayesian deconvolution methods for practitioners working with measurement errors. In the process, we have introduced substantial improvements to several aspects of the original models and algorithms, and have also added options to fit simpler parametric models. Although illustrated on challenges in nutritional epidemiology, the package addresses general deconvolution problems and hence is broadly applicable to other domains as well.
Precision medicine is an innovative approach that aims to customize medical treatments and interventions to patients based on their individual characteristics. Several estimation techniques, including Q-learning, have been developed to determine optimal treatment rules. However, the applicability of these methods depends on the availability of precisely measured variables. This study extends the scope of Q-learning to incorporate compound outcomes, deviating from the commonly assumed univariate outcomes, and further accommodates data with mismeasurement in both binary and continuous covariates. Two methods are described to mitigate the impact of mismeasurement. Numerical studies reveal that mismeasurement in covariates leads to notable estimation bias in parameters indexing the optimal treatment, yet the methods addressing the mismeasured effects yield improved results.
Pub. online:11 Jun 2025Type:Statistical Data ScienceOpen Access
Journal:Journal of Data Science
Volume 23, Issue 3 (2025): Special Issue: 2024 WNAR/IMS/Graybill Annual Meeting, pp. 499–520
Abstract
The rapidly expanding field of metabolomics presents an invaluable resource for understanding the associations between metabolites and various diseases. However, the high dimensionality, presence of missing values, and measurement errors associated with metabolomics data can present challenges in developing reliable and reproducible approaches for disease association studies. Therefore, there is a compelling need for robust statistical analyses that can navigate these complexities to achieve reliable and reproducible disease association studies. In this paper, we construct algorithms to perform variable selection for noisy data and control the False Discovery Rate when selecting mutual metabolomic predictors for multiple disease outcomes. We illustrate the versatility and performance of this procedure in a variety of scenarios, dealing with missing data and measurement errors. As a specific application of this novel methodology, we target two of the most prevalent cancers among US women: breast cancer and colorectal cancer. By applying our method to the Women’s Health Initiative data, we successfully identify metabolites that are associated with either or both of these cancers, demonstrating the practical utility and potential of our method in identifying consistent risk factors and understanding shared mechanisms between diseases.
Abstract: Panel data transcends cross-sectional data by tapping pooled inter- and intra-individual differences, along with between and within individual variation separately. In the present study these micro variations in ill-being are predicted by psychological indicators constructed from the British Household Panel Survey (BHPS). Panel regression effects are corrected for errors-in-variables, which attenuate slopes estimated by traditional panel regressions. These corrections reveal that unhappiness and life dissatisfaction are distinct variables that have different psychological causations.