arXiv ScienceSearch

arXiv · 2509.06779

A nutritionally informed model for Bayesian variable selection with metabolite response variables

Abstract

Understanding the pathways through which diet affects human metabolism is a central task in nutritional epidemiology. This article proposes novel methodology to identify food items associated with blood metabolites in two cohorts of healthcare professionals. We analyze 30 food intake variables that exhibit relationship structure through their correlations and nutritional attributes. The metabolic responses include 244 compounds measured by mass spectrometry, presenting substantial challenges that include missingness, left-censoring, and skewness. While existing methods can address such factors in low-dimensional settings, they are not designed for high-dimensional regression involving strongly correlated predictors and non-normal outcomes. To address these challenges, we propose a novel Bayesian variable selection framework for metabolite response variables based on a skew-normal censored mixture model. To exploit substantive information on the nutritional similarities among dietary factors, we employ a Markov random field prior that encourages joint selection of related predictors, while introducing a new, efficient strategy for its hyperparameter specification. Applying this methodology to the cohort data identifies multiple metabolite-diet associations that are consistent with previous research as well as several potentially novel associations that were not detected using standard methods. The proposed approach is implemented in the R package multimetab, facilitating its use in high-dimensional metabolomic analyses.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Dylan Clark-Boucher, Brent A Coull, Harrison T Reeder, Fenglei Wang, Qi Sun, Jacqueline R Starr, Kyu Ha Lee. 2025-09-08. A nutritionally informed model for Bayesian variable selection with metabolite response variables. https://arxiv.org/abs/2509.06779

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Modelling structural zeros in compositional data via a zero-censored multivariate normal model

We present a new model for analyzing compositional data with structural zeros. Inspired by \cite{butler2008} who suggested a model in the presence of zero values in the data we propose a model that treats the zero values in a different manner. Instead of projecting every zero value towards a vertex, we project them onto their corresponding edge and fit a zero-censored multivariate model.

stat.ME

Inventing a coin-flip classifier: Accounting for multiplicity in machine learning benchmark performance

State-of-the-art (SOTA) performance refers to the highest performance achieved by some model on a test sample, preferably under controlled conditions such as public data (reproducibility) or public challenges (independent sample). Thousands of classifiers are applied, and the highest performance becomes the new reference point for a particular problem. In effect, this set-up is an estimate of the expected best performance among all classifiers applied to a random sample; a sample maximum estimate. In this paper, we argue that SOTA should instead be estimated by the expected performance of the best classifier, which can be done without knowing which classifier it is. Our contribution is the formal distinction between the two, and an investigation into the practical consequences of using the former to estimate the latter. This is done by presenting sample maximum estimator distributions for non-identical and dependent classifiers. We illustrate the impact on real world examples from public challenges.

stat.ME

A computationally efficient multivariate volatility model for many assets

This paper develops a flexible and computationally efficient multivariate volatility model that accommodates dynamic conditional correlations and volatility spillover effects among financial assets. The new model has desirable properties such as identifiability and computational tractability for many assets.A sufficient condition for strict stationarity is derived for the process.Two quasi-maximum likelihood estimation methods are proposed for the new model without and with low-rank constraints on the coefficient matrices, respectively, and the asymptotic properties of both estimators are established. Moreover, a selection-consistent Bayesian information criterion is developed for order selection, and testing for volatility spillover effects is discussed. The finite-sample performance of the proposed methodology is evaluated in simulations for small and moderate dimensions.Its usefulness and inference tools are illustrated by two empirical examples for 5 stock markets and 17 industry portfolios.

stat.ME