arXiv Science⌕ Search

arXiv · 2609.36328

Multilevel regression trees with application to wildfires in the American west

Abstract

We propose a Bayesian regression tree model fit within a multilevel structure and apply it to historic wildfire data in the western United States. Sharing of information between related groups (ecoregions) combined with highly interpretable regression trees allows for better predictions and understanding of climate and land cover variables predictive of wildfires. By doing a simulation study with a range of performance metrics, we demonstrate our method produces tree posteriors most structurally similar to assumed true trees, while simultaneously achieving good out-of-sample predictive performance. Applied to a large wildfire data set, we explore variable splits within regression trees corresponding to each ecoregion in detail, taking into account known features of each location. Shared hyperparameters between trees provide highly useful understanding of both variable and split value importance in predicting wildfires among all ecoregions, with no direct parallel in comparable models. Namely, we highlight potential evaporation, temperature, and evergreen forest land cover as variables most associated with historic wildfires, with some observable patterns in split values most commonly chosen across the groups. We propose a new algorithm based on parallel tempering, conditioning on shared hyperparameters at the true posterior temperature, improving Markov chain mixing, a known bottleneck in Bayesian CART models.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

John Henry V. Gray, Tianjian Zhou, Benjamin A. Shaby. 2026-09-28. Multilevel regression trees with application to wildfires in the American west. https://arxiv.org/abs/2609.36328

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Who Blocks Whom? Probabilistic Pass-Blocking Assignments for Evaluating Blockers and Pass Rushers in American Football

Historically, statistical analysis of offensive lineman has been hindered by the lack of easily measurable quantities. More recently, with the introduction of player tracking data new methodological advances are now possible. Using high-dimensional spatio-temporal data, we adapt the defensive-matchup hidden Markov model of \cite{franks2015characterizing} from basketball to football pass protection, producing frame-by-frame probabilistic assignments of each pass blocker to the rushers. We show how this probabilistic assignment is a usable modeling artifact that augments existing player-evaluation frameworks. We directly quantify the attention a rusher commands, upgrade adjusted plus-minus \citep{Macdonald+2012} from all-or-nothing stints to partial, continuous blocking credit in continuous time, yield block-shedding survival metrics, and measure the space a rusher generates for his teammates. Fit to the first eight weeks of the 2021 NFL season, the resulting metrics recover widely-recognized elite rushers and pass protectors and align with independent charting.

stat.AP↗

Geographic Disparities in Hospice Quality and Family Caregiver Experience: The Roles of Ownership, Social Vulnerability, and Workforce Capacity

Hospice quality should be interpreted in relation to both provider organization and the local conditions under which care is delivered. This study develops a provider-county performance assessment framework by linking national Centers for Medicare & Medicaid Services (CMS) hospice data and Consumer Assessment of Healthcare Providers and Systems (CAHPS) Hospice Survey outcomes with county measures of rurality, social vulnerability, health burden, and workforce and health-resource context. The adjusted analysis includes 2,928 providers in 1,078 counties and combines geographic mapping, blockwise regression, six secondary CAHPS outcomes, and eight sensitivity analyses. Adding county context increased adjusted R-squared from 0.077 to 0.202. After full adjustment, for-profit hospices had overall caregiver ratings 3.832 percentage points lower than nonprofit hospices. This negative association appeared across all six secondary CAHPS domains and remained significant in every sensitivity specification. Higher county social vulnerability was also associated with poorer caregiver experience, although its magnitude depended partly on the specification of community health burden. These findings show that county context materially improves hospice performance assessment but does not eliminate the ownership difference. The framework supports context-aware monitoring, peer comparison, and targeted quality improvement.

stat.AP↗

Feeling Left Behind? Territorial Disadvantage, Well-Being and Cohesion across European Regions

Left-behind places are usually identified using economic, demographic and accessibility indicators, but these may not align with well-being and social cohesion. We combine ten years of territorial indicators with subjective measures derived from georeferenced social-media data for over 1,200 NUTS-3 regions in 28 European countries (2013-2023). Regional profiles and within-between panel models reveal uneven relationships: disadvantaged regions can display contrasting levels of well-being and cohesion, national context alters cross-country associations, and between-region differences diverge from within-region change. The findings show that aggregate measures conceal distinct forms and trajectories of territorial disadvantage.

stat.AP↗