arXiv ScienceSearch

arXiv subjects

Eric Schmitt

Publications and source records attributed to Eric Schmitt.

7 recordsLinked to original sources

Partial recovery of meter-scale surface weather

Near-surface weather varies over tens to hundreds of meters, yet remains unresolved in analyses and forecasts. We test whether this variation can be inferred without resolving atmospheric dynamics. Combining sparse weather stations, high-resolution Earth observation, and coarse atmospheric dynamics, we infer temperature, dewpoint, and wind at 30-m resolution across the contiguous United States. Against measurements held out in space and time, estimates reduce error by 11-28\% relative to the strongest baseline. Within held-out $0.25^\circ$ grid cells, we recover more spatial variance than baselines, explaining nearly half of temperature variability in the median cell. The method captures time-varying differences between locations and produces coherent patterns associated with topography and land cover. Beyond weather, our findings illustrate how sparse observations of a dynamical system can be combined with dense observations of persistent environmental structure to recover otherwise unresolved spatial variability.

cs.LG

Local Off-Grid Weather Forecasting with Multi-Modal Earth Observation Data

Urgent applications like wildfire management and renewable energy generation require precise, localized weather forecasts near the Earth's surface. However, forecasts produced by machine learning models or numerical weather prediction systems are typically generated on large-scale regular grids, where direct downscaling fails to capture fine-grained, near-surface weather patterns. In this work, we propose a multi-modal transformer model trained end-to-end to downscale gridded forecasts to off-grid locations of interest. Our model directly combines local historical weather observations (e.g., wind, temperature, dewpoint) with gridded forecasts to produce locally accurate predictions at various lead times. Multiple data modalities are collected and concatenated at station-level locations, treated as a token at each station. Using self-attention, the token corresponding to the target location aggregates information from its neighboring tokens. Experiments using weather stations across the Northeastern United States show that our model outperforms a range of data-driven and non-data-driven off-grid forecasting methods. They also reveal that direct input of station data provides a phase shift in local weather forecasting accuracy, reducing the prediction error by up to 80% compared to pure gridded data based models. This approach demonstrates how to bridge the gap between large-scale weather models and locally accurate forecasts to support high-stakes, location-sensitive decision-making.

cs.LG

Extending Bayesian structural time-series estimates of causal impact to many-household conservation initiatives

Government agencies offer economic incentives to citizens for conservation actions, such as rebates for installing efficient appliances and compensation for modifications to homes. The intention of these conservation actions is frequently to reduce the consumption of a utility. Measuring the conservation impact of incentives is important for guiding policy, but doing so is technically difficult. However, the methods for estimating the impact of public outreach efforts have seen substantial developments in marketing to consumers in recent years as marketers seek to substantiate the value of their services. One such method uses Bayesian Stuctural Time Series (BSTS) to compare a market exposed to an advertising campaign with control markets identified through a matching procedure. This paper introduces an extension of the matching/BSTS method for impact estimation to make it applicable for general conservation program impact estimation when multi-household data is available. This is accomplished by household matching/BSTS steps to obtain conservation estimates and then aggregating the results using a meta-regression step to aggregate the findings. A case study examining the impact of rebates for household turf removal on water consumption in multiple Californian water districts is conducted to illustrate the work flow of this method.

stat.ME

Transforming how water is managed in the West

California is challenged by its worst drought in 600 years and faces future water uncertainty. Pioneering new data infrastructure to integrate water use data across California's more than a thousand water providers will support water managers in ensuring water reliability. The California Data Collaborative is a coalition of municipal water utilities serving ten percent of California's population who are delivering on that promise by centralizing customer water use data in a recently completed pilot project. This project overview describes tools that have shown promising early results in improving water efficiency programs and optimizing system operations. Longer term, these tools will help navigate future uncertainty and support water managers in ensuring water reliability no matter what the future holds. The uniquely publicly-owned data infrastructure deployed in this project is envisioned to enable the world's first "marketplace of civic analytics" to power the volume of water efficiency measurements water managers require at a radically more cost effective price. More broadly, this data-utility approach is adaptable to domains other than water and shows specific potential for the broader universe of natural resources.

cs.CY

Finite Sample Breakdown of PCS

The Projection Congruent Subset (PCS) is new method for finding multivariate outliers. PCS returns an outlyingness index which can be used to construct affine equivariant estimates of multivariate location and scatter. In this note, we derive the finite sample breakdown point of these estimators.

math.ST

Finding Regression Outliers With FastRCS

The Residual Congruent Subset (RCS) is a new method for finding outliers in the linear regression setting. Like many other outlier detection procedures, RCS searches for a subset which minimizes a criterion. The difference is that the new criterion was designed to be insensitive to the outliers. RCS is supported by FastRCS, a fast regression and affine equivariant algorithm which we also detail. Both an extensive simulation study and two real data applications show that FastRCS performs better than its competitors.

stat.ME

Finding Multivariate Outliers With FastPCS

The Projection Congruent Subset (PCS) Outlyingness is a new index of multivariate outlyingness obtained by considering univariate projections of the data. Like many other outlier detection procedures, PCS searches for a subset which minimizes a criterion. The difference is that the new criterion was designed to be insensitive to the outliers. PCS is supported by FastPCS, a fast and affine equivariant algorithm which we also detail. Both an extensive simulation study and a real data application from the field of engineering show that FastPCS performs better than its competitors.

stat.ME