arXiv ScienceSearch

arXiv subjects

Thorsten Simon

Publications and source records attributed to Thorsten Simon.

9 recordsLinked to original sources

Spatio-seasonal risk assessment of upward lightning at tall objects using meteorological reanalysis data

This study investigates lightning at tall objects and evaluates the risk of upward lightning (UL) over the eastern Alps and its surrounding areas. While uncommon, UL poses a threat, especially to wind turbines, as the long-duration current of UL can cause significant damage. Current risk assessment methods overlook the impact of meteorological conditions, potentially underestimating UL risks. Therefore, this study employs random forests, a machine learning technique, to analyze the relationship between UL measured at Gaisberg Tower (Austria) and $35$ larger-scale meteorological variables. Of these, the larger-scale upward velocity, wind speed and direction at 10 meters and cloud physics variables contribute most information. The random forests predict the risk of UL across the study area at a 1 km$^2$ resolution. Strong near-surface winds combined with upward deflection by elevated terrain increase UL risk. The diurnal cycle of the UL risk as well as high-risk areas shift seasonally. They are concentrated north/northeast of the Alps in winter due to prevailing northerly winds, and expanding southward, impacting northern Italy in the transitional and summer months. The model performs best in winter, with the highest predicted UL risk coinciding with observed peaks in measured lightning at tall objects. The highest concentration is north of the Alps, where most wind turbines are located, leading to an increase in overall lightning activity. Comprehensive meteorological information is essential for UL risk assessment, as lightning densities are a poor indicator of lightning at tall objects.

physics.soc-ph

Confirmatory adaptive group sequential designs for clinical trials with multiple time-to-event outcomes in Markov models

The analysis of multiple time-to-event outcomes in a randomised controlled clinical trial can be accomplished with exisiting methods. However, depending on the characteristics of the disease under investigation and the circumstances in which the study is planned, it may be of interest to conduct interim analyses and adapt the study design if necessary. Due to the expected dependency of the endpoints, the full available information on the involved endpoints may not be used for this purpose. We suggest a solution to this problem by embedding the endpoints in a multi-state model. If this model is Markovian, it is possible to take the disease history of the patients into account and allow for data-dependent design adaptiations. To this end, we introduce a flexible test procedure for a variety of applications, but are particularly concerned with the simultaneous consideration of progression-free survival (PFS) and overall survival (OS). This setting is of key interest in oncological trials. We conduct simulation studies to determine the properties for small sample sizes and demonstrate an application based on data from the NB2004-HR study.

stat.ME

Scalable Estimation for Structured Additive Distributional Regression

Recently, fitting probabilistic models have gained importance in many areas but estimation of such distributional models with very large data sets is a difficult task. In particular, the use of rather complex models can easily lead to memory-related efficiency problems that can make estimation infeasible even on high-performance computers. We therefore propose a novel backfitting algorithm, which is based on the ideas of stochastic gradient descent and can deal virtually with any amount of data on a conventional laptop. The algorithm performs automatic selection of variables and smoothing parameters, and its performance is in most cases superior or at least equivalent to other implementations for structured additive distributional regression, e.g., gradient boosting, while maintaining low computation time. Performance is evaluated using an extensive simulation study and an exceptionally challenging and unique example of lightning count prediction over Austria. A very large dataset with over 9 million observations and 80 covariates is used, so that a prediction model cannot be estimated with standard distributional regression methods but with our new approach.

stat.CO

Upward lightning at wind turbines: Risk assessment from larger-scale meteorology

Upward lightning (UL) has become an increasingly important threat to wind turbines as ever more of them are being installed for renewably producing electricity. The taller the wind turbine the higher the risk that the type of lightning striking the man-made structure is UL. UL can be much more destructive than downward lightning due to its long lasting initial continuous current leading to a large charge transfer within the lightning discharge process. Current standards for the risk assessment of lightning at wind turbines mainly take the summer lightning activity into account, which is inferred from LLS. Ground truth lightning current measurements reveal that less than 50% of UL might be detected by lightning location systems (LLS). This leads to a large underestimation of the proportion of LLS-non-detectable UL at wind turbines, which is the dominant lightning type in the cold season. This study aims to assess the risk of LLS-detectable and LLS-non-detectable UL at wind turbines using direct UL measurements at the Gaisberg Tower (Austria) and S\"antis Tower (Switzerland). Direct UL observations are linked to meteorological reanalysis data and joined by random forests, a powerful machine learning technique. The meteorological drivers for the non-/occurrence of LLS-detectable and LLS-non-detectable UL, respectively, are found from the random forest models trained at the towers and have large predictive skill on independent data. In a second step the results from the tower-trained models are extended to a larger study domain (Central and Northern Germany). The tower-trained models for LLS-detectable lightning is independently verified at wind turbine locations in that domain and found to reliably diagnose that type of UL. Risk maps based on case study events show that high diagnosed probabilities in the study domain coincide with actual UL events.

stat.ML

Identifying Lightning Processes in ERA5 Soundings with Deep Learning

Atmospheric environments favorable for lightning and convection are commonly represented by proxies or parameterizations based on expert knowledge such as CAPE, wind shears, charge separation, or combinations thereof. Recent developments in the field of machine learning, high resolution reanalyses, and accurate lightning observations open possibilities for identifying tailored proxies without prior expert knowledge. To identify vertical profiles favorable for lightning, a deep neural network links ERA5 vertical profiles of cloud physics, mass field variables and wind to lightning location data from the Austrian Lightning Detection and Information System (ALDIS), which has been transformed to a binary target variable labelling the ERA5 cells as cells with lightning activity and cells without lightning activity. The ERA5 parameters are taken on model levels beyond the tropopause forming an input layer of approx. 670 features. The data of 2010-2018 serve as training/validation. On independent test data, 2019, the deep network outperforms a reference with features based on meteorological expertise. SHAP values highlight the atmospheric processes learned by the network which identifies cloud ice and snow content in the upper and mid-troposphere as very relevant features. As these patterns correspond to the separation of charge in thunderstorm cloud, the deep learning model can serve as physically meaningful description of lightning. Depending on the region, the neural network also exploits the vertical wind or mass profiles to correctly classify cells with lightning activity.

physics.ao-ph

Daily-Resolved Lightning Climatology of the Eastern Alpine Region at the Kilometer Scale

Lightning flashes are rare albeit hazardous events. Despite this scarcity, generalized additive models (GAMs) succeed in producing a climatology of lightning occurrence for the eastern Alps and surrounding lowlands at an unprecedented resolution of 1\,km$^2$ for each day of April through September with data from the ALDIS lightning location system. The GAM adds the effects of seasonality, jaggedness of the terrain, and seasonally varying effects of elevation and region, thus combining information from analysis cells sharing similar characteristics. The probability of a cloud-to-ground discharge over 1\,km$^2$ on a given day is typically less than 1\,\% with a rapid increase in spring, followed by a plateau and a gentler tapering-off in fall. Probabilities early are lower at high elevations but increase once their snow cover is gone. Regional patterns of lightning also vary with season with an overall southward shift later in the year but more complex details. Grid cells with jagged topography have a higher probability of lightning.

physics.ao-ph

Upward Lightning at the Gaisberg Tower: Initiation Mechanism and Flash Type and the Atmospheric Influence

Upward lightning is much rarer than downward lightning and requires tall ($100+$~m) structures to initiate. It may be either triggered by other lightning discharges or completely self-initiated. While conventional lightning location systems reliably detect downward lightning, they miss a specific flash type of upward lightning that consists only of a continuous current. Globally, only few specially instrumented towers can detect this flash type. The proliferation of wind turbines in combination with large damage from upward lightning necessitates an improved understanding under which conditions the self-initiated and the undetected subtype of upward lightning occur. To find larger-scale meteorological conditions favorable for self-initiated and undetectable upward lightning, this study uses a random forest machine learning model. It combines direct measurements at the specially instrumented tower at Gaisberg mountain in Austria with explanatory variables from larger-scale atmospheric reanalysis data (ERA5). Atmospheric variables reliably explain whether upward lightning is self-initiated by the tower or triggered by other lightning discharges. The most important variable is the height of the $-10~^\circ$C isotherm above the tall structure: the closer it is the higher is the probability of self-initiated upward lightning. Two-meter temperature and the amount of CAPE are also important. For the occurrence of upward lightning undetectable by lightning location systems, this study finds a strong relationship to the absence of lightning in the vicinity.

physics.ao-ph

Cholesky-based multivariate Gaussian regression

Distributional regression is extended to Gaussian response vectors of dimension greater than two by parameterizing the covariance matrix $\Sigma$ of the response distribution using the entries of its Cholesky decomposition. The more common variance-correlation parameterization limits such regressions to bivariate responses -- higher dimensions require complicated constraints among the correlations to ensure positive definite $\Sigma$ and a well-defined probability density function. In contrast, Cholesky-based parameterizations ensure positive definiteness for all distributional dimensions no matter what values the parameters take, enabling estimation and regularization as for other distributional regression models. In cases where components of the response vector are assumed to be conditionally independent beyond a certain lag $r$, model complexity can be further reduced by setting Cholesky parameters beyond this lag to zero a priori. Cholesky-based multivariate Gaussian regression is first illustrated and assessed on artificial data and subsequently applied to a real-world 10-dimensional weather forecasting problem. There the regression is used to obtain reliable joint probabilities of temperature across ten future times, leveraging temporal correlations over the prediction period to obtain more precise and meteorologically consistent probabilistic forecasts.

stat.ME

bamlss: A Lego Toolbox for Flexible Bayesian Regression (and Beyond)

Over the last decades, the challenges in applied regression and in predictive modeling have been changing considerably: (1) More flexible model specifications are needed as big(ger) data become available, facilitated by more powerful computing infrastructure. (2) Full probabilistic modeling rather than predicting just means or expectations is crucial in many applications. (3) Interest in Bayesian inference has been increasing both as an appealing framework for regularizing or penalizing model estimation as well as a natural alternative to classical frequentist inference. However, while there has been a lot of research in all three areas, also leading to associated software packages, a modular software implementation that allows to easily combine all three aspects has not yet been available. For filling this gap, the R package bamlss is introduced for Bayesian additive models for location, scale, and shape (and beyond). At the core of the package are algorithms for highly-efficient Bayesian estimation and inference that can be applied to generalized additive models (GAMs) or generalized additive models for location, scale, and shape (GAMLSS), also known as distributional regression. However, its building blocks are designed as "Lego bricks" encompassing various distributions (exponential family, Cox, joint models, ...), regression terms (linear, splines, random effects, tensor products, spatial fields, ...), and estimators (MCMC, backfitting, gradient boosting, lasso, ...). It is demonstrated how these can be easily recombined to make classical models more flexible or create new custom models for specific modeling challenges.

stat.CO