arXiv ScienceSearch

arXiv subjects

Gentry White

Publications and source records attributed to Gentry White.

9 recordsLinked to original sources

Estimating Negative Income Distributions via Data Fusion with Vine Copula-based Imputation

Imputation-based data fusion combines datasets by imputing missing variables in one dataset using information from the other. The validity of imputation-based data fusion relies on accurately preserving both the imputed distributions and the underlying multivariate dependence structure across data sources. However, traditional imputation-based data fusion methods often struggle to preserve the dependence structure, particularly when marginal distributions differ or when complex, nonlinear multivariate dependencies exist. Inadequate preservation of these dependencies can result in distorted joint distributions and biased inference in the fused data. To address this challenge, this study introduces a novel imputation-based data fusion framework that utilises conditional sampling from C-vine and D-vine copulas as the imputation mechanism. This approach flexibly models pairwise and higher-order dependencies while accommodating heterogeneous marginal distributions, thereby generating imputations that more accurately maintain the distribution of the multivariate data. The proposed method is validated through a simulation study using survey data and a real-data application that estimates the negative income distribution in survey data using information from administrative tax records. Results indicate that the proposed method outperforms existing approaches in preserving both the imputed distributions and the dependence structure across datasets.

stat.ME

Person-to-person opinion dynamics: an empirical study using an online game

A model needs to make verifiable predictions to have any scientific value. In opinion dynamics, the study of how individuals exchange opinions with one another, there are many theoretical models which attempt to model opinion exchange, one of which is the Martins model, which differs from other models by using a parameter that is easier to control for in an experiment. In this paper, we have designed an experiment to verify the Martins model and contribute to the experimental design in opinion dynamic with our novel method.

physics.soc-ph

Mathematical measures of societal polarisation

In opinion dynamics, as in general usage, polarisation is subjective. To understand polarisation, we need to develop more precise methods to measure the agreement in society. This paper presents four mathematical measures of polarisation derived from graph and network representations of societies and information theoretic divergences or distance metrics. Two of the methods, min-max flow and spectral radius, rely on graph theory and define polarisation in terms of the structural characteristics of networks. The other two methods represent opinions as probability density functions and use the Kullback Leibler divergence and the Hellinger distance as polarisation measures. We present a series of opinion dynamics simulations from two common models to test the effectiveness of the methods. Results show that the four measures provide insight into the different aspects of polarisation and allow real-time monitoring of social networks for indicators of polarisation. The three measures, the spectral radius, Kullback Leibler divergence and Hellinger distance, smoothly delineated between different amounts of polarisation, i.e. how many cluster there were in the simulation, while also measuring with more granularity how close simulations were to consensus. Min-max flow failed to accomplish such nuance.

physics.soc-ph

The effect of biologically mediated decay rates on modelling soil carbon sequestration in agricultural settings

Microbial biomass carbon (MBC), a crucial soil labile carbon fraction, is the most active component of the soil organic carbon (SOC) that regulates bio-geochemical processes in terrestrial ecosystems. Some studies in the literature ignore the effect of microbial population growth on carbon decomposition rates. In reality, we might expect that the decomposition rate should be related to the population of microbes in the soil and have a positive relationship with the size of the microbial biomass pool. In this study, we explore the effect of microbial population growth on the accuracy of modelling soil carbon sequestration by developing and comparing two soil carbon models that consider a carrying capacity and limit to the growth of the microbial pool. We apply our models to three datasets, two small and one large datasets, and we select the best model in terms of having the best predictive performance through two model selection methods. Through this analysis we reveal that commonly used complex soil carbon models can over-fit in the presence of both small and large time-series datasets, and our simpler model can produce more accurate predictions. We conclude that considering the microbial population growth in a soil carbon model improves the accuracy of a model in the presence of a large dataset.

stat.AP

Innovative Approaches in Soil Carbon Sequestration Modelling for Better Prediction with Limited Data

Soil carbon accounting and prediction play a key role in building decision support systems for land managers selling carbon credits, in the spirit of the Paris and Kyoto protocol agreements. Land managers typically rely on computationally complex models fit using sparse datasets to make these accounts and predictions. The model complexity and sparsity of the data can lead to over-fitting, leading to inaccurate results when making predictions with new data. Modellers address over-fitting by simplifying their models and reducing the number of parameters, and in the current context this could involve neglecting some soil organic carbon (SOC) components. In this study, we introduce two novel SOC models and a new RothC-like model and investigate how the SOC components and complexity of the SOC models affect the SOC prediction in the presence of small and sparse time series data. We develop model selection methods that can identify the soil carbon model with the best predictive performance, in light of the available data. Through this analysis we reveal that commonly used complex soil carbon models can over-fit in the presence of sparse time series data, and our simpler models can produce more accurate predictions. The published version of this study is available in Scientific Reports (https://www.nature.com/articles/s41598-024-53516-z/<10.1038/s41598-024-53516-z>)

stat.CO

Direct Sampling of Bayesian Thin-Plate Splines for Spatial Smoothing

Radial basis functions are a common mathematical tool used to construct a smooth interpolating function from a set of data points. A spatial prior based on thin-plate spline radial basis functions can be easily implemented resulting in a posterior that can be sampled directly using Monte Carlo integration, avoiding the computational burden and potential inefficiency of an Monte Carlo Markov Chain (MCMC) sampling scheme. The derivation of the prior and sampling scheme are demonstrated.

stat.ME

Modelling the Proliferation of Terrorism via Diffusion and Contagion

The proliferation of terrorism is a serious concern in national and international security, as its spread is seen as an existential threat to Western liberal democracies. Understanding and effectively modelling the spread of terrorism provides useful insight into formulating effective responses. A mathematical model capturing the theoretical constructs of contagion and diffusion is constructed for explaining the spread of terrorist activity and used to analyse data from the Global Terrorism Database from 2000--2016 for Afghanistan, Iraq, and Israel.

stat.AP

Self-exciting hurdle models for terrorist activity

A predictive model of terrorist activity is developed by examining the daily number of terrorist attacks in Indonesia from 1994 through 2007. The dynamic model employs a shot noise process to explain the self-exciting nature of the terrorist activities. This estimates the probability of future attacks as a function of the times since the past attacks. In addition, the excess of nonattack days coupled with the presence of multiple coordinated attacks on the same day compelled the use of hurdle models to jointly model the probability of an attack day and corresponding number of attacks. A power law distribution with a shot noise driven parameter best modeled the number of attacks on an attack day. Interpretation of the model parameters is discussed and predictive performance of the models is evaluated.

stat.AP