arXiv ScienceSearch

arXiv subjects

Soudeep Deb

Publications and source records attributed to Soudeep Deb.

At least 19 recordsLinked to original sources

Projection Diagnostics for Directional Asymmetry and Tail-Ratio Departure in Multivariate Data

We study projection-based diagnostics for distinguishing directional asymmetry from tail-ratio departure in multivariate data. The procedure reduces the problem to one-dimensional projections and computes two quantile-based summaries: a directional skewness measure evaluated over several quantile levels, and an interquantile tail-ratio evaluated relative to a chosen benchmark. The two summaries lead to a four-regime classification: symmetric benchmark-tail, symmetric tail-departed, skewed benchmark-tail, and skewed tail-departed. The quantile formulation avoids relying on third and fourth moments, which can be unstable in heavy-tailed settings. We establish population properties under central symmetry and ellipticity, uniform finite-sample bounds over the searched directions, and consistency of the threshold classifier under separated regimes. A sparse rank-one calculation is also used to show why coordinate directions can complement random directions in high dimensions. The resulting diagnostic is meant to guide subsequent modelling choices, for example whether a symmetric, skewed, tail-departed, or combined multivariate model is appropriate.

stat.ME

Optimising football transfer strategy under budget constraints: A weighted multi-criteria approach

The football transfer market is a complex, dynamic environment in which clubs compete to acquire players who strengthen their squads. While several frameworks estimate a player's worth, a comprehensive approach that captures both squad optimisation and transfer market dynamics remains limited. In this paper, we propose a quantitative framework for optimising football transfer strategy under budget constraints, integrated with a competitive bidding paradigm. Using data from professional football leagues, we construct player performance and transfer price models using linear mixed-effects frameworks that incorporate player characteristics, recent performance, team context, and league effects. The predicted ratings and estimated transfer prices are then integrated into a weighted multi-criteria constrained optimisation framework that determines a club's transfer activities at the end of the season. Finally, these optimal transfer decisions are embedded within an independent private-value auction model with a random reserve price to analyse market behaviour when multiple teams compete for the same player. We illustrate our approach using the 2018-19 season of the English Premier League to demonstrate its ability to capture transfer-market dynamics.

stat.AP

Nonparametric regression of spatio-temporal data using infinite-dimensional covariates

In spatio-temporal analysis, we often record data at specific time intervals but with varying spatial locations between these timepoints. We propose a conditional model to analyze such spatio-temporal data that accommodates the dependencies alongside second-order stationary explanatory variables, which may be infinite-dimensional and accommodate spatio-temporal covariates. Because of the absence of a mixing-type dependence condition in this case, which is typically required by the existing studies, we consider a weaker polynomially decaying moment contraction (PMC) condition on the covariates. In this paper, we obtain nonparametric point estimates of the mean and covariate functions of such a regression model, which we then show to be statistically consistent. We also obtain a simultaneous confidence interval of the mean function using the central limit theorem for the proposed estimator. Such simultaneous inference tools can be used to test for certain specifications of the mean function. Some simulation studies and two real-data analyses have been illustrated to corroborate the findings.

stat.ME

A nonparametric approach to understand multivariate quantile dynamics in financial time series

Over the last decade, nonparametric methods have gained increasing attention for modeling complex data structures due to their flexibility and minimal structural assumptions. In this paper, we study a general multivariate nonparametric regression framework that encompasses a broad class of parametric models commonly used in financial econometrics. Both the response and the covariate processes are allowed to be multivariate with fixed finite dimensions, and the framework accommodates temporal dependence, thereby introducing additional modeling and theoretical hurdles. To address these challenges, we adopt a functional dependence structure which permits flexible dynamic behavior while maintaining tractable asymptotic analysis. Within this setting, we establish strong and weak convergence results for the estimators of the conditional mean and volatility functions. In addition, we investigate conditional geometric quantiles in the multivariate time series context and prove their consistency under mild regularity conditions. The finite sample performance is examined through comprehensive simulation studies, and the methodology is illustrated by modeling the stock returns of Maersk and Lockheed Martin as a nonparametric function of a geopolitical risk index.

stat.ME

E-STGCN: Extreme Spatiotemporal Graph Convolutional Networks for Air Quality Forecasting

Modeling and forecasting air quality is crucial for effective air pollution management and protecting public health. Air quality data, characterized by nonlinearity, nonstationarity, and spatiotemporal correlations, often include extreme pollutant levels in severely polluted cities (e.g., Delhi, the capital of India). This is ignored by various geometric deep learning models, such as Spatiotemporal Graph Convolutional Networks (STGCN), which are otherwise effective for spatiotemporal forecasting. This study develops an extreme value theory (EVT) guided modified STGCN model (E-STGCN) for air pollution data to incorporate extreme behavior across pollutant concentrations. E-STGCN combines graph convolutional networks for spatial modeling and EVT-guided long short-term memory units for temporal sequence learning. Along with spatial and temporal components, it incorporates a generalized Pareto distribution to capture the extreme behavior of different air pollutants and embed this information into the learning process. The proposal is then applied to analyze air pollution data of 37 monitoring stations across Delhi, India. The forecasting performance for different test horizons is compared to benchmark forecasters (both temporal and spatiotemporal). It is found that E-STGCN has consistent performance across all seasons. The robustness of our results has also been evaluated empirically. Moreover, combined with conformal prediction, E-STGCN can produce probabilistic prediction intervals.

stat.AP

Nonparametric method of structural break detection in stochastic time series regression model

We propose a novel nonparametric test to detect structural breaks in the conditional mean and/or variance of a time series. Our method does not assume any specific parametric form for the dependence structure of the regressor, the time series model, or the distribution of the noise. This flexibility allows our algorithm to be applicable to a wide range of framework. We further apply the proposed test to accurately localize the changepoints and establish theoretical guarantees showing that the estimated structural breaks are consistent, meaning they lie sufficiently close to the true breakpoints when a sufficiently large sample is available. The effectiveness of the proposed algorithm is demonstrated through an extensive simulation study encompassing a diverse range of time series structures, including light, moderately heavy, and heavy tailed distributions. We also show a real-life example, where an application to Bitcoin prices and Google search volume illustrates how the procedure can identify changes in the conditional relationship between market attention and price dynamics.

stat.ME

A divide-and-conquer approach for spatio-temporal analysis of large house price data from Greater London

Statistical research in real estate markets, particularly in understanding the spatio-temporal dynamics of house prices, has garnered significant attention in recent times. Although Bayesian methods are common in spatio-temporal modeling, standard Markov chain Monte Carlo (MCMC) techniques are usually slow for large datasets such as house price data. To tackle this problem, we propose a divide-and-conquer spatio-temporal modeling approach. This method involves partitioning the data into multiple subsets and applying an appropriate Gaussian process model to each subset in parallel. The results from each subset are then combined using the Wasserstein barycenter technique to obtain the global parameters for the original problem. The proposed methodology allows for multiple observations per spatial and time unit, thereby offering added benefits for practitioners. As a real-life application, we analyze house price data of more than 0.6 million transactions from 983 middle layer super output areas in London over a period of eight years. The methodology provides insightful findings about the effects of various amenities, trend patterns, and the relationship between prices and carbon emissions. Furthermore, as demonstrated through a cross-validation study, it shows good predictive accuracy while balancing computational efficiency.

stat.AP

Nonparametric quantile regression for spatio-temporal processes

In this paper, we develop a new and effective approach to nonparametric quantile regression that accommodates ultrahigh-dimensional data arising from spatio-temporal processes. This approach proves advantageous in staving off computational challenges that constitute known hindrances to existing nonparametric quantile regression methods when the number of predictors is much larger than the available sample size. We investigate conditions under which estimation is feasible and of good overall quality and obtain sharp approximations that we employ to devising statistical inference methodology. These include simultaneous confidence intervals and tests of hypotheses, whose asymptotics is borne by a non-trivial functional central limit theorem tailored to martingale differences. Additionally, we provide finite-sample results through various simulations which, accompanied by an illustrative application to real-worldesque data (on electricity demand), offer guarantees on the performance of the proposed methodology.

stat.ME

A Bayesian approach to identify changepoints in spatio-temporal ordered categorical data: An application to COVID-19 data

Although there is substantial literature on identifying structural changes for continuous spatio-temporal processes, the same is not true for categorical spatio-temporal data. This work bridges that gap and proposes a novel spatio-temporal model to identify changepoints in ordered categorical data. The model leverages an additive mean structure with separable Gaussian space-time processes for the latent variable. Our proposed methodology can detect significant changes in the mean structure as well as in the spatio-temporal covariance structures. We implement the model through a Bayesian framework that gives a computational edge over conventional approaches. From an application perspective, our approach's capability to handle ordinal categorical data provides an added advantage in real applications. This is illustrated using county-wise COVID-19 data (converted to categories according to CDC guidelines) from the state of New York in the USA. Our model identifies three changepoints in the transmission levels of COVID-19, which are indeed aligned with the ``waves'' due to specific variants encountered during the pandemic. The findings also provide interesting insights into the effects of vaccination and the extent of spatial and temporal dependence in different phases of the pandemic.

stat.ME

Optimal selection of the starting lineup for a football team

The success of a football team depends on various individual skills and performances of the selected players as well as how cohesively they perform. We propose a two-stage process for selecting optimal playing eleven of a football team from its pool of available players. In the first stage a LASSO-induced modified multinomial logistic regression model is derived to analyse the probabilities of the three possible outcomes. The model considers strengths of the players in the team as well as those of the opponent, home advantage, and also the effects of individual players and player combinations beyond the recorded performances of these players. In the second stage, a GRASP-type meta-heuristic is implemented for the team selection which maximises its probability of winning. The work is illustrated with English Premier League data from 2008/09 to 2015/16. The application demonstrates that the model in the first stage furnishes valuable insights about the deciding factors for different teams whereas the optimisation steps can be effectively used to determine the best possible starting lineup under various circumstances. We propose a measure of efficiency in team selection by the team management and analyse the performance of the teams on this front.

stat.AP

Real-time forecasting within soccer matches through a Bayesian lens

This paper employs a Bayesian methodology to predict the results of soccer matches in real-time. Using sequential data of various events throughout the match, we utilize a multinomial probit regression in a novel framework to estimate the time-varying impact of covariates and to forecast the outcome. English Premier League data from eight seasons are used to evaluate the efficacy of our method. Different evaluation metrics establish that the proposed model outperforms potential competitors inspired by existing statistical or machine learning algorithms. Additionally, we apply robustness checks to demonstrate the model's accuracy across various scenarios.

stat.AP

Effect of influence in voter models and its application in detecting significant interference in political elections

In this article, we study the effect of vector-valued interventions in votes under a binary voter model, where each voter expresses their vote as a $0-1$ valued random variable to choose between two candidates. We assume that the outcome is determined by the majority function, which is true for a democratic system. The term intervention includes cases of counting errors, reporting irregularities, electoral malpractice etc. Our focus is to analyze the effect of the intervention on the final outcome. We construct statistical tests to detect significant irregularities in elections under two scenarios, one where exit poll data is available and more broadly under the assumption of a cost function associated with causing the interventions. Relevant theoretical results on the consistency of the test procedures are also derived. Through a detailed simulation study, we show that the test procedure has good power and is robust across various settings. We also implement our method on three real-life data sets. The applications provide results consistent with existing knowledge and establish that the method can be adopted for crucial problems related to political elections.

stat.AP

A review and recommendations on variable selection methods in regression models for binary data

The selection of essential variables in logistic regression is vital because of its extensive use in medical studies, finance, economics and related fields. In this paper, we explore four main typologies (test-based, penalty-based, screening-based, and tree-based) of frequentist variable selection methods in logistic regression setup. Primary objective of this work is to give a comprehensive overview of the existing literature for practitioners. Underlying assumptions and theory, along with the specifics of their implementations, are detailed as well. Next, we conduct a thorough simulation study to explore the performances of fifteen different methods in terms of variable selection, estimation of coefficients, prediction accuracy as well as time complexity under various settings. We take low, moderate and high dimensional setups and consider different correlation structures for the covariates. A real-life application, using a high-dimensional gene expression data, is also included in this study to further understand the efficacy and consistency of the methods. Finally, based on our findings in the simulated data and in the real data, we provide recommendations for practitioners on the choice of variable selection methods under various contexts.

stat.ME

Nonparametric quantile regression for time series with replicated observations and its application to climate data

This paper proposes a model-free nonparametric estimator of conditional quantile of a time series regression model where the covariate vector is repeated many times for different values of the response. This type of data is abound in climate studies. To tackle such problems, our proposed method exploits the replicated nature of the data and improves on restrictive linear model structure of conventional quantile regression. Relevant asymptotic theory for the nonparametric estimators of the mean and variance function of the model are derived under a very general framework. We provide a detailed simulation study which clearly demonstrates the gain in efficiency of the proposed method over other benchmark models, especially when the true data generating process entails nonlinear mean function and heteroskedastic pattern with time dependent covariates. The predictive accuracy of the non-parametric method is remarkably high compared to other methods when attention is on the higher quantiles of the variable of interest. Usefulness of the proposed method is then illustrated with two climatological applications, one with a well-known tropical cyclone wind-speed data and the other with an air pollution data.

stat.ME

Forecasting Elections from Partial Information Using a Bayesian Model for a Multinomial Sequence of Data

Predicting the winner of an election is of importance to multiple stakeholders. To formulate the problem, we consider an independent sequence of categorical data with a finite number of possible outcomes in each. The data is assumed to be observed in batches, each of which is based on a large number of such trials and can be modeled via multinomial distributions. We postulate that the multinomial probabilities of the categories vary randomly depending on batches. The challenge is to predict accurately on cumulative data based on data up to a few batches as early as possible. On the theoretical front, we first derive sufficient conditions of asymptotic normality of the estimates of the multinomial cell probabilities and present corresponding suitable transformations. Then, in a Bayesian framework, we consider hierarchical priors using multivariate normal and inverse Wishart distributions and establish the posterior convergence. The desired inference is arrived at using these results and ensuing Gibbs sampling. The methodology is demonstrated with election data from two different settings -- one from India and the other from the United States of America. Additional insights of the effectiveness of the proposed methodology are attained through a simulation study.

stat.AP

A mathematical take on the competitive balance of a football league

Competitive balance in a football league is extremely important from the perspective of economic growth of the industry. Many researchers have earlier proposed different measures of competitive balance, which are primarily adapted from the standard economic theory. However, these measures fail to capture the finer nuances of the game. In this work, we discuss a new framework which is more suitable for a football league. First, we present a mathematical proof of an ideal situation where a football league becomes perfectly balanced. Next, a goal based index for competitive balance is developed. We present relevant theoretical results and show how the proposed index can be used to formally test for the presence of imbalance. The methods are implemented on the data from top five European leagues, and it shows that the new approach can better explain the changes in the seasonal competitive balance of the leagues. Further, using appropriate panel data models, we show that the proposed index is more suitable to analyze the variability in total revenues of the football leagues.

stat.AP

Analyzing count data using a time series model with an exponentially decaying covariance structure

Count data appears in various disciplines. In this work, a new method to analyze time series count data has been proposed. The method assumes exponentially decaying covariance structure, a special class of the Mat\'ern covariance function, for the latent variable in a Poisson regression model. It is implemented in a Bayesian framework, with the help of Gibbs sampling and ARMS sampling techniques. The proposed approach provides reliable estimates for the covariate effects and estimates the extent of variability explained by the temporally dependent process and the white noise process. The method is flexible, allows irregular spaced data, and can be extended naturally to bigger datasets. The Bayesian implementation helps us to compute the posterior predictive distribution and hence is more appropriate and attractive for count data forecasting problems. Two real life applications of different flavors are included in the paper. These two examples and a short simulation study establish that the proposed approach has good inferential and predictive abilities and performs better than the other competing models.

stat.ME

A time series method to analyze incidence pattern and estimate reproduction number of COVID-19

The ongoing pandemic of Coronavirus disease (COVID-19) emerged in Wuhan, China in the end of 2019. It has already affected more than 300,000 people, with the number of deaths nearing 13000 across the world. As it has been posing a huge threat to global public health, it is of utmost importance to identify the rate at which the disease is spreading. In this study, we propose a time series model to analyze the trend pattern of the incidence of COVID-19 outbreak. We also incorporate information on total or partial lockdown, wherever available, into the model. The model is concise in structure, and using appropriate diagnostic measures, we showed that a time-dependent quadratic trend successfully captures the incidence pattern of the disease. We also estimate the basic reproduction number across different countries, and find that it is consistent except for the United States of America. The above statistical analysis is able to shed light on understanding the trends of the outbreak, and gives insight on what epidemiological stage a region is in. This has the potential to help in prompting policies to address COVID-19 pandemic in different countries.

stat.AP