arXiv ScienceSearch

arXiv subjects

Harry Joe

Publications and source records attributed to Harry Joe.

15 recordsLinked to original sources

Copula-Based Time Series for Non-Gaussian and Non-Markovian Stationary Processes

In the copula-based approach to univariate time series modeling, the finite dimensional temporal dependence of a stationary time series is captured by a copula. Recent studies investigate how copula-based time series models can be generalized to have long-term autoregressive effects. We study a generalization that comes from a Markov sequence of order p and a q-dependent sequence. We derive the relation of the model to Gaussian-ARMA models and to the Gaussian-GARCH(1,1) model. We investigate distributional properties of the process and discuss the maximum likelihood estimation (MLE). Additionally we analyze the copula moving aggregate process of order one, or MAG(1), as it is a basic building block. Last we test the model in probabilistic forecasting studies on US inflation and German wind energy production.

stat.ME

Extreme Value Inference for CoVaR and Systemic Risk

We develop an extreme value framework for CoVaR centered on $v(q \mid p ; C)$, the copula-adjusted probability level, or equivalently, the CoVaR on the uniform (0,1) scale. We characterize the possible tail regimes of $v(q \mid p ; C)$ through the limit behavior of the copula conditional distribution and show that these regimes are determined by the joint tail expansions of the copula. This leads to tractable conditions for identifying the tail regime and deriving the asymptotic behavior of $v(q | p ; C)$. Building on this characterization, we propose a minimum-distance estimation approach for CoVaR that accommodates multiple tail regimes. The methodology links CoVaR and $\Delta$CoVaR to the underlying joint tail behavior, thereby providing a clear interpretation of these measures in systemic risk analysis. An empirical analysis across U.S. sectors demonstrates the practical value of the approach for assessing systemic risk contributions and exposures with important implications for macroprudential surveillance and risk management.

stat.ME

Margin-closed regime-switching multivariate time series models

A regime-switching multivariate time series model which is closed under margins is built. The model imposes a restriction on all lower-dimensional sub-processes to follow a regime-switching process sharing the same latent regime sequence and having the same Markov order as the original process. The margin-closed regime-switching model is constructed by considering the multivariate margin-closed Gaussian VAR($k$) dependence as a copula within each regime, and builds dependence between observations in different regimes by requiring the first observation in the new regime to depend on the last observation in the previous regime. The property of closure under margins allows inference on the latent regimes based on lower-dimensional selected sub-processes and estimation of univariate parameters from univariate sub-processes, and enables the use of multi-stage estimation procedure for the model. The parsimonious dependence structure of the model also avoids a large number of parameters under the regime-switching setting. The proposed model is applied to a macroeconomic data set to infer the latent business cycle and compared with the relevant benchmark.

stat.ME

Assessing Copula Models for Mixed Continuous-Ordinal Variables

Vine pair-copula constructions exist for a mix of continuous and ordinal variables. In some steps, this can involve estimating a bivariate copula for a pair of mixed continuous-ordinal variables. To assess the adequacy of copula fits for such a pair, diagnostic and visualization methods based on normal score plots and conditional Q-Q plots are proposed. The former utilizes a latent continuous variable for the ordinal variable. Using the Kullback-Leibler divergence, existing probability models for mixed continuous-ordinal variable pair are assessed for the adequacy of fit with simple parametric copula families. The effectiveness of the proposed visualization and diagnostic methods is illustrated on simulated and real datasets.

stat.ME

Vine dependence graphs with latent variables as summaries for gene expression data

The advent of high-throughput sequencing technologies has lead to vast comparative genome sequences. The construction of gene-gene interaction networks or dependence graphs on the genome scale is vital for understanding the regulation of biological processes. Different dependence graphs can provide different information. Some existing methods for dependence graphs based on high-order partial correlations are sparse and not informative when there are latent variables that can explain much of the dependence in groups of genes. Other methods of dependence graphs based on correlations and first-order partial correlations might have dense graphs. When genes can be divided into groups with stronger within group dependence in gene expression than between group dependence, we present a dependence graph based on truncated vines with latent variables that makes use of group information and low-order partial correlations. The graphs are not dense, and the genes that might be more central have more neighbors in the vine dependency graph. We demonstrate the use of our dependence graph construction on two RNA-seq data sets -- yeast and prostate cancer. There is some biological evidence to support the relationship between genes in the resulting dependence graphs. A flexible framework is provided for building dependence graphs via low-order partial correlations and formation of groups, leading to graphs that are not too sparse or dense. We anticipate that this approach will help to identify groups that might be central to different biological functions.

stat.ME

Margin-closed vector autoregressive time series models

Conditions are obtained for a Gaussian vector autoregressive time series of order $k$, VAR($k$), to have univariate margins that are autoregressive of order $k$ or lower-dimensional margins that are also VAR($k$). This can lead to $d$-dimensional VAR($k$) models that are closed with respect to a given partition $\{S_1,\ldots,S_n\}$ of $\{1,\ldots,d\}$ by specifying marginal serial dependence and some cross-sectional dependence parameters. The special closure property allows one to fit the sub-processes of multivariate time series before assembling them by fitting the dependence structure between the sub-processes. We revisit the use of the Gaussian copula of the stationary joint distribution of observations in the VAR($k$) process with non-Gaussian univariate margins but under the constraint of closure under margins. This construction allows more flexibility in handling higher-dimensional time series and a multi-stage estimation procedure can be used. The proposed class of models is applied to a macro-economic data set and compared with the relevant benchmark models.

stat.ME

High-dimensional factor copula models with estimation of latent variables

Factor models are a parsimonious way to explain the dependence of variables using several latent variables. In Gaussian 1-factor and structural factor models (such as bi-factor, oblique factor) and their factor copula counterparts, factor scores or proxies are defined as conditional expectations of latent variables given the observed variables. With mild assumptions, the proxies are consistent for corresponding latent variables as the sample size and the number of observed variables linked to each latent variable go to infinity. When the bivariate copulas linking observed variables to latent variables are not assumed in advance, sequential procedures are used for latent variables estimation, copula family selection and parameter estimation. The use of proxy variables for factor copulas means that approximate log-likelihoods can be used to estimate copula parameters with less computational effort for numerical integration.

stat.ME

Predicting Times to Event Based on Vine Copula Models

In statistics, time-to-event analysis methods traditionally focus on the estimation of hazards. In recent years, machine learning methods have been proposed to directly predict the event times. We propose a method based on vine copula models to make point and interval predictions for a right-censored response variable given mixed discrete-continuous explanatory variables. Extensive experiments on simulated and real datasets show that our proposed vine copula approach provides a decent approximation to other time-to-event analysis models including Cox proportional hazards and Accelerate Failure Time models. When the Cox proportional hazards or Accelerate Failure Time assumptions do not hold, predictions based on vine copulas can significantly outperform other models, depending on the shape of the conditional quantile functions. This shows the flexibility of our proposed vine copula approach for general time-to-event datasets.

stat.ME

On the Selection of Loss Severity Distributions to Model Operational Risk

Accurate modeling of operational risk is important for a bank and the finance industry as a whole to prepare for potentially catastrophic losses. One approach to modeling operational is the loss distribution approach, which requires a bank to group operational losses into risk categories and select a loss frequency and severity distribution for each category. This approach estimates the annual operational loss distribution, and a bank must set aside capital, called regulatory capital, equal to the 0.999 quantile of this estimated distribution. In practice, this approach may produce unstable regulatory capital calculations from year-to-year as selected loss severity distribution families change. This paper presents truncation probability estimates for loss severity data and a consistent quantile scoring function on annual loss data as useful severity distribution selection criteria that may lead to more stable regulatory capital. Additionally, the Sinh-arcSinh distribution is another flexible candidate family for modeling loss severities that can be easily estimated using the maximum likelihood approach. Finally, we recommend that loss frequencies below the minimum reporting threshold be collected so that loss severity data can be treated as censored data.

q-fin.RM

Conditional Inferences Based on Vine Copulas with Applications to Credit Spread Data of Corporate Bonds

Understanding the dependence relationship of credit spreads of corporate bonds is important for risk management. Vine copula models with tail dependence are used to analyze a credit spread dataset of Chinese corporate bonds, understand the dependence among different sectors and perform conditional inferences. It is shown how the effect of tail dependence affects risk transfer, or the conditional distributions given one variable is extreme. Vine copula models also provide more accurate cross prediction results compared with linear regressions. These conditional inference techniques are a statistical contribution for analysis of bond credit spreads of investment portfolios consisting of corporate bonds from various sectors.

stat.ME

Tail Densities of Skew-Elliptical Distributions

Skew-elliptical distributions constitute a large class of multivariate distributions that account for both skewness and a variety of tail properties. This class has simpler representations in terms of densities rather than cumulative distribution functions, and the tail density approach has previously been developed to study tail properties when multivariate densities have more tractable forms. The special skew-elliptical structure allows for derivations of specific forms for the tail densities for those skew-elliptical copulas that admit probability density functions, under heavy and light tail conditions on density generators. The tail densities of skew-elliptical copulas are explicit and depend only on tail properties of the underlying density generator and conditions on the skewness parameters. In the heavy tail case skewness parameters affect tail densities of the skew-elliptical copulas more profoundly than that in the light tail case, whereas in the latter case the tail densities of skew-elliptical copulas are only proportional to the tail densities of symmetrical elliptical copulas. Various examples, including tail densities of skew-normal and skew-t distributions, are given.

math.PR

Prediction based on conditional distributions of vine copulas

Vine copulas are a flexible tool for multivariate non-Gaussian distributions. For data from an observational study where the explanatory variables and response variables are measured together, a proposed vine copula regression method uses regular vines and handles mixed continuous and discrete variables. This method can efficiently compute the conditional distribution of the response variable given the explanatory variables. The performance of the proposed method is evaluated on simulated data sets and a real data set. The experiments demonstrate that the vine copula regression method is superior to linear regression in making inferences with conditional heteroscedasticity.

stat.ME

Model comparison with composite likelihood information criteria

Comparisons are made for the amount of agreement of the composite likelihood information criteria and their full likelihood counterparts when making decisions among the fits of different models, and some properties of penalty term for composite likelihood information criteria are obtained. Asymptotic theory is given for the case when a simpler model is nested within a bigger model, and the bigger model approaches the simpler model under a sequence of local alternatives. Composite likelihood can more or less frequently choose the bigger model, depending on the direction of local alternatives; in the former case, composite likelihood has more "power" to choose the bigger model. The behaviors of the information criteria are illustrated via theory and simulation examples of the Gaussian linear mixed-effects model.

math.ST

Intermediate Tail Dependence: A Review and Some New Results

The concept of intermediate tail dependence is useful if one wants to quantify the degree of positive dependence in the tails when there is no strong evidence of presence of the usual tail dependence. We first review existing studies on intermediate tail dependence, and then we report new results to supplement the review. Intermediate tail dependence for elliptical, extreme value and Archimedean copulas are reviewed and further studied, respectively. For Archimedean copulas, we not only consider the frailty model but also the recently studied scale mixture model; for the latter, conditions leading to upper intermediate tail dependence are presented, and it provides a useful way to simulate copulas with desirable intermediate tail dependence structures.

stat.ME

Simplified Pair Copula Constructions --- Limits and Extensions

So called pair copula constructions (PCCs), specifying multivariate distributions only in terms of bivariate building blocks (pair copulas), constitute a flexible class of dependence models. To keep them tractable for inference and model selection, the simplifying assumption that copulas of conditional distributions do not depend on the values of the variables which they are conditioned on is popular. In this paper, we show for which classes of distributions such a simplification is applicable, significantly extending the discussion of Hob{\ae}k Haff et al. (2010). In particular, we show that the only Archimedean copula in dimension d \geq 4 which is of the simplified type is that based on the gamma Laplace transform or its extension, while the Student-t copula is the only one arising from a scale mixture of Normals. Further, we illustrate how PCCs can be adapted for situations where conditional copulas depend on values which are conditioned on.

stat.ME