arXiv ScienceSearch

arXiv · 2108.00480

Realised Volatility Forecasting: Machine Learning via Financial Word Embedding

Abstract

We examine whether financial news can improve realised volatility forecasting using a parsimonious NLP-based framework that incorporates specialised financial word embeddings alongside general-purpose alternatives. News-only forecasts contain useful predictive information but generally do not outperform strong volatility-history benchmarks. Crucially, combining stock-related news forecasts with a strong volatility-history benchmark lowers forecast losses for several specifications and increases realised utility, providing evidence consistent with forecast complementarity. Performance varies across news types, embedding representations, and volatility regimes. SHAP attributions associate forecast variation with economically interpretable firm-specific and macroeconomic news themes.

Explore related subjects

Keep this discovery

BibTeXRIS

Eghbal Rahimikia, Stefan Zohren, Ser-Huang Poon. 2026-08-31. Realised Volatility Forecasting: Machine Learning via Financial Word Embedding. https://doi.org/10.2139/ssrn.3895272

Cite the original work for its findings. Save a collection to share your selection of sources.

Discover connections

Connections use source metadata and explicit phrase matches, not verified experimental comparisons.

KEEP EXPLORING

Related papers

Latent-Space No-Arbitrage Geometry of Generative Models for Implied Volatility Surfaces

Generative models for implied volatility surfaces must produce outputs that satisfy static no-arbitrage constraints. We study these constraints in latent space. For a fixed generator, we assign each latent code a scalar margin determined by the no-arbitrage conditions of the generated surface. The codes with nonnegative margin form the admissible latent set. We establish conditions under which strictly admissible codes remain admissible under small perturbations and the boundary of the admissible set is characterized by zero margin. For regular boundary components, we formulate a level-set equation whose local dynamics are directed toward the zero-margin set. The analysis treats the generator as a map from latent variables to surfaces and is therefore not restricted to a particular architecture. It applies to variational autoencoders, generative adversarial networks, and other generative models with a deterministic realization map. Numerical tests recover known boundaries in analytic examples. Experiments with a variational autoencoder trained on Heston surfaces show that similar reconstruction errors can correspond to different admissible regions and that the latent prior may be concentrated inside such a region. The computed boundary can also be used to modify latent codes that generate violating surfaces.

q-fin.CP

Risk-Adjusted Harm Scoring for Automated Red Teaming for LLMs in Financial Services

Existing LLM safety evaluations rely on binary attack-success rates and domain-agnostic taxonomies, leaving regulated Banking, Financial Services, and Insurance (BFSI) deployments exposed to failures elicited through legally or professionally plausible framing. We introduce RAHS (Risk-Adjusted Harm Score), a risk-sensitive metric jointly capturing disclosure severity, disclaimer mitigation, and inter-judge agreement, and FinRedTeamBench, a 989-prompt benchmark spanning seven BFSI risk areas and 34 sub-categories mapped to regulatory frameworks. Evaluation uses an ensemble of three heterogeneous LLM judges, validated against human experts, and an adaptive multi-turn red-teaming pipeline. On nine open-weight models, RAHS preserves separation under near-ceiling ASR, ranking is stable under hyperparameter sweeps, and multi-turn pressure drives not only more jailbreaks but more operationally severe disclosures, exposing failure modes that single-turn, domain-agnostic evaluations cannot reveal.

q-fin.CP

Single- and Multilevel Quadrature with Error Control for Fourier Pricing under the Rough Heston Model

Unlike the classical Heston model, Fourier pricing under the rough Heston model requires solving a fractional Riccati equation at every quadrature point. Since the required resolution varies with model parameters and quadrature point, a single uniform time discretization can be inefficient. We develop single- and multilevel Gauss-Laguerre quadrature methods that balance the time discretization and Fourier quadrature errors. Both methods scale the laguerre weight to the estimated Fourier integrand decay. The single-level method allocates a prescribed tolerance between the two errors. The multilevel method splits the integrand into a level-zero term and level differences, selecting quadrature points separately at each level. Suppose that the Fourier integrand discretization error is $O(Δt^p)$, that evaluating the characteristic function once costs $O(Δt^{-β})$, and that the algebraic Gauss-Laguerre quadrature error is $O(N^{-s_{SL}/2})$, where $s_{SL}$ is the smoothness index. Under this estimate and assumptions on the regularity and decay of level differences, we prove that the proposed single-level method requires $O(ε^{-(β/p+2/s_{SL})})$ computational work to achieve accuracy $ε$, whereas the proposed multilevel method requires $O(ε^{-β/p})$ computational work. We also study root-exponential Gauss-Laguerre error models for practical multilevel quadrature allocation. Numerical experiments support the observed fractional Riccati and Fourier integrand convergence rates and root-exponential quadrature behavior, and show substantial reductions in quadrature cost from the proposed scaling. The multilevel method provides clear computational savings over the single-level method. We further benchmark the multilevel fractional Riccati method against the BL2 Markovian approximation and report lower total CPU time in the tested configurations.

q-fin.CP