arXiv ScienceSearch

arXiv subjects

Paolo Barucca

Publications and source records attributed to Paolo Barucca.

At least 19 recordsLinked to original sources

Information-Based Calibration of Uncertainty Quantification in Product-of-Experts Gaussian Process Models

Gaussian process (GP) regression with a single global GP (GP-glo) incurs cubic computational cost, limiting scalability to large datasets. Product-of-experts GP models (GP-pro), which combine local GP models to capture global correlations, alleviate this computational burden. However, training local experts on disjoint data subsets can lead to overestimated posterior variances. We propose GP-pro-c, a product-of-experts GP model that calibrates these variances using an information-based method. The method exploits the monotonicity and submodularity of information gain in GPs to define a calibration ratio that reduces the posterior variance of individual local GP models. We evaluate GP-pro-c using negative log-likelihood (NLL), root mean squared error (RMSE), and expected normalised calibration error (ENCE). Experiments on four synthetic functions and six regression datasets show that GP-pro-c achieves average reductions of 2.3% in NLL and 12.0% in ENCE compared with the uncalibrated GP-pro model. The proposed method mitigates posterior variance overestimation while maintaining predictive accuracy and reducing computational complexity. GP-pro-c provides a promising approach for uncertainty estimation in scalable GP models and may serve as a useful surrogate model for Bayesian optimisation with high-dimensional and large-scale data.

cs.LG

Characterizing Optimizer-Dependent Training Dynamics Through Hessian Eigenvector Displacement and Localization

Hessian spectral properties are a standard tool in analysing neural-network training, with eigenvalues linked to sharpness, generalization, and optimization dynamics. Eigenvalues quantify curvature magnitude, while eigenvectors identify which parameters generate that curvature. In this work, we study how the leading Hessian eigenvectors evolve during training and how they affect the learning trajectories. We track the training dynamics of multilayer perceptrons on a classification problem and measure eigenvector dynamics through two complementary statistics: (i) displacement over time, inspired by analyses of glassy systems, and (ii) localization via the inverse participation ratio. The metrics are compared against a random null model of the Hessian induced by the architecture. Our results reveal clear optimizer-dependent behaviour. SGD leads to progressively more stable leading curvature directions, while Adam exhibits substantially stronger reorganization of eigenvectors throughout training. We also observe a localization phenomenon under Adam, where a small subset of parameters contributes disproportionately to the leading curvature directions. These results suggest that Hessian eigenvector dynamics capture key differences in optimizer behaviour and the resulting training trajectories.

cs.LG

Maximum entropy temporal networks

Temporal networks consist of timestamped directed interactions that may appear continuously in time, yet few studies have directly tackled the continuous-time modeling of networks. Here, we introduce a maximum-entropy approach to temporal networks and with basic assumptions on constraints, the corresponding network ensembles admit a modular and interpretable representation: a set of global time processes and a static maximum-entropy edge, e.g. node pair, probability. This time-edge labels factorization yields closed-form log-likelihoods, degree, clustering and motif expectations, and yields a whole class of effective generative models. We provide the maximum-entropy derivation for the non-homogeneous Poisson Process (NHPP) intensities governing the probability of directed edges in temporal networks via the functional optimization over path entropy, connecting NHPP modeling to maximum-entropy network ensembles. NHPPs consistently improve log-likelihood over generic Poisson processes, while the maximum-entropy edge labels recover strength constraints and reproduce expected unique-degree curves. We discuss the limitations of this framework and how it can be integrated with multivariate Hawkes calibration procedures, renewal theory, and neural kernel estimation in graph neural networks.

cs.SI

Fundamental Limits of Stability Inference in High-Dimensional Complex Systems

Many complex systems, including ecosystems, neural circuits, and financial markets, are inferred to operate close to a threshold of instability, at which a small perturbation can propagate across the entire system. This proximity is often interpreted as functionally advantageous, yet it poses a question common to all these fields: from a finite, noisy recording, how precisely can the distance of a system from that threshold be estimated? Using the multivariate Ornstein-Uhlenbeck process as the canonical linear model of relaxation near a stable fixed point, we show that the attainable precision is governed by three factors: an effective measurement budget, set by the number of samples relative to the system dimension and the sampling interval; the signal-to-noise ratio, given by the magnitude of deterministic interactions relative to stochastic forcing; and the distance to criticality, which simultaneously sets the system's correlation times and degrades both of the preceding factors. As the slowest dynamical mode softens near the threshold, the curvature of the log-likelihood flattens along the direction that determines stability, so that the relative uncertainty on the estimated distance diverges as that distance vanishes. Critically, temporal correlations near instability reduce the effective number of independent observations far below the nominal sample count, and inference breaks down when this effective count falls below the system dimension, even when the raw data volume appears sufficient. A direct consequence is the existence of an optimal sampling interval that diverges as the system approaches criticality, with practical implications for experimental design.

cond-mat.dis-nn

Physics-Informed Neural Networks for Solving Derivative-Constrained PDEs

Physics-Informed Neural Networks (PINNs) recast PDE solving as an optimisation problem in function space by minimising a residual-based objective, yet many applications require additional derivative-based relations that are just as fundamental as the governing equations. In this paper, we present Derivative-Constrained PINNs (DC-PINNs), a general framework that treats constrained PDE solving as an optimisation guided by a minimum objective function criterion where the physics resides in the minimum principle. DC-PINNs embed general nonlinear constraints on states and derivatives, e.g., bounds, monotonicity, convexity, incompressibility, computed efficiently via automatic differentiation, and they employ self-adaptive loss balancing to tune the influence of each objective, reducing reliance on manual hyperparameters and problem-specific architectures. DC-PINNs consistently reduce constraint violations and improve physical fidelity versus baseline PINN variants, representative hard-constraint formulations on benchmarks, including heat diffusion with bounds, financial volatilities with arbitrage-free, and fluid flow with vortices shed. Explicitly encoding derivative constraints stabilises training and steers optimisation toward physically admissible minima even when the PDE residual alone is small, providing reliable solutions of constrained PDEs grounded in energy minimum principles.

cs.LG

Parsimonious Hawkes Processes for temporal networks modelling

Temporal networks are characterised by interdependent link events between nodes, forming ordered sequences of links that may represent specific information flows in the system. Nevertheless, representing temporal networks using discrete snapshots in time partially cancels the effect of time-ordered links on each other, while continuous time models, such as Poisson or Hawkes processes, can describe the full influence between all the potential pairs of links at all times. In this paper, we introduce a continuous Hawkes temporal network model which accounts both for a community structure of the aggregate network and a strong heterogeneity in the activity of individual nodes, thus accounting for the presence of highly heterogeneous clusters with isolated high-activity influencer nodes, communities and low-activity nodes. Our model improves the prediction performance of previously available continuous time network models, and obtains a systematic increase in log-likelihood. Characterising the direct interaction between influencer nodes and communities, we can provide a more detailed description of the system that can better outline the sequence of activations in the components of the systems represented by temporal networks.

cs.SI

How low-cost AI universal approximators reshape market efficiency

The efficient market hypothesis (EMH) famously stated that prices fully reflect the information available to traders. This critically depends on the transfer of information into prices through trading strategies. Traders optimise their strategy with models of increasing complexity that identify the relationship between information and profitable trades more and more accurately. Under specific conditions, the increased availability of low-cost universal approximators, such as AI systems, should be naturally pushing towards more advanced trading strategies, potentially making it harder and harder for inefficient traders to profit. In this paper, we leverage on a generalised notion of market efficiency, based on the definition of an equilibrium price process, that allows us to distinguish different levels of model complexity through investors' beliefs, and trading strategies optimisation, and discuss the relationship between AI-powered trading and the time-evolution of market efficiency. Finally, we outline the need for and the challenge of describing out-of-equilibrium market dynamics in an adaptive multi-agent environment.

q-fin.MF

Granger Causality Detection with Kolmogorov-Arnold Networks

Discovering causal relationships in time series data is central in many scientific areas, ranging from economics to climate science. Granger causality is a powerful tool for causality detection. However, its original formulation is limited by its linear form and only recently nonlinear machine-learning generalizations have been introduced. This study contributes to the definition of neural Granger causality models by investigating the application of Kolmogorov-Arnold networks (KANs) in Granger causality detection and comparing their capabilities against multilayer perceptrons (MLP). In this work, we develop a framework called Granger Causality KAN (GC-KAN) along with a tailored training approach designed specifically for Granger causality detection. We test this framework on both Vector Autoregressive (VAR) models and chaotic Lorenz-96 systems, analysing the ability of KANs to sparsify input features by identifying Granger causal relationships, providing a concise yet accurate model for Granger causality detection. Our findings show the potential of KANs to outperform MLPs in discerning interpretable Granger causal relationships, particularly for the ability of identifying sparse Granger causality patterns in high-dimensional settings, and more generally, the potential of AI in causality discovery for the dynamical laws in physical systems.

cs.LG

How the interplay between power concentration, competition, and propagation affects the resource efficiency of distributed ledgers

Forks in the Bitcoin network result from the natural competition in the blockchain's Proof-of-Work consensus protocol. Their frequency is a critical indicator for the efficiency of a distributed ledger as they can contribute to resource waste and network insecurity. We introduce a model for the estimation of natural fork rates in a network of heterogeneous miners as a function of their number, the distribution of hash rates and the block propagation time over the peer-to-peer infrastructure. Despite relatively simplistic assumptions, such as zero propagation delay within mining pools, the model predicts fork rates which are comparable with the empirical stale blocks rate. In the past decade, we observe a reduction in the number of mining pools approximately by a factor 3, and quantify its consequences for the fork rate, whilst showing the emergence of a truncated power-law distribution in hash rates, justified by a rich-get-richer effect constrained by global energy supply limits. We demonstrate, both empirically and with the aid of our quantitative model, that the ratio between the block propagation time and the mining time is a sufficiently accurate estimator of the fork rate, but also quantify its dependence on the heterogeneity of miner activities. We provide empirical and theoretical evidence that both hash rate concentration and lower block propagation time reduce fork rates in distributed ledgers. Our work introduces a robust mathematical setting for investigating power concentration and competition on a distributed network, for interpreting discrepancies in fork rates -- for example caused by selfish mining practices and asymmetric propagation times -- thus providing an effective tool for designing future and alternative scenarios for existing and new blockchain distributed mining systems.

cs.DC

Whack-a-mole Online Learning: Physics-Informed Neural Network for Intraday Implied Volatility Surface

Calibrating the time-dependent Implied Volatility Surface (IVS) using sparse market data is an essential challenge in computational finance, particularly for real-time applications. This task requires not only fitting market data but also satisfying a specified partial differential equation (PDE) and no-arbitrage conditions modelled by differential inequalities. This paper proposes a novel Physics-Informed Neural Networks (PINNs) approach called Whack-a-mole Online Learning (WamOL) to address this multi-objective optimisation problem. WamOL integrates self-adaptive and auto-balancing processes for each loss term, efficiently reweighting objective functions to ensure smooth surface fitting while adhering to PDE and no-arbitrage constraints and updating for intraday predictions. In our experiments, WamOL demonstrates superior performance in calibrating intraday IVS from uneven and sparse market data, effectively capturing the dynamic evolution of option prices and associated risk profiles. This approach offers an efficient solution for intraday IVS calibration, extending PINNs applications and providing a method for real-time financial modelling.

q-fin.CP

Random matrix ensemble for the covariance matrix of Ornstein-Uhlenbeck processes with heterogeneous temperatures

We introduce a random matrix model for the stationary covariance of multivariate Ornstein-Uhlenbeck processes with heterogeneous temperatures, where the covariance is constrained by the Sylvester-Lyapunov equation. Using the replica method, we compute the spectral density of the equal-time covariance matrix characterizing the stationary states, demonstrating that this model undergoes a transition between stable and unstable states. In the stable regime, the spectral density has a finite and positive support, whereas negative eigenvalues emerge in the unstable regime. We determine the critical line separating these regimes and show that the spectral density exhibits a power-law tail at marginal stability, with an exponent independent of the temperature distribution. Additionally, we compute the spectral density of the lagged covariance matrix characterizing the stationary states of linear transformations of the original dynamical variables. Our random-matrix model is potentially interesting to understand the spectral properties of empirical correlation matrices appearing in the study of complex systems.

cond-mat.dis-nn

Predicting public market behavior from private equity deals

We process private equity transactions to predict public market behavior with a logit model. Specifically, we estimate our model to predict quarterly returns for both the broad market and for individual sectors. Our hypothesis is that private equity investments (in aggregate) carry predictive signal about publicly traded securities. The key source of such predictive signal is the fact that, during their diligence process, private equity fund managers are privy to valuable company information that may not yet be reflected in the public markets at the time of their investment. Thus, we posit that we can discover investors' collective near-term insight via detailed analysis of the timing and nature of the deals they execute. We evaluate the accuracy of the estimated model by applying it to test data where we know the correct output value. Remarkably, our model performs consistently better than a null model simply based on return statistics, while showing a predictive accuracy of up to 70% in sectors such as Consumer Services, Communications, and Non Energy Minerals.

q-fin.CP

No-Arbitrage Deep Calibration for Volatility Smile and Skewness

Volatility smile and skewness are two key properties of option prices that are represented by the implied volatility (IV) surface. However, IV surface calibration through nonlinear interpolation is a complex problem due to several factors, including limited input data, low liquidity, and noise. Additionally, the calibrated surface must obey the fundamental financial principle of the absence of arbitrage, which can be modeled by various differential inequalities over the partial derivatives of the option price with respect to the expiration time and the strike price. To address these challenges, we have introduced a Derivative-Constrained Neural Network (DCNN), which is an enhancement of a multilayer perceptron (MLP) that incorporates derivatives in the objective function. DCNN allows us to generate a smooth surface and incorporate the no-arbitrage condition thanks to the derivative terms in the loss function. In numerical experiments, we train the model using prices generated with the SABR model to produce smile and skewness parameters. We carry out different settings to examine the stability of the calibrated model under different conditions. The results show that DCNNs improve the interpolation of the implied volatility surface with smile and skewness by integrating the computation of the derivatives, which are necessary and sufficient no-arbitrage conditions. The developed algorithm also offers practitioners an effective tool for understanding expected market dynamics and managing risk associated with volatility smile and skewness.

q-fin.CP

Online Learning with Radial Basis Function Networks

Financial time series are characterised by their nonstationarity and autocorrelation. Even if these time series are differenced, technically ensuring their stationarity, they experience regular covariate shifts and concept drifts. Against this backdrop, we combine feature representation transfer with sequential optimisation to provide multi-horizon returns forecasts. Our online learning rbfnet outperforms a random-walk baseline and several powerful batch learners. The rbfnets we formulate are naturally designed to measure the similarity between test samples and continuously updated prototypes that capture the characteristics of the feature space.

cs.CE

Structural importance and evolution: an application to financial transaction networks

A fundamental problem in the study of networks is the identification of important nodes. This is typically achieved using centrality metrics, which rank nodes in terms of their position in the network. This approach works well for static networks, that do not change over time, but does not consider the dynamics of the network. Here we propose instead to measure the importance of a node based on how much a change to its strength will impact the global structure of the network, which we measure in terms of the spectrum of its adjacency matrix. We apply our method to the identification of important nodes in equity transaction networks, and we show that, while it can still be computed from a static network, our measure is a good predictor of nodes subsequently transacting. This implies that static representations of temporal networks can contain information about their dynamics.

cs.CE

Modelling Equity Transaction Networks as Bursty Processes

Trade executions for major stocks come in bursts of activity, which can be partly attributed to the presence of self- and mutual excitations endogenous to the system. In this paper, we study transaction reports for five FTSE 100 stocks. We model the dynamic of transactions between counterparties using both univariate and multivariate Hawkes processes, which we fit to the data using a parametric approach. We find that the frequency of transactions between counterparties increases the likelihood of them to transact in the future, and that univariate and multivariate Hawkes processes show promise as generative models able to reproduce the bursty, hub dominated systems that we observe in the real world. We further show that Hawkes processes perform well when used to model buys and sells through a central clearing counterparty when considered as a bivariate process, but not when these are modelled as individual univariate processes, indicating that mutual excitation between buys and sells is present in these markets.

cs.CE

Deep Reinforcement Learning for Optimal Investment and Saving Strategy Selection in Heterogeneous Profiles: Intelligent Agents working towards retirement

The transition from defined benefit to defined contribution pension plans shifts the responsibility for saving toward retirement from governments and institutions to the individuals. Determining optimal saving and investment strategy for individuals is paramount for stable financial stance and for avoiding poverty during work-life and retirement, and it is a particularly challenging task in a world where form of employment and income trajectory experienced by different occupation groups are highly diversified. We introduce a model in which agents learn optimal portfolio allocation and saving strategies that are suitable for their heterogeneous profiles. We use deep reinforcement learning to train agents. The environment is calibrated with occupation and age dependent income evolution dynamics. The research focuses on heterogeneous income trajectories dependent on agent profiles and incorporates the behavioural parameterisation of agents. The model provides a flexible methodology to estimate lifetime consumption and investment choices for heterogeneous profiles under varying scenarios.

q-fin.PM

Reinforcement Learning for Systematic FX Trading

We explore online inductive transfer learning, with a feature representation transfer from a radial basis function network formed of Gaussian mixture model hidden processing units to a direct, recurrent reinforcement learning agent. This agent is put to work in an experiment, trading the major spot market currency pairs, where we accurately account for transaction and funding costs. These sources of profit and loss, including the price trends that occur in the currency markets, are made available to the agent via a quadratic utility, who learns to target a position directly. We improve upon earlier work by targeting a risk position in an online transfer learning context. Our agent achieves an annualised portfolio information ratio of 0.52 with a compound return of 9.3\%, net of execution and funding cost, over a 7-year test set; this is despite forcing the model to trade at the close of the trading day at 5 pm EST when trading costs are statistically the most expensive.

q-fin.TR