arXiv ScienceSearch

arXiv subjects

Giulia Livieri

Publications and source records attributed to Giulia Livieri.

At least 19 recordsLinked to original sources

Online Learning of Scale Parameters in Score-Driven Filters

A score-driven filter multiplies its scaled log-likelihood score by a scale parameter. We call this coefficient the gain and learn it online. Given the current state and realised scaled score, each admissible gain selects a reachable next state and predictive density. A scalar gain moves along a line; diagonal gains control coordinatewise transmission and may change direction. We evaluate gain selection using a one-step predictive Kullback--Leibler objective. In the scalar unscaled case, the negative consecutive-score product is a stochastic gradient; the positive product used in accelerated recursions is a descent direction. Positive scalar score scaling changes only the effective learning rate. Monotone differentiable gain links induce mirror-descent geometry, while persistence adds a Bregman pull towards a reference gain. Under convexity, compactness, integrability, and schedule conditions, projected and discounted mirror updates satisfy dynamic-regret bounds relative to time-varying, current-information comparators. Simulations isolate score scaling, link geometry, persistence, and coordinatewise gains. Across twelve equity indices, the bounded discounted-logistic gain records a lower out-of-sample mean negative log score than the constant gain in eleven markets, although market-level evidence is mixed. It also avoids the extreme transients of the numerically capped exponential-link benchmark. Improvements are largest in markets spanning multiple crises.

cs.LG

Proper-score observation-driven filters: local geometry, estimation, and continuous-time limits

Observation-driven filters usually use the likelihood score, tying their updates to the logarithmic scoring rule. We study recursions driven instead by the negative parameter derivative of a differentiable proper scoring rule, under a declared working family and predictable scaling. This separates three roles: the rule determines the conditional risk projection and tail response; the scaling converts its derivative into the implemented update; and the autoregressive component determines the composite dynamic centre. We characterize local mean reversion and update noise through risk curvature and innovation variance, which coincide for the log score under the information identity but generally differ. For high-frequency scale models, centred updates converge to diffusions, whereas non-centred updates follow deterministic mean flows. When the rule-specific projection moves over time, the local tracking error admits an Ornstein--Uhlenbeck approximation, valid up to a stopping time and allowing update and target shocks to be correlated. For static parameters, we establish consistency and asymptotic normality under explicit stability and fixed-tuning conditions. Simulations and an international-equity application illustrate how criterion choice affects robustness, adaptation, variance-forecast loss, value-at-risk calibration, and probability-integral-transform diagnostics. The results show that predictive density and updating criterion are distinct design choices whose relative performance depends on the target and disturbance.

math.ST

NeuralChaos: Optimal Adapted Approximation of Square Integrable Predictable Processes

We address fundamental challenges in representing and computing $\mathbb{R}^{d}$-valued predictable square-integrable processes over $[0,T]$, collected in the space $\mathcal{H}^2_T(\mathbb{R}^{d})$. These processes are central to continuous-time stochastic control, reinforcement learning, and mathematical finance. Although Wiener-chaos expansions offer strong theoretical tools, traditional computational methods are hindered by the need for large chaos dictionaries and high-order iterated integrals. To overcome these obstacles, we introduce NeuralChaos -- a neural operator architecture that produces elements of $\mathcal{H}^2_T(\mathbb{R}^{d})$ using only finitely many evaluations of the driving Brownian motion, while preserving predictability and square-integrability. We prove that NeuralChaos is dense in $\mathcal{H}^2_T(\mathbb{R}^{d})$ and achieves the best $N$-term chaoslet approximation rates for compressible and Malliavin--Sobolev regular processes. Moreover, compressibility is shown to be typical for processes from $\mathcal{H}^2_T(\mathbb{R}^{d})$ under non-degenerate sub-Gaussian sampling. In contrast, we show that finite-dimensional Markovian neural SDE models constitute a meagre and Gaussian-null subset in $\mathcal{H}^2_T(\mathbb{R}^{d})$, regardless of discretization, whereas compressible processes are generic. Numerical experiments on a stochastic optimal control problem and dynamic hedging highlight the practical effectiveness of our approach. Our results enable more efficient and expressive modelling in stochastic analysis and mathematical finance.

math.PR

When large trades are not (automatically) news: liquidity tail risk and price discovery

We examine how heavy-tailed liquidity demand changes price discovery in a sequential limit order book with asymmetric information. In our setting, liquidity suppliers observe aggregate order flow, not its decomposition into informed demand and uninformed liquidity shocks. With heavy-tailed uninformed aggregated order flow, large trades remain plausibly uninformed over a wider range of depths, flattening price impact and slowing learning; sufficiently extreme trades can nevertheless become informative. We characterize equilibrium through a non-linear fixed point equation for the marginal-cost schedule; heavy-tailed uninformed aggregated order flow invalidates the monotonicity and compactness arguments available under Gaussianity. Therefore, we establish fixed-point existence within a tail-controlled class, prove posterior consistency for liquidity suppliers in the presence of endogenous dependent order flow, and derive tail asymptotics for marginal costs, informed demand, and aggregate order flow. Additionally, we obtain eventual informed-demand dominance and eventual monotonicity of the book in the far tails. Empirically, using 10-level AAPL data, we document farther-out crossover diagnostics and persistent bid-ask spreads following large heavy-tailed trades.

q-fin.TR

BASIS: Batchwise Advantage Estimation from Single-Rollout Information Sharing for LLM Reasoning

Reinforcement learning with verifiable rewards has become a standard recipe for improving the reasoning abilities of large language models. Existing algorithms face a tradeoff between computational efficiency and sample efficiency in value estimation and policy learning. We introduce BASIS, a critic-free post-training algorithm designed to address this tradeoff. At each online training step, BASIS samples only one rollout per prompt, but leverages rich information across prompts in the entire batch to improve value function estimation. Our experiments demonstrate that BASIS reduces MSE in value function estimation by 69% compared to REINFORCE++, a representative single-rollout baseline, and achieves lower MSE with one rollout than group mean estimators with 8 rollouts. This improvement in value estimation translates to better policy optimization: using substantially less training time, BASIS achieves performance close to multi-rollout GRPO-type baselines and often outperforms single-rollout REINFORCE-type baselines.

cs.LG

READER: Reasoning-Enhanced AI-Generated Text Detection

Recent advances in large language models (LLMs) have made it increasingly difficult to distinguish human-written text from AI-generated content. Many existing detectors train supervised neural classifiers that achieve strong in-distribution performance but are often opaque and can degrade substantially under distribution shift. We present READER, a reasoning-enhanced AI text detector that outputs both a human/AI label and a structured rationale describing the evidence for its decision. A key component of our approach is READ, a curated supervision set of rationales and verdicts. We fine-tune an LLM on READ to build READER, which reasons before detecting at inference time. Despite having only 1.5B parameters, READER consistently outperforms existing detectors as well as prompted, high-capacity LLM baselines (GPT-5.2, Gemini-3-Pro, and DeepSeek-V3.2), which are 100 to 1000 times larger in scale.

cs.CL

Integral equations for the optimal boundary surface of a mean-field game of capacity expansion

We prove that the optimal boundary surface that splits the action and inaction regions in a mean-field game of capacity expansion studied in (Campi et al.,\ Ann.\ Appl.\ Probab.,\ {\bf 32}(5),\, pp.\,3674-3717, 2022) is the unique continuous solution of a nonlinear integral equation of Volterra type. In order to do that, we first establish continuity of the optimal surface. Then we develop an extension of It\^o's formula which weakens assumptions required in the existing literature on the first-order time-derivative and/or second-order space derivative of the value function. The paper also provides an algorithm for the numerical solution of the integral equation and we compute optimal controls numerically for the mean-field game.

math.OC

Statistical Guarantees for Reasoning Probes on Looped Boolean Circuits

We study the statistical behavior of reasoning probes in a stylized model of iterative computation inspired by neural algorithmic reasoning. The underlying computation is given by a looped Boolean circuit whose graph is a perfect $\nu$-ary tree ($\nu\ge 2$), with outputs recursively fed back as inputs across computation rounds. A probe observes a sampled subset of internal nodes and seeks to infer the latent operation at each node, represented as a probability distribution over a finite set of admissible Boolean gates. This partial observability induces a transductive generalization problem on a structured computation graph. We show that when the probe is parameterized by a graph convolutional network and queries $N$ nodes, the worst-case generalization error decays at the optimal rate $\mathcal{O}(\sqrt{\log(2/\delta)}/\sqrt{N})$ with probability at least $1-\delta$. Our analysis combines metric embedding techniques with tools from optimal transport. A key insight is that this rate is achievable independently of the size of the computation graph, enabled by a low-distortion one-dimensional snowflake embedding of the induced graph metric. These results highlight a geometric mechanism underlying statistical efficiency in probing structured, iterative computations.

stat.ML

Competition among seaports through Mean Field Games and real-world data

This paper presents a Mean Field Game (MFG) model for maritime traffic flow, treating the navigation of ships between seaports as a large-scale stochastic control problem. The MFG framework enables the modeling of agents at a microscopic level as rational decision-makers who seek to optimize their utility, thereby translating complex microscopic behaviors into macroscopic models. We build upon this MFG framework to develop a mesoscopic-scale MFG model that defines the payoff and cost functions for a coordinator at each seaport considered in our study. The coordinator determines the routes taken by ships transporting goods between ports by evaluating several key factors: transportation costs, expected profit margins from loading specific goods at the seaports and unloading them at various destinations, and a congestion term that reflects the costs associated with accessing the destination port. We derive an explicit solution for the stationary version of the model under certain approximations and establish conditions necessary to ensure the uniqueness of the corresponding Mean Field Equilibrium (MFE). Furthermore, we introduce a statistical methodology to infer the parameters of the game from real-world data, specifically focusing on costs and the components of expected commercial margins. To validate our model in a real-world context, we analyze the ShipFix dataset of daily ''Dry Coal'' shipments worldwide from 2015 to 2025. Our discussion highlights the influence of empirical traffic flow on various components of costs. We believe that this research represents a significant advancement in the application of MFGs for effective maritime traffic management and offers valuable insights for practitioners in the field.

math.OC

Learning from one graph: transductive learning guarantees via the geometry of small random worlds

Since their introduction by Kipf and Welling in $2017$, a primary use of graph convolutional networks is transductive node classification, where missing labels are inferred within a single observed graph and its feature matrix. Despite the widespread use of the network model, the statistical foundations of transductive learning remain limited, as standard inference frameworks typically rely on multiple independent samples rather than a single graph. In this work, we address these gaps by developing new concentration-of-measure tools that leverage the geometric regularities of large graphs via low-dimensional metric embeddings. The emergent regularities are captured using a random graph model; however, the methods remain applicable to deterministic graphs once observed. We establish two principal learning results. The first concerns arbitrary deterministic $k$-vertex graphs, and the second addresses random graphs that share key geometric properties with an Erd\H{o}s-R\'{e}nyi graph $\mathbf{G}=\mathbf{G}(k,p)$ in the regime $p \in \mathcal{O}((\log (k)/k)^{1/2})$. The first result serves as the basis for and illuminates the second. We then extend these results to the graph convolutional network setting, where additional challenges arise. Lastly, our learning guarantees remain informative even with a few labelled nodes $N$ and achieve the optimal nonparametric rate $\mathcal{O}(N^{-1/2})$ as $N$ grows.

stat.ML

A Mean Field Game approach for pollution regulation of competitive firms

We develop a model based on mean-field games of competitive firms producing similar goods according to a standard AK model with a depreciation rate of capital generating pollution as a byproduct. Our analysis focuses on the widely-used cap-and-trade pollution regulation. Under this regulation, firms have the flexibility to respond by implementing pollution abatement, reducing output, and participating in emission trading, while a regulator dynamically allocates emission allowances to each firm. The resulting mean-field game is of linear quadratic type and equivalent to a mean-field type control problem, i.e., it is a potential game. We find explicit solutions to this problem through the solutions to differential equations of Riccati type. Further, we investigate the carbon emission equilibrium price that satisfies the market clearing condition and find a specific form of FBSDE of McKean-Vlasov type with common noise. The solution to this equation provides an approximate equilibrium price. Additionally, we demonstrate that the degree of competition is vital in determining the economic consequences of pollution regulation.

q-fin.MF

Low-dimensional approximations of the conditional law of Volterra processes: a non-positive curvature approach

Predicting the conditional evolution of Volterra processes with stochastic volatility is a crucial challenge in mathematical finance. While deep neural network models offer promise in approximating the conditional law of such processes, their effectiveness is hindered by the curse of dimensionality caused by the infinite dimensionality and non-smooth nature of these problems. To address this, we propose a two-step solution. Firstly, we develop a stable dimension reduction technique, projecting the law of a reasonably broad class of Volterra process onto a low-dimensional statistical manifold of non-positive sectional curvature. Next, we introduce a sequentially deep learning model tailored to the manifold's geometry, which we show can approximate the projected conditional law of the Volterra process. Our model leverages an auxiliary hypernetwork to dynamically update its internal parameters, allowing it to encode non-stationary dynamics of the Volterra process, and it can be interpreted as a gating mechanism in a mixture of expert models where each expert is specialized at a specific point in time. Our hypernetwork further allows us to achieve approximation rates that would seemingly only be possible with very large networks.

math.NA

Machine-learning regression methods for American-style path-dependent contracts

Evaluating financial products with early-termination clauses, in particular those with path-dependent structures, is challenging. This paper focuses on Asian options, look-back options, and callable certificates. We will compare regression methods for pricing and computing sensitivities, highlighting modern machine learning techniques against traditional polynomial basis functions. Specifically, we will analyze randomized recurrent and feed-forward neural networks, along with a novel approach using signatures of the underlying price process. For option sensitivities like Delta and Gamma, we will incorporate Chebyshev interpolation. Our findings show that machine learning algorithms often match the accuracy and efficiency of traditional methods for Asian and look-back options, while randomized neural networks are best for callable certificates. Furthermore, we apply Chebyshev interpolation for Delta and Gamma calculations for the first time in Asian options and callable certificates.

q-fin.PR

Understanding the householder solar panel consumer: A Markovian model and its societal implications

Household adoption of rooftop photovoltaic (PV) systems is central to the green energy transition, yet diffusion depends on social influence and behavioral biases, as well as payback economics. This study develops a parsimonious Markovian model in which households move sequentially from being unengaged (Carbon) to informed, to planning, and finally to adoption (Green). Transition rates are micro-founded by two mechanisms: (i) social contagion/communication, proxied by the current share of adopters, and (ii) economic profitability, proxied by payback time computed from a Net Present Value framework. Novel to this diffusion setting, bounded rationality is introduced via hyperbolic discounting, creating a procrastination loop that delays adoption even when PV is economically attractive in a long-run perspective. Calibrated on the Italian residential PV diffusion path (2006-2020) and assessed in national and regional applications, the model reproduces observed trajectories and enables forward-looking scenario analysis (2020-2026). Results show that policies yielding similar payback improvements can produce different outcomes once present bias is accounted for and that behaviorally informed intervention are stronger. The findings contribute a micro-to-macro bridge between behavioral economics and technology diffusion modeling and imply that effective policy portfolios (and PV business models) should complement incentives with commitment devices and social-norm peer strategies to accelerate PV uptake and its spillover emissions benefits.

physics.soc-ph

Pricing Transition Risk with a Jump-Diffusion Credit Risk Model: Evidences from the CDS market

Transition risk can be defined as the business-risk related to the enactment of green policies, aimed at driving the society towards a sustainable and low-carbon economy. In particular, the value of certain firms' assets can be lower because they need to transition to a less carbon-intensive business model. In this paper we derive formulas for the pricing of defaultable coupon bonds and Credit Default Swaps to empirically demonstrate that a jump-diffusion credit risk model in which the downward jumps in the firm value are due to tighter green laws can capture, at least partially, the transition risk. The empirical investigation consists in the model calibration on the CDS term-structure, performing a quantile regression to assess the relationship between implied prices and a proxy of the transition risk. Additionally, we show that a model without jumps lacks this property, confirming the jump-like nature of the transition risk.

q-fin.PR

Structural properties in the diffusion of the solar photovoltaic in Italy: individual people/householder vs firms

This paper develops two mathematical models to understand subjects' behavior in response to the urgency of a change and inputs from governments e.g., (subsides) in the context of the diffusion of the solar photovoltaic in Italy. The first model is a Markovian model of interacting particle systems. The second one, instead, is a Mean Field Game model. In both cases, we derive the scaling limit deterministic dynamics, and we compare the latter to the Italian solar photovoltaic data. We identify periods where the first model describes the behavior of domestic data well and a period where the second model captures a particular feature of data corresponding to companies. The comprehensive analysis, integrated with a philosophical inquiry focusing on the conceptual vocabulary and correlative implications, leads to the formulation of hypotheses about the efficacy of different forms of governmental subsidies.

physics.soc-ph

Designing Universal Causal Deep Learning Models: The Case of Infinite-Dimensional Dynamical Systems from Stochastic Analysis

Several non-linear operators in stochastic analysis, such as solution maps to stochastic differential equations, depend on a temporal structure which is not leveraged by contemporary neural operators designed to approximate general maps between Banach space. This paper therefore proposes an operator learning solution to this open problem by introducing a deep learning model-design framework that takes suitable infinite-dimensional linear metric spaces, e.g. Banach spaces, as inputs and returns a universal \textit{sequential} deep learning model adapted to these linear geometries specialized for the approximation of operators encoding a temporal structure. We call these models \textit{Causal Neural Operators}. Our main result states that the models produced by our framework can uniformly approximate on compact sets and across arbitrarily finite-time horizons H\"older or smooth trace class operators, which causally map sequences between given linear metric spaces. Our analysis uncovers new quantitative relationships on the latent state-space dimension of Causal Neural Operators, which even have new implications for (classical) finite-dimensional Recurrent Neural Networks. In addition, our guarantees for recurrent neural networks are tighter than the available results inherited from feedforward neural networks when approximating dynamical systems between finite-dimensional spaces.

math.DS

One-Shot Learning of Stochastic Differential Equations with Data Adapted Kernels

We consider the problem of learning Stochastic Differential Equations of the form $dX_t = f(X_t)dt+\sigma(X_t)dW_t $ from one sample trajectory. This problem is more challenging than learning deterministic dynamical systems because one sample trajectory only provides indirect information on the unknown functions $f$, $\sigma$, and stochastic process $dW_t$ representing the drift, the diffusion, and the stochastic forcing terms, respectively. We propose a method that combines Computational Graph Completion and data adapted kernels learned via a new variant of cross validation. Our approach can be decomposed as follows: (1) Represent the time-increment map $X_t \rightarrow X_{t+dt}$ as a Computational Graph in which $f$, $\sigma$ and $dW_t$ appear as unknown functions and random variables. (2) Complete the graph (approximate unknown functions and random variables) via Maximum a Posteriori Estimation (given the data) with Gaussian Process (GP) priors on the unknown functions. (3) Learn the covariance functions (kernels) of the GP priors from data with randomized cross-validation. Numerical experiments illustrate the efficacy, robustness, and scope of our method.

stat.ML