arXiv Science⌕ Search

arXiv · 2610.03598

When a Correct Reward Is Not Enough: Diagnosing and Guiding PPO in an Analytically Solved Broker-Trader Game

Abstract

Reinforcement learning (RL) is increasingly used for financial optimal-control problems when complex dynamics make analytical strategies difficult to obtain. There are financial mathematics literactures which provides many solved models whose equations and controls could evaluate and guide learning; we ask whether RL can exploit these results. We place a proximal policy optimisation (PPO) agent in an analytically solved continuous-time broker--trader game. PPO replaces the broker and chooses its trading speed while interacting with an informed trader and stochastic uninformed order flow. We derive a finite-step reward from the broker's continuous-time payoff and verify its discrete implementation through grid refinement and an exact one-step identity. With zero uninformed flow, a validation-selected PPO--FFNN approaches the reference action. With stochastic uninformed flow, the tested PPO--FFNN and PPO--LSTM remain inaccurate, although supervised learning confirms that their actors can represent the action. Monte Carlo diagnostics show that their critics do not reliably rank nearby actions; potential-based reward shaping also gives no reliable improvement. Under partial information, a causal certainty-equivalent controller based on the broker's observable history remains close to the reference, while PPO has larger errors and lower payoffs. Finally, we freeze the analytical policy and train PPO to adjust it after the execution cost changes. Halving the cost yields a repeatable improvement that closes \(2.22\%\) of the gap to the changed-cost reference. The analytical solution therefore provides both a benchmark for diagnosing RL and a useful starting policy for adaptation.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Siu Tung Wong, Carlo Campajola. 2026-10-02. When a Correct Reward Is Not Enough: Diagnosing and Guiding PPO in an Analytically Solved Broker-Trader Game. https://arxiv.org/abs/2610.03598

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Adapting the Actor Model of Concurrency for High-Frequency Trading: Synchronous Message Delivery (fast_send) and a Tick-to-Book Latency Study

The actor model - state isolation, data-race freedom, and sequential single-message reasoning - has long been dismissed as unsuitable for high-frequency trading (HFT): actors seem to imply many threads, a mailbox per actor, and a heap-allocated message plus a context switch per interaction, overhead incompatible with a microsecond budget. This paper argues the dismissal is wrong for co-located actors, with a deployed, measured implementation: kaspar-hft, an open-source C++20 framework. Four extensions adapt the model for HFT: fast_send, a synchronous delivery mechanism in which the sending thread runs the receiver's handler inline and returns the reply as a value; actor groups, which co-schedule actors on one thread behind a shared mailbox; per-actor selectable mailbox queues; and a memory pool. fast_send has receiver transparency: the handler is written identically for synchronous and asynchronous delivery and does not depend on which was used or which thread runs it. A grouped synchronous chain runs on one thread, cutting scheduler context switches from O(N) to O(1); a thread-local call-chain test detects cyclic invocation on one thread before any lock is taken, while a cycle spread across threads deadlocks. Microbenchmarks put the synchronous round trip at tens of nanoseconds. On a live CME market-data feed (ES, NQ, ZN futures), when the socket-reader thread decodes each packet and updates the book itself, socket-to-book medians for book updates are 0.8-1.1 microseconds, and a fast_send hop is about 1% of that. The shared-queue group also yields a production/simulation duality: the same actor code runs unchanged in live trading and deterministic backtest.

q-fin.TR↗

Packets, Transactions and Queues: Design Principles for HFT Systems from a Measurement Study of CME Market Data

HFT systems are conventionally built as a single-threaded event loop, on the rule that every thread hop adds latency. We test that rule against a measurement study of more than a year of CME market data for the NQ front-month contract, following every packet and matching-engine transaction through the feed's two exchange timestamps, and checking the results against a live production receiver. Packets arrive in near-critical self-exciting clusters that belong to the matching engine's transactions, not to how the exchange packs them. The engine often processes consecutive transactions within a fraction of a microsecond, while the market-data publisher sends at most one packet per publisher period of about 7.5 microseconds, so a burst reaches the receiver as a train of packets one period apart. This yields design principles for HFT systems. First, a receiver that handles each packet within one publisher period never queues on arrivals, however bursty the market; there one thread is best. Second, above that period a queueing tail appears, driven by the timing of transactions, not by packet rate or size, and two threads can be better than one: splitting the servicing chain into two stages on separate threads removes most of the tail at the cost of one hop on the median. Third, only the slowest stage matters, so a split pays only if it shortens it. Fourth, just under the period, where the production receiver runs, the remaining tail comes from multi-message packets and variable service times, and the levers are cost per message and spread of service, not thread count. An analytic framework, a burst-limit throughput identity and an exact reduction of the tandem to a single bottleneck server, supports these results.

q-fin.TR↗

Latent Continuum of Regimes in Limit Order Book Dynamics

Market-regime models typically assume a finite set of discrete latent states. We examine whether high frequency limit-order-book dynamics exhibit distinct regime separation or apparent regimes result from discretising an underlying continuum, analysing deep limit-order book data for EURO STOXX 50 index futures across 987 clean trading days from 2022 to 2025 using 30, 45 and 60-second aggregation windows. Market states are represented by symmetric positive definite covariance matrices and analysed under Log-Euclidean and affine-invariant geometries; the resulting state cloud has an effective dimension of approximately 1.4. A causal anomaly layer filters statistical, calendar and rollover contamination, while label-free economic validation evaluates market-state separation independently of volatility-based proxy labels. A 17-method zoo spanning conventional and Riemannian representations, reinforced by full-cloud geometric certificates, consistently favours a continuum over a discrete regime structure. A dominant latent coordinate captures between 83.66% and 85.27% of variation in the covariance-state geometry and tracks VSTOXX without regime labels, remaining stable across years although its average level shifts with market conditions, indicating that conventional regimes are coarse quantisations. A divisive hierarchy spanning macro, macro-subregime, micro and micro-subregime tiers is evaluated using expanding-year walk-forward tests, with out-of-sample performance strongest at the macro tier while finer tiers fail to generalise; continuous fine-scale information nonetheless remains predictive, and its discretisation causes the performance loss. Forecasting performance peaks near ten minutes, reaching a maximum pooled R2_OOS of 0.7483. Overall, the latent stress coordinate defines a stable, low-dimensional geometric continuum that captures market stress and its predictive dynamics.

q-fin.TR↗