arXiv ScienceSearch

arXiv · 2509.12456

Reinforcement Learning-Based Market Making as a Stochastic Control on Non-Stationary Limit Order Book Dynamics

Abstract

Reinforcement Learning has emerged as a promising framework for developing adaptive and data-driven strategies, enabling market makers to optimize decision-making policies based on interactions with the limit order book environment. This paper explores the integration of a reinforcement learning agent in a market-making context, where the underlying market dynamics have been explicitly modeled to capture observed stylized facts of real markets, including clustered order arrival times, non-stationary spreads and return drifts, stochastic order quantities and price volatility. These mechanisms aim to enhance stability of the resulting control agent, and serve to incorporate domain-specific knowledge into the agent policy learning process. Our contributions include a practical implementation of a market making agent based on the Proximal-Policy Optimization (PPO) algorithm, alongside a comparative evaluation of the agent's performance under varying market conditions via a simulator-based environment. As evidenced by our analysis of the financial return and risk metrics when compared to a closed-form optimal solution, our results suggest that the reinforcement learning agent can effectively be used under non-stationary market conditions, and that the proposed simulator-based environment can serve as a valuable tool for training and pre-training reinforcement learning agents in market-making scenarios.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Rafael Zimmer, Oswaldo Luiz do Valle Costa. 2026-02-14. Reinforcement Learning-Based Market Making as a Stochastic Control on Non-Stationary Limit Order Book Dynamics. https://arxiv.org/abs/2509.12456

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Event-Time Order-Flow Memory, Operational-Time Impact, and Subordinated Market Observables

We consider two canonical market-microstructure regularities: the long-memory of trade signs and the square-root law of meta-order impact. The point is not to propose new empirical laws, but to separate the clocks on which existing laws are defined. The sign-memory law is an event-time statement about the ordering and fragmentation of hidden orders. The square-root impact law is an operational-time statement about front motion in a locally linear latent order book. Starting from a discrete-time random-walk bid/ask reaction--diffusion order book with a separate event clock, we identify an event-time imbalance reduction, derive the operational-time front and impact equations in the locally linear regime, and then subordinate both sign and impact observables to calendar time. Fractional or tempered clock effects enter through this event-to-calendar projection, not through a different operational-time impact mechanism. This gives a compact clock-aware framework in which event-time sign persistence, operational-time square-root impact, and anomalous calendar-time effects appear as distinct but compatible consequences of a common order-book representation.

q-fin.TR

Auditing Collector-Generated Graduation Labels on Pump.fun: Measurement Error and Temporal Non-Generalization

To examine whether collector-generated terminal labels on the pump.fun platform can be interpreted as platform-side graduation outcomes, and whether a pre-registered logistic association model generalises across a subsequent temporal holdout of the same cohort. A frozen primary cohort of 749,816 unique mints recorded by a V3 off-chain collector between 12 May 2026 and 10 June 2026 is analysed. The design combines a source-code and reconstructed-state measurement audit of the collector's terminal-classification mechanism with a pre-registered locked logistic association model fitted on a 15-day development cohort and evaluated on a subsequent 14-day validation cohort under the pre-registered stability-gate framework. Calendar-day cluster-robust covariance and day-block bootstrap procedures assess uncertainty. Development discrimination is high (AUROC 0.8594), but the model does not generalise to the held-out validation cohort (AUROC 0.4642; 500 of 500 successful bootstrap replicates; 95% percentile interval [0.4112, 0.5196] containing 0.5000). Validation calibration deteriorates substantially (slope 0.013, intercept +2.816); the market-cap functional form is not stable across pre-registered specifications; and of nine automated evaluations, two pass, six fail, and one is not evaluable. The collector's TIMEOUT label does not establish platform-side non-graduation. The paper treats outcome ascertainment as a first-order measurement problem rather than assuming that an off-chain collector's terminal classification measures the platform-side event. It pairs this audit with a pre-registered held-out temporal-generalisation test and reports the resulting negative finding as registered, alongside a fully reproducible correction history.

q-fin.TR

Deep Learning of Robust Market Making under Regime-Switching Order Flow

Classical market-making strategies based on stochastic control, such as the Avellaneda-Stoikov and the Guéant-Lehalle-Fernandez-Tapia (GLFT) extension, provide closed-form quoting rules, but rest on assumptions that break down at realistic microstructure timescales. One of them is that order flow is stationary, while empirical evidence points to the existence of regimes, possibly associated with algorithmic execution of metaorders. In this case, existing methods provide negative PnL. In this paper, we develop a deep reinforcement-learning market maker (RLMM) - a Rainbow-style distributional DQN (C51) which is calibrated and tested in a zero-intelligence limit order book. We find that, in the stationary setting, RLMM outperforms GLFT across the entire observed risk-return frontier. The RLMM is more robust to flow asymmetry than GLFT, but, like any stationarily trained strategy, it still suffers large drawdowns from inventory saturation under persistent directional imbalance. Augmenting the state of RLMM with two auxiliary signals - a Bayesian online change-point filter over the directional flow bias and a queue-adjusted quote-exposure imbalance -restores profitability. A final scenario-bandit step that reweights low-return regime scenarios further improves performance under random-persistence and correlated-direction stress.

q-fin.TR