arXiv Science⌕ Search

arXiv · 2609.38765

Multiperiod bond portfolio optimization with transaction costs using a Markov Decision process

Abstract

Bank treasury portfolios must balance yield, liquidity, and interest-rate risk across bonds of different maturities. Static allocation rules are ill-suited to this task: portfolios concentrated in long-duration securities with no dynamic adjust- ment mechanism can accumulate large mark-to-market losses and liquidity stress under rising interest rates, as illustrated by the failure of Silicon Valley Bank in 2023. We develop a tractable simulation-based framework for multi-period bond port- folio optimization under interest-rate risk and proportional transaction costs. Yield-curve dynamics are modeled using the Dynamic Nelson-Siegel parameter- ization with Vector Autoregressive factor dynamics, from which we construct a time-inhomogeneous discrete-state Markov chain approximating the joint yield process across bond maturities. This chain forms the state space of a finite- horizon Markov Decision Process in which the investor maximizes expected terminal wealth subject to proportional rebalancing costs. The optimal portfolio policy is obtained by backward induction. We also quantify the approximation error introduced by truncating the transition kernel, and show that it leaves mean terminal wealth almost unchanged while substantially distorting drawdown and tail statistics.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Balaji Ramachandran, Srikanth Iyer, Shashi Jain. 2026-09-30. Multiperiod bond portfolio optimization with transaction costs using a Markov Decision process. https://arxiv.org/abs/2609.38765

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

A Spread-Gated Hawkes-Flocking Model for Best Bid and Ask Dynamics, with an Application to Limit Order Placement

We study the joint dynamics of the best bid and ask prices with a spread-gated Hawkes-flocking model. The model tracks four types of best-quote movements: spread-narrowing movements are switched off when the spread is at its one-tick minimum, and a cross-side excitation term, whose activation depends on the prevailing spread, links the two sides of the book. We show that the process is non-explosive on every finite horizon, give an $O(N)$ recursive likelihood, and validate the maximum likelihood estimator by simulation. On real intraday limit order book data for two large-tick stocks, INTC and MSFT, the restriction that removes the cross-side term is rejected, and the full model improves fit substantially by AIC and BIC; the likelihood is multimodal on a single day, so estimation uses a multi-start search. As an application, we derive the closed-form optimal size of a single-period limit order placed at the best or second-best quote, given the model's next-event probabilities and externally supplied execution probabilities.

q-fin.CP↗

The Efficient Frontier from a LASSO Solver

In a recent paper, Schmelzer and Hastie argue that Markowitz's Critical Line Algorithm and the LASSO path trace the same curve. Here we use that identity to compute efficient frontiers with a stock LASSO solver, \texttt{lars\_path} from \texttt{scikit-learn}. It handles long--short portfolios under a leverage cap, fixed leverage with varying risk appetite, and the classical long-only, fully invested frontier. Called naively, the last path stops at the maximum-Sharpe portfolio. One shift of the response, by an amount computed in advance, lets a single call reach the minimum-variance portfolio.

q-fin.CP↗

Information Games: Strategic Crowding and Firm Repositioning in Language-Model Space

Firms follow changing economic opportunities, but rivalry changes their response. We develop ESCAPE, a rational-share game of distribution-valued positioning with heterogeneous capability costs, establish a unique equilibrium, and derive an exact reallocation restriction separating opportunity and crowding contributions. Competition need not push firms apart: rational firms can enter increasingly crowded regions, because the full reallocation gain can be positive even when its crowding component is negative. Rivalry instead changes how differently firms respond to the same opportunity shift. In the benchmark many-firm economy, the relative-response component accounts for 90.5% of the weighted squared composition-response contrast between strategic and independent firms; it is numerically zero under identical capabilities. The contrast remains distinguishable under corpus-scale measurement. Corporate news distributions document persistent but changing peer structure: optimized peer reconstructions retain 86.5% of their aggregate validation advantage one year later even as absolute distance to both frozen reconstructions increases. Peer structure is persistent without being a fixed destination, so current similarity can remain informative without fixing firms' future relative positions.

q-fin.CP↗