arXiv ScienceSearch

arXiv subjects

Suvrorup Mukherjee

Publications and source records attributed to Suvrorup Mukherjee.

2 recordsLinked to original sources

Fear Moves Markets: Sentiment-Augmented POMP for Volatility Modeling of Bitcoin Returns

Cryptocurrency markets exhibit extreme price swings and sentiment-driven regime shifts, which traditional volatility models often fail to capture. To address this, we develop a partially observed Markov processes (POMP) model augmented with sentiment and heavy-tailed distributions. Specifically, we extend Breto's framework by including the Fear and Greed Index (FGI) as an exogenous regressor in the latent volatility dynamics and replacing Gaussian measurement noise with a Student's t distribution. We fit the model via simulation-based inference using daily Bitcoin returns and FGI data from January 2020 to April 2025. Our approach significantly outperforms three benchmark models in log-likelihood and filter stability. Moreover, our model yields interpretable parameters consistent with known features of financial volatility. These results demonstrate that incorporating sentiment signals and heavy-tailed noise improves the modeling of volatility and regime shifts in cryptocurrency markets.

stat.AP

Consistency as a Testable Property: Statistical Methods to Evaluate AI Agent Reliability

This paper establishes a rigorous measurement science for AI agent reliability, providing a foundational framework for quantifying consistency under semantically preserving perturbations. By leveraging $U$-statistics for output-level reliability and kernel-based metrics for trajectory-level stability, we offer a principled approach to evaluating agents across diverse operating conditions. Our proposal highlights the important distinction between the core capability and execution robustness of an agent, showing that minor task-level variations can induce complete strategy breakdowns despite the agent possessing the requisite knowledge for the task. We validate our framework through extensive experiments on three agentic benchmarks, demonstrating that trajectory-level consistency metrics provide far greater diagnostic sensitivity than traditional pass@1 rates. By providing the mathematical tools to isolate where and why agents deviate, we enable the identification and rectification of architectural concerns that hinder the deployment of agents in high-stakes, real-world environments.

cs.AI