arXiv · 2609.20198
Evaluating Financial Sentiment in the Age of AI
Abstract
Financial sentiment measures are widely used in empirical finance, but it remains unclear whether general-purpose large language models (LLMs) improve on existing finance-specific methods. This paper evaluates twelve sentiment models, including dictionary-based methods, finance-specific transformers, and open-source LLMs, using two criteria: linguistic validity and economic validity. We find that general-purpose LLMs achieve classification performance comparable to finance-specific transformer models without task-specific fine-tuning. However, higher classification accuracy does not translate into stronger economic relationships. Several models produce sentiment measures that are significantly associated with earnings surprises, but none is significantly associated with next-day stock returns. Model performance is strongest for announcements with large earnings beats or misses and substantially weaker for announcements with more moderate earnings surprises. These findings suggest that financial sentiment captures information about firms' economic performance but has limited ability to explain short-run market reactions
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Arslan Bisharat, Oudom Hean. 2026-07-29. Evaluating Financial Sentiment in the Age of AI. https://arxiv.org/abs/2609.20198
Cite the original work for its findings. Save a collection to share your selection of sources.