arXiv Science⌕ Search

arXiv · 2609.32746

Self-Evolving Multi-Agent Symbolic Discovery for Financial Fundamental Analysis

Abstract

While symbolic regression (SR) has been successfully used in science to discover new equations, its use in financial valuation is hindered by several limitations. Whereas the natural sciences provide objectively correct relationships, financial valuation constitutes a distinct class of symbolic discovery problems, as it admits multiple valid perspectives, operates under non-stationary market conditions, and involves noisy, continuous performance signals. In this work, we propose Multi-Agent Fundamental Analysis with Symbolic Adaptive learning (MUFASA), a hierarchical multi-agent framework for symbolic discovery in finance. MUFASA introduces (1) disentangled equation discovery via specialized agents representing distinct valuation perspectives, (2) a meta-coordinator that performs hierarchical-level reasoning over market context information, and (3) a memory mechanism that reasons over statistical performance summaries (e.g., accuracy, stability, and tail risk) to guide learning under noisy feedback. Experiments across datasets from multiple countries show that MUFASA achieves state-of-the-art performance on the valuation task compared to classical finance methods, financial large language models, and SR approaches, while simultaneously producing interpretable equations, which we share with the community. We also make publicly available the distilled learnings across evolution iterations and context-dependent strategy weights, which might offer useful insights for future research on financial fundamental analysis.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Kelvin J. L. Koa, Filip Orestav, Shengqiong Wu, Michael J. Wooldridge, Ke-Wei Huang. 2026-09-26. Self-Evolving Multi-Agent Symbolic Discovery for Financial Fundamental Analysis. https://arxiv.org/abs/2609.32746

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

From Evidence to Effect: Authority Semantics and Runtime Infrastructure for Stateful Agents

Stateful agents reuse artifacts after producing executions and permissions change. We formalize authority-sufficient observations and durable effects bound to execution and material identities. WTB implements this interface through runtime adapters, shared evidence, and transactional publication/recovery. Six study families separate the mechanism from its integration. Raw and typed evidence both solve 32/32 authority cases, with model-dependent planning effects. Fixed-intent enforcement blocks six unsafe proposals and executes 12 eligible authorized intents. Complete controls match WTB's capability. Paid integration yields 176/210 accepted benchmark-source stages, including 19/30 publication stages, recovery on 8/8 primary SWE repositories, and the most complete continuous trajectories on each of three source tasks. The findings connect authority information, effect admission, and infrastructure reuse in stateful agents.

cs.MA↗

Scores Alone Do Not Prove Discovery: The Discovery Certification Protocol for Auditing AI Research Agents

AI research agents combine public information and experimental feedback to produce measurable results. The Discovery Certification Protocol (DCP) turns an outcome claim into an executable audit under a registered model, information boundary, and budget. Gate 1 validates useful improvement. Gate 2 tests recovery by matched agents given the starting information and observed Web content, with run history and new measurements withheld. Core requires adequate registered controls, zero recoveries, and a finite-sample recovery bound. Optional Gate 3 compares truthful and neutral feedback from a shared checkpoint; Evidence adds a supported effect and a null-policy equivalence check. Controlled SQLite and virtual catalyst audits pass both decision kernels. On real-data response surfaces, Yacht and Ionosphere pass the Core kernel after zero recoveries in 96 attempts, with an upper bound of 0.0468. Each target combines ten observed utilities and six predictions into a 16-entry data product. Yacht scores 0.7677 on reconstruction of all 32 switch effects, with utility-prediction MAE 0.0315 on its six unmeasured configurations. Fresh truthful continuations recover the target level in 9/30 and 16/30 trials, respectively, separating achieved utility from process repeatability. A deterministic verifier reproduces these local decisions from frozen records.

cs.MA↗

MASTraceBench: Diagnosing Collaboration Gains through Proposal Trajectories in LLM-Based Multi-Agent Systems

LLM-based multi-agent systems (MAS) have shown promise in complex problem solving. As MAS methods diversify, systematic evaluation becomes increasingly challenging. However, existing benchmarks largely focus on final outcomes, leaving unclear how collaboration gains arise, are preserved, or are lost. To address this limitation, we introduce MASTraceBench, a benchmark for diagnosing collaboration gains through proposal trajectories in MAS. Across six cooperative and competitive tasks, MASTraceBench tracks and grades proposal trajectories and provides a multi-layer metric suite covering Task Score, Collaboration Gain, proposal-trajectory indicators, and Token Cost. Using MASTraceBench, we systematically compare representative MAS methods not only by final performance, but also by how agent proposals evolve and are aggregated into the final answer. This analysis reveals a recurring pattern: final MAS answers rarely surpass the strongest initial proposal; interaction often lifts initially weaker proposals toward it, while strong initial proposals are seldom further improved and may regress. To reduce this risk, we propose CLEARS, which replaces whole-proposal exchange with claim-level evaluation across agents to guide reliable synthesis. CLEARS more often preserves or improves upon the strongest initial proposal and achieves the highest Collaboration Gain on five of the six tasks.

cs.MA↗