arXiv Science⌕ Search

arXiv · 2609.36675

Gödel Forest: Balancing Search Depth and Breadth for Data-Centric Recursive Self-Improvement

Abstract

Recursive self-improvement (RSI) aims to achieve compounding gains by having models improve themselves. While most existing RSI systems optimize external agent harnesses or prompts around a frozen base model, data-centric RSI directly updates the model's own parameters by training on agent-generated data. However, because validating data strategies requires expensive model training, existing methods face a fundamental dilemma: a single agent gets trapped in narrow directions and lacks exploration breadth, while naive parallel search or heavy trace sharing sacrifices long-horizon search depth. To address this challenge, we introduce G"odel Forest, a multi-agent framework that organizes recursive self-improvement as an ensemble of co-evolving search trees. In G"odel Forest, each agent autonomously grows a persistent tree, deepening, branching, or pruning data strategies based on model feedback to secure depth, while parallel trees explore distinct regions of the data space to expand breadth. Crucially, rather than leaving trees isolated or flooding them with heavy execution logs, a dynamically co-evolving memory connects the forest: agents continuously distill their successes and failures into compact procedural lessons anchored to a global leaderboard. Through this forest ecosystem, a dead-end in one tree instantly warns the whole forest against unpromising paths, while an empirical breakthrough quickly seeds new exploration branches in neighboring trees. Evaluated on RSIBench-Data across six diverse domains, G"odel Forest outperforms the single-agent baseline by an average of 10.70% while reducing wall-clock time on five tasks. Ablations confirm that co-evolving shared memory yields a +7.00% gain over independent parallel search, demonstrating that collective distillation is key to scalable self-improvement. The code is available at https://github.com/evolvent-ai/Godel-Forest.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Ziqi Zhao, Fanqing Meng, Haocheng Lu, Lingxiao Du, Qiguang Chen, Mengkang Hu, Xiao-Ming Wu. 2026-09-29. Gödel Forest: Balancing Search Depth and Breadth for Data-Centric Recursive Self-Improvement. https://arxiv.org/abs/2609.36675

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

NMIXX: Domain-Adapted Neural Embeddings for Cross-Lingual eXploration of Finance

Financial text embeddings must distinguish changes in event status, perspective, and obligations even when passages share similar wording. NMIXX adapts existing encoders through 18.8k source-linked triplets: paraphrases and Korean-English translations preserve meaning, while targeted financial rewrites introduce semantic contrasts. We examine this recipe across seven backbones on English and Korean financial and general-domain semantic textual similarity (STS), and analyze the composition and passage lengths of KorFinSTS. BGE-M3 attains the highest adapted financial correlations in this comparison, improving from 0.1969 to 0.2967 on FinSTS and from 0.0512 to 0.2732 on KorFinSTS. Its general English and Korean correlations decrease by 0.0391 and 0.0463. Across the seven models, five improve their mean financial correlation, but all reduce their mean general-domain correlation. Per-language comparisons and benchmark-weight sensitivity analysis reveal differences obscured by a single aggregate score. The study contributes a finance-specific supervision design and evidence for evaluating adaptation jointly with retained general semantic capability; direct cross-language retrieval remains outside its evaluation scope.

cs.CL↗

Sequential Bayesian Evaluation of Large Language Model Behavior

It is increasingly important to evaluate the characteristics of systems based on large language models (LLMs). Evaluations in this context often rely on a curated benchmark set of input prompts provided to the LLM, where the output for each prompt may be assigned a binary or ordinal score and the aggregation of scores across prompts is then used as a summary evaluation. In this paper, we develop a Bayesian approach for quantifying the uncertainty that arises in such evaluation metrics as a result of the stochasticity of the LLM-based systems; the same prompt may exhibit different outcomes on repeated runs. Our framework leads naturally to a sequential evaluation, in which we leverage the Bayesian model to preferentially select which prompts in the benchmark to use next, enabling more cost-effective LLM evaluations. We demonstrate this approach through four case studies: pairwise LLM preferences in interactive dialogue (MT-Bench), refusal rates on adversarial inputs (JailbreakBench), embodied agentic reasoning (ALFWorld), and agentic web-search Q&A (GAIA). Our results illustrate how a Bayesian approach can provide insightful uncertainty quantification about the behavior of blackbox LLM-based systems in a manner that is not accessible via deterministic or single-run evaluations.

cs.CL↗

ResidualKV: Residual-Based KV Cache Compression for Efficient Long-Context Inference

Efficient long-context inference faces two coupled bottlenecks: KV-cache memory grows linearly with context length, while attention computation grows quadratically. Existing approaches typically address one at the expense of irreversible token eviction, full-cache retention, or full-history reconstruction, limiting their effectiveness for multi-turn interaction and long-form reasoning. Motivated by two empirical properties, Long-Range Inter-Token Similarity and Smooth Residual Distribution, we propose ResidualKV, which factorizes the KV cache into a sparse set of globally retrieved references and compact, quantized residual codes for the remaining tokens. This representation preserves token-specific information without permanent eviction and, when combined with sparse attention, reconstructs only the selected states on demand. Dynamic-stride scheduling further reduces reference growth from linear to approximately logarithmic at ultra-long contexts. Across Llama, Qwen, LLaVA-OV, and Qwen3-VL backbones, ResidualKV maintains near-full-cache performance using only 13%-16% KV storage and 30% attention computation on LongBench, and 8%-10% storage and 10% computation in matched-budget multimodal evaluation. It also accelerates decoding by up to $1.5\times$ with KV-cache quantization and $3.4\times$ without it. These results show that global cross-token redundancy supports accurate, memory-efficient, and computation-efficient long-context inference. The source code is available at https://github.com/CURRENTF/ResidualKV.

cs.CL↗