arXiv Science⌕ Search

arXiv · 2609.36417

Support-Set Target Leakage in Relational Foundation Models during In-Context Learning: Model Dependence and Evaluation Reliability

Abstract

Relational in-context learning (ICL) uses labeled support examples and their linked relational context to predict labels for new queries. This creates a failure mode when target-derived features are present in the support context but unavailable for the query. We study this setting as support-set target leakage. We construct 20 controlled target-derived features that vary in signal fidelity, representation, semantic transparency, coverage, and zero-, one-, and two-hop relational placement, and evaluate them across 13 RelBench tasks and five relational ICL configurations that vary the ICL head, message-passing depth, pretraining cohort, or relational encoder architecture. We evaluate matched 0-hop, 1-hop, and 2-hop leakage settings, together with a Full leakage condition containing all 20 leaker columns. Within the tested configurations, target-table (0-hop) and Full leakage produce the largest aggregate deviations from clean evaluation, while higher-hop effects are often weaker, consistent with differences in effective exposure associated with temporal reachability, sampling, and aggregation fidelity. Leakage effects are strongly task- and model-dependent and can reverse relative conclusions between model variants even when aggregate changes are small. For leaker detection, we compare an Integrated Gradients (IG)-based screening method with mutual information (MI) and leave-one-column-out (LOCO) on a common Baseline subset. Ranking quality is strongest in the high-impact 0-hop and Full leakage conditions, but detector-based removal does not consistently restore the clean evaluation. A four-task rel-salt case study further shows the same evaluation concern with native-schema leakage candidates from the original relational schema. These results identify the support/query information boundary as an important component of reliable relational ICL evaluation.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Roshan Reddy Upendra, Alexandre Dorais, Joe Meyer, Andrew Pouret, Anastasios Lambrianos Stappas, Dinesh Katupputhur Ramprasath, Viswanath Ganapathy, Tom Palczewski, Minghua Li. 2026-09-29. Support-Set Target Leakage in Relational Foundation Models during In-Context Learning: Model Dependence and Evaluation Reliability. https://arxiv.org/abs/2609.36417

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Generalising from Self-Produced Data: Model Training Beyond Human Constraints

Current large language models (LLMs) are constrained by human-derived training data and limited by a single level of abstraction that impedes definitive truth judgments. This paper introduces a novel framework in which AI models autonomously generate and validate new knowledge through direct interaction with their environment. Central to this approach is an unbounded, ungamable numeric reward - such as annexed disk space or follower count - that guides learning without requiring human benchmarks. AI agents iteratively generate strategies and executable code to maximize this metric, with successful outcomes forming the basis for self-retraining and incremental generalisation. To mitigate model collapse and the warm start problem, the framework emphasizes empirical validation over textual similarity and supports fine-tuning via GRPO. The system architecture employs modular agents for environment analysis, strategy generation, and code synthesis, enabling scalable experimentation. This work outlines a pathway toward self-improving AI systems capable of advancing beyond human-imposed constraints toward autonomous general intelligence.

cs.AI↗

Hierarchical Reasoning Model

Reasoning, the process of devising and executing complex goal-oriented action sequences, remains a critical challenge in AI. Current large language models (LLMs) primarily employ Chain-of-Thought (CoT) techniques, which suffer from brittle task decomposition, extensive data requirements, and high latency. Inspired by the hierarchical and multi-timescale processing in the human brain, we propose the Hierarchical Reasoning Model (HRM), a novel recurrent architecture that attains significant computational depth while maintaining both training stability and efficiency. HRM executes sequential reasoning tasks in a single forward pass without explicit supervision of the intermediate process, through two interdependent recurrent modules: a high-level module responsible for slow, abstract planning, and a low-level module handling rapid, detailed computations. With only 27 million parameters, HRM achieves exceptional performance on complex reasoning tasks using only 1000 training samples. The model operates without pre-training or CoT data, yet achieves nearly perfect performance on challenging tasks including complex Sudoku puzzles and optimal path finding in large mazes. Furthermore, HRM outperforms much larger models with significantly longer context windows on the Abstraction and Reasoning Corpus (ARC), a key benchmark for measuring artificial general intelligence capabilities. These results underscore HRM's potential as a transformative advancement toward universal computation and general-purpose reasoning systems.

cs.AI↗

A memory-based active inference model of DishBrain-like adaptive behaviour

Recent and rapid advances in artificial intelligence (AI) make it increasingly important to understand the foundations of adaptive behaviour in autonomous agents, especially for building safe and efficient systems. While artificial neural networks have dominated the development of AI, recent work has begun to explore living biological neuronal networks as an alternative substrate for computation. These systems promise remarkable data and sample efficiency and rich dynamics, and may also inspire explainable and biologically plausible models. Here, we develop an experiment-informed active inference framework to model decision-making in closed-loop agents that mirror experimental setups using biological neurons. Using a generative model whose dimensions are matched to an experiment protocol, we systematically compare three decision-making schemes within this common generative model. Under matched episode counts (i.e. total data available for learning) to the in-vitro experiment, our simulations show that agents with short memory horizons reach a level of performance close to that of mouse and human cortical cultures (DishBrain platform), whereas longer memory horizons depart from it substantially. Increasing the planning horizon, by contrast, confers no comparable benefit. Because all model parameters are explicit, we can also track the quantities in our generative model that accompany this improvement, such as the risk term and the entropy of the transition and state-action mappings. Together, these results illustrate how active inference offers a formal language for comparing decision-making schemes in similar closed-loop control environments.

cs.AI↗