arXiv Science⌕ Search

arXiv · 2610.04157

Mitigating Over-squashing without Rewiring: A Sheaf Effective Resistance Perspective

Abstract

Graph Neural Networks (GNNs) often struggle to capture long-range dependencies due to over-squashing -- a phenomenon in which the repeated compression of node embeddings into finite-size messages causes representations to collapse. Over-squashing is most often diagnosed as a property of the graph topology, with effective resistance serving as a principled measure of the bottleneck. We provide a complementary view on the matter: building on cellular sheaves, we introduce sheaf effective resistance, a generalization of effective resistance that depends on the sheaf attached to the graph, and we prove that for flat vector bundles, the over-squashing sensitivity in the Jacobian sense is upper bounded by a quantity related to the sheaf effective resistance between the nodes. The bottleneck thus need not lie in the graph itself: it can be relocated, and reduced, by adjusting the sheaf. We instantiate this idea in FlatNSD, a simple message-passing variant of Neural Sheaf Diffusion, and show that it implicitly learns to modulate total sheaf effective resistance, performing well on benchmarks designed to stress over-squashing without altering the original graph topology.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

André Ribeiro, Germano Barcelos, Amauri H. Souza, Diego Mesquita, Ana Luiza Tenório. 2026-10-03. Mitigating Over-squashing without Rewiring: A Sheaf Effective Resistance Perspective. https://arxiv.org/abs/2610.04157

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Convergence of Statistical Estimators via Mutual Information Bounds

Recent advances in statistical learning theory have revealed profound connections between mutual information (MI) bounds, PAC-Bayesian theory, and Bayesian nonparametrics. This work introduces a mutual information bound for statistical models, and derives from it convergence rates for fractional posteriors, for their variational approximations, and for the maximum likelihood estimator. The observations are assumed independent but not identically distributed, and the model is not assumed well-specified, so that the bounds are oracle inequalities and cover regression with a fixed or a conditioned design; the independent and identically distributed, well-specified case is recovered by dropping an index. We illustrate the method on two applications. In the Gaussian sequence model the rate is minimax in both the radius of the Sobolev ball and the sample size, which only appears through its product with the temperature. In logistic regression, where the model is not conjugate and the variational approximation is computed by a stochastic gradient method, the bound applies to the output of the algorithm rather than to an idealized minimizer, and attains the parametric order in the Renyi risk with no logarithmic factor, the statistical and the optimization error being separated.

stat.ML↗

Risk-Calibrated Proposal Transport for Finite-Particle Diffusion Steering

Inference-time steering combines pretrained diffusion experts or rewards without retraining by changing the dynamics that transport noise to data. Feynman-Kac correction compensates for proposal mismatch through importance-weighted sequential Monte Carlo (SMC), whose finite-particle behavior depends on the proposal. Variance-controlling guidance (VCG) improves that proposal by fitting a linear drift correction to minimize empirical log-weight-rate variance. Although its population optimum cannot worsen residual variance, finite-particle VCG can nearly eliminate its fitting residual while increasing residual risk on new states by orders of magnitude. The resulting update can degrade unweighted generation or accelerate particle collapse. We show that the centered Feynman-Kac rate is the normalized transport residual and that expected out-of-fit benefit is exactly population headroom minus coefficient-estimation penalty. Under regularity assumptions, a Wasserstein analysis bounds the unweighted proposal's terminal error using this residual. These results motivate Risk-Calibrated Proposal Transport (RCPT), which uses deletion leave-one-out residuals to calibrate the retained fraction of the VCG update, adding no model calls and only small linear-algebra overhead. Experiments on 2D checker distributions, scaffold decoration, molecular property optimization, and class-conditional CIFAR-10 generation demonstrate recovery from harmful fitted updates. Across molecular and image domains, RCPT mitigates harmful fitted updates and improves a broad range of terminal metrics relative to uncalibrated VCG.

stat.ML↗

Local Fisher Information Enables Sparse Causal Discovery

Sparse causal discovery calls for methods that exploit graph structure without estimating high-dimensional densities. We introduce Fisher Information Completion Search (FiCS), a source-first algorithm for additive noise models that uses one local Fisher score for both ordering and parent selection. Under regularity and nonconstant-parent conditions, we prove that a node's local Fisher information equals the noise Fisher information exactly when the conditioning set contains all parents, provided that it contains no descendants. This Fisher parent completion identifies the parent set as the unique minimal Fisher completion. With a maximum conditioning set size $q$ at least the maximum indegree $d$, population FiCS queries marginals of at most $q+1$ variables and recovers the true directed acyclic graph under a positive ordering margin. Bounded conditioning also has a population advantage: reducing $q$ toward $d$ cannot decrease, and can strictly increase, the ordering margin. A growing non-Gaussian family separates local Fisher selection from conditional-variance and leaf-first Fisher ordering. For the regularized kernel Stein estimator, we establish high-dimensional DAG consistency under $q\{1+\log(p/q)\}+\log p=o(n)$, uniform Fisher separation, local approximation, and compatible ridge and parent penalty parameters. Experiments show the strongest gains when $n$ is small relative to $p$, quantify the effect of the conditioning size, and demonstrate competitive reference-graph recovery on three real-data benchmarks.

stat.ML↗