arXiv Science⌕ Search

arXiv subjects

Search papers

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

At least 1,531 records · Page 85Linked to original sources

FedSAP: Federated Learning with Structured Adaptive Partitioning for Multi-Domain Heterogeneous Edge Devices

Federated learning (FL) on heterogeneous edge devices must jointly accommodate unequal resource budgets and domain-shifted local data. Existing resource-adaptive methods decide how much of a model each client trains but not where retained capacity should reside or how it should be shared, whereas federated domain-generalization methods usually assume a shared full architecture. Uniform compression can therefore discard high-utility channels, and a single aggregation path can mix transferable features with domain-sensitive updates. We propose FedSAP, a domain-aware heterogeneous FL framework that casts structured pruning as budget-constrained tri-state channel allocation. FedSAP converts each keep ratio into non-uniform layer budgets, assigns stable channels to a Global pool, useful domain-sensitive channels to pseudo-domain-specific Private pools, and low-utility channels to a Dropped state. This partition lets broadly useful features benefit from cross-client pooling while isolating domain-sensitive updates from incompatible clients. Domain-Guided Assignment infers pseudo-domains from shallow-gradient similarity, while Type-Matched Aggregation restricts each channel to its intended sharing scope. Across three random seeds, FedSAP reaches 76.00% and 72.67% mean global accuracy on Digits and Office-Caltech, exceeding the strongest baseline by 1.70 and 4.92 percentage points while supporting client pruning ratios of up to 80% across heterogeneous clients.

cs.LG↗

Not All Error Yields to Scale: Where Scaling Stops in Vision-Language Inference

Vision-language models (VLMs) face a fixed-budget trade-off between processing more visual information for fine-grained perception and using a larger language backbone for complex reasoning. Existing studies do not tell us which combination of backbone size and input resolution to deploy, especially in high-resolution deployments. To address this gap, we propose the Separable Law that describes how VLM performance changes with language backbone size and visual token count. We fit the law to measurements from 26 InternVL and QwenVL models, with language backbone sizes from 1B to 72B, on four high-resolution benchmarks with image sizes from 224 pixels to 8K. We find that the questions responding to scaling can be predicted from the skill they require, while a substantial fraction never responds at all. We also find that the two model families gain similarly from a larger backbone, while their gains from more visual tokens differ sharply. Combined with a cost law, the Separable Law gives a closed-form rule for allocating compute between backbone size and visual tokens. When deployment is limited to available configurations, the law identifies model and image sizes that perform close to the best feasible choice under the same budget. We hope our work offers a principled way to decide how much a model should be allowed to see at high resolution, given what it must reason about.

cs.CV↗

MCIR: A Feature Dependence-Aware Explainability Method with Reliability Guarantees

Modern machine-learning models often contain strongly dependent or redundant features, making feature attribution difficult because shared predictive information can be distributed across correlated predictors. Existing methods such as SHAP, LIME, HSIC, MI/CMI, and SAGE may therefore produce unstable rankings under multicollinearity or near-duplicate predictors. We propose the Mutual Correlation Impact Ratio Method (MCIR-M), a dependence-aware global feature-importance approach that quantifies the unique predictive information contributed by each feature beyond a selected dependence neighbourhood. MCIR-M introduces the Mutual Correlation Impact Ratio (MCIR), which conditions each feature on strongly dependent neighbours and computes a normalized ratio of conditional to block-level information. The population score lies in [0,1] and equals zero under exact conditional redundancy. We also introduce a lightweight estimation procedure that computes MCIR using a fraction of the available data and evaluates agreement with full-data explanations. Across controlled synthetic redundancy experiments and the UCI HAR benchmark, MCIR shows dependence-aware ranking behaviour, with its clearest advantage under injected near-duplicate predictors. Comparisons with independent and conditional SHAP, SAGE, HSIC, MI-based scores, and CIR-family baselines are mixed across real-data criteria. Reduced explanation samples lower computational burden in the evaluated configurations, while agreement with full-data explanations is assessed separately through ranking, head-set, and faithfulness diagnostics. Overall, MCIR-M provides a practical dependence-aware diagnostic for global explanation under strong feature dependence.

cs.LG↗

Linear toral endomorphisms, exactness, Bernoulliness and generators

Let $T$ be a non-invertible surjective continuous endomorphism of the group $\ttf^d\,:= \rrf^d/\zzf^d$. Then $T$ preserves the Haar measure, is associated to some $d \times d$ matrix $A$ with integer coefficients and is $r$-to-one, where $r = |\det A| \ge 2$. In~\cite{Krzyzewski}, Krzyżewski gives a necessary and sufficient condition for $T$ to be exact, and shows that $T$ -- viewed as a measure-preserving map -- is isomorphic to the direct product of some (possibly trivial) toral automorphism by some exact toral endomorphism. We re-prove these results by more elementary and constructive methods, whereas Krzyżewski actually works with endomorphisms of compact Abelian groups and relies on non-constructive theorems. Krzyżewski's exactness condition and Mihailescu's papers ~\cite{Mihailescu 2012} and~\cite{Mihailescu 2013} show that $T$ is isomorphic to the one-sided uniform Bernoulli shift on $\odc 0,r-1 \fdc^{\zzf\_+}$ if and only if $A$ is expanding (i.e. the modulus of each eigenvalue is $>1$). When $A$ is not expanding and hyperbolic (i.e. all eigenvalues have modulus different from $1$), Mihailescu~\cite{Mihailescu 2012} shows that $T$ cannot have a generating Rokhlin partition. More generally, when $A$ is not expanding (whether hyperbolic or not), we show that the endomorphism $T$ cannot have a smooth finite generator (whether independent or not).

math.DS↗

Finite Time Singularities of the Kähler Ricci Flow on Compact Kähler Surfaces are of Type I

We establish a Type I curvature bound for finite-time volume-collapsing Kähler--Ricci flows on compact Kähler surfaces. No symmetry assumption is imposed, and the initial Kähler class need not be rational. The proof combines local symplectic topology with the classification of gradient Kähler--Ricci shrinking solitons. Bamler's compactness and structure theory provides the shrinking limits used in the argument.

math.DG↗

SonarVoxNet: Diver Detection in 3D Bounding Box using 3D Sonar

Autonomous underwater vehicles (AUVs) assisting human divers must continuously track not only the diver's 3D position but also their full-body orientation. However, vision-based perception is unreliable underwater, and forward-looking sonar -- despite being widely used -- discards the elevation information needed for orientation estimation, posing a fundamental limitation. Recently commercialized 3D sonar preserves elevation but produces sparse, noisy returns, and existing detectors are built for dense LiDAR data and for targets that remain upright and rotate only about the yaw axis (e.g., vehicles, pedestrians), making them unable to represent a freely pitching and rolling diver. To address this gap, we present two contributions. First, SonarVoxNet adapts a voxel-based encoder and an anchor-free center-based detection head to 3D sonar data, replacing the conventional yaw-only rotation representation with a continuous 6D rotation parameterization to predict full 9-DoF oriented bounding boxes -- to our knowledge, the first 3D sonar diver detector to do so. Second, Diver3D is the first public 3D sonar dataset with full 3D orientation labels for divers in diverse, non-upright poses, collected at a natural cave-diving site. Through controlled ablations over the backbone and detection head, we show that the dominant factor behind accurate 3D sonar-based diver detection is the transition from yaw-only rotation to full-SO(3) rotation. This transition substantially improves detection accuracy and reduces orientation error. These results demonstrate that full-body diver orientation is recoverable from 3D sonar alone, laying the groundwork for future work on diver pose estimation and diver-robot interaction.

cs.RO↗

Beyond Demographic Balance: Multi-Metric and Intersectional Evaluation of Fairness in MIMIC-IV Mortality Prediction

Fairness conclusions in clinical prediction can depend strongly on both the metrics reported and the demographic resolution at which performance is evaluated. We revisit these evaluation choices for ICU mortality prediction on MIMIC-IV, comparing predictive-utility and subgroup-error metrics across several fairness interventions. As a complementary case study, we introduce a lightweight adaptation strategy that jointly balances ethnicity--gender--insurance representation without conditioning on mortality outcomes, allowing demographic representation balancing to be examined separately from outcome-conditioned or direct error-rate interventions. We evaluate its behavior at both marginal and corresponding three-way intersectional subgroup levels, while accounting for the statistical support of finer-grained estimates. The results show that interventions can receive substantially different assessments across accuracy/AUROC, sensitivity, and false-positive rate, and that marginal demographic summaries can conceal heterogeneous error profiles within their constituent intersections, including among larger subgroups. These findings highlight the importance of evaluating fairness interventions at both complementary metric and subgroup resolutions, while accounting for the intervention target and the reliability of subgroup estimates.

cs.LG↗

DIADA: Automatic Data Composition in Data Lakes

Data lakes contain a plethora of attributes scattered across many tables that, when combined, provide enhanced assets for data analysis. Nonetheless, deciding which attributes belong together in meaningful relations remains a manual, per-task effort. Merging by joinability alone provides no guarantees regarding attribute relevance, while selecting features against a single target discards attributes useful to other tasks. To address this gap, we introduce the data composition problem: organizing a fragmented, heterogeneous lake into meaningful relations, agnostic of any particular analytical task so that the resulting organization can serve as a common foundation for diverse downstream analyses. We propose DIADA, a composition system that employs multivariate dependence as the criterion for assessing the meaningfulness of a relation and approximates it by hypothesizing independence among attributes and identifying those sets that violate this hypothesis. To do so, we map the attributes to a predicate space, forming a lattice under inclusion and mining those predicate sets that exhibit dependence among their constituents. We contribute a dedicated and scalable algorithm to effectively explore this space, outscaling classical algorithms for mining relationships, thus discovering dependencies that would otherwise be impractical to identify. We demonstrate that applying a single data composition process benefits diverse potential downstream tasks. This is the result of providing a subset of low-noise, statistically relevant attributes that increases the confidence that detected patterns are grounded in real relationships, thus preventing common modeling issues in large-scale environments.

cs.DB↗

Towards a Cloud Fog Edge System for Smart Building

In this article, we present our vision and recent advancements toward creating a decentralized system capable of learning from real-time data within buildings to support sustainable and privacy-preserving smart environments. Our approach promotes the concept of the building itself as the data center, aligning with the principles of edge computing to safeguard confidentiality and reduce reliance on external cloud infrastructure. This is particularly valuable in humanitarian contexts, where data sovereignty, energy efficiency, and infrastructure constraints are critical. We detail a lightweight, "Kubernetes-like" orchestration framework for deploying AI services within such environments and demonstrate our progress in implementing AI algorithms on low-power, cost-effective microcontrollers such as those in the Arduino ecosystem. By enabling in-situ learning directly on sensors or microcontrollers, our work aims to bring intelligent services to resource-limited settings, fostering autonomy, resilience, and sustainable development in vulnerable or underserved communities. The contributions in this article are related, firstly, to our project "Online Machine Learning Algorithms for Embedded Systems" and the evaluation of two new online algorithms. Secondly, we envision a cloud-fog-edge architecture based on the KOptim and FIWARE components, and we propose a methodology for coupling them. Experimental results of the online algorithms are also presented, showcasing real-world traces.

cs.DC↗

Exact Locality Gaps for Matchable Semi-Matchings

An assignment of tasks to servers can resist every small improvement and still make tasks wait longer than necessary. We determine exactly how inefficient such an assignment can be when each task requires one unit of service and the eligibility constraints permit all tasks to use distinct servers. For every move size $r$ and maximum current server load $K$, we give a closed formula for the worst ratio between locally optimal and globally optimal total completion time. Local optimality here allows every feasible reassignment changing at most $r$ tasks. Every finite-cap bound is attained on a tree where each task has at most two eligible servers. Thus the worst behavior already occurs under simple eligibility constraints. At load cap two, the exact ratio is $1+1/(r+2)$, attained on a path with $r+2$ tasks. Without a load cap, the worst-case supremum is $3/2$ for single-task moves and approximately $1.294503159$ for two-task moves; its excess above one is $1/(r+2)+O(2^{-r}/r)$ as $r$ grows. The proof uses an explicit rational potential on a comparison graph and matching extremal constructions. These results give sharp guarantees for bounded-size local search on matchable semi-matchings, including exact guarantees under degree bounds.

cs.DS↗

CrossGMN: Graph Metanetworks for Cross-Architecture Weight-Space Transformations

Weight-space networks operate directly on parameters of other neural networks, enabling tasks such as predicting model properties, editing trained models, and generating weights. Weight-space symmetries such as neuron permutations make equivariance a key design principle. However, existing equivariant weight-space architectures have primarily been studied for transformations that preserve the network architecture. In contrast, many practical transformations, including model compression and upscaling, map a trained source network into a target network with a different architecture. In this setting, the source and target permutation symmetries act on different parameter spaces, making equivariance less straightforward to formulate. Our key idea for addressing this mismatch is to reformulate cross-architecture operators with two inputs: a trained source network and an initialization of the target network. This lets us define equivariant cross-architecture operators that refine the initialization of the target network using information from the source network, while being invariant to source-network permutations and equivariant to target-network permutations. Based on this formulation, we introduce CrossGMN, a graph metanetwork that jointly processes both networks through symmetry-preserving cross-network message passing. We prove CrossGMN is universal for continuous cross-architecture operators on compact sets under a general-position assumption. We evaluate CrossGMN for model compression, predicting a smaller network's parameters to accelerate subsequent knowledge distillation. Across 2-D and 3-D INRs and image classification with MLPs, CNNs, and Vision Transformers, CrossGMN speeds up distillation by up to 8.89x, transfers across datasets without retraining (3.78x), and a single model can accelerate compression from heterogeneous source architectures into a common target architecture.

cs.LG↗

Combining Homomorphic Encryption and Differential Privacy in Federated Learning for Model Inspection and Availability

The increasing prevalence of decentralized data has led to a growing interest in federated learning, which enables collaborative model training without clients sharing their sensitive local data. However, FL alone does not sufficiently protect sensitive training data and is generally coupled with privacy-preserving techniques, such as differential privacy and homomorphic encryption. Although powerful, these techniques address separate concerns via different mechanisms, so relying on just one might prove insufficient or impractical for addressing challenges associated with federated learning. In this work, we propose a privacy-preserving federated learning framework that combines homomorphic encryption-based training with differential privacy-based model inspection and release. We adopt a Markov chain Monte Carlo-based Bayesian privacy estimation method to estimate the privacy of our proposed framework. Our results show that this method improves both model utility and estimated privacy over the baseline method that relies solely on differential privacy for training. In our experiments with the FEMNIST dataset, by the end of training, our method reaches a test loss of $1.09$, compared to $2.37$ for the differential privacy-only approach, while providing stronger estimated privacy protection, with the estimated posterior mean of the privacy parameter $ε$ of $4.32$, compared to $7.26$ for the differential privacy-only approach. We also show that intermittent model monitoring can preserve the encrypted training trajectory while, under our evaluated experimental setting, providing estimated privacy comparable to or stronger than the differential privacy-only approach.

cs.CR↗

Theoretical guarantees for stochastic gradient Langevin dynamics

We prove asymptotic bias bounds for stochastic gradient Langevin dynamics in Wasserstein distance of order two. We assume that the negative log-density is strongly convex with a Lipschitz gradient, and that the stochastic gradient estimator is unbiased with an error satisfying a mean-square Lipschitz condition. The bounds are of order $h$ under a fourth moment assumption on the stochastic gradient error and of order $h^{1/2}$ under only a second moment assumption, where $h$ is the stepsize. A spiked-noise example shows that a second moment assumption alone is insufficient for a bound of order $h$ that is uniform over noise distributions with a fixed variance.

math.PR↗

Iterative Policy Refinement through Semantic Rollout Analysis

Structured policies improve efficiency, robustness, and interpretability in imitation learning by introducing task-specific inductive bias, but existing structure generation methods rely either on extensive human input or on static domain knowledge encoded in LLMs, which may be inconsistent with the expert demonstrations. We propose a closed-loop framework that iteratively refines structured policies using LLM-guided analysis of policy rollouts. By logging rollouts as semantically meaningful tabular data and prompting the LLM to generate diagnostic analysis code, our method identifies suboptimalities in the policy structure and iteratively corrects them without requiring human instruction. Experiments on car racing and door opening tasks show that our approach improves imitation learning performance by up to 15% over zero-shot LLM-generated structures and requires 75% less compute to achieve the same reinforcement learning performance. These results demonstrate that tabular rollout analysis provides an effective feedback signal to align LLM-generated policy structures with expert demonstrations, and we can utilize it to generate good policy structures automatically.

cs.LG↗

The isoperimetric problem for four-dimensional parallelohedra

We show that the regular $24$-cell uniquely minimizes surface area among four-dimensional parallelohedra of fixed volume. As a consequence, it also minimizes surface area among parallelohedra containing a fixed ball. Moreover, we establish quadratic Hausdorff stability, and we derive a second-variation formula valid across changes of Voronoi type, which identifies which root-lattice Voronoi cells are strict local minima under metric and affine perturbations.

math.MG↗

Identifying Panel Conditioning with Refreshment Samples: Sharp Bounds and Design Assumptions

Refreshment samples are the standard remedy for panel attrition, and the identification results behind them maintain that participation does not change measurement. We characterize what a refreshment sample identifies about panel conditioning, modelled as a deterministic monotone map at reinterview, when attrition is unrestricted. A candidate map is consistent with the data if and only if the retention-scaled distribution of the stayers' implied latent outcomes is setwise dominated by the refreshment distribution; every such map is rationalized by an explicit attrition process. Without attrition the map is identified on the latent-outcome support; with attrition, a density-ratio condition governs the identified set, and tail behaviour alone does not determine it. For an unrestricted map the survivors' mean effect has the familiar trimming bounds; for an item with all categories reported, the model reduces to a test of no conditioning. Within a cohort, other waves, dropout patterns and entry-wave items leave the set unchanged unless restrictions link selection across waves. Under explicit selection restrictions, survival matching, symmetric matching and entry-wave correction identify survivor effects. We derive their biases and give rank conditions under which refreshment schedules identify curvature in the conditioning path. A Japanese panel illustrates the results.

stat.ME↗

MDIRNET: Multi-Degradation Image Restoration Network via Deep Unfolding

Real images often exhibit unknown and mixed degradations, making restoration substantially more challenging than single-task image restoration because multiple distortion types interact within the same observation. Consequently, existing methods often rely on prior knowledge of the degradation type or separate task-specific models, which may oversmooth fine structures or leave residual artifacts, motivating a compact model-driven alternative. We propose the Multi-Degradation Image Restoration Network (MDIRNET), a unified framework that combines a model-driven low-rank prior with end-to-end learning. Here, unified refers to joint training on three degradation types: noise, rain, and blur. A single MDIRNET model restores all three without requiring task-specific models, modules, or branches at inference. The low-rank prior exploits the redundancy and compact structure of natural image patches. To identify this underlying low-dimensional representation, we formalize restoration via Orthogonal Variational PCA (OVPCA) and translate its iterative inference into a deep unfolding network. To handle spatially non-uniform corruption and local content variability, we further introduce a learnable patch-partitioning strategy and a lightweight dynamic rank-allocation module that predicts the appropriate subspace dimension for each region. Spatially adaptive reconstruction refinement is performed using a supervised attention module. Extensive experiments on standard denoising, deblurring, and deraining benchmarks show that MDIRNET achieves competitive or superior performance over strong baselines across most metrics, while controlled mixed-degradation experiments demonstrate consistent performance across the evaluated synthetic degradation combinations. The code is available at https://github.com/ScholarForge/mdirnet.git.

eess.IV↗

On regularity of averaged Green's functions in homogenization of elliptic PDE

This paper is concerned with establishing estimates on averaged Green's functions for a uniformly elliptic divergence form partial difference equation with random coefficients on the $d$ dimensional integer lattice $\mathbb{Z}^d$. It has previously been shown that the averaged Green's function is point-wise well approximated by the homogenized Green's function at large length scales. Corresponding results also hold for (fractional) derivatives of the averaged Green's function up to but not including the second derivative. Second derivative results have been established when the elliptic equation generates a positive semi-group, as in the diagonal case. Here the second derivative result is shown more generally by using an insight of Bourgain.

math.AP↗