arXiv Science⌕ Search

SEARCH · arXiv Science

Search arXiv Science

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,333 records · Page 74Linked to original sources

MWOP: Modality-aware Width-wise Operation Pruning for Efficient MLLMs

Multimodal large language models (MLLMs) incur substantial inference costs when processing long visual-textual sequences. While existing operation compression methods exploit modality-level redundancy, they largely treat computation within attention heads and shared feed-forward network (FFN) channels as unified units, leaving finer-grained redundancy underexplored. We find that redundancy varies both across modality-interaction paths within the same attention head and across visual and textual executions of the same FFN channel. Based on these findings, we propose Modality-aware Width-wise Operation Pruning (MWOP), which independently prunes visual-to-visual (V2V), text-to-visual (T2V), and text-to-text (T2T) attention paths within each layer, and separately selects FFN channels for visual and textual inputs. A first-order Taylor criterion guides the pruning process, with FFN importance re-evaluated after attention pruning and LoRA-based recovery training. To translate the resulting fine-grained sparsity into practical acceleration, we further develop path-sparse Triton attention kernels and compact visual-side FFN execution. MWOP preserves the token sequence while reducing attention and FFN computation, making it complementary to token compression and enabling simultaneous reduction of sequence length and per-token computation. On LLaVA-OneVision-7B, MWOP alone achieves a $1.6\times$ prefill speedup with 99.7\% average performance retention across 12 benchmarks. Combined with two representative token compression methods, it further increases their prefill speedups from $2.0\times$ and $1.9\times$ to $2.9\times$ and $2.7\times$, respectively. Results on Qwen2.5-VL-7B further demonstrate its applicability across architectures. The code is available at https://github.com/EIT-NLP/MWOP.

cs.CV↗

Distillation of Tabular Foundation Models into Efficient Predictors

Tabular foundation models (TFMs) achieve strong predictive performance through in-context learning, yet repeatedly conditioning on labeled data makes inference expensive. Knowledge distillation can reduce this cost by transferring their predictive ability to lightweight, dataset-specific students. However, the dependence of TFM predictions on both a labeled context and a query introduces two design questions: how to construct teacher supervision and whether expanding query coverage improves distillation. We examine these questions across two TFMs and both neural and tree-based students, and derive an effective distillation recipe. The recipe uses the full labeled training set as teacher context and trains students solely on teacher predictions for observed and synthetic queries. On TabArena, the resulting students outperform their supervised trained tuned-and-ensembled counterparts by 57-98 Elo points. Applied unchanged to TALENT, the same recipe improves matched default students on 236-258 of 300 datasets and reduces median primary error by 4.0-6.4%. The distilled students also achieve median inference speedups of 3.0-21.6 times over their teachers, offering a practical trade-off between predictive performance and repeated inference cost. Code is available at https://github.com/nums-ai/TFM_Distillation .

cs.LG↗

A Deterministic and Auditable AI Security Risk Assessment Framework with ATLAS Aligned Executable Rules and Formal Verification

Artificial intelligence systems are increasingly deployed in high impact and safety critical settings, yet security assessment remains difficult to reproduce and defend under audit. Existing approaches often rely on narrative checklists or assessor driven scoring, and they lack an explicit, machine evaluable mapping from observable engineering artefacts to stable technique level outcomes. We present an evidence driven AI security assessment framework that operationalises assessment as a deterministic decision function. The framework normalises heterogeneous artefacts into a project independent Control ID taxonomy scored on a bounded four level ordinal scale, compiles technique level predicates from a pinned MITRE ATLAS snapshot via an explicit mitigation to control mapping, and outputs technique indexed feasibility and impact levels with traceable links back to the triggering evidence. We package all normative choices as a versioned assessment policy object to support repeatable reassessment across snapshots. To ensure semantic correctness, we formally verify boundedness, totality, ordered semantic consistency, and monotonicity of the compiled evaluator over the full declared score domain. We evaluate the framework on five public open source AI projects pinned to explicit repository snapshots, quantify before and after changes under a unified hardening intervention, and validate responsiveness to real engineering changes through fork based implementations of Software Bill of Materials (SBOM) generation and Continuous integration (CI) security scanning gates. Results show consistent downward shifts in feasibility profiles under strengthened observable controls, while worst case residual feasibility persists when technique specific core controls remain absent from the evidence scope.

cs.AI↗

Agentic Wireless Digital Twin Construction and Calibration Using Real-World Measurements

Wireless digital twins (WDTs) are promising enablers for developing and evaluating AI-native radio access networks, yet constructing a high-fidelity WDT typically requires substantial manual effort to integrate heterogeneous information and infer unknown propagation-related parameters. This paper proposes Agentic WDT (AWDT), an end-to-end agentic framework for autonomous WDT construction and calibration using readily available environmental information and measurements readily obtainable from commercial smartphones. AWDT comprises three agents: EnvAgent constructs the propagation environment, OpAgent infers BS and sector configurations, and MatAgent calibrates radio material properties. The agents iteratively refine the WDT using discrepancies between LTE/NR reference signal received power (RSRP) measurements and ray-tracing predictions. Real-world experiments show that AWDT reduces the RSRP prediction MAE from 11.46 to 5.25 dB, demonstrating a substantial improvement in ray-tracing fidelity. Evaluation with an independent measurement system further demonstrates cross-device transferability with lightweight device-specific bias adaptation, highlighting the potential of agentic AI for scalable WDT construction and calibration with limited prior knowledge.

cs.IT↗

DRelay: Global Draft Context for Prefix-Aware Parallel Speculative Decoding Repair

Parallel drafting reduces the drafting overhead of speculative decoding for large language models (LLMs), but its gains remain limited by the accepted prefix length. Even when the correct token is present in the candidate pool, a single early selection error prevents subsequent predictions from being used. We propose DRelay, which uses global information from the entire draft block to perform prefix-aware selective repair of candidate selections before target-model verification. DRelay bases its decisions on candidate correlations and the selected path: a global reader extracts predictive information across positions for each candidate. While a causal selector combines candidate-level information extracted by the global read with the tokens selected at preceding positions to determine whether the native choice at the current position is consistent with the global evidence and the selected prefix. It then decides whether to retain or replace the token, thereby repairing early errors and extending the accepted prefix. We further jointly train the draft backbone and the selector, combining candidate-support learning with a repair objective, while weighting the repair loss according to each block position's potential contribution to the consecutive accepted prefix. Across eight diverse benchmarks on an H800 GPU, DRelay consistently improves both average acceptance length and end-to-end decoding performance over DFlash, Domino, and DSpark. Under SGLang serving, DRelay improves average end-to-end speedup over DFlash, Domino, and DSpark by 14.7%-16.8%, 8.7%-9.3%, and 8.1%-9.3%, respectively.

cs.AI↗

Integer reachability in VASS with transfers: a refined complexity analysis

Integer reachability is NP-complete for vector addition systems with states (VASS), but becomes PSPACE-complete in the presence of transfer operations. We refine this complexity gap for single-transfer VASS by identifying structural features of transfers responsible for the increase in complexity. Each system induces a transfer graph whose vertices are counters and whose edges represent possible transfers. We classify its vertices as good or bad, according to the branching and cyclic structure of their reachable subgraphs. Let $b$ be the number of bad vertices. We show that every positive instance admits a polynomially verifiable certificate of size $|I|^{O(b+1)}$, where $|I|$ is the input size. Consequently, integer reachability for single-transfer VASS can be decided in nondeterministic time $|I|^{O(b+1)}$; in particular, it belongs to NP for every class with a bounded number of bad counters. Conversely, we show that bad counters provide sufficient structural power to encode space-bounded computation. For every transfer graph with $b$ bad vertices, we construct a single-transfer VASS that encodes the acceptance of a Turing machine using $b^{O(1)}$ tape cells. This yields PSPACE-hardness for every polynomial-time constructible family of transfer graphs containing linearly many bad vertices. Our results isolate the transfer patterns responsible for the complexity of integer reachability.

cs.FL↗

On the Classical and Parameterized Complexity of Strong Odd Coloring

A strong odd $k$-coloring of a graph $G$ is a proper $k$-coloring such that every color appearing in the neighborhood of a non-isolated vertex appears an odd number of times. The minimum $k$ for which $G$ admits a strong odd $k$-coloring is the \emph{strong odd chromatic number}, denoted by $χ_{\text{so}}(G)$, of $G$. Given a graph $G$ and an integer $k$, \textsc{strong odd $k$-colorability} problem asks whether $G$ admits a strong odd $k$-coloring. It is known that STRONG ODD $k$-COLORABILITY is NP-complete in general graphs. In this paper, we prove that the problem is NP-complete on perfect elimination bipartite graphs for $k\geq3$, which is a subclass of bipartite graphs. Furthermore, we show that $χ_{\text{so}}(G)$ is inapproximable within a factor of $O(n^{\frac{1}{2}-\varepsilon})$ for every $\varepsilon>0$. On the positive side, we obtain a linear time algorithm to compute an optimal strong odd coloring for block graphs. From a parameterized perspective, we present an FPT algorithm for STRONG ODD $k$-COLORABILITY when parameterized by treewidth. Moreover, we show that the problem cannot be solved in time $(k-\varepsilon)^{\texttt{tw}}n^{O(1)}$ for every $k\geq3$ and $\varepsilon>0$ when parameterized by treewidth under SETH. Furthermore, we show that STRONG ODD $k$-COLORABILITY does not admit a polynomial kernel when parameterized by feedback vertex set. Lastly, we prove that STRONG ODD $k$-COLORABILITY is W[1]-hard when parameterized by clique-width.

cs.DM↗

An improved lower and upper bound of the k-limited domination number

We continue the study of $k$-limited domination in graphs, a domination variant in which each vertex of a dominating set may dominate at most $k$ vertices outside the set. This concept extends classical domination by incorporating capacity constraints on dominating vertices. We improve the general lower bound for the $k$-limited domination number $γ_k^L(G)$ by employing the concept of $k$-capacitated domination. As a consequence, we characterize graphs satisfying $γ_k^L(G)=\lceil \frac{n}{k+1} \rceil$. We also refine the known upper bound for graphs with $k<δ(G)$ using the $k$-limited packing number. Under this condition, we describe graphs attaining $γ_k^L(G)=n-k$. These results extend and unify previous investigations for the case $k=1$ and provide a complete characterization of graphs attaining the extreme values of the $k$-limited domination number under the considered assumptions.

math.CO↗

GridSMR: Causal Compression for Sharded Blockchains

We present GridSMR, a sharded blockchain that scales execution horizontally while allowing dependent cross-shard operations to progress within a single block. Existing sharded systems typically place coordination between dependent cross-shard steps, making latency grow with causal depth. GridSMR localizes atomicity to individual accounts and executes cross-account work asynchronously. Using an execute-before-agree architecture, dependent operations execute across shards as they become available, while consensus later validates and commits the resulting schedule. This enables Causal Compression: cross-shard latency need not grow with the causal depth of a computation. GridSMR scales single-validator execution to 1.07M requests/s and four-validator execution to 193K committed requests/s, while reducing 16-hop causal-chain latency by 8.2x versus deferred execution.

cs.DC↗

Detection of kilosecond hard lags in the new pulsating ULX candidate NGC 7456 ULX-1

Context. Ultraluminous X-ray sources (ULXs) are thought to be powered, in many cases, by super-Eddington accretion onto compact objects. While soft X-ray lags have been detected in several ULXs, hard lags remain rare and poorly understood. Aims. We investigate the temporal and energy dependence of X-ray lags in NGC 7456 ULX-1 to constrain their physical origin and probe the super-Eddington accretion flow. Methods. We analyzed the two deepest XMM-Newton observations, taken in 2018 and 2023. Hard (1-10 keV) and soft (0.3-1 keV) light curves were cross-correlated using adaptive-binning techniques optimized for Poissonian low-count data. We measured lags over consecutive 10 ks intervals and investigated their energy dependence. Spectra were modeled using thermal and Comptonization models. Results. We detect significant hard X-ray lags in both observations, with the hard emission delayed by $\sim 10^3$ s during phases of rapid flux variability. The delays are primarily driven by the lowest-energy photons. Spectral modeling indicates a Comptonization-dominated flow comprising a cooler, extended outer region and a hotter, compact inner flow embedded in an optically thick wind. We interpret the delays as the combined effect of inward propagation of accretion-rate fluctuations and photon diffusion within the dense outflow. Fluctuations first enhance the soft-emitting outer regions and then propagate toward the hotter inner flow, where photons undergo stronger Comptonization before escaping with a kilosecond delay. The small inferred inner emitting radius disfavors an intermediate-mass black hole accretor. Conclusions. The sign, amplitude, and energy dependence of the delays disfavor standard reverberation. Propagation-driven variability coupled with radiative transfer in optically thick winds appears to play a major role in shaping the timing properties of super-Eddington accretion flows.

astro-ph.HE↗

ibUMAP: Coherent and Scalable Field Evaluation for UMAP Optimization

UMAP achieves scalable layout optimization through stochastic negative sampling. However, this stochasticity can lead to unstable embeddings across reruns and downstream reuse, as the estimated repulsive forces depend on the ordering of sampling events. We present ibUMAP, a coherent field-based alternative that evaluates attraction and repulsion from a shared embedding snapshot and applies them synchronously. Its degree-weighted repulsive field is motivated by the conditional expectation of negative sampling for a fixed embedding and represented by three scalar moments, which are evaluated efficiently on CPUs and GPUs using an interpolation-based FFT scheme. This formulation avoids explicit all-pairs computations while inducing optimization dynamics that differ from those of standard online UMAP. Controlled experiments show that synchrony and kernel capping alter the local-global fidelity trade-off, whereas FFT evaluation produces small average changes in final quality. End-to-end benchmarks show median speedups of 3.29x unseeded and 5.79x seeded over umap-learn on CPU, and 1.44x over cuML on million-scale datasets under unseeded GPU execution. These gains accompany greater run-to-run stability and measurable fidelity trade-offs.

cs.LG↗

Unflattening by Flattening -- How Input Distributions Shape Output Variance in Angle-Encoded Circuits

Barren plateaus hinder training of parameterized quantum circuits by making loss gradients exponentially small in the number of qubits. For angle-encoded product states, we show how the input distribution affects output variation through algebraic input purity, i.e. the input's overlap with the circuit's dynamical Lie algebra. With the specified readouts and Haar or exact group 2-design sampling, matchgate circuits on $n$ qubits retain output variance of order $1/n$ for every pure product input. The off-diagonal family instead has zero output on computational-basis inputs at any depth and for every parameter choice. Independent uniform angles yield mean variance of the same order as matchgates. A count of diagonal Pauli strings identifies these zero-output endpoints. For any fixed dataset of nonzero inputs with at least $2n$ coordinates, we construct a classical preprocessing certificate. A shared random rotation is accepted only when the dataset's mean purity passes a computable threshold. This certifies mean output variance of order $1/n$ for the off-diagonal family before circuit execution, with a constant expected number of trials. Binary and ternary weighted encodings also recover the independent-angle mean purity from one uniform scalar input. Numerical experiments examine training behavior and extensions to larger algebras. The guarantee concerns output variance averaged over inputs, while successful learning also depends on the task and information preserved by the encoding.

quant-ph↗

An $O(4^{\log^* n})$ Bound for the KLS Constant

The Kannan--Lovász--Simonovits (KLS) conjecture asks whether every isotropic log-concave probability measure on $\mathbb R^n$ has a Cheeger constant bounded below by a universal positive constant. The best previous upper bound is $ψ_n\lesssim\log^{1/4}n$, due to Letwin [Let26]. We prove that $ψ_n\le C \cdot 4^{\log^*(n+2)}$ for a universal constant $C$, where $\log^*x$ is the least number of successive natural logarithms needed to bring $x$ to at most one. We also prove that $C_P(μ)\le C'16^{\log^*(n+2)}$ for every isotropic log-concave probability measure $μ$ on $\mathbb R^n$, with a universal constant $C'>0$.

math.PR↗

On skew-symmetric distributions and their use in Monte Carlo sampling algorithms: coordinate-free, Gibbs-style and manifold versions of the Barker proposal

Skew-symmetric probability distributions provide a principled mechanism for incorporating gradient information into Markov chain Monte Carlo algorithms. Here we review the (preconditioned) Barker proposal, a Metropolis--Hastings algorithm built on skew-symmetric distributions, and motivate its design. We then introduce three natural extensions. First, we propose coordinate-free variants of the Barker algorithm. Second, we introduce a Gibbs-style Barker algorithm that re-evaluates the gradient at each partially updated coordinate. Third, we derive a simplified manifold Barker algorithm, producing a manifold sampler with enhanced robustness compared to natural comparators. Numerical experiments demonstrate that the Gibbs-style variant improves raw sampling efficiency on correlated targets, that the coordinate-free variants offer limited practical advantage over the standard Barker proposal once computational costs are accounted for, and that the simplified manifold Barker algorithm can achieve significant advantages over simplified manifold MALA when the local geometric structure of the target is irregular or unreliable.

stat.CO↗

Fixed-Time Voltage Regulation in Distribution Networks with Impedance Awareness

This letter introduces an optimization-based fixed time control algorithm for solving the voltage regulation problem of a radial and balanced power distribution network. The proposed algorithm requires no prior knowledge of the exact network impedance (resistance and reactance), yet guarantees voltage convergence to the predefined safe limit within a fixed time-window. We embrace results from Fixed-time stability (FxTs) and Control Lyapunov function (CLF) to analyze the stability and robustness of the underlying algorithm and then synthesize it leveraging the Quadratic Programming (QP) approach. We first analytically provide sufficient conditions on the control gains and the design parameters, that can ensure voltage regulation in fixed-time. Thereafter we transform the existing voltage regulation problem into an equivalent QP-based optimization framework and translate the aforesaid conditions to find feasible solutions for voltage stability. We empirically verify the efficacy of our proposed algorithm on the IEEE-33 bus distribution network and compare its performance against other existing robust control methods.

eess.SY↗

Code-Switching Spoken Language Identification as Multi-Label Set Prediction

Code-switched (CS) speech leaks through the monolingual language identification (LID) filters used to curate massive speech corpora, calling for CS-aware LID (CS-LID). We formulate utterance-level CS-LID as multi-label language-set prediction and propose a set generator that directly outputs the languages in an utterance, comparing it against atomic-pair and score-based classification baselines. Oracle Top-k is the strongest baseline, but thresholding fails because no single threshold separates CS from monolingual speech. Our set generator predicts the correct language count on unseen pairs without assuming the number of languages, but underperforms oracle Top-k in exact set accuracy. Our analysis identifies the key obstacles to robust CS-LID: oracle cardinality, threshold instability, language bias in CS training data, and the synthetic-to-real gap.

eess.AS↗

A Multi-Agent LLM Framework for Personalized Health Checkup Interpretation and Guidance

Personalized interpretation of health checkup results requires reasoning across longitudinal records, medical knowledge, lifestyle guidance, and healthcare navigation. We present a multi-agent large language model (LLM) system that identifies multiple intents, maps each to a task-specific agent, executes them in parallel, and synthesizes their outputs. We compared answers generated in Single Agent and Multi Agent settings on 120 Korean compound queries combining two to four requirements, using synthetic health checkup records. The Multi Agent improved the weighted LLM-judge score from 1.695 to 1.797 (p = 0.027), and three additional LLM judges showed consistent improvements ($Δ$ = +0.111 to +0.186, all p < 0.05). The gains came from usefulness, consistency, and the handling of every requirement in compound queries, whereas numerical accuracy and grounding improved significantly under only one of the four judges and medical safety did not differ, and critical failures occurred at similar rates (Single Agent 15.0% vs. Multi Agent 13.3%). Two human evaluators preferred Multi Agent in 66.7% and 68.3% of pairwise comparisons. Multi Agent execution increased latency and cost by 1.31$\times$ and 2.02$\times$, respectively. In exploratory subgroup analyses, the improvement was concentrated in queries involving personal-record lookup.

cs.AI↗

Uncertainty-Guided Handshake: Efficient Human-in-the-Loop Refinement for Surgical-Grade Glioma Segmentation

While state-of-the-art automated models for medical image segmentation achieve high mean performance, they frequently suffer from localized, catastrophic failures that preclude safe clinical deployment, particularly in neuro-oncology. Interactive segmentation frameworks mitigate this by incorporating human oversight, but traditionally impose prohibitive cognitive and temporal workloads by requiring clinicians to manually search for errors. In this project, we present an efficient, Hybrid Structural-Aleatoric Human-in-the-Loop framework for glioma segmentation that bridges the gap between automated baseline performance and surgical-grade precision, achieving sub-2.0 mm HD95 on curated benchmarks while providing safety-net routing for structural failures across real-world clinical data. By extracting voxel-wise Test-Time Augmentation (TTA) uncertainty and applying hierarchical topological filtering, our method proactively isolates high-risk structural anomalies. We comprehensively evaluated our approach on a challenging out-of-distribution clinical stress-test cohort (N = 362). Operating under a simulated Human Oracle, the framework improved the Whole Tumor (WT) Dice score from 0.891 to 0.914 and reduced the 95th percentile Hausdorff Distance (HD95) from 5.82 mm to 4.76 mm. Critically for surgical safety, the system rescued severe boundary failures in the Tumor Core, reducing mean HD95 from 17.96 mm to 14.83 mm (improving absolute TC Dice to 0.356). These spatial rescues were achieved while demanding a median interactive workload of just 11.3% of the target volume. Acknowledging this as a simulated upper bound lacking real-world cognitive friction, the framework nevertheless demonstrates a highly Pareto-efficient pathway for safely deploying clinical AI.

cs.CV↗