arXiv ScienceSearch

SEARCH · arXiv Science

Search arXiv Science

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7Linked to original sources

SPID: Distilled Protein Backbone Generation

Diffusion- and flow-based generative models have recently demonstrated strong performance in protein backbone generation tasks, offering unprecedented capabilities for de novo protein design. However, despite their generation quality, these models are constrained by slow sampling, often requiring hundreds of iterative steps. This computational bottleneck limits their practical utility in large-scale protein discovery, where thousands to millions of candidate structures are needed. To address this challenge, we explore the techniques of score distillation, which has shown great success in reducing the number of sampling steps in the vision domain while maintaining high generation quality. However, a straightforward adaptation of these methods results in unacceptably low designability. We introduce Score Protein identity Distillation (SPID), which resolves this incompatibility by combining few-step generation with inference-time noise scaling. SPID adapts the Score identity Distillation (SiD) framework to both diffusion- and flow-based models without requiring access to pretraining data. Applied to the Proteina flow-matching model, our 16-step generator achieves 94.4% designability, matching the 400-step teacher, while delivering more than a 20-fold reduction in effective backbone-generation time and maintaining comparable diversity and novelty. SPID generalizes across unconditional generation, fold-class conditional generation, and motif scaffolding, and extends to equivariant diffusion architectures, achieving significant reduction in generation time with comparable generation quality to the teacher in all tasks. The resulting reduction in inference cost could facilitate large-scale in silico protein design, thereby advancing diffusion-based models toward real-world protein engineering applications. The PyTorch implementation is available at https://github.com/LY-Xie/SiD_Protein

cs.LG

Random algebraic constructions for extremal and Ramsey problems

Building on Bukh's random algebraic method, we develop a framework for extremal and Ramsey problems involving apex hypergraphs. If $\mathcal{H}$ is a $(d-1)$-partite $(d-1)$-uniform hypergraph with $S$ edges and $\mathcal{H}(t)$ is obtained by adjoining $t$ vertices with common link $\mathcal{H}$, we prove that $\operatorname{ex}(n,\mathcal{H}(t))=Ω_{\mathcal{H}}(n^{d-1/S})$ for $t>9^{S+o_d(S)}$, which is best possible when $\mathcal{H}$ is Sidorenko. Our framework also yields sharper sided Zarankiewicz bounds, quantitative generalized Tur'an bounds, and diagonal multicolor Ramsey constructions. For each fixed $s\geq 2$ and $K\geq 3$, we further prove $\operatorname{r}_K(\mathcal K_{s,t};\mathcal K_n) =Θ_{s,t,K}((n/\log n)^s)$ for $t>9^{s+o(s)}$ for $t>9^{s+o(s)}$, extending a theorem of Alon and Rödl from factorial to exponential $t$. The main ingredients are interpolation on $m$-independent varieties, control of the dependencies imposed by symmetry, and linear spaces of forms whose nonzero members remain regular after a common algebraic slice. Limited edge independence then gives the spectral and local-density estimates needed for the Ramsey application.

math.CO

Find Your Optimal Teacher: Personalized Data Synthesis via Router-Guided Multi-Teacher Distillation

Training student models on synthetic data generated by strong teacher models is a promising way to distilling the capabilities of teachers. However, recent studies show that stronger models are not always optimal teachers, revealing a mismatch between teacher outputs and student learnability. To address this issue, we propose PerSyn (Personalized data Synthesis), a novel synthesis strategy that operates under a new ``Route then Generate'' paradigm to create data tailored to each student model, enabling it to learn more effectively. Specifically, PerSyn first assigns each prompt to its optimal teacher via a query-level router that jointly considers student learnability and teacher response quality. Each teacher then synthesizes data only for its assigned prompts, making the process more efficient than the conventional ``Generate then Select'' paradigm, where all teachers must generate parallel responses for the entire prompt set before constructing the final dataset. Extensive experiments across different model families and scales demonstrate that PerSyn consistently achieves superior or comparable performance to all baselines in instruct tuning and math reasoning settings. Further analysis verifies the effectiveness of PerSyn and offers extra insights to propel future research.

cs.LG

An optimal two-step estimation approach for two-phase studies

Two-phase sampling is commonly adopted for reducing cost and improving estimation efficiency. In many two-phase studies, the outcome and some cheap covariates are observed for a large sample in Phase I, and expensive covariates are obtained for a selected subset of the sample in Phase II. As a result, the analysis of the association between the outcome and covariates faces a missing data problem. Complete-case analysis, which relies solely on the Phase II sample, is generally inefficient. In this paper, we study a two-step estimation approach, which first obtains an estimator using the complete data, and then updates it using an asymptotically mean-zero estimator obtained from a working model between the outcome and cheap covariates using the full data. This two-step estimator is asymptotically at least as efficient as the complete-data estimator and is robust to misspecification of the working model. We propose a kernel-based method to construct a two-step estimator that achieves optimal efficiency. Additionally, we develop a simple joint update approach based on multiple working models to approximate the optimal estimator when a fully nonparametric kernel approach is infeasible. We illustrate the proposed methods with various outcome models. We demonstrate their advantages over existing approaches through simulation studies and provide an application to a major cancer genomics study.

stat.ME

Accelerated stochastic first-order method for convex optimization under heavy-tailed noise

We study convex composite optimization problems, where the objective function is given by the sum of a prox-friendly function and a convex function whose subgradients are estimated under heavy-tailed noise. Existing work often employs gradient clipping or normalization techniques in stochastic first-order methods to address heavy-tailed noise. %In this paper, we demonstrate that a vanilla stochastic algorithm---without additional modifications such as clipping or normalization---can achieve optimal complexity for these problems. In this paper, we analyze the first-order oracle complexity of vanilla stochastic algorithms---without additional modifications such as clipping or normalization---for solving these problems. In particular, we establish that an accelerated stochastic proximal subgradient method achieves a first-order oracle complexity for finding an approximate optimal solution in expectation that is universally optimal for smooth, weakly smooth, and nonsmooth convex optimization, as well as for stochastic convex optimization under heavy-tailed noise. Moreover, we derive high-probability first-order oracle complexity bounds for the accelerated stochastic proximal subgradient method under heavy-tailed and sub-Weibull noise, respectively. Numerical experiments are further provided to illustrate the numerical behavior of the methods.

math.OC

High-Resolution Modelling of Coronae and Winds in Solar-type Stars with Varying Rotation Rates I. X-ray Coronae

Stellar coronae are believed to be the main birthplace of various stellar magnetic activities. However, the structures and properties of stellar coronae remain poorly understood. Using the Space Weather Modelling Framework with the Alfvén Wave Solar Model (SWMF-AWSoM) and dynamo-generated surface magnetic maps, here we model the coronae of four solar-type stars. By incorporating the Sun, our work covers a range of stars with the rotation varying from 1.0 to 23.3 $Ω_\odot$ (periods of 25 to 1 days). Guided by observations, we scale the magnetic field strength with increasing rotation, covering a range between 6.0 G to 1200 G approximately. In our models, energy release associated with small-scale magnetic flux is a key source of coronal heating and is essential for reproducing realistic coronal structures. Our models capture dense (1$-$2 orders of magnitude higher than solar values) and ultra-hot ($\sim 10\,\mathrm{MK}$) coronae dominated by closed field structures. Using the CHIANTI atomic database, we also compute synthetic X-ray spectra and derive the corresponding X-ray luminosities $(L_X)$, which follow a scaling law to magnetic field $L_X \propto \langle|\mathbf{B}|\rangle^{1.75}$. Furthermore, the coronal X-ray emission is found to be rotationally modulated by the alternating presence of bright active regions and dark coronal holes. These results provide new insights into the extremely high-energy coronae of rapidly rotating solar-type stars, which differ markedly from the Sun.

astro-ph.SR

Rigidity of one-dimensional point processes via optimal transport

We investigate rigidity phenomena in one-dimensional point processes. We show that the existence of an $L^1$ transport map from a stationary lattice or the Lebesgue measure to a point process is sufficient to guarantee the properties of Number-Rigidity and Cyclic-Factor. We then apply this result to non-singular Riesz gases with parameter $s\in(-2,-1]$, defined in infinite volume as accumulation points of stationarized finite-volume Riesz gases. This includes, for $s=-1$, the well-known one-dimensional Coulomb gas (also called Jellium plasma, or the one-component 1D plasma).

math.PR

Three-dimensional symmetric designs of propriety 3

We define symmetric designs of dimension $n$ and propriety $d$, providing a unifying generalization of several classes of higher-dimensional symmetric designs previously studied. We focus on the case $n=d=3$, which leads to the following question: Can we fill the $v^3$ cells of a $v\times v\times v$ cube with $\{0,1\}$ in such a way that each layer parallel to each face contains a fixed number $k$ of ones, and that for every two parallel layers there are exactly $λ$ positions where they have matching ones? We establish necessary conditions on the parameters $(v,k,λ)$, introduce notions of difference sets and multipliers for these objects, and enumerate small examples up to equivalence. Furthermore, we construct infinite families of these objects using difference sets, symmetric designs, doubly regular tournaments, Hadamard matrices, Latin cubes, and association schemes on triples.

math.CO

Gaillard-Zumino non-invertible symmetries

We uncover an infinite class of novel zero-form non-invertible symmetries in a broad family of four-dimensional models, studied years ago by Gaillard and Zumino (GZ), which includes several extended supergravities as particular subcases. The GZ models consist of abelian gauge fields coupled to a neutral sector, typically including a set of scalars, whose equations of motion are classically invariant under a continuous group $\mathscr{G}$ acting on the electric and magnetic field strengths via symplectic transformations. The standard lore holds that, at the quantum level, these symmetries are broken to an integral subgroup $\mathscr{G}_\mathbb{Z}$. We show that, in fact, a much larger subgroup $\mathscr{G}_\mathbb{Q}$ survives, albeit through non-invertible topological defects. We explicitly construct these defects and compute some of their fusion rules. As illustrative examples, we consider the axion-dilaton-Maxwell model and the bosonic sector of a class of $\mathcal{N}=2$ supergravities of the kind that appear in type II Calabi-Yau compactifications. Finally, we comment on how (part of) these non-invertible zero-form symmetries can be broken by gauging the $\mathscr{G}_\mathbb{Z}$ subgroup of invertible symmetries.

hep-th

Memory-Free Continual Learning with Null Space Adaptation for Zero-Shot Vision-Language Models

Pre-trained vision-language models (VLMs), such as CLIP, have demonstrated remarkable zero-shot generalization, enabling deployment in a wide range of real-world tasks without additional task-specific training. However, in real deployment scenarios with evolving environments or emerging classes, these models inevitably face distributional shifts and novel tasks. In such contexts, static zero-shot capabilities are insufficient, and there is a growing need for continual learning methods that allow models to adapt over time while avoiding catastrophic forgetting. We introduce NuSA-CL (Null Space Adaptation for Continual Learning), a lightweight memory-free continual learning framework designed to address this challenge. NuSA-CL employs low-rank adaptation and constrains task-specific weight updates to lie within an approximate null space of the model's current parameters. This strategy minimizes interference with previously acquired knowledge, effectively preserving the zero-shot capabilities of the original model. Unlike methods relying on replay buffers or costly distillation, NuSA-CL imposes minimal computational and memory overhead, making it practical for deployment in resource-constrained, real-world continual learning environments. Experiments show that our framework not only effectively preserves zero-shot transfer capabilities but also achieves highly competitive performance on continual learning benchmarks. These results position NuSA-CL as a practical and scalable solution for continually evolving zero-shot VLMs in real-world applications.

cs.AI

Streaming Generation for Music Accompaniment

Music generation models can produce high-fidelity coherent accompaniment given complete audio input, but are limited to editing and loop-based workflows. We study real-time audio-to-audio accompaniment: as a model hears an input audio stream (e.g., a singer singing), it has to also simultaneously generate in real-time a coherent accompanying stream (e.g., a guitar accompaniment). In this work, we propose a model design considering inevitable system delays in practical deployment with two design variables: future visibility $t_f$, the offset between the output playback time and the latest input time used for conditioning, and output chunk duration $k$, the number of frames emitted per call. We train Transformer decoders across a grid of $(t_f,k)$ and show two consistent trade-offs: increasing effective $t_f$ improves coherence by reducing the recency gap, but requires faster inference to stay within the latency budget; increasing $k$ improves throughput but results in degraded accompaniment due to a reduced update rate. Finally, we observe that naive maximum-likelihood streaming training is insufficient for coherent accompaniment where future context is not available, motivating advanced anticipatory and agentic objectives for live jamming.

cs.SD

M-CIF: Multi-Scale Alignment For CIF-Based Non-Autoregressive ASR

The Continuous Integrate-and-Fire (CIF) mechanism provides effective alignment for non-autoregressive (NAR) speech recognition. This mechanism creates a smooth and monotonic mapping from acoustic features to target tokens, achieving performance on Mandarin competitive with other NAR approaches. However, without finer-grained guidance, its stability degrades in some languages such as English and French. In this paper, we propose Multi-scale CIF (M-CIF), which performs multi-level alignment by integrating character and phoneme level supervision progressively distilled into subword representations, thereby enhancing robust acoustic-text alignment. Experiments show that M-CIF reduces WER compared to the Paraformer baseline, especially on CommonVoice by 4.21% in German and 3.05% in French. To further investigate these gains, we define phonetic confusion errors (PE) and space-related segmentation errors (SE) as evaluation metrics. Analysis of these metrics across different M-CIF settings reveals that the phoneme and character layers are essential for enhancing progressive CIF alignment.

cs.SD

VividCam: Learning Unconventional Camera Motions from Virtual Synthetic Videos

Although recent video generative models are getting more capable of following external camera controls, imposed by either text descriptions or camera trajectories, they still struggle to generalize to unconventional camera motions, which is crucial in creating truly original and artistic videos. The challenge lies in finding sufficient training videos with the intended uncommon camera motions. To this end, we propose VividCam, a training paradigm that enables diffusion models to learn complex camera motions from synthetic videos, releasing the reliance on collecting realistic training videos. VividCam incorporates multiple disentanglement strategies that isolate camera motion learning from synthetic appearance artifacts, ensuring more robust motion representation and mitigating domain shift. We show that our design synthesizes a wide range of precisely controlled camera motions using surprisingly simple synthetic data. Notably, this synthetic data often consists of basic geometries within a low-poly 3D scene and can be efficiently rendered by engines like Unity. Our video results can be found in https://wuqiuche.github.io/VividCamDemoPage/ .

cs.CV

Monetary Regimes and Trade before the Classical Gold Standard: Evidence from the Latin Monetary Union

This paper reexamines the trade effects of the Latin Monetary Union (LMU), a 19th century agreement to standardize gold and silver coinage among several European countries. The LMU provides a useful setting for studying whether monetary arrangements fostered trade before the classical gold standard, when gold, silver, bimetallic, and paper regimes coexisted. Because some countries already shared other monetary standards, treating all non-member pairs as a single control group mixes pairs with and without alternative forms of monetary coordination. I classify pairs by standard and estimate the LMU effect relative to pairs without a common standard, bringing the comparison closer to those used in the literature on the gold standard and contemporary currency unions. The results suggest that the LMU increased trade between its members by approximately 30\% during its early years, when bimetallism was still credible. These effects subsequently faded, converging to zero by the end of the 1870s. More broadly, these findings also highlight the importance of accounting for the existing monetary regimes when estimating the trade effects of other international policies.

econ.GN

Generalized Additive Decompositions of Symmetric Tensors

This article addresses the Generalized Additive Decomposition (GAD) of symmetric tensors, that is, degree-$d$ forms $f \in \mathcal{S}_d$. From a geometric perspective, a GAD corresponds to representing a point on a secant of osculating varieties to the Veronese variety, providing a compact and structured description of a tensor that captures its intrinsic algebraic properties. We provide a linear algebra method for measuring the GAD size and prove that the minimal achievable size, which we call the GAD-rank of the considered tensor, coincides with the rank of suitable Catalecticant matrices, under certain regularity assumptions. We provide a new explicit description of the apolar scheme associated with a GAD as the annihilator of a polynomial-exponential series. We show that if the Castelnuovo-Mumford regularity of this scheme is sufficiently small, then both the GAD and the associated apolar scheme are minimal and unique. Leveraging these results, we develop a numerical GAD algorithm for symmetric tensors that effectively exploits the underlying algebraic structure, extending existing algebraic approaches based on eigen computation to the treatment of multiple points. We illustrate the effectiveness and numerical stability of such an algorithm through several examples, including Waring and tangential decompositions.

math.AC

Sufficient conditions for a digraph to contain: a pre-Hamiltonian cycle and cycles of lengths 3 and 4

Let $D$ be a digraph of order $p\geq5$ with minimum degree at least $p-1$ and with minimum semi-degree at least $p/2-1$. In his excellent and renowned paper, ``Long Cycles in Digraphs" (Proc. London Mathematical Society (3), 42 (1981), Thomassen fully characterized the following for $p=2n+1$: (i) $D$ has a cycle of length at least $2n$; and (ii) $D$ is Hamiltonian. Motivated by this result, and building on some of the ideas in Thomassen's paper, we investigated the Hamiltonicity (when $p$ is even) and pancyclcity (when $p$ is arbitrary) such digraphs. We have given a complete description of whether such digraphs are Hamiltonian ($p$ is even), are pancyclic ($p$ is arbitrary). Since the proof is very long, we have divided it into three parts. In this paper, we provide a full description of the following: (iii) for $k=3$ and $k=4$, the digraph $D$ contains a cycle of length $k$; and (iv) the digraph $D$ contains a pre-Hamiltonian cycle, i.e. a cycle of length $p-1$.

math.CO

Thor: Towards Human-Inspired Whole-Body Reactions for Intense Contact-Rich Environments

Maintaining whole-body stability and motion tracking under large interaction forces remains challenging for humanoids. We present Thor, a reinforcement learning framework for forceful humanoid loco-manipulation. Thor jointly trains lower-body, waist, and upper-body policies with shared whole-body observations and body-specific rewards to coordinate locomotion and force adaptation, waist posture regulation, and upper-body motion tracking. We further introduce a force-adaptive torso-tilt (FAT2) objective that derives a load-dependent horizontal center-of-mass offset reference from quasi-static moment balance. Capacity-matched simulation blations show that the three-policy architecture improves tracking under large external force disturbances, while real-world ablations demonstrate that FAT2 increases peak pulling capability. On the Unitree G1, Thor achieves mean peak dual-hand pulling forces of 167.7 N and 145.5 N during backward and forward locomotion, exceeding the best-performing baseline by 68.9% and 74.7%, respectively. Real-world demonstrations include opening a fire door with one hand using approximately 60 N of pulling force and towing a 1.7-ton passenger car.

cs.RO

Natural methods of unsupervised topological alignment

In this paper, we consider methods for the diagonal multi-omics integration of heterogeneous datasets. Several approaches to the nature of biological heterogeneity are analyzed and developed to comprehend more clearly the generated differences. Specifically, the extremal trace problems for the coupled Laplacian on sets homeomorphic to the Stiefel manifold embedded in the complex Euclidean space are investigated. The gradient ascent method for the maximization problem is elaborated in the classical terms of functional analysis, which is of significant interest in itself. On this basis, we introduce a novel characteristic of dataset heterogeneity by employing the norm of the difference between the maximum and minimum points.

math.FA