arXiv ScienceSearch

arXiv subjects

Jingyi Zhang

Publications and source records attributed to Jingyi Zhang.

At least 19 recordsLinked to original sources

Uniform-in-Time Boltzmann Mean-Field Limits for Anchored Binary Opinion Dynamics

We study a continuous-time opinion model in which agents interact in pairs and each agent has a fixed anchor. At each interaction epoch, both opinions are updated according to a possibly nonlinear rule with random inputs. Under suitable stability and moment assumptions, we establish uniform-in-time propagation of chaos: as the population grows, the joint law of any fixed number of agents converges, uniformly over time, to the corresponding product of the law solving a nonlinear Boltzmann equation. We first obtain a graphical estimate in total variation on finite time intervals, assuming only that the interaction rule is measurable. We extend this approximation uniformly over time by showing that both the finite system and its nonlinear limit converge exponentially fast to their respective stationary laws. We also prove uniform-in-time convergence in Wasserstein distance of order two under a contraction condition in mean square. For this result, we assume either a bounded state space or a drift condition that keeps moments of some order greater than two uniformly bounded. We examine two applications. For an anchored Friedkin-Johnsen model with random coefficients, we derive an affine stochastic recurrence for the stationary deviation of an opinion from its anchor. We use this recurrence to give conditions for sub-Gaussian or power-law tails and to characterize the exponent in the power-law case. We also study a model of biased assimilation, in which agents favor information that agrees with their current opinions. With the anchor and incoming evidence held at neutral values, we determine when an interaction moves an opinion farther from the neutral value, even if it starts close to it. Simulations illustrate the tail predictions and show how a single peak in the opinion distribution can split into two.

math.PR

AdaptVPR: Route-Aware Hard Positive Generation for Robust Visual Place Recognition

Visual Place Recognition (VPR) localizes a query image by retrieving database images of the same or nearby place, yet its robustness is often degraded by domain shifts arising from illumination, weather, seasonal changes, and dynamic occlusions. One contributing factor is the limited appearance diversity of the same place in existing training data. To address this issue, we propose AdaptVPR, a route-aware generative augmentation framework that constructs same-place hard positives for robust VPR training. AdaptVPR first uses a vision language model to parse scene attributes and estimate editing feasibility, while a rule-based scheduler determines the generation route according to editability scores and risk constraints. The generation process is decomposed into three complementary routes: the Global Appearance Route introduces global scene changes in weather, illumination, and time of day; the Local Occlusion Route inserts plausible dynamic occluders; and the Dual Route combines both types of perturbations to produce more challenging appearance shifts. Each generated candidate is evaluated using a VPR-oriented verification scheme based on geometric consistency and appearance diversity, reducing the risk of structural drift while ensuring sufficient appearance variation. Global candidates are generated once and rejected if verification fails, while Local Occlusion and Dual candidates use verification feedback for limited prompt refinement and regeneration. Using this framework, we construct AdaptCities, containing 160K verified synthetic same-place hard positives. Experiments across multiple VPR baselines and vision foundation backbones show consistent gains on standard benchmarks and substantial improvements under challenging domain shifts, with R@1 gains of up to 9.2%. The source code and data resources are publicly available at https://github.com/chenshunpeng/AdaptVPR.

cs.CV

POSEIDON III: The Aligned Orbit of the Hot Neptune Around the Hot Star WASP-195

Stellar obliquities provide important clues as to the formation and migration histories of planetary systems, but measurements remain scarce for Neptune-mass planets, especially those orbiting hot stars (above the Kraft break). Here we present observations of the Rossiter-McLaughlin effect in the hot-star/hot-Neptune system WASP-195 ($T_{\rm eff}=6470\pm100$ K, $v\sin{i_\star}=10.5\pm1.1$ km s$^{-1}$) obtained with the Keck Planet Finder and NEID spectrographs. A joint analysis of these observations, archival photometry, and archival radial velocities yields a sky-projected stellar obliquity of $\lambda=-10\pm7^\circ$, consistent with spin-orbit alignment. This makes WASP-195 one of the few hot-star/hot-Neptune systems with a measured obliquity. Archival radial velocities from SOPHIE exclude Jupiter-mass planets within approximately 3 au at $5\sigma$ confidence. The aligned and nearly circular orbit is naturally consistent with a history of disk-driven migration, although coplanar high-eccentricity migration or Roche-lobe overflow cannot be ruled out. We also investigate why so few Neptunes around hot stars have measured obliquities. Their scarcity likely reflects a combination of the lower intrinsic occurrence of short-period Neptunes around hot stars and the difficulty of confirming planet candidates in this regime, where rapid stellar rotation broadens spectral lines and hampers conventional radial-velocity confirmation. Rapid rotation also increases the detectability of the Rossiter-McLaughlin effect, a feature that could help to widen the planet confirmation bottleneck while expanding the obliquity census of small planets around hot stars.

astro-ph.EP

Not all generalisation failures can be bought back: four boundaries in affective audio modelling

Models mapping acoustic properties onto affective response underpin applications from music recommendation to sound design, yet are evaluated almost entirely within the corpus they were fitted on. When one fails outside it, the standard response -- more data, or a larger model -- assumes every failure is a shortage of resources. We show it is not, and that the alternative calls for the opposite remedy. Using four corpora of rated sound, four pretrained representations and three corpora of physiological recording, we pushed one mapping across four boundaries an application must cross: to new material, to edited audio, to a sensor in place of a self-report, and to an individual listener. At each we report the ceiling the target permits, the fraction surviving the crossing, and the price in target-side observations of closing the gap. Within a corpus, prediction reaches 84% of the ceiling set by inter-listener agreement. A same-domain corpus swap costs a fifth of that, and a hundred target labels return two-thirds of the loss. Crossing between music and environmental sound costs four-fifths to all of it, and four pretrained representations recover none of it. Against physiological response no information source we constructed exceeds a third of the attainable ceiling. "The model does not generalise" is therefore two diagnoses, not one, with mutually exclusive remedies; treating the second as the first is the more expensive mistake.

eess.AS

Analysis of Error Propagation in Autoencoder-Based Reduced-Order Neural Ordinary Differential Equations

Neural ODE reduced-order models often achieve comparable local prediction accuracy, yet their long-horizon extrapolation behavior can differ substantially. To analyze this discrepancy, we develop a path-integral identity that separates local discrepancy injection from amplification in the learned latent dynamics. The associated multi-step Jacobian norms quantify transport sensitivity and distinguish different propagation regimes. Experiments on the Burgers and Gray--Scott systems exhibit two distinct patterns of error evolution. In Burgers systems, prediction errors remain bounded and are primarily associated with persistent local discrepancies. In contrast, Gray--Scott systems exhibit pronounced amplification during extrapolation, where Jacobian norms serve as sensitivity diagnostics rather than direct indicators of physical prediction accuracy.

math.NA

Adaptive Matrix Multiplication for Dynamic Shapes on Ascend NPUs

Matrix Multiplication (MatMul) faces a "generalization crisis" driven by highly dynamic tensor shapes. This crisis is particularly acute on Ascend NPUs, where explicitly controlled architectures and strict physical constraints render existing GPU-centric optimizations ineffective. To resolve this, we propose AdaptCore, an adaptive framework for universally high-performance MatMul on Ascend NPUs. AdaptCore systematically decouples operator optimization into spatial tiling and instruction orchestration. It first maps dynamic shapes into a hardware-aware 2D tiling taxonomy to balance on-chip capacity limits and multi-core parallelism. Furthermore, it integrates a composable optimization library with a deterministic analytical performance model. By mathematically evaluating hardware state mutations, AdaptCore proactively selects and caches optimal implementations, enabling O(1) overhead runtime dispatching. Evaluations demonstrate that AdaptCore delivers a remarkable 1.85x mean speedup across 80,000 input shapes, and achieves up to a 1.48x acceleration in representative end-to-end models over the highly-tuned native vendor library (ACLNN).

cs.AR

OneEmo: A Unified Multimodal Reasoning Model for Emotion Perception, Understanding, and Interaction

Multimodal Large Language Models (MLLMs) have demonstrated remarkable capabilities in emotional intelligence. However, prevailing research predominantly focuses on task-specific specialization, often neglecting inter-task synergy and leaving latent reasoning potential underexplored. To bridge this gap, we introduce OneEmo, a unified affective generalist capable of mastering emotion perception, comprehension, and interaction. For this purpose, we first construct EmoWorld-130K, a comprehensive dataset that distills specialized affective knowledge into explicit reasoning trajectories via a human-in-the-loop workflow. Supervised fine-tuning on this corpus reveals significant mutual benefits derived from multi-task learning. Second, to fully unlock the latent reasoning potential, we propose Emo-Chord, a novel reinforcement learning strategy that stabilizes optimization through unified multi-task reward allocation. Extensive experiments demonstrate that OneEmo achieves state-of-the-art performance against similarly sized baselines across most benchmarks. Notably, despite having significantly fewer parameters than commercial models, OneEmo delivers highly competitive results. This paper paves the way for more reliable and interpretable affective computing. The code is available at https://github.com/waHAHJIAHAO/OneEmo.

cs.HC

Physical Properties of 6.7 Million Galaxies from the DESI Bright Galaxy Survey: Spectral Fitting and Systematic Tests with Mock Spectra

We present a comprehensive analysis of the physical properties of galaxies in the Dark Energy Spectroscopic Instrument (DESI) Data Release 1 (DR1) Bright Galaxy Survey (BGS), based on full spectral fitting of $\sim 6.7$ million galaxy spectra. Using a customized spectral fitting pipeline, we derive key physical parameters including stellar mass, stellar velocity dispersion, stellar population age, dust attenuation, and emission-line properties. To quantify the reliability and systematic uncertainties of our measurements, we construct a large set of mock spectra that closely reproduce the observed properties of DESI data, including realistic noise and spectral features. By comparing the recovered parameters with the known inputs, we assess the performance of the spectral fitting as a function of stellar continuum signal-to-noise ratio (S/N, defined as the ratio of the median continuum flux to its associated error) and redshift. We find that stellar masses can be robustly recovered with negligible bias for spectra with $\mathrm{S/N} \gtrsim 5$, while low-S/N spectra ($\mathrm{S/N} \lesssim 5$) show a mild systematic overestimation of $\sim 0.1$ dex and increased scatter. Similar trends are observed for stellar population parameters, while emission-line fluxes are recovered with high accuracy and minimal bias. We further validate our stellar mass estimates by comparison with independent measurements from photometric spectral energy distribution fitting, finding good overall consistency within the expected systematic uncertainties. The value-added catalog presented in this work enables a wide range of statistical studies of galaxy evolution with DESI, and provides a foundation for future analyses.

astro-ph.GA

Scaling Limits of Constant-Stepsize SGD at Flat Minima

For stochastic gradient descent (SGD) with a constant stepsize $\alpha$, the invariant law of the iterates, centered at a minimizer, describes the behavior of the algorithm over long time horizons. In the strongly convex case, this invariant law has the familiar $\sqrt{\alpha}$ scaling and a Gaussian limit as $\alpha\downarrow 0$. We show that this behavior changes fundamentally for convex objectives $H$ with flat minima and (sub)quadratic tails. More specifically, we study SGD with Markovian noise generated by a contractive driving chain. For every sufficiently small constant stepsize $\alpha$, we prove existence, uniqueness, and geometric convergence to an augmented invariant law in a Wasserstein distance induced by an $\alpha$-dependent metric. When the minimizer $x_\star$ has local flatness exponent $m\ge2$, meaning that $\nabla^2 H(x)\asymp \lVert x-x_\star\rVert^{m-2} I_d$ as $x\to x_\star$, we obtain a contraction bound with factor $1-c\alpha^{m-1}$, where $c>0$ is a constant. This recovers the factor $1-c\alpha$ in the quadratic case $m=2$. We then analyze the small-stepsize scaling limit. We show that the invariant law concentrates on the scale $\alpha^{1/m}$ and that the rescaled iterates converge weakly to the stationary distribution of the stochastic differential equation $$ dY_t=-h_0(Y_t)\,dt+\Sigma^{1/2}\,dB_t , $$ where $h_0$ is the limiting drift at the minimizer and $\Sigma$ denotes the asymptotic covariance. This recovers the Gaussian limit when $m=2$ and gives generally non-Gaussian stationary limits in the flat case $m>2$. Finally, we give corresponding results for coordinate-separable objectives with unequal flatness exponents.

cs.LG

MobileSAM2: Lightweight Segment Anything for Spatial Intelligence

The recent large video foundation model, SAM2, enables segment anything in both images and videos, serving as a powerful base model for various applications. However, many of such use cases require to operate on resource-constrained devices like mobile phones and laptops. In this work, we aim to make SAM2 more mobile-friendly by distilling the heavyweight SAM2 into a lightweight model, facilitating segment anything in both images and videos on mobile devices. To this end, we propose Hypergraphical Knowledge Distill (HyperKD), which introduces the idea of hypergraph into knowledge distillation, aiming to effectively model and transfer SAM2's generalizable and comprehensive knowledge. HyperKD consists of Temporal HyperKD and Granularity HyperKD that construct hypergraphs to explicitly model and extract the generalizable temporal knowledge and the comprehensive multi-granularity knowledge from SAM2 respectively, which are then distilled into the lightweight student model by aligning it with the constructed hypergraphs. Besides, we present MobileSAM2, a new family of lightweight SAM2 that balances efficiency and effectiveness via searching the best model architectures with HyperKD during model size reduction. Extensive experiments validate MobileSAM2 across multiple benchmarks and show promising generalization performance on embodied AI tasks.

cs.CV

Environmental Imprints on the Assembly of the Cool Gas around Bright Cluster Galaxies

Galaxy clusters represent extreme cosmic laboratories where environmental processes dramatically reshape their constituent galaxies, yet their effect on the gaseous halos of central galaxies remains poorly constrained. Here we present the first statistical mapping of cool gas around massive brightest cluster galaxies (BCGs) at $z\approx0.55$. Using Mg II absorption in stacked sight-line spectra from over a million background quasars observed by the Dark Energy Spectroscopic Instrument, we compare BCGs to a matched sample of field galaxies and trace the radial profile from 40 kpc to 15 Mpc. Our analysis reveals a striking dual environmental signature: within 200 kpc, the circumgalactic medium (CGM) around BCGs is significantly suppressed compared to that of field galaxies, while at larger radii (200 kpc to 10 Mpc) a pronounced excess of cool gas emerges. This clear transition from suppression in the core to enhancement on such large scales delineates a novel observed pattern for gas regulation by the dense environment. It suggests that clusters may not only strip gas in the core but also facilitate its accumulation in the outskirts. Our results provide key observational constraints on theoretical models of environmental processing in and around the most massive dark matter halos.

astro-ph.GA

IB-Flow: Information Bottleneck-Guided CFG Distillation for Few-Step Text-to-Image Generation

While large-scale text-to-image generative models have achieved unprecedented visual performance, their inherent reliance on multi-step iterative solvers incurs severe inference latency. Few-step distillation targeting the Classifier-Free Guidance (CFG) trajectory has emerged as the prevalent dual-dimensional compression paradigm. However, existing frameworks remain subjugated by a coarse-grained blind injection paradigm that perpetually enforces a globally static guidance strength while indiscriminately sampling the supervisor timestep. This state-agnostic design completely disregards the intrinsic nature of image generation as a dynamic evolutionary process characterized by progressive entropy reduction, which not only restricts the performance boundary of few-step compression but also precipitates severe CFG over-conditioning artifacts. To transcend these limitations, we re-examine the distillation procedure through the theoretical lens of Information Theory, formally modeling it as a dynamic mutual information game constrained by the Information Bottleneck (IB) principle. Specifically, we dismantle traditional blind assumptions via a dual-track adaptive framework. To determine the injection target, we propose an instance-aware selection mechanism that transmutes the intractable KL divergence constraint into a zero-overhead closed-form solution predicated on the local vector field norm. To regulate the injection strength, we introduce an entropy-aware schedule that dynamically decays alongside the SNR, applying maximal thrust for initial structural anchoring before smoothly reverting to the natural manifold to refine micro-details. Extensive empirical evaluations corroborate that our framework fundamentally eradicates over-conditioning artifacts, shattering the performance ceiling to achieve SOTA generative fidelity under extremely stringent 2-step configurations.

cs.CV

CoDMD: Copula-aware Distribution Matching Distillation for Fast Video Generation

Few-step distillation for video diffusion models has attracted significant attention, driven by the urgent demand for efficient deployment in real-world scenarios. However, Distribution Matching Distillation (DMD), a leading paradigm, tends to degrade under limited NFE budgets, manifesting in video generation as layout instability, oversaturation, and broken motion dynamics. We trace this failure to a structural limitation: standard DMD is an intra-sample distribution-matching objective with coordinate-wise gradients, and thus imposes no explicit constraint on the relational geometry across batch elements or temporal frames, leaving the underlying copula largely unregulated. Combined with the mode-seeking tendency of its reverse-KL objective, this absence of relational guidance makes DMD prone to collapsing into local optima in the few-step regime. Motivated by this insight, we propose Copula-aware DMD (CoDMD), a lightweight relational regularizer that reuses score estimates already produced by the frozen teacher and the online fake model to construct pairwise relation matrices across samples and frames. These are matched through a supplementary distributional objective that requires no additional networks, datasets, or sampling trajectories. On the Wan-2.1-T2V model series at 1.3B & 14B scales, CoDMD distills 50-step teachers into 4-step students, achieving an approximate 25$\times$ speed-up while attaining VBench scores of 84.46 & 84.87, outperforming prior trajectory-based (rCM 82.81 & 84.05) and distribution-based (DMD 83.38 & 83.81) methods.

cs.CV

OpenClaw-Skill: Collective Skill Tree Search for Agentic Large Language Models

Equipping Large Language Model (LLM) agents with effective skills is crucial for solving complex tasks in real-world systems like OpenClaw. In this work, we aim to develop a framework that automatically constructs such reusable skills to enhance LLMs in tool use, multi-step reasoning, and dynamic environment interaction. To this end, we propose Collective Skill Tree Search (CSTS), a novel tree-search-based skill construction framework that constructs structured, diverse and generalizable tree of skills. The core idea of CSTS is to leverage collective intelligence to jointly search, identify and compose effective skills via two iterative phases: Collective Skill Node Generation (CSN-Gen) and Collective Skill Node Assessment (CSN-Assess). CSN-Gen exploits collective knowledge from multiple models to explore diverse candidate skills for each subtask, enabling comprehensive skill exploration. CSN-Assess employs multiple models as judges to evaluate and select skill nodes with two scoring mechanisms: (1) collective quality scoring that aggregates independent evaluations to produce a robust estimate of skill effectiveness, and (2) collective transferability scoring that explicitly verifies whether a skill generalizes well across different models. With CSTS, we construct a set of comprehensive tree of skills along with skill-augmented training data, enabling models to effectively learn and utilize skills. Besides, we introduce Collective Skill Reinforcement Learning, which actively selects multiple relevant skills from the tree to broaden solution-space exploration, avoid being trapped by a single skill and its resulting homogeneous or suboptimal solutions. As a result, our trained model, OpenClaw-Skill, exhibits outstanding agentic capabilities in long-horizon planning, tool use and generalization over challenging benchmarks.

cs.AI

Detector Development for HUBS I: Initial Testing of Small-Area TES Microcalorimeters

We report progress on the ongoing development of microcalorimeter detector technology for the Hot Universe Baryon Surveyor (HUBS) mission. We show the results from testing and characterizing selected pixels in a 10$\times$10 microcalorimeter array. The microcalorimeter is based on a Mo/Cu transition-edge sensor (TES) coupled to an Au absorber. To better understand the properties of the devices, we have first measured the energy resolution of a selected pixel in a TES array of the same design with a pulsed laser system that produces 3 eV photons, and found that individual photon peaks are easily resolved with the TES, indicating good performance. We have then exposed the microcalorimeter array to radiation from a $^{55}$Fe source, and found that the pixels tested show energy resolutions as good as 3.7$\pm$0.1 eV at 5.9 keV. The energy resolution is found to vary monotonically with the bias point for all the devices, showing little evidence for the presence of the so-called excess noise. This is consistent with the results from modeling the measured noise spectrum. The effects of thermal crosstalk are evident, leading to the degradation of energy resolution.

astro-ph.IM

Wan-Image: Pushing the Boundaries of Generative Visual Intelligence

We present Wan-Image, a unified visual generation system explicitly engineered to paradigm-shift image generation models from casual synthesizers into professional-grade productivity tools. While contemporary diffusion models excel at aesthetic generation, they frequently encounter critical bottlenecks in rigorous design workflows that demand absolute controllability, complex typography rendering, and strict identity preservation. To address these challenges, Wan-Image features a natively unified multi-modal architecture by synergizing the cognitive capabilities of large language models with the high-fidelity pixel synthesis of diffusion transformers, which seamlessly translates highly nuanced user intents into precise visual outputs. It is fundamentally powered by large-scale multi-modal data scaling, a systematic fine-grained annotation engine, and curated reinforcement learning data to surpass basic instruction following and unlock expert-level professional capabilities. These include ultra-long complex text rendering, hyper-diverse portrait generation, palette-guided generation, multi-subject identity preservation, coherent sequential visual generation, precise multi-modal interactive editing, native alpha-channel generation, and high-efficiency 4K synthesis. Across diverse human evaluations, Wan-Image exceeds Seedream 5.0 Lite and GPT Image 1.5 in overall performance, reaching parity with Nano Banana Pro in challenging tasks. Ultimately, Wan-Image revolutionizes visual content creation across e-commerce, entertainment, education, and personal productivity, redefining the boundaries of professional visual synthesis.

cs.CV

Discovery of low-redshift analogues to "Little Red Dots" in DESI: A later evolutionary stage of compact LRDs?

The James Webb Space Telescope (JWST) has recently discovered a population of compact, red sources at z > 4 known as "Little Red Dots" (LRDs). They are characterized by their V-shaped continuum spectra and prominent broad Balmer emission lines. As their underlying physical nature remains debated and direct study at high-redshift is challenging; therefore, we seek to identify and characterize LRD analogues in the low-redshift universe to constrain their properties and potential evolutionary pathways. We identified five candidates at z = 0.2-0.4 from the Dark Energy Spectroscopic Instrument (DESI) that exhibit spectral energy distributions (SEDs) and broad Balmer emission lines closely resembling their high-redshift counterparts. However, we find significant differences: our low-redshift sample occupies a different region on the Baldwin, Phillips \& Terlevich (BPT) diagram, and their stellar masses are significantly higher, suggesting a more substantial host galaxy contribution. These sources are not necessarily direct local analogues of high-redshift LRDs, but may represent later evolutionary stages of compact, rapidly accreting systems, or systems with related observational properties arising under different physical conditions. This sample provides a valuable laboratory for detailed follow-up studies to elucidate the nature of LRD-like phenomena.

astro-ph.GA

HyperLiDAR: Adaptive Post-Deployment LiDAR Segmentation via Hyperdimensional Computing

LiDAR semantic segmentation plays a pivotal role in 3D scene understanding for edge applications such as autonomous driving. However, significant challenges remain for real-world deployments, particularly for on-device post-deployment adaptation. Real-world environments can shift as the system navigates through different locations, leading to substantial performance degradation without effective and timely model adaptation. Furthermore, edge systems operate under strict computational and energy constraints, making it infeasible to adapt conventional segmentation models (based on large neural networks) directly on-device. To address the above challenges, we introduce HyperLiDAR, the first lightweight, post-deployment LiDAR segmentation framework based on Hyperdimensional Computing (HDC). The design of HyperLiDAR fully leverages the fast learning and high efficiency of HDC, inspired by how the human brain processes information. To further improve the adaptation efficiency, we identify the high data volume per scan as a key bottleneck and introduce a buffer selection strategy that focuses learning on the most informative points. We conduct extensive evaluations on two state-of-the-art LiDAR segmentation benchmarks and two representative devices. Our results show that HyperLiDAR outperforms or achieves comparable adaptation performance to state-of-the-art segmentation methods, while achieving up to a 13.8x speedup in retraining.

cs.CV