arXiv Science⌕ Search

arXiv subjects

Ming Sun

Publications and source records attributed to Ming Sun.

At least 55 records · Page 3Linked to original sources

X-ray line diagnostics of the multi-phase gas in the Centaurus cluster core with XRISM/Resolve

We report the multi-temperature structure of the intracluster medium (ICM) in the Centaurus cluster core observed with XRISM/Resolve. Thanks to its high energy resolution, Resolve enables us to measure fine structures of highly ionized emission lines from Si to Fe and to directly determine the excitation temperature and the ionization temperature from the emission line ratio diagnostics. The observed spectrum in the Centaurus core is well-represented by a double-temperature thermal plasma at collisional ionization equilibrium state rather than an isothermal one. The line ratio diagnostics also support this biphasic temperature structure. Particularly, the observed line ratios show a trend of increasing ionization temperature with atomic mass, while the ionization and excitation temperatures of Fe show nearly the same temperature. The resultant line ratios, which are well-represented by the two temperatures ICM, ~ 1.6 and ~ 3 keV, are also fairly consistent with the expected numbers when assuming the radial single-temperature ICM was projected in the cluster core along the line of sight. Due to the limited low-energy sensitivity of the Resolve with the gate valve closed, we investigated the effect of the cool component using the XMM-Newton/RGS spectrum, but it ultimately did not affect our results. The observed flux ratio between the Fe XXV He alpha resonance and forbidden lines shows an about 20% reduction, suggesting the presence of resonant scattering.

astro-ph.HE↗

Equipping LLM with Directional Multi-Talker Speech Understanding Capabilities

Recent studies have demonstrated that prompting large language models (LLM) with audio encodings enables effective speech understanding capabilities. However, most speech LLMs are trained on single-channel, single-talker data, which makes it challenging to directly apply them to multi-talker and multi-channel speech understanding task. In this work, we present a comprehensive investigation on how to enable directional multi-talker speech understanding capabilities for LLMs, specifically in smart glasses usecase. We propose two novel approaches to integrate directivity into LLMs: (1) a cascaded system that leverages a source separation front-end module, and (2) an end-to-end system that utilizes serialized output training. All of the approaches utilize a multi-microphone array embedded in smart glasses to optimize directivity interpretation and processing in a streaming manner. Experimental results demonstrate the efficacy of our proposed methods in endowing LLMs with directional speech understanding capabilities, achieving strong performance in both speech recognition and speech translation tasks.

cs.CL↗

SA-VLA: Spatially-Aware Flow-Matching for Vision-Language-Action Reinforcement Learning

Vision-Language-Action (VLA) models exhibit strong generalization in robotic manipulation, yet reinforcement learning (RL) fine-tuning often degrades robustness under spatial distribution shifts. For flow-matching VLA policies, this degradation is closely associated with the erosion of spatial inductive bias during RL adaptation, as sparse rewards and spatially agnostic exploration increasingly favor short-horizon visual cues. To address this issue, we propose \textbf{SA-VLA}, a spatially-aware RL adaptation framework that preserves spatial grounding during policy optimization by aligning representation learning, reward design, and exploration with task geometry. SA-VLA fuses implicit spatial representations with visual tokens, provides dense rewards that reflect geometric progress, and employs \textbf{SCAN}, a spatially-conditioned annealed exploration strategy tailored to flow-matching dynamics. Across challenging multi-object and cluttered manipulation benchmarks, SA-VLA enables stable RL fine-tuning and improves zero-shot spatial generalization, yielding more robust and transferable behaviors. Code and project page are available at https://xupan.top/Projects/savla.

cs.RO↗

SLM-TTA: A Framework for Test-Time Adaptation of Generative Spoken Language Models

Spoken Language Models (SLMs) are increasingly central to modern speech-driven applications, but performance degrades under acoustic shift - real-world noise, reverberation, and microphone variation. Prior solutions rely on offline domain adaptation, which is post-hoc, data-intensive, and slow. We introduce the first test-time adaptation (TTA) framework for generative SLMs that process interleaved audio-text prompts. Our method updates a small, targeted subset of parameters during inference using only the incoming utterance, requiring no source data or labels. This stabilizes token distributions and improves robustness to acoustic variability without degrading core task accuracy. Evaluated on automatic speech recognition, speech translation, and 19 audio understanding tasks from AIR-Bench, our approach yields consistent gains under diverse corruptions. Because adaptation touches only a small fraction of weights, it is both compute- and memory-efficient, supporting deployment on resource-constrained platforms. This work enhances the robustness and adaptability of generative SLMs for real-world speech-driven applications.

cs.SD↗

InstantViR: Real-Time Video Inverse Problem Solver with Distilled Diffusion Prior

Video inverse problems are fundamental to streaming, telepresence, and AR/VR, where high perceptual quality must coexist with tight latency constraints. Diffusion-based priors currently deliver state-of-the-art reconstructions, but existing approaches either adapt image diffusion models with ad hoc temporal regularizers - leading to temporal artifacts - or rely on native video diffusion models whose iterative posterior sampling is far too slow for real-time use. We introduce InstantViR, an amortized inference framework for ultra-fast video reconstruction powered by a pre-trained video diffusion prior. We distill a powerful bidirectional video diffusion model (teacher) into a causal autoregressive student that maps a degraded video directly to its restored version in a single forward pass, inheriting the teacher's strong temporal modeling while completely removing iterative test-time optimization. The distillation is prior-driven: it only requires the teacher diffusion model and known degradation operators, and does not rely on externally paired clean/noisy video data. To further boost throughput, we replace the video-diffusion backbone VAE with a high-efficiency LeanVAE via an innovative teacher-space regularized distillation scheme, enabling low-latency latent-space processing. Across streaming random inpainting, Gaussian deblurring and super-resolution, InstantViR matches or surpasses the reconstruction quality of diffusion-based baselines while running at over 35 FPS on NVIDIA A100 GPUs, achieving up to 100 times speedups over iterative video diffusion solvers. These results show that diffusion-based video reconstruction is compatible with real-time, interactive, editable, streaming scenarios, turning high-quality video restoration into a practical component of modern vision systems.

cs.CV↗

Unveiling Chemical Enrichment in the Abell 2029 Core with XRISM, XMM-Newton, and Chandra

We present new measurements of the chemical abundance pattern in the core of the nearby galaxy cluster Abell~2029, based on XRISM observations with Resolve (37 ks) and Xtend (500 ks), combined with archival data from XMM-Newton (EPIC, RGS) and Chandra. Fe abundances derived from Resolve, Xtend, and EPIC are broadly consistent, while RGS gives systematically lower values. Because the XRISM gate valve remained closed during these observations, Resolve spectral fitting is restricted to the 2--10 keV band, providing reliable constraints only for elements with strong lines in this band (S, Ar, Ca, Fe, Ni). Abundances of the $α$-elements are therefore derived using complementary observations from Xtend, EPIC, RGS, and Chandra. We construct an average X/Fe pattern in the cluster core by using Resolve exclusively for S/Fe, Ar/Fe, Ca/Fe, and Ni/Fe, and RGS + Xtend for O/Fe. The Ne/Fe ratio is averaged from Xtend, EPIC, RGS, and Chandra measurements; Mg/Fe from EPIC and Chandra measurements; and Si/Fe from Xtend, EPIC, and Chandra. Comparison with the supernovae yield models indicates that the observed abundance pattern in A2029 core is best reproduced by a combination of core-collapsed yields from low-metallicity progenitors ($Z_{\rm init}=0.001$) and a sub-Chandrasekhar-mass, double-degenerate Type Ia model. Additionally, we find an excess in Ca abundance in the core of A2029 that cannot be reproduced by the standard supernovae yield models.

astro-ph.GA↗

The relationship between warm and hot gas-phase metallicity in massive elliptical galaxies and the influence of AGN feedback

Warm ionized gas is ubiquitous at the centers of X-ray bright elliptical galaxies. While it is believed to play a key role in the feeding and feedback processes of supermassive black holes, its origins remain under debate. Existing studies have primarily focused on the morphology and kinematics of warm ionized gas. This work aims to provide a new perspective on warm (10,000 K) ionized gas and its connection to X-ray-emitting hot gas (>10^6 K) by measuring and comparing their metallicities. We conducted a joint analysis of 13 massive elliptical galaxies using MUSE/VLT and Chandra observations. Emission-line ratios were measured for the warm ionized gas using MUSE observation, and used to infer the ionization mechanisms and derive metallicities of the warm ionized gas using HII, and LIN(E)R calibrations. We also computed the warm phase metallicity using X-ray/EUV, and pAGB stars models. For two sources at higher redshift, direct Te method was also used to measure warm gas metallicities. Our observations reveal that most sources exhibit composite ionization, with contributions from both star formation and LINER-like emission. A positive linear correlation was found between the gas-phase metallicities of the warm and hot phases, ranging from 0.3 to 1.5 Zsun, and suggest the intimate connection between the two gas phases, likely driven by gas cooling and/or mixing. In some sources the warm gas metallicity shows a central drop. A similar radial trend has been reported for the hot gas metallicity in some galaxy clusters. The ionization mechanisms of cooling flow elliptical galaxies are diverse, suggesting multiple channels for powering the warm ionized gas. The large variation in the warm gas metallicity further suggests that cold gas mass derived under the assumption of solar metallicity for the CO-to-H2 conversion factor needs to be revised by approximately an order of magnitude.

astro-ph.GA↗

Constraining star formation in M87 using deep HST UV data

We analyzed the deepest Hubble Space Telescope (HST) F275W ultraviolet (UV) imaging of M87 to obtain the most robust constraints on its star formation rate (SFR) and star formation history (SFH). After removing the galaxy continuum and globular clusters, we detected an excess of UV point sources near the center. By comparing their colors to young stellar source (YSS) colors generated by stochastically simulated star formation (SF) for various SFRs and SFHs, we ruled out their origin as a UV-upturn population and identified them as YSS. We found an extremely low SFR of $\sim 2\times10^{-5}$ M$_\odot$ yr$^{-1}$ in M87, with evidence of a weak starburst $\sim$125 Myr ago that formed $\sim 1000$ M$_\odot$ of stars. Unlike other cool-core clusters where SF is stronger and directly linked to cooling gas, we found no spatial correlation between YSS and H$α$ filaments. Comparing SF activity with M87's AGN outburst history suggests that recent AGN feedback events ($\lesssim$12 Myr ago) neither triggered nor were associated with any detectable SF, however, earlier outbursts may have triggered weak starbursts. We detected UV filaments co-spatial with H$α$ filaments with similar lengths and widths, though they are obscured by dust near the center. These filaments are likely powered by metal-line emission from collisional ionization, suggesting ongoing low-level precipitation of the intracluster medium. Our results indicate that AGN feedback has quenched SF significantly in M87 for at least 200 Myr, even though some precipitation persists. Additionally, we identified a hotspot created by the counterjet, with the spectral index also constrained.

astro-ph.GA↗

Introducing the Descriptive Parametric Model: Gaseous Profiles for Galaxies, Groups, and Clusters

We develop and present the Descriptive Parametric Model (DPM), a tool for generating profiles of gaseous halos (pressure, electron density, and metallicity) as functions of radius, halo mass, and redshift. The model assumes single-phase, spherically symmetric, volume-filling warm/hot gas. The DPM framework enables mock observations of the circumgalactic medium (CGM), group halos, and clusters across a number of wavebands including X-ray, sub-millimeter/millimeter, radio, and ultraviolet (UV). We introduce three model families calibrated to reproduce cluster profiles while having different extrapolations to the CGM -- (i) self-similar halos, (ii) a reduced gas model for lower halo masses, and (iii) a model with shallower radial slopes at lower masses. We demonstrate how our z=0.0-0.6 models perform when applied to stacked and individual X-ray emission profiles, measurements of the thermal and kinetic Sunyaev-Zel'dovich Effect, electron dispersion measures from fast radio bursts, O VI absorption, and UV-derived pressures. Our investigation supports models that remove baryons from halos more effectively and have shallower profiles at lower halo mass. We discuss biases and systematics when modelling observables using consistent hot gaseous halo models for all wavebands explored. We release the DPMhalo code to encourage the use of our framework and new formulations in future investigations. Included with the DPMhalo distribution is a set of recent observations that allow the reproduction of most plots in this paper.

astro-ph.GA↗

Electron-Ion Equilibration in the Merging Galaxy Cluster A665

Galaxy cluster mergers drive powerful shock fronts that heat the intracluster medium (ICM) and accelerate particles, redistributing the energy in a merger. A665 is one of only a few clusters with such a powerful shock ($\mathcal{M}\sim$3), and it provides a unique opportunity to study the thermalization timescale of the ICM, particularly the electron-ion equilibration timescale. Understanding this timescale is crucial for determining how the energy from the merger is distributed between thermal and nonthermal particle populations. Using $\sim$200 ks of NuSTAR observations, we measure the temperature distribution across the shock to distinguish between two heating models: (1) an instant collisionless model, where ions and electrons are immediately heated at the shock front; and (2) a collisional model, where electrons are initially adiabatically compressed at the shock and subsequently equilibrate with the ions over $\sim$100 Myr. Our measurements favor the delayed-equilibration model, suggesting that electrons do not immediately reach thermal equilibrium with the ions at the shock front and instead equilibrate over $t_{eq} = (4.0 \pm 3.4) \times 10^8$ yr. Additionally, our temperature measurements indicate that the Mach number may be lower than previously estimated ($\mathcal{M} = 2.8 \pm 0.7$), suggesting that the shock strength has been overestimated in past studies. These results add to our understanding of the microphysics governing how thermal energy is distributed in diffuse plasmas like the ICM, with implications for galaxy cluster evolution, large-scale structure formation, and cosmology.

astro-ph.CO↗

Multi-Channel Differential ASR for Robust Wearer Speech Recognition on Smart Glasses

With the growing adoption of wearable devices such as smart glasses for AI assistants, wearer speech recognition (WSR) is becoming increasingly critical to next-generation human-computer interfaces. However, in real environments, interference from side-talk speech remains a significant challenge to WSR and may cause accumulated errors for downstream tasks such as natural language processing. In this work, we introduce a novel multi-channel differential automatic speech recognition (ASR) method for robust WSR on smart glasses. The proposed system takes differential inputs from different frontends that complement each other to improve the robustness of WSR, including a beamformer, microphone selection, and a lightweight side-talk detection model. Evaluations on both simulated and real datasets demonstrate that the proposed system outperforms the traditional approach, achieving up to an 18.0% relative reduction in word error rate.

eess.AS↗

XRISM Reveals Complex Multi-Temperature Structures in the Abell 2029 Galaxy Cluster

We present $\sim$500 ks XRISM observations covering the central and two northern regions of the Abell 2029 galaxy cluster. Resolve enables us to distinguish multiple emission lines from hydrogen-like and helium-like iron (Fe) ions. This study focuses on the multi-temperature structure of Abell 2029 using line-ratio diagnostics. Using a single-temperature collisionally ionized equilibrium model, we measure average plasma temperatures of 6.73 keV, 7.61 keV, and 8.14 keV in the central, inner northern, and outer northern regions, respectively, spanning a radial range up to 700 kpc. To further investigate thermal structure, we derive excitation and ionization temperatures by comparing observed emission-line flux ratios with atomic database predictions. Significant deviations from the single-temperature CIE model in the central and inner northern regions indicate the presence of multi-phase gas. The excitation and ionization temperatures range from 2.85 keV to 8.5 keV in the central region, 4.3 keV to 9.8 keV in the inner northern region, and 8.3 keV to 10.4 keV in the outer northern region. These temperature distributions are largely consistent with the previously observed temperature gradient of A2029. However, Resolve detects two notably cooler components--3.42 keV in the central region and $\sim$4.3 keV in the inner northern region--likely associated with displaced cool gas due to gas sloshing. Additionally, we thermally resolve a 2.85 keV gas component at the core of A2029--potentially a significant development in our understanding of gas cooling. We propose that this cooler gas is a direct product of ongoing cooling processes in A2029, having already cooled to its present temperature. If this temperature structure is stable and no heating mechanism is present, this reservoir is likely to cool to even lower temperatures and form stars.

astro-ph.HE↗

Generating Query-Relevant Document Summaries via Reinforcement Learning

E-commerce search engines often rely solely on product titles as input for ranking models with latency constraints. However, this approach can result in suboptimal relevance predictions, as product titles often lack sufficient detail to capture query intent. While product descriptions provide richer information, their verbosity and length make them unsuitable for real-time ranking, particularly for computationally expensive architectures like cross-encoder ranking models. To address this challenge, we propose ReLSum, a novel reinforcement learning framework designed to generate concise, query-relevant summaries of product descriptions optimized for search relevance. ReLSum leverages relevance scores as rewards to align the objectives of summarization and ranking, effectively overcoming limitations of prior methods, such as misaligned learning targets. The framework employs a trainable large language model (LLM) to produce summaries, which are then used as input for a cross-encoder ranking model. Experimental results demonstrate significant improvements in offline metrics, including recall and NDCG, as well as online user engagement metrics. ReLSum provides a scalable and efficient solution for enhancing search relevance in large-scale e-commerce systems.

cs.IR↗

Thinking in Directivity: Speech Large Language Model for Multi-Talker Directional Speech Recognition

Recent studies have demonstrated that prompting large language models (LLM) with audio encodings enables effective speech recognition capabilities. However, the ability of Speech LLMs to comprehend and process multi-channel audio with spatial cues remains a relatively uninvestigated area of research. In this work, we present directional-SpeechLlama, a novel approach that leverages the microphone array of smart glasses to achieve directional speech recognition, source localization, and bystander cross-talk suppression. To enhance the model's ability to understand directivity, we propose two key techniques: serialized directional output training (S-DOT) and contrastive direction data augmentation (CDDA). Experimental results show that our proposed directional-SpeechLlama effectively captures the relationship between textual cues and spatial audio, yielding strong performance in both speech recognition and source localization tasks.

eess.AS↗

Directional Source Separation for Robust Speech Recognition on Smart Glasses

Modern smart glasses leverage advanced audio sensing and machine learning technologies to offer real-time transcribing and captioning services, considerably enriching human experiences in daily communications. However, such systems frequently encounter challenges related to environmental noises, resulting in degradation to speech recognition and speaker change detection. To improve voice quality, this work investigates directional source separation using the multi-microphone array. We first explore multiple beamformers to assist source separation modeling by strengthening the directional properties of speech signals. In addition to relying on predetermined beamformers, we investigate neural beamforming in multi-channel source separation, demonstrating that automatic learning directional characteristics effectively improves separation quality. We further compare the ASR performance leveraging separated outputs to noisy inputs. Our results show that directional source separation benefits ASR for the wearer but not for the conversation partner. Lastly, we perform the joint training of the directional source separation and ASR model, achieving the best overall ASR performance.

cs.SD↗

ALMA-JELLY I: High Resolution CO(2-1) Observations of Ongoing Ram Pressure Stripping in NGC 4858 Reveal Asymmetrical Gas Tail Formation and Fallback

We present new CO(2-1) observations (resolution $\sim1" = 460$pc) of the Coma cluster jellyfish galaxy NGC 4858 obtained from the ALMA-JELLY large program. Analyzing this data alongside complimentary Subaru H$α$ and HST (F600LP / F350LP) observations, we find numerous structural and kinematic features indicative of the effects from strong, inclined ram pressure, including an asymmetric inner gas tail. We estimate a highly-inclined disk-wind angle of $ϕ_{DW} = 75^{+10}_{-27}$. By subtracting a simple circular velocity model, we find (1): gas clumps that are being accelerated by ram pressure, and (2): signatures of gas clumps that had been previously pushed out of the disk but are now falling inwards. We also discuss head-tail morphologies in star complexes within the stellar disk that appear to be RPS-influenced. Lastly, we compare this galaxy to state-of-the-art galaxy ``wind tunnel'' simulations. We find that this galaxy is one of the best nearby examples of strong and inclined ram pressure gas stripping, and of gas that is perturbed by ram pressure but not fully stripped and falls back. We emphasize the importance of torques due to ram pressure in highly-inclined interactions, which help drive gas inwards on the side rotating against the wind, contributing to the formation of asymmetric inner RPS tails.

astro-ph.GA↗

Rethinking Diffusion Posterior Sampling: From Conditional Score Estimator to Maximizing a Posterior

Recent advancements in diffusion models have been leveraged to address inverse problems without additional training, and Diffusion Posterior Sampling (DPS) (Chung et al., 2022a) is among the most popular approaches. Previous analyses suggest that DPS accomplishes posterior sampling by approximating the conditional score. While in this paper, we demonstrate that the conditional score approximation employed by DPS is not as effective as previously assumed, but rather aligns more closely with the principle of maximizing a posterior (MAP). This assertion is substantiated through an examination of DPS on 512x512 ImageNet images, revealing that: 1) DPS's conditional score estimation significantly diverges from the score of a well-trained conditional diffusion model and is even inferior to the unconditional score; 2) The mean of DPS's conditional score estimation deviates significantly from zero, rendering it an invalid score estimation; 3) DPS generates high-quality samples with significantly lower diversity. In light of the above findings, we posit that DPS more closely resembles MAP than a conditional score estimator, and accordingly propose the following enhancements to DPS: 1) we explicitly maximize the posterior through multi-step gradient ascent and projection; 2) we utilize a light-weighted conditional score estimator trained with only 100 images and 8 GPU hours. Extensive experimental results indicate that these proposed improvements significantly enhance DPS's performance. The source code for these improvements is provided in https://github.com/tongdaxu/Rethinking-Diffusion-Posterior-Sampling-From-Conditional-Score-Estimator-to-Maximizing-a-Posterior.

cs.CV↗

Accelerating Diffusion-based Super-Resolution with Dynamic Time-Spatial Sampling

Diffusion models have gained attention for their success in modeling complex distributions, achieving impressive perceptual quality in SR tasks. However, existing diffusion-based SR methods often suffer from high computational costs, requiring numerous iterative steps for training and inference. Existing acceleration techniques, such as distillation and solver optimization, are generally task-agnostic and do not fully leverage the specific characteristics of low-level tasks like super-resolution (SR). In this study, we analyze the frequency- and spatial-domain properties of diffusion-based SR methods, revealing key insights into the temporal and spatial dependencies of high-frequency signal recovery. Specifically, high-frequency details benefit from concentrated optimization during early and late diffusion iterations, while spatially textured regions demand adaptive denoising strategies. Building on these observations, we propose the Time-Spatial-aware Sampling strategy (TSS) for the acceleration of Diffusion SR without any extra training cost. TSS combines Time Dynamic Sampling (TDS), which allocates more iterations to refining textures, and Spatial Dynamic Sampling (SDS), which dynamically adjusts strategies based on image content. Extensive evaluations across multiple benchmarks demonstrate that TSS achieves state-of-the-art (SOTA) performance with significantly fewer iterations, improving MUSIQ scores by 0.2 - 3.0 and outperforming the current acceleration methods with only half the number of steps.

cs.CV↗