arXiv Science⌕ Search

arXiv subjects

Ming Sun

Publications and source records attributed to Ming Sun.

At least 37 records · Page 2Linked to original sources

MMG2Skill: Can Agents Distill In-the-Wild Guides into Self-Evolving Skills?

Abundant procedural knowledge on the Web holds great potential for helping agents solve long-horizon tasks. However, such knowledge is often multimodal, heterogeneous, noisy, and implicitly assumes human executors, making it difficult to use directly as the skills required by agents. To bridge the gap between human-oriented guides and agent-executable skills, we formalize this problem as guide-to-skill learning: converting in-the-wild guides into executable skills and continuously improving them from trajectories observable to the agent. To evaluate the capability of existing agents on this task, we introduce MMG2Skill-Bench, the first benchmark designed for this problem. We further propose MMG2Skill, a closed-loop framework that compiles guides into editable skills, conditions a fixed vision-language model (VLM) agent on these skills during execution, and revises the skills from trajectory-level root-cause feedback without using benchmark scores. Across GUI control, open-ended gameplay, and strategic card play with six VLM backbones, MMG2Skill consistently outperforms vanilla baseline agents in every model-domain setting, achieving macro-average gains of +12.8 to +25.3 percentage points across backbones. Ablation studies show that directly prompting agents with raw guides can degrade performance, while both structured skill construction and trajectory-driven revision are necessary for the observed improvements. On success-inferable tasks, analyzer-based early stopping further prevents late-stage performance regressions and saves 25%-53% of attempts when the success signal is properly calibrated.

cs.CL↗

A Post-starburst Galaxy Undergoing Ram-pressure Stripping at Redshift 3.06

Understanding how galaxies ignite and extinguish their star formation remains a cornerstone question in modern astrophysics. Recent JWST surveys have revealed an overabundance of massive quiescent galaxies in the first billion years of the Universe, challenging current models of galaxy evolution. In the nearby Universe, ram pressure stripping (RPS) is a major environmental mechanism capable of rapidly shutting down star formation, yet direct observation remains scarce at redshift $z\gtrsim1$, and its role at $z>2$ is even poorly constrained by simulations. Here, we utilize JWST and ALMA observations to present direct evidence of RPS in the post-starburst galaxy A2744-JF-z3, residing in a galaxy group at redshift 3.06, the earliest such detection to date. Spectroscopic diagnostics and spectral energy distribution modeling reveal the ongoing removal of cold gas and dust, coincident with the abrupt cessation of star formation. Contrary to hydrodynamical simulations that predict a reduced incidence of RPS at high redshift, our results instead imply that RPS can operate at $z>3$, suggesting a highly stochastic and impulsive stripping within a clumpy, filamentary intra-group and circumgalactic medium. These observations extend environmental quenching well into the epoch of galaxy assembly, highlighting RPS as a previously overlooked decisive pathway to rapid quenching in nascent groups and protoclusters in the early Universe.

astro-ph.GA↗

Solving the cooling flow problem with combined jet-wind AGN feedback

Active galactic nucleus (AGN) feedback is widely viewed as the most promising solution to the long-standing cooling flow problem in galaxy clusters, yet previous models prescribe jet properties inconsistent with accretion physics. We perform an idealized hydrodynamic simulation of a galaxy cluster with no merger history and a relaxed state, with its other properties similar to the Perseus cluster using the MACER framework, incorporating both jets and winds whose properties are constrained by general relativistic magnetohydrodynamic simulations of black hole accretion and observations. The combined feedback reproduces key observables, including cold gas mass, star formation rate, thermodynamic radial profiles, and black hole growth, while jet-only or wind-only models fail. The success arises from turbulence driven by jet-wind shear that enhances kinetic-to-thermal energy conversion, boosting heating efficiency by factors of three and six relative to wind-only and jet-only cases, respectively.

astro-ph.GA↗

Structure-Centric Graph Foundation Model via Geometric Bases

Graph foundation models (GFMs) seek transferable representations across graph domains but are limited by structural heterogeneity and incompatible node feature spaces. We propose Structure-Centric Graph Foundation Models (SCGFM), which treat graph topology as the primary source of transferable knowledge. Modeling graphs as metric measure spaces, SCGFM introduces learnable geometric bases that define a shared structural coordinate system. Graphs are aligned to these bases via Gromov-Wasserstein distances, yielding structure-aligned latent representations that accommodate heterogeneous graph topologies. To address feature incompatibility, SCGFM employs a structure-aware feature re-encoding mechanism that unifies node representations without assuming a fixed feature dimensionality or requiring dataset-specific preprocessing. Experiments on graph- and node-level tasks demonstrate strong in-domain and cross-domain generalization, outperforming existing GFM approaches.

cs.LG↗

Towards Backdoor-Based Ownership Verification for Vision-Language-Action Models

Vision-Language-Action models (VLAs) support generalist robotic control by enabling end-to-end decision policies directly from multi-modal inputs. As trained VLAs are increasingly shared and adapted, protecting model ownership becomes essential for secure deployment and responsible open-source usage. In this paper, we present GuardVLA, the first backdoor-based ownership verification framework specifically designed for VLAs. GuardVLA embeds a stealthy and harmless backdoor watermark into the protected model during training by injecting secret messages into embodied visual data. For post-release verification, we propose a swap-and-detect mechanism, in which the trigger projector and an external classifier head are used to activate and detect the embedded backdoor based on prediction probabilities. Extensive experiments across multiple datasets, model architectures, and adaptation settings demonstrate that GuardVLA enables reliable ownership verification while preserving benign task performance. Further results show that the embedded watermark remains detectable under post-release model adaptation.

cs.RO↗

WebCompass: Towards Multimodal Web Coding Evaluation for Code Language Models

Large language models are rapidly evolving into interactive coding agents capable of end-to-end web coding, yet existing benchmarks evaluate only narrow slices of this capability, typically text-conditioned generation with static-correctness metrics, leaving visual fidelity, interaction quality, and codebase-level reasoning largely unmeasured. We introduce WebCompass, a multimodal benchmark that provides unified lifecycle evaluation of web engineering capability. Recognizing that real-world web coding is an iterative cycle of generation, editing, and repair, WebCompass spans three input modalities (text, image, video) and three task types (generation, editing, repair), yielding seven task categories that mirror professional workflows. Through a multi-stage, human-in-the-loop pipeline, we curate instances covering 15 generation domains, 16 editing operation types, and 11 repair defect types, each annotated at Easy/Medium/Hard levels. For evaluation, we adopt a checklist-guided LLM-as-a-Judge protocol for editing and repair, and propose a novel Agent-as-a-Judge paradigm for generation that autonomously executes generated websites in a real browser, explores interactive behaviors via the Model Context Protocol (MCP), and iteratively synthesizes targeted test cases, closely approximating human acceptance testing. We evaluate representative closed-source and open-source models and observe that: (1) closed-source models remain substantially stronger and more balanced; (2) editing and repair exhibit distinct difficulty profiles, with repair preserving interactivity better but remaining execution-challenging; (3) aesthetics is the most persistent bottleneck, especially for open-source models; and (4) framework choice materially affects outcomes, with Vue consistently challenging while React and Vanilla/HTML perform more strongly depending on task type.

cs.SE↗

CodeTracer: Towards Traceable Agent States

Code agents are advancing rapidly, but debugging them is becoming increasingly difficult. As frameworks orchestrate parallel tool calls and multi-stage workflows over complex tasks, making the agent's state transitions and error propagation hard to observe. In these runs, an early misstep can trap the agent in unproductive loops or even cascade into fundamental errors, forming hidden error chains that make it hard to tell when the agent goes off track and why. Existing agent tracing analyses either focus on simple interaction or rely on small-scale manual inspection, which limits their scalability and usefulness for real coding workflows. We present CodeTracer, a tracing architecture that parses heterogeneous run artifacts through evolving extractors, reconstructs the full state transition history as a hierarchical trace tree with persistent memory, and performs failure onset localization to pinpoint the failure origin and its downstream chain. To enable systematic evaluation, we construct CodeTraceBench from a large collection of executed trajectories generated by four widely used code agent frameworks on diverse code tasks (e.g., bug fixing, refactoring, and terminal interaction), with supervision at both the stage and step levels for failure localization. Experiments show that CodeTracer substantially outperforms direct prompting and lightweight baselines, and that replaying its diagnostic signals consistently recovers originally failed runs under matched budgets. Our code and data are publicly available.

cs.SE↗

KAT-Coder-V2 Technical Report

We present KAT-Coder-V2, an agentic coding model developed by the KwaiKAT team at Kuaishou. KAT-Coder-V2 adopts a "Specialize-then-Unify" paradigm that decomposes agentic coding into five expert domains - SWE, WebCoding, Terminal, WebSearch, and General - each undergoing independent supervised fine-tuning and reinforcement learning, before being consolidated into a single model via on-policy distillation. We develop KwaiEnv, a modular infrastructure sustaining tens of thousands of concurrent sandbox instances, and scale RL training along task complexity, intent alignment, and scaffold generalization. We further propose MCLA for stabilizing MoE RL training and Tree Training for eliminating redundant computation over tree-structured trajectories with up to 6.2x speedup. KAT-Coder-V2 achieves 79.6% on SWE-bench Verified (vs. Claude Opus 4.6 at 80.8%), 88.7 on PinchBench (surpassing GLM-5 and MiniMax M2.7), ranks first across all three frontend aesthetics scenarios, and maintains strong generalist scores on Terminal-Bench Hard (46.8) and tau^2-Bench (93.9). Our model is publicly available at https://streamlake.com/product/kat-coder.

cs.CL↗

QPT V2: Masked Image Modeling Advances Visual Scoring

Quality assessment and aesthetics assessment aim to evaluate the perceived quality and aesthetics of visual content. Current learning-based methods suffer greatly from the scarcity of labeled data and usually perform sub-optimally in terms of generalization. Although masked image modeling (MIM) has achieved noteworthy advancements across various high-level tasks (e.g., classification, detection etc.). In this work, we take on a novel perspective to investigate its capabilities in terms of quality- and aesthetics-awareness. To this end, we propose Quality- and aesthetics-aware pretraining (QPT V2), the first pretraining framework based on MIM that offers a unified solution to quality and aesthetics assessment. To perceive the high-level semantics and fine-grained details, pretraining data is curated. To comprehensively encompass quality- and aesthetics-related factors, degradation is introduced. To capture multi-scale quality and aesthetic information, model structure is modified. Extensive experimental results on 11 downstream benchmarks clearly show the superior performance of QPT V2 in comparison with current state-of-the-art approaches and other pretraining paradigms.

cs.CV↗

Score2Instruct: Scaling Up Video Quality-Centric Instructions via Automated Dimension Scoring

Classical video quality assessment methods generate a numerical score to judge a video's perceived visual fidelity and clarity. Yet, a score fails to describe the video's complex quality dimensions, restricting its applicability. Benefiting from the human-friendly linguistic output, adapting video large multimodal models to VQA via instruction tuning has the potential to address this issue. The core of the approach lies in the video quality-centric instruction data. Previous explorations mainly focus on the image domain, and their data generation processes heavily rely on human quality annotations and proprietary systems, limiting data scalability and effectiveness. To address these challenges, we propose the Score-based Instruction Generation pipeline. Specifically, SIG first scores multiple quality dimensions of an unlabeled video and maps scores to text-defined levels. It then explicitly incorporates a hierarchical Chain-of-Thought to model the correlation between specific dimensions and overall quality, mimicking the human visual system's reasoning process. The automated pipeline eliminates the reliance on expert-written quality descriptions and proprietary systems, ensuring data scalability and generation efficiency. To this end, the resulting Score2Instruct dataset contains over 320K diverse instruction-response pairs, laying the basis for instruction tuning. Moreover, to advance video LMMs' quality scoring and justification abilities simultaneously, we devise a progressive tuning strategy to fully unleash the power of S2I. Built upon SIG, we further curate a benchmark termed S2I-Bench with 400 open-ended questions to better evaluate the quality justification capacity of video LMMs. Experimental results on the S2I-Bench and existing benchmarks indicate that our method consistently improves quality scoring and justification capabilities across multiple video LMMs.

cs.CV↗

Versatile Recompression-Aware Perceptual Image Super-Resolution

Perceptual image super-resolution (SR) methods restore degraded images and produce sharp outputs. In practice, those outputs are usually recompressed for storage and transmission. Ignoring recompression is suboptimal as the downstream codec might add additional artifacts to restored images. However, jointly optimizing SR and recompression is challenging, as the codecs are not differentiable and vary in configuration. In this paper, we present \textbf{Versatile Recompression-Aware Perceptual Super-Resolution (VRPSR)}, which makes existing perceptual SR aware of versatile compression. First, we formulate compression as conditional text-to-image generation and utilize a pre-trained diffusion model to build a generalizable codec simulator. Next, we propose a set of training techniques tailored for perceptual SR, including optimizing the simulator using perceptual targets and adopting slightly compressed images as the training target. Empirically, our VRPSR achieves 10% - 40% bitrate savings based on Real-ESRGAN and S3Diff under H.264/H.265/H.266 single-picture (intra) compression. Besides, our VRPSR facilitates joint optimization of SR and the post-processing model after recompression.

cs.CV↗

Pioneering Perceptual Video Fluency Assessment: A Novel Task with Benchmark Dataset and Baseline

Accurately estimating humans' subjective feedback on video fluency, e.g., motion consistency and frame continuity, is crucial for various applications like streaming and gaming. Yet, it has long been overlooked, as prior arts have focused on solving it in the video quality assessment (VQA) task, merely as a sub-dimension of overall quality. In this work, we conduct pilot experiments and reveal that current VQA predictions largely underrepresent fluency, thereby limiting their applicability. To this end, we pioneer Video Fluency Assessment (VFA) as a standalone perceptual task focused on the temporal dimension. To advance VFA research, 1) we construct a fluency-oriented dataset, FluVid, comprising 4,606 in-the-wild videos with balanced fluency distribution, featuring the first-ever scoring criteria and human study for VFA. 2) We develop a large-scale benchmark of 23 methods, the most comprehensive one thus far on FluVid, gathering insights for VFA-tailored model designs. 3) We propose a baseline model called FluNet, which deploys temporal permuted self-attention (T-PSA) to enrich input fluency information and enhance long-range inter-frame interactions. Our work not only achieves state-of-the-art performance but, more importantly, offers the community a roadmap to explore solutions for VFA.

cs.CV↗

HST view of NGC 5044: Constraints on Filament Widths, Magnetic Support, Multiphase Structure, and Comparison with Cluster Environments

We present new Hubble Space Telescope (HST) imaging of ionised filaments in the brightest group galaxy NGC 5044. These filaments extend several kiloparsecs and have widths of $\sim$50--120 pc, with some as narrow as those in cluster cores and others broader, reflecting the lower confining pressure in groups. Filament width ($W$) scales with ambient pressure ($P$) as $W \propto P^{-0.4}$. Combining HST, ALMA, and MUSE data, we measure column densities and magnetic field strengths. Equipartition fields decline from $\sim$40 $μ$G at the centre to $\sim$20 $μ$G at 5 kpc, about 2--3 times weaker than in clusters. Dynamical stability requires stronger radial fields ($\sim$10$^2$ $μ$G), consistent with simulations and magnetic draping, though such high values exceed Faraday Rotation Measure limits. Turbulence and cosmic rays also contribute support. Group and cluster filaments are stable against gravitational collapse, and ultraviolet imaging reveals no star formation in NGC 5044 ($<$10$^{-3}$ M$_\odot$ yr$^{-1}$). NGC 5044 hosts an ionised gas core within its Bondi radius with $n_e \propto r^{-1}$ and filling factor $f \gtrsim 3 \times 10^{-3}$, that is connected to the extended filaments, suggesting a channel for gas inflow toward the black hole. Group and cluster filaments likely share a common origin, with magnetic fields and AGN feedback preserving their structure. Ambient pressure and dust survival regulate molecular gas formation. Lower-pressure groups favour broader, more diffuse filaments with sporadic molecular clumps and weaker dust shielding, whereas higher-pressure clusters host narrower strands with stronger molecular-ionised gas alignment. We predict that (i) filament width scales with ambient pressure, (ii) filament-coincident Faraday rotation structures emerge at $\leq 0.1$ kpc resolution, and (iii) molecular/ionised gas co-spatiality is weaker in groups than in clusters.

astro-ph.GA↗

Tuning Real-World Image Restoration at Inference: A Test-Time Scaling Paradigm for Flow Matching Models

Although diffusion-based real-world image restoration (Real-IR) has achieved remarkable progress, efficiently leveraging ultra-large-scale pre-trained text-to-image (T2I) models and fully exploiting their potential remain significant challenges. To address this issue, we propose ResFlow-Tuner, an image restoration framework based on the state-of-the-art flow matching model, FLUX.1-dev, which integrates unified multi-modal fusion (UMMF) with test-time scaling (TTS) to achieve unprecedented restoration performance. Our approach fully leverages the advantages of the Multi-Modal Diffusion Transformer (MM-DiT) architecture by encoding multi-modal conditions into a unified sequence that guides the synthesis of high-quality images. Furthermore, we introduce a training-free test-time scaling paradigm tailored for image restoration. During inference, this technique dynamically steers the denoising direction through feedback from a reward model (RM), thereby achieving significant performance gains with controllable computational overhead. Extensive experiments demonstrate that our method achieves state-of-the-art performance across multiple standard benchmarks. This work not only validates the powerful capabilities of the flow matching model in low-level vision tasks but, more importantly, proposes a novel and efficient inference-time scaling paradigm suitable for large pre-trained models.

cs.CV↗

The shocking features in the closest rich galaxy cluster Norma

The merger shocks generated by the collision of galaxy clusters elevate the pressure within the intracluster medium, significantly influencing the evolution of embedded cluster galaxies. We detect a merger shock (Mach number $\sim 1.3$) on the northwest side of the closest rich galaxy cluster Norma (A3627), using XMM-Newton and Chandra data. The textbook ram pressure stripping (RPS) galaxy ESO 137-001 appears to be located in the post-shock region. The shock boosts RPS and may induce the formation of the brightest known X-ray tail behind a cluster late-type galaxy. Another prominent head-tail radio galaxy ESO 137-007, with one of the longest radio continuum tails ($> 500$ kpc), is also likely in the post-shock region. The shock may have reversed the upstream jet to a one-sided radio head-tail morphology. Moreover, the shock can strip and roll the jet cocoon into a vortex ring structure like a `smoke ring' behind the end of the jet as observed by the ASKAP data. Therefore, the cluster merger shock can remarkably change cluster galaxies. Furthermore, Norma is the second brightest non-cool-core cluster following the Coma cluster, with a cool core remnant on its southeast side. Its original cool core may be disrupted by cluster mergers and/or active galactic nuclei.

astro-ph.GA↗

ShiftLUT: Spatial Shift Enhanced Look-Up Tables for Efficient Image Restoration

Look-Up Table based methods have emerged as a promising direction for efficient image restoration tasks. Recent LUT-based methods focus on improving their performance by expanding the receptive field. However, they inevitably introduce extra computational and storage overhead, which hinders their deployment in edge devices. To address this issue, we propose ShiftLUT, a novel framework that attains the largest receptive field among all LUT-based methods while maintaining high efficiency. Our key insight lies in three complementary components. First, Learnable Spatial Shift module (LSS) is introduced to expand the receptive field by applying learnable, channel-wise spatial offsets on feature maps. Second, we propose an asymmetric dual-branch architecture that allocates more computation to the information-dense branch, substantially reducing inference latency without compromising restoration quality. Finally, we incorporate a feature-level LUT compression strategy called Error-bounded Adaptive Sampling (EAS) to minimize the storage overhead. Compared to the previous state-of-the-art method TinyLUT, ShiftLUT achieves a 3.8$\times$ larger receptive field and improves an average PSNR by over 0.21 dB across multiple standard benchmarks, while maintaining a small storage size and inference time. The code is available at: https://github.com/Sailor-t/ShiftLUT .

cs.CV↗

Early Results from the Coma Legacy IFU Survey (CLIFS): Ram Pressure Induced Shocks and Ionization in Jellyfish Tails

Jellyfish galaxies, which exhibit tails of gas opposite to their direction of motion, are a galaxy population showcasing the most extreme effects of ram pressure stripping (RPS). We present the emission line properties of a preliminary sample of five jellyfish galaxies in the Coma cluster, observed with the WEAVE Large-IFU as part of the Coma Legacy IFU Survey (CLIFS). When complete, CLIFS will form a sample of 29 jellyfish galaxies in Coma, selected based on the presence of one-sided tails in the radio continuum, enabling a comprehensive picture of the effects of ram pressure on galaxies in the Coma cluster. We extract emission line properties and confirm consistency between disk fluxes measured from WEAVE and MaNGA for galaxies with overlapping disk coverage between surveys. Comparing resolved radio and H$α$-based star formation rates, we find that, in contrast to the disk, the dominant source of tail emission is not star formation. We find evidence for diffuse ionized gas excited by RPS-driven shocks in the tails, as indicated by: (1) LINER-like tail emission with the [OI]/H$α$ BPT diagnostic; (2) enhanced [OII]/H$α$ ratios in the tails relative to the disks; and (3) similarly elevated emission line velocities and velocity dispersions in the tails with respect to the disks. These results demonstrate that ram-pressure-driven shocks dominate the ionized emission in jellyfish galaxy tails.

astro-ph.GA↗

X-ray line diagnostics of the multi-phase gas in the Centaurus cluster core with XRISM/Resolve

We report the multi-temperature structure of the intracluster medium (ICM) in the Centaurus cluster core observed with XRISM/Resolve. Thanks to its high energy resolution, Resolve enables us to measure fine structures of highly ionized emission lines from Si to Fe and to directly determine the excitation temperature and the ionization temperature from the emission line ratio diagnostics. The observed spectrum in the Centaurus core is well-represented by a double-temperature thermal plasma at collisional ionization equilibrium state rather than an isothermal one. The line ratio diagnostics also support this biphasic temperature structure. Particularly, the observed line ratios show a trend of increasing ionization temperature with atomic mass, while the ionization and excitation temperatures of Fe show nearly the same temperature. The resultant line ratios, which are well-represented by the two temperatures ICM, ~ 1.6 and ~ 3 keV, are also fairly consistent with the expected numbers when assuming the radial single-temperature ICM was projected in the cluster core along the line of sight. Due to the limited low-energy sensitivity of the Resolve with the gate valve closed, we investigated the effect of the cool component using the XMM-Newton/RGS spectrum, but it ultimately did not affect our results. The observed flux ratio between the Fe XXV He alpha resonance and forbidden lines shows an about 20% reduction, suggesting the presence of resonant scattering.

astro-ph.HE↗