arXiv ScienceSearch

arXiv subjects

Shuo Cao

Publications and source records attributed to Shuo Cao.

At least 19 recordsLinked to original sources

ProcArena: A Multi-Scenario Benchmark for LLMs on Direct and Interactive PL/SQL Development from Natural Language

Large language models (LLMs) have shown strong potential for translating natural-language (NL) requirements into PL/SQL programs, attracting increasing attention from the database community. However, existing NL-to-PL/SQL efforts primarily focus on directly generating PL/SQL from complete NL requirements. In practice, PL/SQL development involves diverse scenarios, such as from-scratch development, code modification, debugging, and optimization, and may require either direct generation or multi-turn interaction. Yet, no comprehensive benchmark evaluates multi-scenario, direct and interactive, and multi-dialect NL-to-PL/SQL development. In this paper, we present ProcArena, an execution-based benchmark covering both Direct and Interactive modes. ProcArena comprises 3,998 executable tasks over 157 databases, spanning nine development subscenarios in PostgreSQL and Oracle. We construct challenging Direct tasks through Iterative Logic Enhancement and scenario-specific adapters, and derive paired Interactive tasks through Knowledge Integration and Requirement Perturbation while preserving executable targets. We further design a controlled Solver-User Simulator protocol that allows models to clarify user intent and inspect the database environment without exposing hidden execution feedback. Evaluating seven language models, we find that the best average scores are only 62.2% and 57.8% in Direct and Interactive, respectively, demonstrating that realistic NL-to-PL/SQL development remains challenging, particularly in interactive settings.

cs.CL

Anisotropic Tensile Strength and Fracture Mechanism of $\theta$-TaN: A Machine-Learning Potential Molecular Dynamics Study

theta-phase tantalum nitride (theta-TaN) combines metallic conductivity with exceptionally high thermal conductivity, making it a potential material for device thermal management and interconnect applications. However, its tensile strength and fracture behavior remain unclear. Here, we investigate the anisotropic tensile response and fracture mechanism of theta-TaN using neuroevolution-potential molecular dynamics simulations. Size-convergence tests show that a 20 nm long model is sufficient for reliable prediction, and the mechanical parameters vary by less than 3.5% over the strain-rate range of 10^7 to 10^9 s^-1. The results reveal strong tensile anisotropy. The c-axis direction ([0001]) shows a higher strength of 80.10 GPa and modulus of 748.63 GPa, but a lower fracture strain of 15.02%. In contrast, the a-axis direction ([2-1-10]) shows a lower strength of 56.87 GPa and modulus of 570.74 GPa, but a higher fracture strain of 17.71%. From 300 to 900 K, the mechanical properties decrease nearly linearly, while more than 73% of the 300 K strength is retained at 900 K. Fracture occurs without observable dislocation activity and is governed by cleavage-plane selection: {10-10} prismatic planes under a-axis tension and the (0001) basal plane under c-axis tension. Atomic displacement analysis shows that local separation and microvoid formation precede macroscopic crack growth, indicating a brittle fracture process driven by local bond-network instability. These results provide atomic-scale mechanical data for assessing the reliability of theta-TaN in thermal management applications.

cond-mat.mtrl-sci

See2Think: Do Multimodal Models Really Use Intermediate Visual States?

Multimodal large language models increasingly use sketches, annotations, tools, and intermediate images during reasoning, but it remains unclear whether they truly rely on these visual states. Existing benchmarks are limited both by task collections with narrow coverage or partially text-solvable samples and by evaluations that emphasize final answers without diagnosing how intermediate visual states are generated, rendered, and used. We introduce See2Think, a unified evaluation framework comprising See2ThinkBench and Visual Action-of-Thought (VAoT). See2ThinkBench contains 1,200 open-ended, visually dependent problems across 12 task categories spanning 2D structured, 3D scene, and real-world reasoning. VAoT records textual thoughts, visual actions, rendered states, and subsequent reasoning under four controlled inference settings. Evaluating representative proprietary and open-source multimodal models, we find that visual reasoning is strongly model- and environment-dependent, with no single setting consistently dominating across tasks. Process analysis further shows that models usually select relevant visual operations, while faithful rendering remains the clearest bottleneck and high feedback uptake does not necessarily translate into accuracy gains. Under task-relevant corrupted feedback, models exhibit behavioral dependence on visual states, with accuracy dropping by over 10 percentage points in controlled interventions.

cs.CV

No Strong Evidence for Plasma Lensing in FRB 20240114A

FRB~20240114A is an extremely active repeating fast radio burst for which plasma lensing has been proposed to explain its burst-rate variations, spectral evolution, and apparently ``carbon-copy'' burst pairs. Using FAST data and publicly available Parkes observations, we test this interpretation with a one-dimensional Gaussian plasma-lens model. Although the burst-rate enhancements can be fitted separately, the corresponding magnification peaks and demagnification troughs are offset by far more than predicted and show no consistent periodicity. Moreover, with more than 10,000 bursts detected, a few apparently ``carbon-copy'' pairs can readily occur by chance. The burst bandwidth is not systematically narrower during the proposed lensing interval, nor are the burst energies significantly enhanced during the predicted magnification interval. These results provide no compelling evidence that a single Gaussian plasma lens explains the observed variability, which is more likely dominated by intrinsic source activity.

astro-ph.HE

Extragalactic test of General Relativity with time-delay gravitational lenses

Strong gravitational lensing, a key prediction of General Relativity (GR), offers a unique environment for examining alternative modified gravity theories. In this Letter, we employ a model-independent approach to estimate the parameterized post-Newtonian parameter $\gamma_{\rm PPN}$ using the time-delay measurements from H0LiCOW strong lensing systems. To minimize potential biases from cosmological models in testing GR, we use Gaussian Process regression (GPR) to reconstruct angular diameter distances ($D_{\rm A}$) from the newest baryon acoustic oscillation (BAO) measurements, provided by the Dark Energy Spectroscopic Instrument (DESI) DR2 data. Based on the reconstructed angular diameter distances and four H0LiCOW lenses, we directly estimate the post-Newtonian parameter $\gamma_{\rm PPN}=0.93^{+0.16}_{-0.17}$ and the sound horizon scale $r_{\rm d}=136.36^{+5.14}_{-3.20}~{\rm Mpc}$. This is the first simultaneous measurement of $\gamma_{\rm PPN}$ and $r_{\rm d}$ without any assumptions about the contents of the universe or the theory of gravity. In the new framework of distance ratio $D_{\Delta t}/D_{\rm l}$ which avoids the bias introduced by $r_{\rm d}$, the $\gamma_{\rm PPN}$ constraint can be further improved to $\gamma_{\rm PPN}=0.89^{+0.19}_{-0.15}$. Our results provide a direct test of GR at the extragalactic scale, which is well consistent with the prediction of GR within $1\sigma$.

astro-ph.CO

A Pilot Study of Mildly Recycled Pulsars: A Case Study of PSR J2338+4818

Mildly recycled pulsars are neutron stars partially spun up through relatively short mass-transfer phases, typically with massive carbon-oxygen (CO) or oxygen-neon-magnesium (ONeMg) white dwarf companions. PSR J2338+4818, a mildly recycled pulsar, was discovered with the Five-hundred-meter Aperture Spherical Telescope (FAST). As a pilot study on the formation and evolutionary pathways of mildly recycled pulsars, we present the updated timing solution for PSR J2338+4818 and examine its single pulses and scintillation properties. Aided by the sensitivity of FAST, the single pulses of PSR J2338+4818 were systematically studied. 27,228 single pulses with S/N > 7 have been detected in our observations. For the FAST ultra-wideband observation on MJD 61045, the receiver was still in the technical commissioning phase, and then only a preliminary single-pulse search was performed. Pulse nulling was examined using a Markov Chain Monte Carlo (MCMC) method, but no evidence for nulling was found. The possible long-term nulling reported by previous studies did not occur in any of our observations in either the 1.0 to 1.5 GHz band or the 300 to 600 MHz band. Interstellar scintillation is evident in our observations. The measured scintillation timescales and bandwidths range from 2.93 to 25.26 minutes and 1.68 to 27.41 MHz, respectively. In all observations, no clear scintillation arc was found in the secondary spectra of PSR J2338+4818.

astro-ph.HE

StableI2I: Spotting Unintended Changes in Image-to-Image Transition

In most real-world image-to-image (I2I) scenarios, existing evaluations primarily focus on instruction following and the perceptual quality or aesthetics of the generated images. However, they largely fail to assess whether the output image preserves the semantic correspondence and spatial structure of the input image. To address this limitation, we propose StableI2I, a unified and dynamic evaluation framework that explicitly measures content fidelity and pre--post consistency across a wide range of I2I tasks without requiring reference images, including image editing and image restoration. In addition, we construct StableI2I-Bench, a benchmark designed to systematically evaluate the accuracy of MLLMs on such fidelity and consistency assessment tasks. Extensive experimental results demonstrate that StableI2I provides accurate, fine-grained, and interpretable evaluations of content fidelity and consistency, with strong correlations to human subjective judgments. Our framework serves as a practical and reliable evaluation tool for diagnosing content consistency and benchmarking model performance in real-world I2I systems.

cs.CV

Inflation driven by a bare cosmological constant and its graceful exit

Vacuum energy, a prediction of quantum field theory, manifests itself as a cosmological constant in general relativity. In this Letter, we propose a novel inflationary scenario driven by a bare cosmological constant $\Lambda$, which terminates naturally through a self-tuning mechanism. Within Fab-Four gravity, self-tuning destabilizes the de Sitter state and drives the system toward a stiff-fluid attractor, thereby yielding a graceful exit. We construct two explicit models in which the slow-roll parameter evolves exponentially or as a power law. We show that the latter model, derived from center-manifold dynamics, significantly relaxes the required tuning of initial conditions. Our results establish, for the first time, that bare-vacuum-energy inflation with natural termination constitutes a viable dynamical possibility.

gr-qc

FAST Polarization Catalog of FRB 20240114A

Polarization measurements of fast radio bursts (FRBs) probe the magnetized plasma surrounding their central engines. FRB~20240114A is an exceptionally active repeating source, with 17,356 bursts detected between 2024 January 28 and 2025 May 30 by FAST, enabling time-resolved polarimetric studies. In this work, we present a polarimetric catalog of 6,131 bright bursts (with a signal-to-noise ratio S/N $\geq$ 20, 35.3% of the total sample), including arrival time (MJD$_{\text{topo}}$), dispersion measure (DM), burst width (W$_{\text{eff}}$), bandwidth, Faraday rotation measure (RM), linear and circular polarization degrees (DOL, DOC), and intrinsic polarization angle (PA$_0$). We detect a clear temporal evolution of RM: after an initial stable phase, it decreases linearly by $\sim$200 $\rm rad\ m^{-2}$ over 200 days, forming a bimodal distribution, whereas DM remains stable at 528.9 $\rm pc\ cm^{-3}$. The linear polarization fraction is generally high, with the 3$\sigma$ lower bound around 76%, while circular polarization is low, with 1,157 of 17,356 bursts (6.67%) having DOC $\geq$10%. We perform a power-law fit between $|\textrm{V}|$/I and $|\textrm{RM}|$, which yields an index of $-2.98 \pm 0.80$. It is found that the combined 2D distribution of L/I versus V/I remains stable, implying that the emission mechanism is largely invariant. Our PA$_0$ measurements show a broad, non-uniform distribution, implying a complex emission geometry. These results suggest that FRB~20240114A resides in a dynamically evolving magneto-ionic environment. This catalog provides a foundation for studies of repeating FRB progenitors and their environments.

astro-ph.HE

Testing Screened Modified Gravity with Strongly Lensed Gravitational Waves

Screening mechanisms are essential components in many modified gravity theories, which satisfy local tests of General Relativity (GR) and address cosmic acceleration on cosmological scales. The strong gravitational lensing of gravitational waves (GWs) offers a unique observational probe into cosmology and fundamental physics. In this paper, we investigate the possibility of testing screened modified gravity theories with strongly lensed gravitational waves. Specially, we develop the refined theoretical and statistical framework, in order to measure the post-Newtonian parameter $\gamma_{\text{PN}}$ in the presence of screening effects. Specially, the mass-truncated power-law and Navarro-Frenk-White (NFW) models are introduced to quantify the modified lensing potential. Our analysis also addresses the mass-sheet degeneracy (MSD) problem, by incorporating the absolute magnification and time delay measurements accessible through strongly lensed GW systems. We find that individual lensed GW system detected by next-generation GW detectors can provide stringent constraints on the PPN parameter ($\gamma_{\text{PN}}$) across different screening scales ($\Lambda$). Therefore, future measurements of strongly lensed GWs have great promise to seek departures from GR on kpc-Mpc scales, due to more precise time delay from lensed GW signals.

astro-ph.GA

Hierarchical cosmological constraints through strong lensing distance ratio

Strong gravitational lensing provides an independent and powerful probe of cosmic expansion by directly linking observables to cosmological distances. Upcoming surveys such as LSST will discover large number of galaxy-galaxy strong lensing systems, offering a new route to precise cosmological constraints. In this paper, we propose a Fisher-like sensitivity factor to map how the cosmological information of strong-lensing distances changes across the lens-source redshift plane. Applying such factor to the distance ratio $D_{ls}/D_s$, the time-delay distance $D_{\Delta t}$, and the double-source-plane ratio, we determine the ``sensitivity valleys'' where an observable becomes insensitive to a given parameter. The realistically simulated LSST lens population, which largely lies outside the distance-ratio valleys, covers the most sensitive region for $(w_0,w_a)$ parameter space. We then develop a new hierarchical framework, which could calibrate the redshift evolution of lens mass-density slopes and constrain cosmological parameters simultaneously. Focusing on the LSST mock data, we demonstrate that ignoring mass-profile evolution can bias $\Omega_m$ by up to $\sim 10\sigma$, while modeling the lens evolution could perfectly recovers the fiducial cosmology and yield stringent cosmological constraints (e.g., $\Delta\Omega_m \simeq 0.01$ and $\Delta w \simeq 0.1$ for $\sim 10^4$ lenses).

astro-ph.CO

Accelerating Masked Image Generation by Learning Controlled Latent Dynamics

Masked Image Generation Models (MIGMs) have achieved great success, yet their efficiency is hampered by the multiple steps of bi-directional attention. In fact, there exists notable redundancy in their computation: when sampling discrete tokens, the rich semantics contained in the continuous features are lost. Some existing works attempt to cache the features to approximate future features. However, they exhibit considerable approximation error under aggressive acceleration settings. We attribute this to their limited expressivity and the failure to account for sampling information. To fill this gap, we propose learning a lightweight model that incorporates both previous features and sampled tokens, and regresses the average velocity field of feature evolution. The model has moderate complexity that suffices to capture the subtle dynamics while keeping lightweight compared to the original base model. We apply our method to two representative MIGMs and tasks. In particular, on the state-of-the-art Lumina-DiMOO, it achieves over 4x acceleration of text-to-image generation while maintaining quality, significantly pushing the Pareto frontier of masked image generation. The code and model weights are available at https://github.com/Kaiwen-Zhu/MIGM-Shortcut.

cs.CV

Toward Generalizable Deblurring: Leveraging Massive Blur Priors with Linear Attention for Real-World Scenarios

Image deblurring has advanced rapidly with deep learning, yet most methods exhibit poor generalization beyond their training datasets, with performance dropping significantly in real-world scenarios. Our analysis shows this limitation stems from two factors: datasets face an inherent trade-off between realism and coverage of diverse blur patterns, and algorithmic designs remain restrictive, as pixel-wise losses drive models toward local detail recovery while overlooking structural and semantic consistency, whereas diffusion-based approaches, though perceptually strong, still fail to generalize when trained on narrow datasets with simplistic strategies. Through systematic investigation, we identify blur pattern diversity as the decisive factor for robust generalization and propose Blur Pattern Pretraining (BPP), which acquires blur priors from simulation datasets and transfers them through joint fine-tuning on real data. We further introduce Motion and Semantic Guidance (MoSeG) to strengthen blur priors under severe degradation, and integrate it into GLOWDeblur, a Generalizable reaL-wOrld lightWeight Deblur model that combines convolution-based pre-reconstruction & domain alignment module with a lightweight diffusion backbone. Extensive experiments on six widely-used benchmarks and two real-world datasets validate our approach, confirming the importance of blur priors for robust generalization and demonstrating that the lightweight design of GLOWDeblur ensures practicality in real-world applications. The project page is available at https://vegdog007.github.io/GLOWDeblur_Website/.

cs.CV

UniPercept: Towards Unified Perceptual-Level Image Understanding across Aesthetics, Quality, Structure, and Texture

Multimodal large language models (MLLMs) have achieved remarkable progress in visual understanding tasks such as visual grounding, segmentation, and captioning. However, their ability to perceive perceptual-level image features remains limited. In this work, we present UniPercept-Bench, a unified framework for perceptual-level image understanding across three key domains: Aesthetics, Quality, Structure and Texture. We establish a hierarchical definition system and construct large-scale datasets to evaluate perceptual-level image understanding. Based on this foundation, we develop a strong baseline UniPercept trained via Domain-Adaptive Pre-Training and Task-Aligned RL, enabling robust generalization across both Visual Rating (VR) and Visual Question Answering (VQA) tasks. UniPercept outperforms existing MLLMs on perceptual-level image understanding and can serve as a plug-and-play reward model for text-to-image generation. This work defines Perceptual-Level Image Understanding in the era of MLLMs and, through the introduction of a comprehensive benchmark together with a strong baseline, provides a solid foundation for advancing perceptual-level multimodal image understanding.

cs.CV

dMLLM-TTS: Self-Verified and Efficient Test-Time Scaling for Diffusion Multi-Modal Large Language Models

Diffusion Multi-modal Large Language Models (dMLLMs) have recently emerged as a novel architecture unifying image generation and understanding. However, developing effective and efficient Test-Time Scaling (TTS) methods to unlock their full generative potential remains an underexplored challenge. To address this, we propose dMLLM-TTS, a novel framework operating on two complementary scaling axes: (1) trajectory exploration scaling to enhance the diversity of generated hypotheses, and (2) iterative refinement scaling for stable generation. Conventional TTS approaches typically perform linear search across these two dimensions, incurring substantial computational costs of O(NT) and requiring an external verifier for best-of-N selection. To overcome these limitations, we propose two innovations. First, we design an efficient hierarchical search algorithm with O(N+T) complexity that adaptively expands and prunes sampling trajectories. Second, we introduce a self-verified feedback mechanism that leverages the dMLLMs' intrinsic image understanding capabilities to assess text-image alignment, eliminating the need for external verifier. Extensive experiments on the GenEval benchmark across three representative dMLLMs (e.g., Lumina-DiMOO, MMaDA, Muddit) show that our framework substantially improves generation quality while achieving up to 6x greater efficiency than linear search. Project page: https://github.com/Alpha-VLLM/Lumina-DiMOO.

cs.CV

PICABench: How Far Are We from Physically Realistic Image Editing?

Image editing has achieved remarkable progress recently. Modern editing models could already follow complex instructions to manipulate the original content. However, beyond completing the editing instructions, the accompanying physical effects are the key to the generation realism. For example, removing an object should also remove its shadow, reflections, and interactions with nearby objects. Unfortunately, existing models and benchmarks mainly focus on instruction completion but overlook these physical effects. So, at this moment, how far are we from physically realistic image editing? To answer this, we introduce PICABench, which systematically evaluates physical realism across eight sub-dimension (spanning optics, mechanics, and state transitions) for most of the common editing operations (add, remove, attribute change, etc.). We further propose the PICAEval, a reliable evaluation protocol that uses VLM-as-a-judge with per-case, region-level human annotations and questions. Beyond benchmarking, we also explore effective solutions by learning physics from videos and construct a training dataset PICA-100K. After evaluating most of the mainstream models, we observe that physical realism remains a challenging problem with large rooms to explore. We hope that our benchmark and proposed solutions can serve as a foundation for future work moving from naive content editing toward physically consistent realism.

cs.CV

LinearSR: Unlocking Linear Attention for Stable and Efficient Image Super-Resolution

Generative models for Image Super-Resolution (SR) are increasingly powerful, yet their reliance on self-attention's quadratic complexity (O(N^2)) creates a major computational bottleneck. Linear Attention offers an O(N) solution, but its promise for photorealistic SR has remained largely untapped, historically hindered by a cascade of interrelated and previously unsolved challenges. This paper introduces LinearSR, a holistic framework that, for the first time, systematically overcomes these critical hurdles. Specifically, we resolve a fundamental, training instability that causes catastrophic model divergence using our novel "knee point"-based Early-Stopping Guided Fine-tuning (ESGF) strategy. Furthermore, we mitigate the classic perception-distortion trade-off with a dedicated SNR-based Mixture of Experts (MoE) architecture. Finally, we establish an effective and lightweight guidance paradigm, TAG, derived from our "precision-over-volume" principle. Our resulting LinearSR model simultaneously delivers state-of-the-art perceptual quality with exceptional efficiency. Its core diffusion forward pass (1-NFE) achieves SOTA-level speed, while its overall multi-step inference time remains highly competitive. This work provides the first robust methodology for applying Linear Attention in the photorealistic SR domain, establishing a foundational paradigm for future research in efficient generative super-resolution.

cs.CV

Probing potential redshift-dependent systematics in the Hubble tension: Model-independent $H_0$ constraints from DESI R2

We present a determination of the Hubble constant ($H_0$) using the latest observational data from multiple cosmological probes, providing an independent geometric calibration of the SN Ia distance scale. By combining baryon acoustic oscillation (BAO) measurements from the second data release of the Dark Energy Spectroscopic Instrument (DESI DR2), cosmic chronometer $H(z)$ data, and the Pantheon Plus Type Ia supernova (SN Ia) sample, we reconstruct the cosmic expansion history through Gaussian process regression without assuming a specific cosmological model. Our analysis fully incorporates the complete covariance structure and yields $H_0$ constraints at five distinct redshifts: $65.72 \pm 1.99$ (z=0.51), $67.78 \pm 1.75$ (z=0.706), $70.74 \pm 1.39$ (z=0.934), $71.04 \pm 1.93$ (z=1.321), and $68.37 \pm 3.95~\mathrm{km~s^{-1}~Mpc^{-1}}$ (z=1.484). The Bayesian combination of these measurements gives $\hat{H}_0 = 69.29 \pm 0.81~\mathrm{km~s^{-1}~Mpc^{-1}}$ with 1.2\% precision, which occupies an intermediate position between the Planck CMB result and the SH0ES local measurement. While we observe a non-monotonic pattern in $H_0$ values across redshifts, statistical tests show this apparent evolution is not significant (p = 0.208). Our approach delivers independent constraints at multiple redshifts, enabling investigation of potential redshift-dependent systematic effects in the Hubble tension. The results demonstrate that an independent geometric method yields an $H_0$ value consistent with the intermediate range of current measurements, providing a crucial cross-check of distance ladder determinations.

astro-ph.CO