arXiv ScienceSearch

arXiv subjects

Dong Xu

Publications and source records attributed to Dong Xu.

At least 19 recordsLinked to original sources

GaussVid: Sparse-View Gaussian Splatting with 3D-Aware Video Diffusion Priors

3D Gaussian Splatting (3DGS) has achieved remarkable success in novel view synthesis; however, reconstructions under sparse views often exhibit noticeable artifacts. While recent video diffusion models provide strong spatio-temporal priors for 3DGS restoration, directly fine-tuning them for restoration is suboptimal, as they lack awareness of the underlying multi-camera geometry, resulting in multi-view inconsistencies. In this work, we propose a novel 3D-aware video restoration framework designed to enhance the quality of sparse 3DGS reconstruction. Specifically, we construct a large-scale 3DGS video dataset to enable specialized fine-tuning. To bridge the gap between 2D video generation and 3D multi-view constraints, we introduce a camera-conditioned geometric prior. By using the first and last frames as boundary anchors and encoding the corresponding camera relationships, we explicitly inject spatial structure into the video generation pipeline. This boundary-anchored, camera-aware prior guides the network toward geometrically grounded restoration that remains coherent across viewpoints. Extensive experiments show that, among video-prior restoration methods, our approach attains the best pixel- and structure-level fidelity (PSNR/SSIM) and improves multi-view consistency, while remaining competitive in perceptual quality (LPIPS).

cs.CV

Minutes-long soft X-ray prompt emission from a compact object merger

Compact object mergers are multi-messenger sources and known progenitors of some gamma-ray bursts, bright flashes of high-energy radiation powered by a central engine, either an accreting black hole or a neutron star. Our understanding of these events has so far been shaped primarily by observations in the gamma-ray band, leaving their prompt phase poorly constrained at lower energies. A long-lasting ($\approx$100 s) engine-driven X-ray emission was discussed to explain rapidly fading X-ray afterglows following several ($\approx$30%) bursts of short ($\lesssim$2 s) duration. However, this prompt X-ray component was not directly observed and past candidates were not confirmed. Here we report the discovery of EP250704a containing a minutes-long ($\sim$560 s) flash of soft (0.5--4 keV) X-rays immediately following the short ($\sim$0.4 s) GRB 250704B. The variability and spectral shape of this emission are inconsistent with the canonical picture of a hard, accretion-powered spike followed by a standard external-shock afterglow. Instead, the long-soft bump points to a distinct phase of prompt emission in X-rays, which would not have been detected without the soft X-ray coverage of Einstein Probe. The detection of a prompt soft X-ray counterpart in an otherwise ordinary short GRB shows that long-lasting X-ray emission is likely a common feature of merger-driven bursts and a promising electromagnetic counterpart to gravitational wave sources.

astro-ph.HE

GRB 260310A / SN 2026fgk: A Multi-Wavelength Study of a Nearby Underluminous Long GRB and SN with a Complex Afterglow

We present a comprehensive multi-wavelength study of GRB 260310A / SN 2026fgk, a nearby ($z=0.153$), long-duration gamma-ray burst (GRB) with an exceptionally underluminous prompt $γ$-ray emission and a Comptonized spectrum. The burst occurred at the edge of a blue host galaxy at a projected distance of 15 kpc, which is one of the largest offsets reported for a long GRB. The bright optical afterglow, with dense coverage from COLIBRÍ, likely peaked at a few to several hours post-burst, followed by a shallow decay not expected from canonical afterglow models. Both the optical and X-ray light curves show a brief chromatic plateau from $4-7$ days. We show that the subsequent rebrightening observed at $\sim20$ days is best explained by the combined contribution of the associated Type Ic-BL supernova, identified in GTC spectra, and a late-time refreshed shock. The broadband optical to X-ray spectral energy distribution is well described by synchrotron emission from the forward shock, while the radio observations demand an additional emission component. We model the afterglow using (a) an on-axis uniform jet from a dirty fireball with late-time energy injection and (b) a misaligned jet with power-law angular structure, both having material emitting along our line-of-sight (LOS) moving with an initial Lorentz factor of $Γ_0\sim20-35$. We conclude that at more typical GRB distances ($z\gtrsim0.5$) the prompt $γ$-ray emission from this source would likely have escaped detection, whereas its optical afterglow would have remained observable, making the event appear as an orphan afterglow or a gamma-ray quiet fast X-ray transient.

astro-ph.HE

READ: A Retrieval-Alignment Diffusion Framework for Structure-based Drug Design

Structure-based drug design (SBDD) models are central to modern pharmaceutical research, enabling the rational exploration of protein-ligand interactions at atomic resolution. However, most existing approaches frame molecular generation as an isolated optimization or a one-to-one matching task, overlooking the shared binding patterns and intrinsic similarities among protein-ligand complexes. This fragmented perspective constrains their ability to capture the fundamental principles governing molecular recognition and binding specificity. Moreover, the limited availability of high-quality experimental data further hampers model generalization and real-world applicability. To address these challenges, we present READ, a retrieval-alignment molecular generation framework that conditions the generative process on small molecules targeting homologous proteins. Retrieved ligands are aligned with a diffusion model across multiple representational spaces and integrated as conditional guidance throughout successive stages of generation. Under a standardized docking-based evaluation protocol, READ achieves consistently strong performance against state-of-the-art SBDD methods. More importantly, it introduces a retrieval-alignment paradigm for structure-based molecular generation, offering a practical framework for early-stage computational hit generation while leaving prospective experimental validation as future work.

q-bio.BM

DegradeQuery: Counterfactual Tuple Pretraining for Context-Aware PROTAC Degradation Prediction

Proteolysis-targeting chimeras (PROTACs) induce protein degradation by recruiting a target protein to an E3 ubiquitin ligase, making degradation a joint outcome of the degrader molecule and its biological context. Although public databases contain thousands of structured molecule-target-E3 records, degradation measurements are available for only a small fraction of them. Existing supervised approaches therefore leave most recorded chemical-biological relationships unused. We introduce DegradeQuery, a context-aware prediction framework that converts these label-missing records into a pretraining signal. Its counterfactual tuple pretraining objective contrasts recorded tuples with alternatives formed by replacing the target, the E3 ligase, or both, enabling the model to learn contextual associations without assigning activity pseudo-labels. The resulting representation is then fine-tuned to predict degradation from the complete molecule-target-E3 context. On the official PROTAC-8K benchmark, DegradeQuery achieves an area under the receiver operating characteristic curve of 0.9065 and an accuracy of 0.8500, outperforming the compared methods. Controlled analyses further show that the improvement is primarily attributable to tuple-level pretraining, can be recovered using only label-missing records, and remains complementary to protein language model representations. These findings demonstrate that incompletely labeled PROTAC databases contain useful relational supervision and provide a practical route for learning context-aware degradation predictors from scarce experimental labels.

q-bio.BM

Versatile Video Representation via Feed-Forward 2D Gaussian Splatting Tokenization

Recent video representation methods that rely on fixed-grid, patch-wise tokenization often exhibit limited versatility.Spatially, uniformly allocating a fixed number of tokens often leads to over-encoding in low-information regions. Temporally, reducing redundancy remains challenging without explicitly distinguishing between static and dynamic content. In this work, we introduce the Gaussian Video Transformer (GVT), a versatile video representation framework built on a feed-forward 2D Gaussian Splatting (2DGS) tokenization scheme. We first extract latent rigid features from a video clip and represent them with a set of 2D Gaussians generated by our proposed Spatio-Temporal Gaussian Embedding (STGE) mechanism in a feed-forward manner. Such 2D Gaussians not only enhance spatial adaptability by assigning higher (resp., lower) rendering weights to regions with higher (resp., lower) information content during rasterization, but also improve generalization by avoiding per-video optimization. To enhance the temporal versatility, we introduce a Gaussian Set Partitioning (GSP) strategy that separates the 2D Gaussians into static and dynamic sets, which explicitly model static content shared across different time-steps and dynamic content specific to each time-step, enabling a compact representation. We evaluate GVT across four tasks: video reconstruction, video action recognition, video compression, and video generation, on the UCF101, Kinetics, and DAVIS datasets. The results demonstrate state-of-the-art reconstruction and compression performance, improved action recognition, and video generation performance comparable to the baseline MAGVIT-v2.

cs.CV

OmniCustom: Sync Audio-Video Customization Via Joint Audio-Video Generation Model

Existing mainstream video customization methods focus on generating identity-consistent videos based on given reference images and textual prompts. Benefiting from the rapid advancement of joint audio-video generation, this paper proposes a more compelling new task: sync audio-video customization, which aims to synchronously customize both video identity and audio timbre. Specifically, given a reference image $I^{r}$ and a reference audio $A^{r}$, this novel task requires generating videos that maintain the identity of the reference image while imitating the timbre of the reference audio, with spoken content freely specifiable through user-provided textual prompts. To this end, we propose OmniCustom, a powerful DiT-based audio-video customization framework that can synthesize a video following reference image identity, audio timbre, and text prompts all at once in a zero-shot manner. Our framework is built on three key contributions. First, identity and audio timbre control are achieved through separate reference identity and audio LoRA modules that operate through self-attention layers within the base audio-video generation model. Second, we introduce a contrastive learning objective alongside the standard flow matching objective. It uses predicted flows conditioned on reference inputs as positive examples and those without reference conditions as negative examples, thereby enhancing the model ability to preserve identity and timbre. Third, we train OmniCustom on our constructed large-scale, high-quality audio-visual human dataset. Extensive experiments demonstrate that OmniCustom outperforms existing methods in generating audio-video content with consistent identity and timbre fidelity. Project page: https://omnicustom-project.github.io/page/.

cs.SD

Reverse to Advance: Teleoperation-Cost Effective Hard Policy Learning from Reversed Easy Tasks

High-quality teleoperation datasets are costly to collect, particularly for hard tasks. We observe that many tasks exhibit directional asymmetry: completing the forward hard task is difficult, whereas reversing it by relaxing or disrupting the environment is comparatively easy. This suggests that reversed easy-task trajectories can serve as a scalable supervision signal for the hard task, reducing the cost of manual demonstration collection. However, reversed data can be noisy, and directly training on it may yield suboptimal policies. To enable largely automated acquisition and effective use of reversed data, we propose a teleoperation-cost effective framework for hard policy learning via temporal reversal of easy tasks, consisting of three key components: a closed-loop data collection pipeline that alternates between hard-task and easy-task policies to autonomously reset the environment and generate diverse trajectories; a hierarchical data refinement pipeline that temporally inverts easy-task rollouts and filters low-quality motion using kinematic priors and a critic-guided advantage filter; and an iterative policy learning method that trains the hard-task policy using both initial reversed easy-task demonstrations and the filtered reversed data in a continuous online learning loop. By combining automated collection, hierarchical refinement, and iterative learning, our method enables scalable, reliable training of complex, high-precision manipulation tasks. Across two simulated benchmarks and real-robot experiments, we demonstrate that our method improves hard-task success rates with higher data efficiency and more stable training compared to reversal-based and reinforcement-learning baselines, without requiring extensive hard-task teleoperation.

cs.RO

Large Language Models Versus Physicians in Traditional Chinese Medicine: A Real-World Clinical Case Evaluation

Large language models (LLMs) are increasingly being explored for clinical applications, yet their assessment for real-world traditional Chinese medicine (TCM) practice remains limited We constructed a clinical case library comprising 349 de-identified outpatient cases from 62 hospitals and evaluated 16 LLMs and a comparator cohort of 60 practicing TCM physicians using 60 representative cases selected from this library. Model outputs and physician reports were anonymized and scored by five senior TCM experts across nine diagnostic and therapeutic dimensions. Cutting-edge general-purpose LLMs achieved higher expert scores than the physician comparators, particularly for medical advice, treatment principles and selected diagnostic tasks. However, prescription-level analyses revealed discrepancies in herb selection, dosage, and treatment strategy, and qualitative safety review identified hallucinations and undesirable template-driven outputs. These findings highlight the potential of LLMs for TCM decision support while underscoring the need for physician oversight, safety constraints and prospective clinical evaluation.

cs.CL

Render-in-the-Loop: Vector Graphics Generation via Visual Self-Feedback

Multimodal Large Language Models (MLLMs) have shown promising capabilities in generating Scalable Vector Graphics (SVG) via direct code synthesis. However, existing paradigms typically adopt an open-loop "blind drawing" approach, where models generate symbolic code sequences without perceiving intermediate visual outcomes. This methodology severely underutilizes the powerful visual priors embedded in MLLMs vision encoders, treating SVG generation as a disjointed textual sequence modeling task rather than an integrated visuo-spatial one. Consequently, models struggle to reason about partial canvas states and implicit occlusion relationships, which are visually explicit but textually ambiguous. To bridge this gap, we propose Render-in-the-Loop, a novel generation paradigm that reformulates SVG synthesis as a step-wise, visual-context-aware process. By rendering intermediate code states into a cumulative canvas, the model explicitly observes the evolving visual context at each step, leveraging on-the-fly feedback to guide subsequent generation. However, we demonstrate that applying this visual loop naively to off-the-shelf models is suboptimal due to their inability to leverage incremental visual-code mappings. To address this, we first utilize fine-grained path decomposition to construct dense multi-step visual trajectories, and then introduce a Visual Self-Feedback (VSF) training strategy to condition the next primitive generation on intermediate visual states. Furthermore, a Render-and-Verify (RaV) inference mechanism is proposed to effectively filter degenerate and redundant primitives. Our framework, instantiated on a multimodal foundation model, outperforms strong open-weight baselines on the standard MMSVGBench. This result highlights the remarkable data efficiency and generalization capability of our Render-in-the-Loop paradigm for both Text-to-SVG and Image-to-SVG tasks.

cs.CV

X-rays breaking out of pre-explosion ejecta mark a supernova's first light

Massive stars die as core-collapse supernovae, whose optical light emerges days after the implosion. Theory predicts that the initial collapse-driven shock, upon breaking through the star and dense circumstellar medium, emits a brief thermal flash of soft X-rays and ultraviolet. Yet these elusive first signals have remained largely undetected, owing to limited wide-field soft X-ray monitoring. Here we report the discovery of a soft X-ray flash, EP260321a, followed days later by a broad-lined supernova from an envelope-stripped progenitor. Its X-ray spectrum, best modeled with blackbody, establishes it as the long-sought archetypal shock breakout. The burst's duration and energetics place the breakout at a radius of 300 solar radii, tracing a dense surrounding shell and revealing abrupt mass ejection within the final month before collapse.

astro-ph.HE

ELBO-T2IAlign: A Generic ELBO-Based Method for Calibrating Pixel-level Text-Image Alignment in Diffusion Models

Diffusion models excel at image generation. Recent studies have shown that these models not only generate high-quality images but also encode text-image alignment information through attention maps or loss functions. This information is valuable for various downstream tasks, including segmentation, text-guided image editing, and compositional image generation. However, current methods heavily rely on the assumption of perfect text-image alignment in diffusion models, which is not the case. In this paper, we propose using zero-shot referring image segmentation as a proxy task to evaluate the pixel-level image and class-level text alignment of popular diffusion models. We conduct an in-depth analysis of pixel-text misalignment in diffusion models from the perspective of training data bias. We find that misalignment occurs in images with small-sized, occluded, or rare object classes. Therefore, we propose ELBO-T2IAlign--a simple yet effective method to calibrate pixel-text alignment in diffusion models based on the evidence lower bound (ELBO) of likelihood. ELBO-T2IAlign is training-free and generic: it requires no additional annotations, model retraining, or architectural modifications, and it can be directly applied to different diffusion backbones. Extensive experiments on zero-shot referring image segmentation, text-guided image editing, and compositional image generation verify that the proposed calibration improves pixel-text alignment across complementary downstream tasks.

cs.CV

EP251023a: A fast X-ray transient featuring a magnetar-powered optical internal plateau followed by a steep decay

EP251023a is an extragalactic fast X-ray transient (eFXT) detected solely by EP without a gamma-ray counterpart. The prompt emission consists of a main emission with a duration $T_{90}=292\pm19$ s, followed by a long-lasting tail emission that persists until the observation ends at $T_0+1571$ s. With the upper limit of Konus--Wind, we derived a conservative upper limit on the isotropic gamma-ray energy $E_{γ,\rm{iso}}$ of $5.7 \times 10^{52}$ erg for the main emission phase. A redshift of $z = 2.232\pm0.001$ is identified from strong absorption features in the Keck spectrum, which also indicate a relatively low host-galaxy HI column density. Based on the broadband spectral energy distribution, the late-time light curves show an achromatic plateau, followed by an extremely steep decay with a slope of 3.99 after a break at about 49 ks, which is consistent with a rapidly spinning millisecond magnetar engine. Under the isotropic wind scenario, we obtain the initial period $P_0<2.27$~ms and the magnetic field strength $B_p<8.33\times10^{14}$~G for the magnetar; whereas considering a jet collimation with a typical opening angle of 0.1 rad relaxes these constraints to $P_0<32.15$~ms and $B_p<1.18\times10^{16}$~G. Together with GRB\,070707, EP251023a may represent a rare class of optical magnetar-powered internal plateaus with little external-shock contamination, unlike previous examples detected primarily in X-rays. Future discoveries of similar events will help clarify the relationship between magnetar-powered internal emission observed in the optical band and that detected only in X-rays.

astro-ph.HE

GRB 250424A: A Case Study of Energy Injection with Multiwavelength Observations

We present a comprehensive multiwavelength analysis of the long-duration gamma-ray burst (GRB) 250424A. Our dataset spans from the prompt gamma-ray emission to late-time optical monitoring, including spectra obtained with the Keck 10\,m telescope. We find that the afterglow light curves display a prominent, simultaneous shallow decay phase in both X-ray and optical bands, followed by an achromatic transition to a standard decay regime. The broadband spectral energy distributions are well-modeled by a single power-law function, indicating a common synchrotron origin for the emission across frequencies. We interpret the afterglow evolution within the framework of a relativistic forward shock refreshed by continuous energy injection. This scenario successfully reproduces the observed temporal and spectral behavior, yielding an isotropic equivalent kinetic energy of $E_{\rm K,iso} \approx 5.5 \times 10^{52}$ erg and an injection index of $q\approx 0.34$ in a constant-density circumburst environment. The shallow decay phase is consistent with sustained energy injection lasting $\sim$ 9 ks. Despite the relatively low redshift, late-time optical observations reveal no distinct supernova component; however, our derived upper limits do not strictly rule out the presence of a typical GRB-associated supernova.

astro-ph.HE

Failed jet breakout in the metal-poor broad-lined type Ic supernova 2026gzf

A long-standing question in the death of massive stars is the role of relativistic jets. While many gamma-ray bursts and some fast X-ray transients seem to be associated with broad-lined type Ic supernovae, the opposite is not true. The lack of observable jet emission in those Ic-BL SNe can be explained by invoking off-axis jets, choked jets that inject all their energy into the stellar envelope, baryon-loaded jets for which the prompt high-energy emission is strongly suppressed, or non-jetted SNe. The lack of exact explosion time in the majority of SNe presents an obstacle to distinguish between these scenarios. Here we report the properties of SN 2026gzf associated with the X-ray thermal Einstein Probe shock-breakout EP260321a at z=0.0343. The absence of compelling shocked cocoon and radio emission up to 54 days, combined with initial expansion velocities of ~30,000 km/s and a circumstellar shell of ~0.07 M$_\odot$, favour a scenario for SN 2026gzf in which a jet was choked in the circumstellar shell. Our high-spatial resolution images of the SN environment show that the progenitor was located between two highly star-forming regions with a metallicity lower than any previously known Ic-BL SN. As the first case of a Ic-BL SN associated with high-energy prompt emission without the signature of a jet, SN 2026gzf provides a unique perspective to understand the successful launch of relativistic jets during the deaths of massive stars.

astro-ph.HE

General framework for incoherent topological structured light and optical information encoding

Topology provides a powerful language for describing global invariants in physical systems, yet optical topology has been explored predominantly with fully coherent light. Recent studies have shown that incoherent light can host topological structures mediated by coherence singularities; however, a general framework for their construction and control has been lacking. Here, we introduce an incoherent Milnor polynomial, which establishes a theoretical framework for real-space incoherent topological structured light, in which topology and statistical coherence emerge as independent and jointly addressable degrees of freedom. This framework overcomes a fundamental limitation of coherent topological structured light, enabling arbitrary intensity engineering without altering the underlying topological configuration. Experimentally, we realize incoherent Hopf-linked and trefoil-knotted coherence singularities with programmable statistical coherence. We further demonstrate a robust optical information-encoding scheme inspired by Rubik's-cube-like rotations, where statistical coherence determines far-field intensity patterns associated with the cube's initial states, and topological structures govern controlled rotations acting as encryption keys. Our results advance incoherent topological structured light from a physical curiosity to a programmable photonic platform, opening new avenues for optical information encoding, statistical photonics, and coherence-engineered functionalities beyond coherent optical topology.

physics.optics

EP250827b/SN 2025wkm: An X-ray Flash-Supernova Powered by a Central Engine and Circumstellar Interaction

We present the discovery of EP250827b/SN 2025wkm, an X-ray Flash (XRF) discovered by the Einstein Probe (EP), accompanied by a broad-line Type Ic supernova (SN Ic-BL) at $z = 0.1194$. EP250827b possesses a prompt X-ray luminosity of $\sim 10^{45} \, \rm{erg \, s^{-1}}$, lasts over 1000 seconds, and has a peak energy $E_{\rm{p}} < 1.5$ keV at 90\% confidence. SN 2025wkm possesses a double-peaked optical light curve (LC), though its bolometric luminosity plateaus after its initial peak for $\sim 20$ days, consistent with a central engine injecting additional energy into the explosion. Its spectrum transitions from a blue to red continuum with clear blueshifted broad absorption features consistent with a SN Ic-BL classification. We do not detect any transient radio emission and rule out the existence of an on-axis, energetic jet $\gtrsim 10^{50}~$erg assuming a typical LGRB circumburst constant density ($n \approx 10^{-3}$--$10^{-1}~{\rm cm}^{-3}$) and microphysical parameters ($ε_{\rm e} = 0.1$ and $ε_{\rm B} = 0.01$). In the model we invoke, the collapse gives rise to a long-lived magnetar, potentially surrounded by an accretion disk. Magnetically--driven winds from the magnetar and the disk mix together and break out with a velocity $\sim 0.35c$ and interact with an extended circumstellar medium with radius $\sim 10^{13}$ cm, generating X-ray breakout emission through non-thermal free-free processes. The disk outflows and magnetar winds power blackbody photospheric emission as they cool adiabatically and thermalize, producing the first SN peak. The spin-down luminosity of the magnetar and radioactive decay of $^{56}$Ni powers the late-time emission. We end by discussing the landscape of XRF-SNe within the context of EP's recent discoveries.

astro-ph.HE

Z-Order Transformer for Feed-Forward Gaussian Splatting

Recent advances in 3D Gaussian Splatting (3DGS) have enabled significant progress in photorealistic novel view synthesis. However, traditional 3DGS relies on a slow, iterative optimization process, which limits its use in scenarios demanding real-time results. To overcome this bottleneck, recent feed-forward methods aim to predict Gaussian attributes directly from images, but they often struggle with the redundancy of Gaussian primitives and rendering quality. In this work, we introduce a transformer-based architecture specifically designed for feed-forward Gaussian Splatting. Our key insight is that spatial and semantic relationships among Gaussians can be effectively captured through a sparse attention mechanism, enabled by a Z-order strategy that organizes the unstructured Gaussian set into a spatially coherent sequence. Furthermore, we incorporate this Z-order strategy to adaptively suppress redundancy while preserving critical structural details. This allows the transformer to efficiently model context, compress Gaussian primitives, and predict Gaussian attributes in a single forward pass. Comprehensive experiments demonstrate that our method achieves fast and high-quality novel view synthesis with fewer Gaussian primitives.

cs.CV