arXiv ScienceSearch

arXiv subjects

Hongfei Zhang

Publications and source records attributed to Hongfei Zhang.

At least 19 recordsLinked to original sources

WFST Follow-up of S251112cm: Searching for an Optical Counterpart to a Subsolar-mass Compact-binary Merger Candidate

Electromagnetic counterparts to gravitational-wave (GW) sources probe the properties, environments, and evolution of compact-object mergers. S251112cm belongs to an emerging class of GW candidates with possible subsolar-mass components. We present an optical counterpart search for S251112cm with the Wide Field Survey Telescope (WFST). The observations began 20.2 hr after the GW trigger and continued for three nights, covering approximately $780~\mathrm{deg}^{2}$ and $51\%$ of the localization probability in the updated skymap. We searched for newly emerging, rapidly evolving, off-nuclear optical transients and identified four candidates whose host-galaxy distances are broadly consistent with the GW distance estimate. However, all four evolve substantially more slowly than AT 2017gfo, ruling out an AT 2017gfo-like origin and disfavoring their association with S251112cm as rapidly evolving kilonova counterparts. Combining WFST and DECam observations increases the covered localization probability to approximately 67%. Conditional on the counterpart lying within this footprint, we constrain a grid of binary-neutron-star kilonova models. For viewing angles $θ_{\rm obs}<60^{\circ}$, with the other parameters fixed to the best-fitting values for AT 2017gfo, models with dynamical ejecta mass $M_{\rm dyn}\gtrsim0.01\,M_{\odot}$ or wind ejecta mass $M_{\rm wind}\gtrsim0.05\,M_{\odot}$ are disfavored over most of the covered GW probability. Beyond kilonova models, our phenomenological analysis constrains rapidly evolving optical transients with $T_{\rm rise,1/2}\leq2$ days to peak absolute magnitudes fainter than approximately -13 mag. To our knowledge, the coordinated WFST and DECam observations provide the strongest optical constraints to date on possible rapidly evolving counterparts to this new class of GW candidates involving subsolar-mass compact objects.

astro-ph.HE

FlashDrive: Flash Vision-Language-Action Inference for Autonomous Driving

Vision-Language-Action (VLA) models promise to bring end-to-end reasoning to autonomous driving, but their computational cost remains far too high for real-time control. The core challenge is structural: VLA inference is not a single bottleneck but a cascade of four. Visual encoding wastes compute on overlapping video frames; language-model prefill recomputes context that could be carried over from the previous timestep; reasoning tokens are generated serially despite low entropy; and flow-matching denoising applies uniform compute to a non-uniform velocity field. Addressing any one stage in isolation leaves the others untouched. We propose FlashDrive, an algorithm-system co-design framework that targets all four stages simultaneously. Our key insight is that each bottleneck admits a distinct, lightweight algorithmic shortcut: temporal overlap enables streaming KV-cache reuse across frames; the low per-token entropy and strong intra-block correlations of driving-domain reasoning make a non-autoregressive diffusion drafter highly effective for speculative decoding; and the velocity field's structure---sharp at the endpoints, flat in the middle---permits adaptive step caching that concentrates compute where it matters. Layered on system-level CUDA Graph compilation and kernel fusion, these techniques compound. Applied to Alpamayo 1.5-10B with W4A8 quantization, FlashDrive reduces end-to-end latency from 717ms to 151ms (4.7x) while leaving accuracy essentially unchanged: minADE6@6.4s shifts by only 0.08m, minADE1 improves, and closed-loop collision and off-road rates improve in simulation. By raising a 10B-parameter reasoning VLA from 1.4~Hz to 6.6~Hz on a single GPU, FlashDrive moves end-to-end autonomous driving substantially closer to real-time deployment.

cs.AI

Search for L4 Earth Trojan asteroids with the 2.5-meter Wide Field Survey Telescope

Earth Trojan asteroids (ETAs) are a mysterious population, and dynamically stable ETAs, if primordial, could be "living fossils" of the early solar system. To date, there are only two known ETAs, but both are temporary ETAs. The aim of our survey is to discover new temporary or stable ETAs; in the absence of detections, we derive upper limits on the population of stable ETAs. We conducted the largest wide-area survey of the Earth's L4 Lagrange point region so far using the Wide Field Survey Telescope, covering about 236.74 deg^2, corresponding to 33.24% of the probability coverage for sky regions where dynamically stable L4 ETAs are likely to reside. No new ETAs were detected in our survey. We place a cumulative upper limit of N(H < 19.1) < 19 on the stable population of objects larger than ~520 m (for an assumed albedo of 0.15). This represents the most stringent constraint on the ETA population to date.

astro-ph.EP

Cross-Modal Iteration Distillation for Robust IHD Screening: The IDNet Framework and A New Benchmark

Color Fundus Photography (CFP) offers a low-cost and non-invasive route for ischemic heart disease (IHD) screening, but current studies are limited by scarce public benchmarks and ineffective fusion of retinal images with sparse clinical variables. We propose IDNet, a multimodal framework with a Cross-Modal Distillation Aggregator (CDA) that uses learnable queries to sequentially integrate left-eye, right-eye, and clinical features, mitigating the imbalance between high-dimensional visual features and low-dimensional tabular inputs. We also construct a reproducible UK Biobank benchmark with open-source curation and quality-control pipelines, yielding 50,410 images from 25,205 subjects. On this benchmark, IDNet outperforms image-only, clinical-only, and several multimodal baselines, and CDA consistently improves multiple visual encoders as a plug-in fusion module.

cs.CV

OmniCoT: A Benchmark for Global and Multi-Step Panoramic Reasoning

Multimodal Large Language Models (MLLMs) have demonstrated promising spatial reasoning capabilities, while these abilities remain underexplored in the emerging visual modality of panoramic imagery. The full 360°$\times$180° field of view of panoramas essentially supports complex global multi-step reasoning, which is also the fundamental advantage of panoramas in applications such as embodied intelligence. However, existing panoramic benchmarks largely focus on simplistic queries that rely on local cues or single-/few-step reasoning, thereby ignoring the fundamental advantage of panoramas and failing to fully exploit their potential. To address this gap, we introduce OmniCoT, a panoramic spatial reasoning suite designed to enable MLLMs to use global evidence and perform multi-step inference across viewpoints. It includes OmniCoT-B (6.7K data) for evaluation, which measures both answer accuracy and reasoning quality, OmniCoT-Real (1K data) as a manually annotated real-world subset to quantify the Sim-to-Real gap. For training, OmniCoT-T (14.3K data) is purpose-built with structured stepwise Chain-of-Thought annotations that explicitly link intermediate reasoning steps to panoramic evidence. Based on OmniCoT-T, we introduce OmniCoT-R1 and adopt a two-stage training strategy tailored to the geometrically complex panoramic space, where Supervised Fine-tuning (SFT) anchors reasoning to panoramic evidence (e.g., bearings, proximity) and GRPO penalizes geometrically incoherent paths to consolidate global 360° spatial consistency. Through OmniCoT, we aim to recalibrate the difficulty of panoramic spatial reasoning to better align with the intrinsic capabilities of panoramic imagery, thereby fostering meaningful progress in this research area.

cs.CV

Discovery of a Featureless Tidal Disruption Event at z~1 with the Wide Field Survey Telescope

We report the discovery of tidal disruption event (TDE) WFST250820mmsw/AT2025wet by the 2.5-meter Wide Field Survey Telescope (WFST). It exhibits a blue nuclear flare throughout the observed evolution with a g-band peak magnitude ~22, which is about 3 magnitudes brighter than its host galaxy. A Keck/LRIS spectrum taken near the optical peak reveals a featureless blue continuum, with no discernible emission lines. However, its redshift can be accurately determined to be 1.037 by its host galaxy absorption lines. Blackbody fits to the multiband spectral energy distribution (SED) of AT2025wet yield a constant temperature of ~19,000K and a peak luminosity of (8.27 +0.92 -0.71)*10^44 erg s^-1 while actually the SED likely peaks at a much shorter wavelength than a 19,000K blackbody. The SED modeling of the host galaxy implies a stellar mass of ~10^11.2 M_odot and an estimated central black hole mass of ~10^8 M_odot, with no evidence of significant active galactic nucleus activity prior to the flare. All of these observations are well consistent with a featureless TDE scenario, making it the highest-redshift non-jetted TDE known to date. TDEs at such high redshift provide us a unique opportunity to explore the intrinsic SEDs of TDEs, particularly to test whether they peak in the extreme-UV regime, thereby addressing the missing energy puzzle and the origin of optical emission in TDEs. Ongoing surveys represented by WFST and the Legacy Survey of Space and Time (LSST) are expected to discover an increasing number of TDEs at higher redshifts, which will extend our census of SMBHs across redshift space and help unravel the mysteries of optical TDEs through direct probes of their UV emission.

astro-ph.HE

WildFeedback: Aligning LLMs With In-situ User Interactions And Feedback

As large language models (LLMs) continue to advance, aligning these models with human preferences has emerged as a critical challenge. Traditional alignment methods, relying on human or LLM annotated datasets, are limited by their resource-intensive nature, inherent subjectivity, misalignment with real-world user preferences, and the risk of feedback loops that amplify model biases. To overcome these limitations, we introduce WildFeedback, a novel framework that leverages in-situ user feedback during conversations with LLMs to create preference datasets automatically. Given a corpus of multi-turn user-LLM conversation, WildFeedback identifies and classifies user feedback to LLM responses between conversation turns. The user feedback is then used to create examples of preferred and dispreferred responses according to users' preference. Our experiments demonstrate that LLMs fine-tuned on WildFeedback dataset exhibit significantly improved alignment with user preferences, as evidenced by both traditional benchmarks and our proposed checklist-guided evaluation. By incorporating in-situ feedback from actual users, WildFeedback addresses the scalability, subjectivity, and bias challenges that plague existing approaches, marking a significant step toward developing LLMs that are more responsive to the diverse and evolving needs of their users.

cs.CL

Panoramic Affordance Prediction

Affordance prediction serves as a critical bridge between perception and action in embodied AI. However, existing research is confined to pinhole camera models, which suffer from narrow Fields of View (FoV) and fragmented observations, often missing critical holistic environmental context. In this paper, we present the first exploration into Panoramic Affordance Prediction, utilizing 360-degree imagery to capture global spatial relationships and holistic scene understanding. To facilitate this novel task, we first introduce PAP-12K, a large-scale benchmark dataset containing over 1,000 ultra-high-resolution (12k, 11904 x 5952) panoramic images with over 12k carefully annotated QA pairs and affordance masks. Furthermore, we propose PAP, a training-free, coarse-to-fine pipeline inspired by the human foveal visual system to tackle the ultra-high resolution and severe distortion inherent in panoramic images. PAP employs recursive visual routing via grid prompting to progressively locate targets, applies an adaptive gaze mechanism to rectify local geometric distortions, and utilizes a cascaded grounding pipeline to extract precise instance-level masks. Experimental results on PAP-12K reveal that existing affordance prediction methods designed for standard perspective images suffer severe performance degradation and fail due to the unique challenges of panoramic vision. In contrast, PAP framework effectively overcomes these obstacles, significantly outperforming state-of-the-art baselines and highlighting the immense potential of panoramic perception for robust embodied intelligence.

cs.CV

DVD: Deterministic Video Depth Estimation with Generative Priors

Existing video depth estimation faces a fundamental trade-off: generative models suffer from stochastic geometric hallucinations and scale drift, while discriminative models demand massive labeled datasets to resolve semantic ambiguities. To break this impasse, we present DVD, the first framework to deterministically adapt pre-trained video diffusion models into single-pass depth regressors. Specifically, DVD features three core designs: (i) repurposing the diffusion timestep as a structural anchor to balance global stability with high-frequency details; (ii) latent manifold rectification (LMR) to mitigate regression-induced over-smoothing, enforcing differential constraints to restore sharp boundaries and coherent motion; and (iii) global affine coherence, an inherent property bounding inter-window divergence, which enables seamless long-video inference without requiring complex temporal alignment. Extensive experiments demonstrate that DVD achieves state-of-the-art zero-shot performance across benchmarks. Furthermore, DVD successfully unlocks the profound geometric priors implicit in video foundation models using 163x less task-specific data than leading baselines. Notably, we fully release our pipeline, providing the whole training suite for SOTA video depth estimation to benefit the open-source community.

cs.CV

WFST Astrometric Calibration -- I. Modeling Global Geometric Distortion with Zernike Polynomials

Accurate modeling of geometric distortion is essential for precise astrometric calibration in wide-field imaging surveys. We present a self-calibration method based on Zernike polynomials, applied to imaging data from the Wide Field Survey Telescope (WFST). Our approach constructs a global geometric distortion (GD) model from the position offsets of stars in the WFST r-band relative to Gaia DR3, achieving a median systematic uncertainty of below 10 mas for individual exposures. The correspondence between Zernike polynomials and optical aberrations reveals that the global GD of WFST is dominated by coma, inherent to the optical design, while rapid variations are likely attributed to the atmospheric dispersion corrector. Applying this method to 82 exposures from a single night (20250218), we find that the relative positions of the WFST CCDs remain stable, with standard deviations of less than 0.1 pixel in translation and 1.8 arcsec in rotation. The corrected WFST astrometric system is thereby tied to the Gaia DR3 coordinate frame, with further refinements to be presented in future work.

astro-ph.IM

WFST Supernovae in the First Year: I. Statistical Study of 16 Early-phase Type Ia Supernovae from the Pilot Survey

In this paper we present 16 early-phase type Ia supernovae (SNe Ia) discovered during the pilot survey of the 2.5-meter Wide Field Survey Telescope (WFST-PS) from March 4 to July 10, 2024, including three SNe Ia with early-excess emission features (EExSNe Ia). The discovery magnitude of the 16 WFST-PS early-phase SNe is at least 3 mag fainter than their peak brightness. A large scatter of color indices is found in approximately the first 10 days of supernova explosions, indicating diverse photometric behaviors in the early phase. Three EExSNe Ia show relatively brighter peak luminosities and longer rise time compared to those of non-EExSNe Ia. The results indicate that current theoretical models require further refinement to fully capture the early photometric evolution of SNe Ia. Based on the initial high-cadence ugr-band data from the WFST-PS survey, we emphasize that early near-ultraviolet (NUV) observations are indispensable for placing tight constraints on the explosion mechanisms and progenitor systems of SNe Ia.

astro-ph.HE

WFST Supernovae in the First Year: II. SN 2024aedt: Systematical Study of a Transitional Type Ia Supernova

We present comprehensive photometric and spectroscopic observations of a transitional type Ia SN 2024aedt, discovered by the 2.5-meter Wide Field Survey Telescope (WFST) within one day of the explosion. Its light curve is characterized by a peak absolute magnitude of $M_B = -18.49 \pm 0.03$ mag and a decline rate of $Δm_{15}(B) = 1.53 \pm 0.36$ mag, placing the object on the $Δm_{15}(B)$--$M_B$ diagram in the transition region between normal and subluminous SNe Ia. Furthermore, the early-color evolution and host galaxy environment of SN 2024aedt underscore its transitional nature, sharing properties with both normal and 91bg-like SNe Ia. Light-curve modeling with MOSFiT yields a synthesized $^{56}\mathrm{Ni}$ mass of $0.414 \pm 0.042\,M_{\odot}$ and a total ejecta mass of $0.548 \pm 0.108\,M_{\odot}$. A comparison with theoretical models suggests that the evolutionary trend can be broadly explained by both delayed-detonation (DDT) and double-detonation (DDet) scenarios while possible early-excess emissions predicted by DDet cannot be identified given the limited detections soon after the SN explosion. Although the overall spectral evolution of SN 2024aedt is similar to that of other transitional SNe Ia, the spectroscopic comparison reveals diversity in the early-phase blue-end features, which becomes more homogeneous at later phases. The result indicates the importance of early-time observations in understanding the origin of SN Ia diversity.

astro-ph.HE

WFST Supernovae in the First Year: III. Systematical Study of the Photometric Behavior of Early-phase Core-collapse Supernovae

We investigate the multiband photometric properties of seven supernovae (SNe) showing double-peaked light-curve evolution and prominent shock-cooling emission, observed by the Wide Field Survey Telescope (WFST) during its first year of operation. By jointly employing an analytic early shock-cooling model and the Arnett radioactive-diffusion model, we fit the bolometric light curves and infer ejecta masses in the range $1.1$-$2.6 M_\odot$, consistent with a transitional population between ultra-stripped supernovae (USSNe) and normal stripped-envelope supernovae (SESNe). The envelope masses are estimated to be $M_{\rm env}=0.1$-$0.4 M_\odot$, while the progenitors are constrained to be yellow or blue supergiants (YSGs/BSGs) with radii of $R=120$-$300 R_\odot$. Using empirical relations, we estimate progenitor luminosities of $L=10^{4.6}$-$10^{4.9} L_\odot$, corresponding to zero-age main-sequence (ZAMS) masses of $8$-$20 M_\odot$. Theoretical models suggest that such progenitors are more naturally produced through binary evolution channels, as single-star evolutionary pathways are unable to yield ejecta masses this low.

astro-ph.HE

Illuminating the Mass Gap Through Deep Optical Constraint on a Neutron Star Merger Candidate S250206dm

The gravitational wave (GW) event S250206dm, as the first well-localized neutron star merger candidate potentially located in the mass gap, presented a unique opportunity to probe the electromagnetic signatures from such a system. Here we report a deep, multiband search with the new 2.5-meter Wide Field Survey Telescope (WFST), covering about 64% of the localization region up to a 5-sigma limiting magnitude of 23 mag. In total, 12 potential candidates have been identified while none of them are likely related to S250206dm. This non-detection provides the most stringent constraint to date on any associated kilonova. Crucially, an AT 2017gfo-like event at 269 Mpc can be excluded by WFST observations alone. Based on ejecta mass limits, a neutron star-black hole with a large mass ratio (Q >= 3.2) is disfavored. This optical-derived constraint on the mass ratio reaches, for the first time, a precision comparable to that inferred from the GW signal. This work presents the best observation of this type of events until now, and demonstrates the power of rapid, deep follow-up observations to constrain the properties of compact binary progenitors, offering key insights into the constituents of the mass gap.

astro-ph.HE

Show, Don't Tell: Morphing Latent Reasoning into Image Generation

Text-to-image (T2I) generation has achieved remarkable progress, yet existing methods often lack the ability to dynamically reason and refine during generation--a hallmark of human creativity. Current reasoning-augmented paradigms most rely on explicit thought processes, where intermediate reasoning is decoded into discrete text at fixed steps with frequent image decoding and re-encoding, leading to inefficiencies, information loss, and cognitive mismatches. To bridge this gap, we introduce LatentMorph, a novel framework that seamlessly integrates implicit latent reasoning into the T2I generation process. At its core, LatentMorph introduces four lightweight components: (i) a condenser for summarizing intermediate generation states into compact visual memory, (ii) a translator for converting latent thoughts into actionable guidance, (iii) a shaper for dynamically steering next image token predictions, and (iv) an RL-trained invoker for adaptively determining when to invoke reasoning. By performing reasoning entirely in continuous latent spaces, LatentMorph avoids the bottlenecks of explicit reasoning and enables more adaptive self-refinement. Extensive experiments demonstrate that LatentMorph (I) enhances the base model Janus-Pro by $16\%$ on GenEval and $25\%$ on T2I-CompBench; (II) outperforms explicit paradigms (e.g., TwiG) by $15\%$ and $11\%$ on abstract reasoning tasks like WISE and IPV-Txt, (III) while reducing inference time by $44\%$ and token consumption by $51\%$; and (IV) exhibits $71\%$ cognitive alignment with human intuition on reasoning invocation.

cs.CV

Asymptotic Behavior of the Principal Eigenvalue Problems with Large Divergence-Free Drifts

In this paper, we consider the following principal eigenvalue problem with a large divergence-free drift: \begin{equation}\label{0.1} -\varepsilonΔϕ-2α\nabla m(x)\cdot\nabla ϕ+V(x)ϕ=λ_αϕ \,\ \text{in}\, \ H_0^1(Ω),\tag{0.1} \end{equation} where the domain $Ω\subset \mathbb{R}^N (N\ge 1)$ is bounded with smooth boundary $\partialΩ$, the constants $\varepsilon>0$ and $α>0$ are the diffusion and drift coefficients, respectively, and $m(x)\in C^{2}(\barΩ)$, $V (x)\in C^γ(\barΩ)~(0<γ<1)$ are given functions. For a class of divergence-free drifts where $m$ is a harmonic function in $Ω$ and has no first integral in $H_{0}^{1}(Ω)$, we prove the convergence of the principal eigenpair $(λ_α, ϕ)$ for (0.1) as $α\rightarrow+\infty$, which addresses a special case of the open question proposed in [H. Berestycki, F. Hamel and N. Nadirashvili, CMP, 2005]. Moreover, we further investigate the refined limiting profiles of the principal eigenpair $(λ_α, ϕ)$ for (0.1) as $α\rightarrow+\infty$, which display the visible effects of the large divergence-free drifts on the principal eigenpair $(λ_α, ϕ)$.

math.AP

Artificial intelligence pioneers the double-strangeness factory

Artificial intelligence (AI) is transforming not only our daily experiences but also the technological development landscape and scientific research. In this study, we pioneered the application of AI in double-strangeness hypernuclear studies. These studies which investigate quantum systems with strangeness via hyperon interactions provide insights into fundamental baryon-baryon interactions and contribute to our understanding of the nuclear force and composition of neutron star cores. Specifically, we report the observation of a double hypernucleus in nuclear emulsion achieved via innovative integration of machine learning techniques. The proposed methodology leverages generative AI and Monte Carlo simulations to produce training datasets combined with object detection AI for effective event identification. Based on the kinematic analysis and charge identification, the observed event was uniquely identified as the production and decay of resulting from Ξ- capture by 14N in the nuclear emulsion. Assuming capture in the atomic 3D state, the binding energy of the two Λ hyperons in 13BΛΛ, BΛΛ, was determined as 25.57 +- 1.18(stat.) +- 0.07(syst.) MeV. The ΛΛ interaction energy obtained was 2.83 +- 1.18(stat.) +- 0.14(syst.) MeV. This study marks a new era in double-strangeness research.

nucl-ex

Refined Limiting Profiles of the Principal Eigenvalue Problems with Large Advection

In this paper, we are concerned with the following eigenvalue problem with an advection term: \begin{equation}\label{0.1} \left\{ \begin{split} -εΔϕ-2α\nabla m(x)\cdot\nabla ϕ+V(x)ϕ&=λϕ \ \text{in}\ \ Ω,\\ ϕ&=0\ \ \hbox{on}\ \ \partialΩ, ~~~\text{(0.1)} \end{split} \right. \end{equation} where $Ω\subset\mathbb{R}^N~(N\geq1)$ satisfying $\partialΩ\in C^{2}$ is a bounded domain and contains the origin as an interior point, the constants $ε>0$ and $α>0$ are the diffusive and advection coefficients, respectively, and $m(x)\in C^{2}(\barΩ)$, $V (x)\in C^γ(\barΩ)~(0<γ<1)$ are given functions. We analyze the refined limiting profiles of the principal eigenpair $(λ, ϕ)$ for (0.1) as $α\rightarrow\infty$, which display the visible effect of the large advection on $(λ, ϕ)$. It expects that our argument is applicable to investigating the refined expansions of the general principal eigenvalue problems.

math.AP