arXiv ScienceSearch

arXiv subjects

Hui Sun

Publications and source records attributed to Hui Sun.

At least 19 recordsLinked to original sources

CIT-CAD: Constraint Intent Tree-based CAD Code Generation and Verification

Natural-language Computer-Aided Design (CAD) code generation aims to turn design intent into executable and editable parametric programs. Large language models (LLMs) make this goal increasingly practical, but useful systems must preserve the construction process behind the rendered geometry. Existing benchmarks and methods mostly focus on how closely the generated CAD model matches the reference geometry, often using metrics such as Intersection over Union (IoU). Such metrics can miss errors in part decomposition, construction hierarchy, Boolean operations, sketch structure, and geometric relations. This gap calls for a representation that makes design intent explicit and lets a system check generated code against that intent. We propose CIT-CAD, a framework that infers a Constraint Intent Tree (CIT) from the input description to represent the intended entities, hierarchy, operations, and relations. The tree has two roles: it guides CAD code generation and defines expected constraints for verification. The framework extracts actual constraints from the generated program, compares them with the expected constraints, and uses mismatches to localize and repair design violations. Experiments show that the framework improves CAD generation performance, with larger gains on more complex multi-entity designs. By turning design intent into an explicit and checkable object, this work is the first attempt to move text-to-CAD generation beyond rendered-geometry matching toward construction-aware synthesis, verification, and repair.

cs.AI

Evaluating Language Models on Cross-Language Code Functional Equivalence

Background: Large Language Models (LLMs) have demonstrated strong performance across a variety of code-understanding tasks, leading many to believe that they can reason about program semantics. However, existing evaluations primarily focus on single-language settings or rely on synthetically generated code, raising concerns about whether current results reflect true semantic understanding. Aims: We investigate whether LLMs can accurately judge functional equivalence across different programming languages in human-written code, a setting that requires deeper reasoning beyond superficial similarity. Method: We introduce PolyHuman, a dataset of human-written programs in CPP, Java, and Python. Using this dataset, we evaluate intra- and inter-language equivalence detection across open-weight and proprietary LLMs, selecting GPT-o4-mini as a representative model to assess stability. We then manually analyze 81 cases of systematic disagreement in which models incorrectly judge functional equivalence, examining the code logic and the generated Chain-of-Thought reasoning. Finally, we categorize these failures and compare them across GPT-o4-mini, Claude-Opus-4.7, and Gemini-3-Flash to determine whether they reflect model-specific issues or broader limitations of state-of-the-art LLMs. Results: We identify a difficulty-dependent breakdown in equivalence judgment (harder problems make the model increasingly prone to misclassifying non-equivalent code as equivalent), a model-specific sensitivity to programming language for the best-performing model (particularly a more conservative behavior on Python), and a partial reliance on similarity-based cues. GPT-o4-mini also shows substantial run-to-run instability under identical settings, indicating inconsistent rather than absent capability. Conclusions: Current LLMs do not reliably capture functional equivalence within or across languages.

cs.SE

Nonparametric Schr\"odinger Bridge Time Series Generator: Algorithm, Convergence Analysis and Applications

We conduct a convergence analysis for the Schr\"odinger Bridge Time Series (SBTS) data generator. Starting from a regularized formulation in which the data ensemble is mixed with a standard multivariate Gaussian distribution with a prescribed probability, we prove that the Euler-Maruyama discretization converges to the mixed target distribution with half-order convergence rate, provided that the ensemble size and kernel bandwidth are chosen appropriately. We further show that the regularized distribution converges to the original target distribution as the mixing probability tends to zero. The analysis simultaneously accounts for the ensemble approximation error, kernel approximation error, and time-discretization error, and therefore provides a full distributional convergence result for the SBTS generator. Empirically, we further examine the flexibility of the method by replacing the Wiener reference measure with the path measure induced by a more general SDE. The numerical experiments show that the schemes based on both the original Wiener reference measure and the SDE-induced reference measure achieve comparable performance, demonstrating the robustness and stability of the SBTS framework.

math.NA

REATS: LLM Reasoning-based Ensemble Learning for Adaptive Time Series Forecasting

Due to the diversity of real-world time series, no single forecasting model consistently dominates across all samples. Ensemble learning addresses this by combining complementary model strengths, yet existing methods rely on fixed rules or black-box models based solely on numerical inputs, failing to leverage LLM reasoning for interpretable weighting decisions. We propose REATS, which leverages LLM reasoning capabilities as an intelligent ensemble router that jointly processes textual temporal pattern descriptions and numerical features to produce interpretable, sample-adaptive ensemble weights through chain-of-thought reasoning. To enable effective LLM-based ensembling, we study its key design choices and propose: (i) a structured input pipeline that transforms raw time series into hybrid textual--numerical representations with fixed token cost, enabling rule-based chain-of-thought construction without API dependency, augmented with retrieved similar-sample priors; (ii) a diverse multi-row weight supervision scheme coupled with a token-efficient percentage-table format that reduces numerical complexity and mitigates LLM hallucinations; and (iii) a two-stage fine-tuning framework combining SFT with GRPO, where a reciprocal reward mapping transforms the continuous unbounded MSE gap into bounded signals with amplified near-oracle sensitivity, addressing the uniform sensitivity and outlier-dominated advantage compression inherent in naive reward designs for regression-based GRPO. Experiments on eight benchmarks demonstrate that REATS outperforms competitive ensemble baselines while providing natural language explanations and demonstrating strong transfer learning and out-of-domain generalization to unseen candidate models.

cs.LG

A JWST redshift for the host galaxy of EP250207b of z=3.2: a collapsar origin is viable

We present James Webb Space Telescope (JWST) and Hubble Space Telescope (HST) observations of the field of the fast X-ray transient (FXT) detected by Einstein Probe, EP250207b, to resolve any ambiguity about the host galaxy and redshift of the FXT. EP250207b was originally associated with a nearby galaxy at z=0.082, based on its low chance alignment probability, and a binary neutron star merger origin was proposed. However, we report the detection of a background galaxy at z=3.2 at the location of EP250207b. Assuming this galaxy is the actual host galaxy, the rest-frame energetics and timescales of the event change. Furthermore, the available data are not able to rule out the presence of a supernova associated with EP250207b if at this redshift. We model the X-ray, optical, near-infrared and radio light curves using a tophat jet model implemented in Redback and find that they are consistent with an on-axis gamma ray burst afterglow. The energetics and host galaxy properties do not allow us to distinguish between a collapsar and a merger driven event.

astro-ph.HE

Auxiliary-Channel-Assisted Cross-Talk Noise Removal in LISA Pathfinder

LISA Pathfinder (LPF) is the technology demonstration mission for the future Laser Interferometer Space Antenna (LISA). Besides the science interferometric channel, LPF is equipped with numerous auxiliary channels that monitor the instrument and environment disturbances, some of which contain information correlated with noise in the science channel. In this work, we investigate cross-talk noise subtraction in LPF using a frequency-domain transfer-function approach, in which the coherence between the science channel and auxiliary channels is used to estimate and remove correlated noise. The effectiveness of this method is first validated using simulated science and auxiliary channel data with a common disturbance. The simulation shows that the subtraction performance strongly depends on the auxiliary-channel noise level, providing a practical framework for determining the auxiliary-channel noise requirement needed to achieve a desired subtraction performance. This method is then used to analyze 9 auxiliary channels in publicly available LPF telemetry data. In the cross-talk dominated frequency band from $1\times10^{-2}$ to $6\times10^{-2}\,\mathrm{Hz}$, this method achieves cross-talk noise suppression comparable to the standard fitting based subtraction. Above $6\times10^{-2}\,\mathrm{Hz}$, it further suppresses the residual noise by avoiding the introduction of auxiliary-channel readout noise associated with the fitting procedure, resulting in better performance than the standard pipeline.

astro-ph.IM

M-EPDet: Real-Time Real-Bogus Classification and Transient Candidate Judgement for the EP-WXT Pipeline via Multi-Modal Data

The Wide-field X-ray Telescope (WXT) onboard the Einstein Probe (EP) produces a large post-detection candidate stream in which genuine astrophysical sources coexist with instrumental artifacts and Cosmic Ray events. We present M-EPDet, a three-step post-detection framework for real-time candidate vetting in EP-WXT lobster-eye Micro-pore Optics (MPO) data. The framework combines a ResNet-based Arm filter, a dual-branch temporal-spectral Cosmic Ray filter, and a background-aware Bayesian Blocks module for single-exposure variability screening. Using on-orbit EP-WXT observations, we report decoupled metrics for the cascading system. M-EPDet achieves a Real-Bogus Recall of 98.31\% ($98.53\% \times 99.78\%$) for genuine astrophysical sources, together with rejection rates of 92.99\% for instrumental artifacts and 98.18\% for Cosmic Ray events. In the final step, the Bayesian Blocks module flags 0.75\% of the post-filtration observations, corresponding to a 99.25\% reduction in candidate volume. The system is deployed in the EP-WXT pipeline as a lightweight real-time service, reducing the manual-inspection burden in candidate vetting.

astro-ph.IM

X-rays breaking out of pre-explosion ejecta mark a supernova's first light

Massive stars die as core-collapse supernovae, whose optical light emerges days after the implosion. Theory predicts that the initial collapse-driven shock, upon breaking through the star and dense circumstellar medium, emits a brief thermal flash of soft X-rays and ultraviolet. Yet these elusive first signals have remained largely undetected, owing to limited wide-field soft X-ray monitoring. Here we report the discovery of a soft X-ray flash, EP260321a, followed days later by a broad-lined supernova from an envelope-stripped progenitor. Its X-ray spectrum, best modeled with blackbody, establishes it as the long-sought archetypal shock breakout. The burst's duration and energetics place the breakout at a radius of 300 solar radii, tracing a dense surrounding shell and revealing abrupt mass ejection within the final month before collapse.

astro-ph.HE

Exploring Pass-Rate Reward in Reinforcement Learning for Code Generation

Reinforcement learning (RL) from unit-test feedback has become a standard post-training recipe for improving large language models (LLMs) on code generation. However, the pass-all-tests binary reward can be sparse, yielding no learning signal on challenging problems where none of the sampled solutions passes all tests. A common remedy is to use the test-case pass rate as a surrogate reward. In this work, we study pass-rate rewards in critic-free RL for code generation (e.g., GRPO and RLOO) and report a consistent pattern across base models and algorithms: despite alleviating reward sparsity, pass-rate rewards do not reliably improve final performance over binary rewards in rigorous controlled experiments. To understand this discrepancy, we analyze reward density and the resulting gradient directions. We find that pass-rate rewards are denser, but the induced gradient updates do not consistently move probability mass toward full-pass solutions. This arises because test-case pass rate is a miscalibrated surrogate for progress toward full correctness, and partial-pass solutions within the same group can induce conflicting gradient directions that cancel out. Overall, our results suggest that, in critic-free RL, pass-rate rewards are insufficient to improve code generation and motivate reward designs that better align optimization with the goal of full correctness.

cs.LG

Multi-wavelength study of EP250416a / GRB 250416C: An Optically Dark Long GRB with a Late Jet Break

We present multi-wavelength study of the $\gamma$/X-ray transient EP250416a (also designated GRB 250416C), triggered by the Einstein Probe (EP) Wide-field X-ray Telescope and also by SVOM and Konus-Wind. Observations spanning the gamma-ray, X-ray, and optical bands facilitated detailed analysis of the burst's prompt emission, afterglow evolution, and physical origin. EP250416a exhibits a burst duration of 30 s in X-ray and 17.7 s in gamma-rays, with joint spectral fitting of 0.5-5000 keV data gives $E\rm_{peak}=342_{-232}^{+90}$ keV. Optical spectroscopy of the afterglow, acquired with the Gemini Multi-Object Spectrograph (GMOS) on Gemini South, yielded a redshift of $z=0.963$. Accounting for the measured redshift, the isotropic energies are $E\rm_{X,iso}=2.7_{-0.5}^{+0.9}\times10^{50}$ erg and $E\rm_{\gamma,iso}=7.34_{-2.1}^{+5.1}\times10^{51}$ erg, aligning with the Amati relation for long GRBs. The fluence ratio $\rm S(25-50~keV)/S(50-100~keV)=0.78_{-0.15}^{+0.1}$ classifies EP250416a as an X-ray rich (XRR) GRB. The X-ray afterglow shows an initial shallow decay ($\alpha \approx -0.5$) transitioning to a canonical decay phase ($\alpha \approx -1$), with a very late jet break at $t\sim 1.5\times 10^6$ s, corresponding to a jet half-opening angle of $\theta _j=10.6_{-1.8}^{+1.9}$ degrees. EP250416a is optically dark, as it shows only a faint $r$-band detection ($r=24.16$ mag) from Gemini South-GMOS and a low optical-to-X-ray spectral index $\beta_{\rm OX} = 0.3$. This may be attributed to significant host-galaxy extinction, with a required $A_V^{\text{host}}=5.5\ \text{mag}$ derived from the extinction curve model.

astro-ph.HE

A fast X-ray transient with chromatic flares: signatures of violent collisions induced by late-time central engine reactivation

Extragalactic Fast X-ray Transients (EFXTs) represent an emerging class of high-energy phenomena characterized by X-ray outbursts lasting from tens to hundreds of seconds. However, for more than half of the EFXTs, their physical origins remain elusive. In this Letter, we report the discovery of EP250302a, a luminous EFXT detected by the Einstein Probe (EP) at a redshift of $z = 1.131$. The multi-wavelength light curves of EP250302a reveal remarkable temporal features that distinguish it from the previously known EP-detected EFXT population, most notably a needle-like X-ray flare accompanied by smooth optical rebrightening during the afterglow phase. We suggest that the distinct X-ray and optical behaviors constitute the first observed instance of late-time violent collision of two relativistic shells in an EFXT. Drawing on insights from GRB studies, such a collision process strongly indicates the reactivation of a central engine, making EP250302a-like transients a unique laboratory for probing the late-time activity and jet physics of EFXT central engines.

astro-ph.HE

A Sample-Wise Adjoint Regression Framework for Mean-Field Control with Connections to Adjoint Matching

This work proposes a novel numerical approach for solving mean-field control (MFC) problems using an adjoint-based optimization framework motivated by the stochastic maximum principle (SMP). Rather than solving the adjoint processes $(Y_t,Z_t)$ explicitly, we construct sample-wise unbiased estimators of their discretized counterparts and use them to approximate the Hamiltonian control gradient through recursive regression. The control is then updated by a gradient-descent scheme. This approach differs from conventional deep-learning methods, which typically follow a ``discretize-then-optimize'' paradigm and directly optimize a globally parameterized control through the discretized objective. Numerical experiments demonstrate competitive, and in many cases improved performance compared with direct deep-learning approaches. The sample-wise and regression-based structure also supports scalable implementation in high-dimensional settings and makes the method attractive for generative modeling problems involving particle-based representations and distributional objectives. On the theoretical side, we establish a connection between a particular class of MFC problems and the \emph{adjoint matching} framework. Using the SMP under the mean field control setting, we further show that the self-consistent critical point of the resulting mean-field adjoint matching loss coincides with the optimal control.

math.OC

Efficient Learned Data Compression via Dual-Stream Feature Decoupling

While Learned Data Compression (LDC) has achieved superior compression ratios, balancing precise probability modeling with system efficiency remains challenging. Crucially, uniform single-stream architectures struggle to simultaneously capture micro-syntactic and macro-semantic features, necessitating deep serial stacking that exacerbates latency. Compounding this, heterogeneous systems are constrained by device speed mismatches, where throughput is capped by Amdahl's Law due to serial processing. To this end, we propose a Dual-Stream Multi-Scale Decoupler that disentangles local and global contexts to replace deep serial processing with shallow parallel streams, and incorporate a Hierarchical Gated Refiner for adaptive feature refinement and precise probability modeling. Furthermore, we design a Concurrent Stream-Parallel Pipeline, which overcomes systemic bottlenecks to achieve full-pipeline parallelism. Extensive experiments demonstrate that our method achieves state-of-the-art performance in both compression ratio and throughput, while maintaining the lowest latency and memory usage. The code is available at https://github.com/huidong-ma/FADE.

cs.CL

ACES: Who Tests the Tests? Leave-One-Out AUC Consistency for Code Generation

Selecting LLM-generated code candidates using LLM-generated tests is challenging because the tests themselves may be incorrect. Existing methods either treat all tests equally or rely on ad-hoc heuristics to filter unreliable tests. Yet determining test correctness requires knowing which codes are correct, creating a \emph{circular dependency}. Our key insight is that we need not determine test correctness at all: \emph{test votes should rank, not merely count}. What matters is not how many codes pass a test, but whether the test can \emph{distinguish} correct from incorrect code. We break the circular dependency via leave-one-out evaluation: hold out one test, rank codes by their aggregate scores on all remaining tests, and measure whether the held-out test's pass/fail pattern agrees with this ranking. We formalize this agreement as the leave-one-out AUC~(LOO-AUC) and prove that the expected LOO-AUC is proportional to each test's ability to separate correct code from incorrect code. Building on this, we propose \textbf{ACES}~(\textbf{A}UC \textbf{C}onsist\textbf{E}ncy \textbf{S}coring) with two complementary variants: ACES-C provides closed-form weights that provably approximate the oracle in expectation under a mild assumption on average test quality; ACES-O drops this assumption and iteratively optimizes a differentiable LOO-AUC objective. Both operate solely on the binary pass matrix with negligible overhead, and achieve state-of-the-art Pass@$k$ on multiple code generation benchmarks.

cs.LG

Project Imaging-X: A Survey of 1000+ Open-Access Medical Imaging Datasets for Foundation Model Development

Foundation models have demonstrated remarkable success across diverse domains and tasks, primarily due to the thrive of large-scale, diverse, and high-quality datasets. However, in the field of medical imaging, the curation and assembling of such medical datasets are highly challenging due to the reliance on clinical expertise and strict ethical and privacy constraints, resulting in a scarcity of large-scale unified medical datasets and hindering the development of powerful medical foundation models. In this work, we present the largest survey to date of medical image datasets, covering over 1,000 open-access datasets with a systematic catalog of their modalities, tasks, anatomies, annotations, limitations, and potential for integration. Our analysis exposes a landscape that is modest in scale, fragmented across narrowly scoped tasks, and unevenly distributed across organs and modalities, which in turn limits the utility of existing medical image datasets for developing versatile and robust medical foundation models. To turn fragmentation into scale, we propose a metadata-driven fusion paradigm (MDFP) that integrates public datasets with shared modalities or tasks, thereby transforming multiple small data silos into larger, more coherent resources. Building on MDFP, we release an interactive discovery portal that enables end-to-end, automated medical image dataset integration, and compile all surveyed datasets into a unified, structured table that clearly summarizes their key characteristics and provides reference links, offering the community an accessible and comprehensive repository. By charting the current terrain and offering a principled path to dataset consolidation, our survey provides a practical roadmap for scaling medical imaging corpora, supporting faster data discovery, more principled dataset creation, and more capable medical foundation models.

cs.CV

Long Distance Daylight Drone-based Quantum Key Distribution under Relative Motion

Low-altitude drones can serve as dynamic nodes apparently mitigating terrain-induced impacts for quantum networks. However, it is extremely hard to establish a sable quantum link in a drone-based dynamic platform, which requires centimeter-level positioning techniques and high-precision time synchronization technologies. In this paper, we develop a single-ended polarization adaptive correction technology at both the transmitting and receiving ends. Based on this, we present the world's first kilometer-scale drone-based QKD network, achieving an 1.2 km free-space QKD link with a secure key rate of 2.76 kbps, suitable for urban quantum network deployment. We validate the feasibility of QKD between dynamic drone and ground unmanned vehicle at a relative speed of 1 m/s over a distance of 100 m, attaining a secure key rate of 70.94 kbps. This work advances drone-based QKD from static demonstrations to practical dynamic network, boasting great development potential for an airborne quantum internet.

quant-ph

Test Code Review in the Era of GitHub Actions: A Replication Study

Test code is indispensable in software development, ensuring the correctness of production code and supporting maintainability. Nonetheless, errors or omissions in the test code can conceal production defects. While code review is widely adopted to assess code quality and correctness, little research has examined how test code is reviewed. Spadini et al.'s research on Gerrit (a pre-commit review model) found that test code receives significantly less discussion than production code. However, the most popular review model is currently based on pull requests (PRs), in which contributors propose changes for discussion and approval, a more negotiable and flexible model compared to Gerrit. Furthermore, GitHub Actions (GHA) has become widely used to automate pre-checks and testing, potentially impacting review practices. This leads us to explore whether Spadini et al.'s findings still hold for the PR model in the era of GHA? Our work replicates and extends their work. We focus on GitHub PRs and analyze six open-source projects. We investigate the impact of the PR model and GHA on test code review. Our results show that GitHub's PR model fosters more balanced discussions between test and production files than Gerrit, albeit with lower overall comment density. However, despite cross-project heterogeneity, GHA adoption triggered a sharp pivot toward production code. Post-GHA, for PRs involving tests, both review probability and comment density reached a median of zero. These findings reveal how evolving continuous integration pipelines can marginalize test code review. The observed decline in test-centric discussion under GHA warrants concern regarding long-term software quality. Our work also presents recommendations for stakeholders involved in the software development life cycle.

cs.SE

Design-Specification Tiling for ICL-based CAD Code Generation

Large language models~(LLMs) have demonstrated remarkable capabilities in code generation, yet their performance remains limited on domain-specific tasks such as Computer-Aided Design~(CAD) code generation, largely due to the scarcity of high-quality training data. In-Context Learning~(ICL) provides a training-free alternative by prompting LLMs with task-specific exemplars, but its effectiveness critically depends on how exemplars are selected. Existing selection strategies mainly rely on similarity or point-wise diversity, often overlooking the compositional nature of CAD design specifications, where a query may involve multiple functional requirements, geometric constraints, and design primitives. As a result, selected exemplars can be individually relevant but collectively redundant, providing insufficient coverage for complex design requirements. In this work, we propose \emph{knowledge sufficiency} as a principled objective for exemplar selection, aiming to select a compact set of exemplars that maximally satisfies the requirements contained in a target design specification. To instantiate this objective, we introduce \emph{Design-Specification Tiling~(DST)}, which estimates knowledge sufficiency through a surrogate tiling ratio by decomposing design specifications into multi-granular components and measuring the proportion of query components covered by selected exemplars. We further show that optimizing this objective can be formulated as a submodular maximization problem, and develop a polynomial-time greedy algorithm tailored to this setting with a $(1-1/e)$-approximation guarantee. Extensive experiments across multiple LLMs demonstrate that DST substantially improves CAD code generation quality and consistently outperforms existing ICL exemplar selection strategies, highlighting the importance of requirement-level knowledge coverage for domain-specific code generation.

cs.SE