arXiv ScienceSearch

arXiv subjects

Shiyu Wang

Publications and source records attributed to Shiyu Wang.

At least 19 recordsLinked to original sources

Low-leakage superconducting-qubit measurement with sub-100-ns total duration

Fast, accurate, and low-leakage qubit measurement is a key requirement for quantum error correction. Here, we demonstrate measurement of a superconducting transmon qubit with a total duration of 97(1) ns, defined as the time from the start of the measurement pulse until the measurement-induced error on a subsequent $π$-pulse operation falls below $10^{-4}$. By combining a large state-averaged resonator decay rate of $κ_\mathrm{eff}/2π$ = 30.8 MHz with a dispersive shift close to the optimal SNR-per-photon condition, we achieve an assignment error of 0.17(1)% using a 58-ns measurement pulse, with residual readout photons depleting passively in tens of nanoseconds without an active depletion pulse. Using a repeated-measurement sequence together with a leakage-sensitive measurement, we benchmark the measurement-induced state transitions, finding a per-measurement leakage rate of $2.7(2) \times 10^{-5}$, only twice the background rate and two orders of magnitude below the measurement-induced relaxation rate, which dominates the assignment error. Floquet simulations indicate that the multiphoton resonances present at the operating point are weakly coupled and traversed diabatically, without causing leakage. These results demonstrate that a large resonator decay rate, combined with a dispersive shift close to the optimal SNR-per-photon condition, can enable fast, high-fidelity, low-leakage dispersive readout at small qubit-resonator detuning.

quant-ph

When Do Corrective Features Help? An Agent for Corrective Feature Discovery on Black-Box Forecasters

Frozen pretrained forecasters often fail in structured, recurring ways that are costly to repair through fine-tuning. We study corrective feature discovery: mining interpretable features of a frozen forecaster's residual to drive a lightweight post-hoc corrector. Prior automated feature engineering models the data-generating process; corrective features instead model the model-failure process. We present CRAFTER (Corrective Residual Agent with Feature-based Temporal Exploration and Reasoning), which keeps the backbone frozen and mines its residual with two complementary generators: a compositional search over the raw input channels, and a large language model (LLM) that proposes named feature combinations, binary flags, and short executable code. A single validation-grounded gate accepts or rejects every candidate regardless of its origin, and a validation-selected corrector applies the accepted features or leaves the forecast unchanged. This source-agnostic pipeline also allows prior feature-engineering systems to be evaluated under identical conditions, making CRAFTER an instrument for attributing forecast improvements to the feature source alone. Across six public datasets and six frozen backbones, CRAFTER surpasses every dedicated feature-engineering system at every feature budget, roughly doubling the improvement achieved by the corrector alone and reducing the error of the weakest backbones by up to 27%. These gains are robust across different LLM backends and persist even when applied on top of fine-tuned backbones.

cs.LG

QuBE/Qubex: an integrated hardware-software system for superconducting qubit experiments with broadband control

Achieving high-fidelity operation in large-scale superconducting qubit systems requires not only control hardware with broad frequency coverage, low crosstalk, and tight synchronization but also software that coordinates system configuration, experiment execution, and data analysis. Here we present an integrated qubit-control system that combines broadband microwave hardware with a pulse-level software stack for scalable superconducting qubit experiments. The hardware provides broadband microwave coverage, including an instantaneous span of up to 1.6 GHz from a control output, while the software reduces setup and calibration overhead through automated configuration and built-in experiment workflows. We validate the system on a 64-qubit fixed-frequency transmon chip through full-chip frequency identification and representative demonstrations, including multi-unit far-detuned cross-resonance calibration and benchmarking that yields a measured two-qubit gate fidelity of 98.34%, and multilevel readout beyond the computational subspace. By disclosing the hardware architecture and releasing the software stack as open source, this work provides an inspectable hardware-software foundation for scalable superconducting qubit control experiments.

quant-ph

Sonar-TS: Search-Then-Verify Natural Language Querying for Time Series Databases

Natural Language Querying for Time Series Databases (NLQ4TSDB) aims to assist non-expert users retrieve meaningful events, intervals, and summaries from massive temporal records. However, existing Text-to-SQL methods are not designed for continuous morphological intents such as shapes or anomalies, while time series models struggle to handle ultra-long histories. To address these challenges, we propose Sonar-TS, a neuro-symbolic framework that tackles NLQ4TSDB via a Search-Then-Verify pipeline. Analogous to active sonar, it utilizes a feature index to ping candidate windows via SQL, followed by generated Python programs to lock on and verify candidates against raw signals. To enable effective evaluation, we introduce NLQTSBench, the first large-scale benchmark designed for NLQ over TSDB-scale histories. Our experiments highlight the unique challenges within this domain and demonstrate that Sonar-TS effectively navigates complex temporal queries where traditional methods fail. This work presents the first systematic study of NLQ4TSDB, offering a general framework and evaluation standard to facilitate future research.

cs.AI

Machine learning study on single production of a singlet vectorlike lepton at the Large Hadron Collider

Vectorlike leptons are nonchiral, colorless fermions from new physics beyond the Standard Model, appearing in many theoretical extensions. We investigate the prospect for detecting the single production of a singlet vectorlike lepton that mixes with the $τ$ lepton at the Large Hadron Collider. The corresponding final states are classified as the three- and four-lepton search channels. The machine learning algorithm XGBoost is employed to enhance signal-background discrimination. Our analysis indicates that, at $\sqrt{s} = 14~\mathrm{TeV}$ with an integrated luminosity of $3000~\mathrm{fb}^{-1}$ under the assumption of negligible systematic uncertainties, the expected $2σ$ exclusion limits in the three- and four-lepton channels can reach vectorlike lepton masses up to $500$ and $405~\mathrm{GeV}$ in the parameter region allowed by the electroweak oblique parameter constraint, respectively. These findings demonstrate that machine learning techniques can substantially improve the sensitivity of collider searches for vectorlike leptons.

hep-ph

InsightVQA: High-Dimensional Emotion-Cognitive Visual Question Answering Benchmark

Visual emotion understanding requires models not only to recognize emotional states, but also to why they arise and perform higher-level cognitive reasoning. However, existing benchmarks mainly focus on emotion recognition, offering limited support for grounded understanding and response-oriented analysis. To address this gap, we introduce \textbf{InsightVQA}, a large-scale dataset for hierarchical visual question answering on emotion understanding and cognitive reasoning. Building from 351K images collected from six public sources, we apply a rigorous multi-stage filtering pipeline to curate 138K high-confidence images. Each image is annotated at three hierarchical levels: perception QA for emotion and valence recognition, grounded understanding QA constructed from visual trigger extraction through constraint-guided generation, and cognition QA centered on response intent prediction and sequential insight reasoning. In total, InsightVQA contains 725K QA pairs. We further present \textbf{InsightVQA-Bench}, a high-quality evaluation benchmark comprising 30K samples for fine-grained evaluation. To support evaluation, we introduce \textbf{InsightNet}, an emotion-tuned baseline for MLLMs. Results demonstrate that InsightVQA poses significant challenges for grounded emotion understanding and reasoning.

cs.CV

Nested Spatio-Temporal Time Series Forecasting

Spatiotemporal forecasting is critical for real-world applications like traffic management, yet capturing reliable interactions remains challenging under noisy and non-stationary conditions. Existing methods primarily rely on historical spatial priors, often failing to account for evolving temporal correlations and suffering from systematic errors. In this work, we propose a nested forecasting framework that couples future macro-level regional trends with micro-level historical observations, enabling top-down guidance from abstract future representations for fine-grained forecasting. Specifically, we employ a spectral clustering-based approach to construct semantically coherent regions, providing both theoretical and empirical evidence that this representation effectively filters systematic noise while preserving essential trends. Building on this, we develop a progressive coarse-to-fine predictor to integrate these representative features into the inference process. This enables the model to leverage trend predictions to anticipate dynamic anomalies, such as periodic offsets, in advance. Furthermore, extensive experiments on multiple high-dimensional datasets demonstrate that our method consistently outperforms state-of-the-art baselines, validating the effectiveness of future macro-guided nested forecasting.

cs.LG

UFT: Unifying Fine-Tuning of SFT and RLHF/DPO/UNA through a Generalized Implicit Reward Function

By pretraining on trillions of tokens, an LLM gains the capability of text generation. However, to enhance its utility and reduce potential harm, SFT and alignment are applied sequentially to the pretrained model. Because SFT and alignment have different objectives and underlying processes, performance on certain tasks can decline. To address this, we seamlessly introduce Unified Fine-Tuning (UFT), which integrates SFT and alignment into a single training stage using the same objective and loss functions through an implicit reward function. Our experimental results demonstrate that UFT outperforms SFT on instruction-tuning data alone. Moreover, when combining instruction-tuning data with alignment data, UFT effectively prevents the degradation on some tasks across these two stages and shows a clear advantage over sequentially applying SFT and alignment. This is evident in the significant improvements observed in the \textbf{ifeval} task for instruction-following and the \textbf{truthful} task for factuality. The proposed general fine-tuning framework UFT establishes an effective and efficient paradigm for LLM post-training.

cs.CL

Exploring Accuracy Law for Deep Time Series Forecasters: An Empirical Study

Deep time series forecasting has emerged as a rapidly growing field in recent years. Despite the exponential growth of community interests, progress on standard benchmarks is often limited to marginal improvements. A common consensus of the community is that time series forecasting inherently faces a non-zero error lower bound due to its partially observable and uncertain nature. However, a fundamental question arises: how to estimate the performance upper bound of deep time series forecasters? We delve into univariate time series forecasting, a prevalent forecasting paradigm spanning traditional statistical models to advanced time series foundation models. Going beyond classical series-wise predictability metrics, we realize that the forecasting performance is highly related to window-wise properties due to the sequence-to-sequence forecasting paradigm of deep time series models and introduce a quantitative measurement of window-wise pattern complexity. Through rigorous statistical analyses over more than 4700 newly trained deep forecasting models, we discover a consistent empirical relationship between the minimum attainable forecasting error of deep models and the complexity of window-wise series patterns, which is termed the accuracy law. We further demonstrate that this empirical finding successfully guides us to identify saturated tasks from widely used benchmarks and derive an effective training strategy for time series foundation models, offering valuable insights for future research.

cs.LG

STReasoner: Empowering LLMs for Spatio-Temporal Reasoning in Time Series via Spatial-Aware Reinforcement Learning

Spatio-temporal reasoning in time series involves the explicit synthesis of temporal dynamics, spatial dependencies, and textual context. This capability is vital for high-stakes decision-making in systems such as traffic networks, power grids, and disease propagation. However, the field remains underdeveloped because most existing works prioritize predictive accuracy over reasoning. To address the gap, we introduce ST-Bench, a benchmark consisting of four core tasks, including etiological reasoning, entity identification, correlation reasoning, and in-context forecasting, developed via a network SDE-based multi-agent data synthesis pipeline. We then propose STReasoner, which empowers LLM to integrate time series, graph structure, and text for explicit reasoning. To promote spatially grounded logic, we introduce S-GRPO, a reinforcement learning algorithm that rewards performance gains specifically attributable to spatial information. Experiments show that STReasoner achieves average accuracy gains between 17% and 135% at only 0.004X the cost of proprietary models and generalizes robustly to real-world data.

cs.CL

Wave--particle transition and quantum Zeno effect in which-way experiments with a superconducting quantum processor

Wave--particle duality demonstrates the peculiar nature of quantum mechanics. In which-way experiments, depending on the measurement scheme, a particle exhibits either wave-like or particle-like properties, as summarized by Bohr's principle of complementarity. In this work, we implement Mach-Zehnder (MZ) interferometry on a two-dimensional (2D) superconducting quantum processor. With precise control of the which-way measurement strength, we demonstrate the transition of a photon from wave-like to particle-like behavior. Furthermore, by performing quantum state tomography on two qubits located in the two paths, we demonstrate that which-way measurements break the entanglement and coherence between the two paths and cause information leakage from the quantum system to the environment. To capture this behavior quantitatively, we derive complementarity relations between the entropy and the fringe visibility. By applying a continuous which-way measurement during the evolution, we also observe the quantum Zeno effect that partially obstructs the interferometer path, giving rise to nonmonotonic behavior of purity and von Neumann entropy. Our experiments provide a detailed characterization of the full interferometer dynamics, reveal the relation between wave--particle duality and quantum information, and demonstrate the potential of superconducting quantum processors for testing quantum foundations under high precision and controllability.

quant-ph

Webscale-RL: Automated Data Pipeline for Scaling RL Data to Pretraining Levels

Large Language Models (LLMs) have achieved remarkable success through imitation learning on vast text corpora, but this paradigm creates a training-generation gap and limits robust reasoning. Reinforcement learning (RL) offers a more data-efficient solution capable of bridging this gap, yet its application has been constrained by a critical data bottleneck: existing RL datasets are orders of magnitude smaller and less diverse than web-scale pre-training corpora. To address this, we introduce the Webscale-RL pipeline, a scalable data engine that systematically converts large-scale pre-training documents into millions of diverse, verifiable question-answer pairs for RL. Using this pipeline, we construct the Webscale-RL dataset, containing 1.2 million examples across more than 9 domains. Our experiments show that the model trained on this dataset significantly outperforms continual pretraining and strong data refinement baselines across a suite of benchmarks. Notably, RL training with our dataset proves substantially more efficient, achieving the performance of continual pre-training with up to 100$\times$ fewer tokens. Our work presents a viable path toward scaling RL to pre-training levels, enabling more capable and efficient language models.

cs.CL

Timer-S1: A Billion-Scale Time Series Foundation Model with Serial Scaling

We introduce Timer-S1, a strong Mixture-of-Experts (MoE) time series foundation model with 8.3B total parameters, 0.75B activated parameters for each token, and a context length of 11.5K. To overcome the scalability bottleneck in existing pre-trained time series foundation models, we perform Serial Scaling in three dimensions: model architecture, dataset, and training pipeline. Timer-S1 integrates sparse TimeMoE blocks and generic TimeSTP blocks for Serial-Token Prediction (STP), a generic training objective that adheres to the serial nature of forecasting. The proposed paradigm introduces serial computations to improve long-term predictions while avoiding costly rolling-style inference and pronounced error accumulation in the standard next-token prediction. Pursuing a high-quality and unbiased training dataset, we curate TimeBench, a corpus with one trillion time points, and apply meticulous data augmentation to mitigate predictive bias. We further pioneer a post-training stage, including continued pre-training and long-context extension, to enhance short-term and long-context performance. Evaluated on the large-scale GIFT-Eval leaderboard, Timer-S1 achieves state-of-the-art forecasting performance, attaining the best MASE and CRPS scores as a pre-trained model. Timer-S1 is released to facilitate further research.

cs.AI

Ultrafast Non-Volatile Weyl LuminoMem for Mid-Infrared In-Memory Computing

Integrated optoelectronic systems strive to combine the logic/memory density of electronics with the bandwidth of photonics, but monolithic realization is impeded by the inefficient electronic-to-photonic interface. Current architectures rely on separate readout circuitry and modulators, creating bottlenecks in energy and latency, while existing direct transduction methods often compromise on switching speed or non-volatility. Here, we report an ultrafast, non-volatile optoelectronic memory, named LuminoMem, that integrates electrical storage and mid-infrared light emission in a single device. The device utilizes a floating-gate architecture, in which the Weyl semiconductor tellurium serves simultaneously as a charge-trapping storage layer and an emissive medium. This design enables nanosecond-scale electrical programming of non-volatile photoluminescence at 3.4 um, allowing direct optical access to stored states without external modulation. We demonstrate 4-bit (16-level) optical storage capacity and validate the device's performance through neural network simulations that achieve high accuracy on the Fashion-MNIST dataset. By effectively bridging the gap between electronic storage and mid-infrared photonics, the demonstrated mid-infrared LuminoMem provides a hardware foundation for promoting current computation efficiency and potential intelligent platforms that co-integrate computing, memory, and sensing capabilities.

cond-mat.mtrl-sci

Disentangled World Models: Learning to Transfer Semantic Knowledge from Distracting Videos for Reinforcement Learning

Training visual reinforcement learning (RL) in practical scenarios presents a significant challenge, $\textit{i.e.,}$ RL agents suffer from low sample efficiency in environments with variations. While various approaches have attempted to alleviate this issue by disentangled representation learning, these methods usually start learning from scratch without prior knowledge of the world. This paper, in contrast, tries to learn and understand underlying semantic variations from distracting videos via offline-to-online latent distillation and flexible disentanglement constraints. To enable effective cross-domain semantic knowledge transfer, we introduce an interpretable model-based RL framework, dubbed Disentangled World Models (DisWM). Specifically, we pretrain the action-free video prediction model offline with disentanglement regularization to extract semantic knowledge from distracting videos. The disentanglement capability of the pretrained model is then transferred to the world model through latent distillation. For finetuning in the online environment, we exploit the knowledge from the pretrained model and introduce a disentanglement constraint to the world model. During the adaptation phase, the incorporation of actions and rewards from online environment interactions enriches the diversity of the data, which in turn strengthens the disentangled representation learning. Experimental results validate the superiority of our approach on various benchmarks.

cs.CV

Gate-Tunable Mid-Infrared Electroluminescence from Te/MoS2 p-n Heterojunctions

Mid-infrared (MIR) emitters are critical components in advanced photonic systems, driving progress in fields such as chemical sensing, environmental monitoring, medical diagnostics, thermal imaging and free-space communications. Conventional MIR emitters based on III-V heterostructures rely on complex epitaxial growth on rigid lattice-matched substrates and suffer from limited integration compatibility with CMOS or flexible platforms. The recent development of novel MIR emitters based on two-dimensional (2D) materials such as black phosphorus (BP) is more suitable for on-chip applications but faces challenges related to stability and emission efficiency. Based on the recently discovered highly efficient photoluminescence of Te, we demonstrate a gate-tunable midinfrared light-emitting diode based on a van der Waals heterojunction formed by multilayer transition metal dichalcogenide (TMD) MoS2 and tellurium (Te). The device emits polarized electroluminescence (EL) centered at 3.5 $μ$m under forward bias at 25 K, and the EL persists up to 80 K with reduced intensity. Gate control of the MoS2 Fermi level modulates the band alignment and injection efficiency, enabling dynamic tuning of the EL intensity. The emission remains spectrally stable under varying bias and gating, indicating robust band-edge recombination. These results establish the Te/TMD heterostructure as a promising platform for integrated polarized mid-infrared optoelectronics.

cond-mat.mes-hall

Position: Vector Prompt Interfaces Should Be Exposed to Enable Customization of Large Language Models

As large language models (LLMs) transition from research prototypes to real-world systems, customization has emerged as a central bottleneck. While text prompts can already customize LLM behavior, we argue that text-only prompting does not constitute a suitable control interface for scalable, stable, and inference-only customization. This position paper argues that model providers should expose \emph{vector prompt inputs} as part of the public interface for customizing LLMs. We support this position with diagnostic evidence showing that vector prompt tuning continues to improve with increasing supervision whereas text-based prompt optimization saturates early, and that vector prompts exhibit dense, global attention patterns indicative of a distinct control mechanism. We further discuss why inference-only customization is increasingly important under realistic deployment constraints, and why exposing vector prompts need not fundamentally increase model leakage risk under a standard black-box threat model. We conclude with a call to action for the community to rethink prompt interfaces as a core component of LLM customization.

cs.CL

Select, then Balance: Exploring Exogenous Variable Modeling of Spatio-Temporal Forecasting

Spatio-temporal (ST) forecasting is critical for dynamic systems, yet existing methods predominantly rely on modeling a limited set of observed target variables. In this paper, we present the first systematic exploration of exogenous variable modeling for ST forecasting, a topic long overlooked in this field. We identify two core challenges in integrating exogenous variables: the inconsistent effects of distinct variables on the target system and the imbalance effects between historical and future data. To address these, we propose ExoST, a simple yet effective exogenous variable modeling general framework highly compatible with existing ST backbones that follows a "select, then balance" paradigm. Specifically, we design a latent space gated expert module to dynamically select and recompose salient signals from fused exogenous information. Furthermore, a siamese dual-branch backbone architecture captures dynamic patterns from the recomposed past and future representations, integrating them via a context-aware weighting mechanism to ensure dynamic balance. Extensive experiments on real-world datasets demonstrate the ExoST's effectiveness, universality, robustness, and efficiency.

cs.LG