arXiv ScienceSearch

arXiv subjects

Yuchen Zhang

Publications and source records attributed to Yuchen Zhang.

At least 19 recordsLinked to original sources

Reconfigurable Antennas for Next-Generation Wireless Communications: Technologies, Prototypes, Architectures, and Signal Processing

The transition to sixth-generation (6G) networks calls for wireless transceivers with enhanced adaptability and efficiency. Reconfigurable antennas (RAs) have emerged as a promising solution, enabling dynamic control over the electromagnetic properties of individual antenna elements. Their integration into antenna arrays is particularly attractive due to their adaptability and potential for improved energy efficiency. This article provides a comprehensive overview of RA technologies for advanced communication systems, encompassing hardware advancements, early experimental studies, novel system architectures, and key signal processing challenges. From the combined perspective of antenna design and communication operation, we highlight the potential of RAs to enable next-generation wireless communications, while also identifying key challenges and promising research opportunities.

eess.SP

PlannerForge: LLM Agents for Scenario-Based Testing of Motion Planners in Autonomous Driving

Ensuring the safety of autonomous driving is a critical challenge. Scenario-based testing is a systematic process used to validate Autonomous Driving Systems (ADSs), but it remains a fragmented modular pipeline in which scenario generation, retrieval, modification, ADS execution, and results analysis are performed by separate tools with little interaction. Large Language Model (LLM) agents have shown promise across ADS sub-systems such as perception, planning, and control. However, no prior work covers the whole scenario-based testing pipeline for ADSs with a unified LLM-agent framework. We present PlannerForge, an LLM-agent framework that extends all scenario-based testing stages (from Scenario Generation to ADS Assessment) and adds two further LLM-enhanced stages: ADS Enhancement and ADS Benchmarking. We evaluate PlannerForge with 10 off-the-shelf LLMs across all tasks (Generation, Selection, Modification, Module Routing, Planner Testing, and Enhancement) under 5 prompt conditions. Best-per-task scores range from 0.88 to 1.00, and open-source 20-35B backends match commercial APIs on most tasks. Open-source models such as Qwen3.6:35B match commercial APIs on three of the five tasks. Chaining the modules end-to-end retains 83% / 78% of seed queries (commercial / open). It outperforms Scenario Factory 2.0 (Finkeldei et al., 2025) on natural-language generation (193 vs. 144 executable of 200) and realises 92-96% of requested city, road and vehicle attributes. It outperforms BM25 (Robertson and Zaragoza, 2009) at rank 1 selection (92.0% vs. 67.5%) and From-Words-to-Collisions (Gao et al., 2025) on physically valid edits (>=94% vs. 31%). At N=400, cost-tuning lifts planner success from 50.4% to 70.2% and cuts collisions from 19.0% to 8.4%, without domain-specific fine-tuning.

cs.AI

Integrated Positioning and Communication via LEO Satellites: A Contemporary Overview

Low Earth orbit (LEO) satellites, as a prominent technology in the sixth generation non-terrestrial network, offer both positioning and communication capabilities. While these two applications have each been extensively studied and have achieved substantial progress in recent years, the potential synergistic benefits of integrating them remain an underexplored yet promising avenue. This article provides a contemporary overview of integrated positioning and communication (IPAC) systems enabled by LEO satellites. Leveraging the distinct characteristics of LEO satellites, we examine how communication systems can enhance positioning accuracy and, conversely, how positioning information can be exploited to improve communication efficiency. A representative case study is included to illustrate enhanced LEO positioning signal acquisition with the assistance of communication-side information. Finally, several key open research challenges for LEO-based IPAC systems are discussed.

eess.SP

Orbital Errors in LEO Satellite Positioning: Modeling, Analysis, and Bayesian Calibration

Low-Earth orbit (LEO) satellites offer a promising alternative to global navigation satellite systems for precise positioning. However, their relatively low altitudes make them more susceptible to orbital perturbations, which in turn degrade positioning accuracy. In this work, we study (i) the mechanisms through which orbital errors affect positioning performance and (ii) the fundamental calibration strategies for mitigating low-Earth orbital errors. We first model the satellite orbit and analyze historical two-line element data to characterize the statistical behavior of low-Earth orbital errors. Based on these real-data statistics, we derive the misspecified Cramér-Rao bound to quantify the impact of orbital errors on positioning performance. Subsequently, we develop a Bayesian orbit calibration framework that leverages the learned orbital error statistics as prior information to enhance calibration accuracy. This work sheds light on orbital error modeling, real-data-based evaluation, theoretical impact analysis, and calibration methodologies. Extensive simulations show that stale orbital information can severely degrade positioning performance, whereas learning and incorporating orbital error statistics can effectively enhance orbit calibration and ultimately improve positioning accuracy.

eess.SP

VERPO: Verified Evidence Regularized Policy Optimization

Verifiable outcome rewards guide language-model post-training, but sequence-level advantages do not identify which token-level decisions should be preserved or revised. Evidence-conditioned Teachers provide denser supervision by replaying sampled trajectories with privileged feedback. Yet indiscriminate imitation risks transferring formatting or reasoning-style shifts that do not support task success. We introduce VERPO, a Verified Evidence Regularized Policy Optimization framework that treats evidence as a proposal for policy correction while retaining the outcome objective. It separates evidence-free reference restoration from signed token-level evidence corrections. Fisher Evidence Contrast attenuates corrections along an estimated evidence-presence direction. A stopped token-wise ZPD controller scales acceptance according to local reward alignment and Fisher movement cost, while the reference channel remains independent of acceptance. Across five scientific-reasoning and tool-use tasks, the best variant on each backbone exceeds the strongest compared baseline in average score. The averages rise from 0.6826 to 0.6857 on Qwen3-4B, from 0.6895 to 0.7058 on Qwen3-8B, and from 0.4751 to 0.5657 on Llama-3.2-1B.

cs.LG

Quantization-Aware EE Optimization and SE-EE Tradeoff for MiLAC-Aided MU-MISO Beamforming

In large antenna arrays, hardware power consumption becomes a dominant design constraint, making energy efficiency (EE) a first-class objective alongside spectral efficiency (SE). Microwave linear analog computer (MiLAC)-aided beamforming, whose front end is a passive reciprocal stream-to-antenna network, addresses this tension by reducing the active radio-frequency chain count to the stream number, at a moderate SE cost. Despite this promise, no EE optimization framework has been established for MiLAC-aided beamforming that accounts for digital-to-analog converter quantization noise and post-quantized transmit power. We fill this gap for downlink multiuser multiple-input single-output (MU-MISO) systems by formulating quantization-aware EE maximization over the MiLAC-feasible beamformer and characterizing the resulting SE-EE tradeoff. Three contributions follow. First, we prove a row-space optimality property of the effective MiLAC-aided beamformer, yielding an equivalent reduced-dimension reformulation whose complexity scales with the stream number rather than the antenna number. Second, we develop a low-complexity Dinkelbach-weighted minimum mean-square error algorithm aided by projected gradient descent that is guaranteed to converge to a stationary point. Third, we cast the SE-EE tradeoff as a multi-objective problem and trace its Pareto boundary via a weighted-sum method that combines an alternative reduced-dimension coordinate with auxiliary-variable successive convex approximation, yielding convex per-iteration subproblems with guaranteed convergence. Numerical results on a DeepMIMO v4 deployment show MiLAC-aided beamforming substantially improves EE over digital and hybrid benchmarks at a moderate SE cost and significantly expands the achievable SE-EE operating region.

eess.SP

SUPER ODOMETRY 2.0: Resilient Odometry via Hierarchical Adaptation

Resilient and robust odometry is crucial for autonomous systems operating in complex and dynamic environments. Existing odometry systems often struggle with severe sensory degradations and extreme conditions such as smoke, sandstorms, snow, or low-light conditions, threatening both the safety and functionality of robots. To address these challenges, we present Super Odometry, a sensor fusion framework that dynamically adapts to varying levels of environmental degradation. Super Odometry employs a hierarchical structure to integrate four core modules from lower-level to higher-level adaptability including adaptive feature selection, adaptive state direction selection, adaptive engine selection, and a novel learning- based inertial odometry. The inertial odometry, trained on over 100 hours of heterogeneous robotic platforms, captures comprehensive motion dynamics. Super Odometry elevates the inertial measurement unit (IMU) to equal importance with camera and LiDAR within the sensor fusion framework, providing a reliable fallback when exteroceptive sensors fail. Super Odometry has been validated across 200 kilometers and 800 operational hours on a fleet of aerial, wheeled, and legged robots, under diverse sensor configurations, environmental degradation, and aggressive motion profiles. It marks an important step towards safe and long-term robotic autonomy in all-degraded environments.

cs.RO

Constraining the Baryon Fraction in Extragalactic Diffuse Ionized Gas with 124 Localized Fast Radio Bursts

Fast radio bursts (FRBs) are increasingly recognized as powerful cosmological tools for constraining the baryon fraction in extragalactic diffuse ionized gas, presenting a promising approach to address the missing baryon problem. In this paper, we constrain the baryon fraction in extragalactic diffuse ionized gas ($f_\mathrm{d}$) utilizing the latest sample of 124 localized FRBs across three different cosmological models. Our analysis models the probability distribution of the extragalactic diffuse ionized gas dispersion measure with a form that accurately reproduces mock observations. For a constant $f_\mathrm{d}$ model, we find that more than 90\% of baryons reside in the diffuse ionized gas phase. This result is robust against the choice of dark-energy parametrization under the current combination of datasets, although the fitted cosmological parameters shift accordingly. We also find that the inferred $f_\mathrm{d}$ is sensitive to the assumed dispersion measure distributions of both the Milky Way halo and the FRB host galaxies. Furthermore, the current data do not show statistically significant evidence for redshift evolution in $f_\mathrm{d}$, but the constraints are limited by the redshift distribution of the sample. Our conclusions are insensitive to the adopted baryonic feedback parameters and to the dispersion measure selection effect. These results provide strong evidence that the majority of the missing baryons reside in the diffuse ionized intergalactic medium.

astro-ph.CO

Tri-Hybrid Multi-User Precoding Using Pattern-Reconfigurable Antennas: Fundamental Models and Practical Algorithms

The integration of pattern-reconfigurable antennas into hybrid multiple-input multiple-output (MIMO) architectures presents a promising path toward high-efficiency and low-cost transceiver solutions. Pattern-reconfigurable antennas can dynamically steer per-antenna radiation patterns, enabling more efficient power utilization and interference suppression. In this work, we study a tri-hybrid MIMO architecture for multi-user communications that integrates digital, analog, and antenna-domain precoding using pattern-reconfigurable antennas. For characterizing the reconfigurability of antenna radiation patterns, we develop two models---Model~I and Model~II. Model~I captures realistic hardware constraints through limited pattern selection, while Model~II explores the performance upper bound by assuming arbitrary pattern generation. Based on these models, we develop two corresponding tri-hybrid precoding algorithms grounded in the weighted minimum mean square error (WMMSE) framework, which alternately optimize the digital, analog, and antenna precoders under practical per-antenna power constraints. Realistic simulations conducted in ray-tracing generated environments are utilized to evaluate the proposed system and algorithms. The results demonstrate the significant potential of the considered tri-hybrid architecture in enhancing communication performance and hardware efficiency. However, they also reveal that the existing hardware is not yet capable of fully realizing these performance gains, underscoring the need for joint progress in antenna design and communication theory development.

eess.SP

CLARA: Clip-Level Multimodal Alignment with VLM-Derived Rationales for Hateful Video Detection

Hateful video detection has become increasingly important with the rapid growth of video-centric social media platforms, given the serious risks that hate speech poses to both individual well-being and social cohesion. Compared with text or static multimodal content, hateful video detection remains underexplored and significantly more challenging, as hateful meaning often arises from complex interactions among multimodal cues, including speech, audio, and visual content. Moreover, such signals are often brief, implicit, and temporally dependent, making them difficult to capture using conventional video-level representations. In this work, we propose CLARA, a clip-level multimodal framework for hateful video detection. Instead of treating a video as a single instance, CLARA models it as a sequence of fine-grained clips, enabling more precise capture of temporally localized hateful signals. We introduce a Mixture-of-Experts clip encoder for adaptive multimodal alignment, a local-global segment contrastive objective to jointly model short-term cues and long-range temporal dependencies, and VLM-derived rationales integrated via a gated Transformer to provide high-level semantic guidance. Extensive experiments on three hateful video datasets demonstrate that CLARA consistently outperforms state-of-the-art methods. Further ablation studies and parameter analyses validate the effectiveness of each component.

cs.CV

Language Family Matters: Evaluating LLM-Based ASR Across Linguistic Boundaries

Large Language Model (LLM)-powered Automatic Speech Recognition (ASR) systems achieve strong performance with limited resources by linking a frozen speech encoder to a pretrained LLM via a lightweight connector. Prior work trains a separate connector per language, overlooking linguistic relatedness. We propose an efficient and novel connector-sharing strategy based on linguistic family membership, enabling one connector per family, and empirically validate its effectiveness across two multilingual LLMs and two real-world corpora spanning curated and crowd-sourced speech. Our results show that family-based connectors reduce parameter count while improving generalization across domains, offering a practical and scalable strategy for multilingual ASR deployment.

cs.CL

Speak in Context: Multilingual ASR with Speech Context Alignment via Contrastive Learning

Automatic speech recognition (ASR) has benefited from advances in pretrained speech and language models, yet most systems remain constrained to monolingual settings and short, isolated utterances. While recent efforts in context-aware ASR show promise, two key challenges persist: limited multilingual support and the absence of principled alignment between speech and contextual representations. In this paper, we introduce a context-aware multilingual ASR framework that supports diverse languages and accents while preserving the modularity of pretrained models. Our approach combines a frozen speech encoder and a decoder-only language model via a lightweight projection module, allowing structured context prompts, including dialogue history and biasing words, to guide transcription. To improve interaction between speech and context, we employ a contrastive learning objective that aligns their representations in a shared embedding space. Evaluations on over 1,500 hours of real-world conversational speech across 11 languages and 5 English dialects show that contextual input consistently improves recognition quality. Contrastive alignment provides additional gains when applied to different context types, with an overall performance gain of over 5%. These results highlight the importance of both contextual modeling and cross-modal alignment in multilingual ASR.

cs.CL

ReMP: Low-Downtime Runtime Model-Parallelism Reconfiguration for LLM Serving

Current large language model (LLM) inference systems universally deploy ultra-large-scale models using a combination of Tensor Parallelism (TP) and Pipeline Parallelism (PP). However, existing systems treat the model parallelism topology as a static configuration that cannot be flexibly adjusted at runtime. This rigid design creates a fundamental contradiction with the dynamically changing inference workloads in real-world scenarios. State-of-the-art systems lack online reconfiguration capabilities and can only switch configurations by restarting the service, resulting in several minutes of service interruption, KV cache loss, and prohibitive recomputation overhead. To address this problem, this paper presents ReMP, a runtime model parallelism reconfiguration framework that supports low downtime. ReMP achieves dynamic adjustment through three key techniques: (1) decoupling the model parallelism topology from runtime state to avoid full service reconstruction; (2) designing a two-dimensional KV cache migration mechanism to preserve reusable cache states after TP/PP changes; and (3) implementing end-to-end online reconfiguration. Experiments demonstrate that ReMP can complete most topology switches within 1-7 seconds on models ranging from 7B to 70B parameters, achieving speedups of tens to over a hundred times compared to the restart approach. Moreover, ReMP significantly outperforms fixed configurations under dynamic workloads, delivering superior performance in terms of TTFT, TPOT, and output throughput.

cs.DC

VOLA: Improving Open-World Driving by VLM-Based Semantic Attribute Prediction

Driving in the real world is open-world: a car may encounter a fallen mattress, a deer, or other objects outside its training data. Naming them is not enough. The system must know how to treat each region: can it drive over it, and how severe would a collision be? We therefore shift scene perception from category labels to dense action-relevant attributes, where each pixel is labeled by how it should affect motion rather than by object name. We instantiate this general formulation with two ordered attributes: 7-rank drivability and 5-rank vulnerability. We read Qwen3.5 image-token hidden states directly as a spatial semantic representation. A lightweight boundary-aware decoder then turns this coarse token grid into sharp full-resolution attribute maps. The whole process requires neither autoregressive text generation nor an external mask model such as SAM. We train on dense attribute labels built in CARLA and test transfer to real scenes and to novel obstacles never seen in training. We compare with vision-only segmenters trained on the same attributes and prompted VLM segmenters. Our model matches strong vision-only segmenters on familiar categories and improves transfer to real open-world anomalies, reaching 69.4% mean vulnerability-rank recall versus 57.1% for the best vision-only baseline and 53.9% for the best prompted VLM baseline. These results show that VLM image tokens provide useful semantic cues for transferring driving attributes to objects outside the training vocabulary.

cs.CV

Observational constraints on fractional holographic dark energy in the light of DESI DR2

Based on the fractional entropy from fractional quantum mechanics, fractional holographic dark energy (FHDE) has been proposed with the Hubble horizon as the IR cutoff (FHDEH). We extend this framework by adopting the future event horizon and the particle horizon as the IR cutoff, proposing the FHDEF and FHDEP models. Using the SN+OHD+DESI DR2 dataset to constrain these models, we find that all three models provide a marginally lower $χ^{2}_{min}$ compared to $Λ$CDM but without significant preference according to AIC and BIC. When CMB distance priors are included, the FHDEH and FHDEP models are strongly ruled out. We further analyze the cosmological evolution for these models, and find that only the FHDEF model predicts nearly identical evolutions of $Ω_{m}$ and $Ω_{de}$ to those of the $Λ$CDM model across cosmic history, but its deceleration parameter $q$ deviate from the $Λ$CDM model in the future, indicating richer late time dynamics beyond the standard $Λ$CDM cosmology.

gr-qc

AirFlow: Context Preserving and Multi-Rate State Modeling for Air Quality Forecasting

Accurate air quality forecasting is essential for public health and urban environmental management, but remains challenging because pollutant channels differ in periodicity and distribution drift, while their concentration trajectories contain both multi-scale dependencies and rapid changes. Recent methods have improved spatial dependency learning and meteorological covariate modeling. However, pollutant channels are still passed through the same normalization rule and temporal backbone, using a shared latent representation for channel-specific distributions and changes at different rates. To address this limitation, we propose AirFlow, a pollutant-aware dual-stream framework that operates on station multivariate observations without additional graph propagation or predefined signal decomposition. Specifically, AirFlow designs two novel blocks: (1) a statistic-guided normalization routing mechanism that selects a normalization path for each pollutant according to its 24-hour autocorrelation and distribution drift; and (2) a hierarchical dual-stream state model that combines multi-scale state space propagation with learnable response coefficients, where gated bidirectional cross-attention exchanges information and adaptively fuses the resulting representations. Experiments on real-world data from multiple cities show that AirFlow achieves the best performance in 34 of 36 metrics comparisons, with reductions of up to 11.11% root mean square error over the state-of-the-art baseline. AirFlow also requires only 0.0483M parameters and 0.0215G FLOPs, achieving high forecasting accuracy with low computational overhead.

cs.AI

Memorization Dynamics in Knowledge Distillation for Language Models

Knowledge Distillation (KD) is increasingly adopted to transfer capabilities from large language models to smaller ones, offering significant improvements in efficiency and utility while often surpassing standard fine-tuning. Beyond performance, KD is also explored as a privacy-preserving mechanism to mitigate the risk of training data leakage. While training data memorization has been extensively studied in standard pre-training and fine-tuning settings, its dynamics in a knowledge distillation setup remain poorly understood. In this work, we study memorization across the KD pipeline using three large language model (LLM) families (Pythia, OLMo-2, Qwen-3) and three datasets (FineWeb, Wikitext, Nemotron-CC-v2). We find: (1) distilled models memorize significantly less training data than standard fine-tuning (reducing memorization by more than 50%); (2) some examples are inherently easier to memorize and account for a large fraction of memorization during distillation (over ~95%); (3) student memorization is predictable prior to distillation using features based on zlib entropy, KL divergence, and perplexity; and (4) while soft and hard distillation have similar overall memorization rates, hard distillation poses a greater risk: it inherits $2.7\times$ more teacher-specific examples than soft distillation. Overall, we demonstrate that distillation can provide both improved generalization and reduced memorization risks compared to standard fine-tuning.

cs.CL

Science Edge Evaluation: SEE the Missing Step Toward Real Scientific Discovery

Large language models (LLMs) are increasingly involved in scientific discovery, yet it remains unclear whether they can support complex real laboratory science. Here we introduce Science Edge Evaluation (SEE), a multimodal benchmark of expert-curated questions grounded in peer-reviewed literature and experimental practice in chemistry, biology, and materials science. Evaluation of 19 multimodal large language models (MLLMs) shows that even the best-performing model reaches only 48.7% accuracy. Moreover, general-purpose models outperform science-specialized models on average. In the visual-agent evaluation, the use of tools increases the best accuracy to 52.7%. Tool use can expand the information available to models, but more information does not necessarily lead to reliable scientific reasoning. The key challenge is whether models can manage tool-derived information within the boundaries of the original experimental evidence. Together, these findings reveal that current MLLMs still cannot reliably make justified and evidence-bounded inferences from experimental results, which is an essential capability in real scientific discovery. Bridging this gap requires MLLMs to transition from explaining established scientific concepts to deriving novel and evidence-based insights from experimental data.

cs.AI