arXiv ScienceSearch

arXiv subjects

Xiang Guo

Publications and source records attributed to Xiang Guo.

At least 19 recordsLinked to original sources

Weakly Driven and Finite Detuning Boundary Time Crystals Enabled by Low-Dissipation Dynamical Channels

Spontaneous breaking of continuous time-translation symmetry in driven-dissipative systems gives rise to boundary time crystals (BTCs), characterized by persistent oscillations sustained by coherent driving and collective dissipation. Conventional BTCs, however, typically require strong driving and exact atom-drive resonance, imposing stringent constraints on their realization. Here we consider two atomic ensembles coupled to a common Markovian reservoir and show that shared dissipation organizes dissipation-free and low-dissipation modes into dynamically accessible low-dissipation channels, enabling BTCs under weak driving and finite detuning. Finite detuning further selects a unique stable limit cycle from an initial-state-dependent family of oscillatory trajectories. Our results establish low-dissipation dynamical channels as a route to robust BTCs under relaxed driving and resonance conditions.

quant-ph

Kimi K3: Open Frontier Intelligence

We introduce Kimi K3, a 2.8T parameter Mixture-of-Experts model with 104 billion activated parameters, native vision capabilities, and a 1-million-token context window. Kimi K3 is built on Kimi Delta Attention and Attention Residuals, which improve information flow across sequence length and model depth. Together with Stable LatentMoE, which effectively activates 16 of 896 routed experts per token, and refined training and data recipes, these advances yield an approximately 2.5x improvement in overall scaling efficiency over Kimi K2. Post-training highlights reinforcement learning across general, agentic, and coding domains and multiple reasoning-effort levels, enabling compositional generalization and robust long-horizon execution. At 2.8T scale, Kimi K3 is supported by infrastructure advances in multiple areas: algorithm-system co-design for KDA, perfectly balanced expert-parallel training with efficient memory management, million-token agentic RL with persistent rollout and sandbox states, and deployment innovations. Extensive evaluations show that Kimi K3 achieves frontier-level performance across long-horizon coding, agentic, knowledge, reasoning, and vision tasks. While its overall performance still trails the most powerful proprietary models, namely Claude Fable 5 and GPT-5.6 Sol, Kimi K3 consistently outperforms other open and proprietary models evaluated in our suite. We release the full Kimi K3 model weights to facilitate future research and accelerate the broader deployment and adoption of frontier intelligence.

cs.CL

Answer First, Reason Later: Aligning Search Relevance via Mode-Balanced Reinforcement Learning

Building a search relevance model that achieves both low latency and high performance is a long-standing challenge in the search industry. To satisfy the millisecond-level response requirements of online systems while retaining the interpretable reasoning traces of Large Language Models (LLMs), we propose a novel \textbf{Answer-First, Reason Later (AFRL)} paradigm. This paradigm requires the model to output the definitive relevance score in the very first token, followed by a structured logical explanation. Inspired by the success of reasoning models, we adopt a "Supervised Fine-Tuning (SFT) + Reinforcement Learning (RL)" pipeline to achieve AFRL. However, directly applying existing RL training often leads to \textbf{mode collapse} in the search relevance task, where the model forgets complex long-tail rules in pursuit of high rewards. From an information theory perspective: RL inherently minimizes the \textbf{Reverse KL divergence}, which tends to seek probability peaks (mode-seeking) and is prone to "reward hacking." On the other hand, SFT minimizes the \textbf{Forward KL divergence}, forcing the model to cover the data distribution (mode-covering) and effectively anchoring expert rules. Based on this insight, we propose a \textbf{Mode-Balanced Optimization} strategy, incorporating an SFT auxiliary loss into Stepwise-GRPO training to balance these two properties. Furthermore, we construct an automated instruction evolution system and a multi-stage curriculum to ensure expert-level data quality. Extensive experiments demonstrate that our 32B teacher model achieves state-of-the-art performance. Moreover, the AFRL architecture enables efficient knowledge distillation, successfully transferring expert-level logic to a 0.6B model, thereby reconciling reasoning depth with deployment latency.

cs.LG

ETR: Outcome-Guided Elastic Trust Regions for Policy Optimization

Reinforcement Learning with Verifiable Rewards (RLVR) has emerged as an important paradigm for unlocking reasoning capabilities in large language models, exemplified by the success of OpenAI o1 and DeepSeek-R1. Currently, Group Relative Policy Optimization (GRPO) stands as the dominant algorithm in this domain due to its stable training and critic-free efficiency. However, we argue that GRPO suffers from a structural limitation: it imposes a uniform, static trust region constraint across all samples. This design implicitly assumes signal homogeneity, a premise misaligned with the heterogeneous nature of outcome-driven learning, where advantage magnitudes and variances fluctuate significantly. Consequently, static constraints fail to fully exploit high-quality signals while insufficiently suppressing noise, often precipitating rapid entropy collapse. To address this, we propose \textbf{E}lastic \textbf{T}rust \textbf{R}egions (\textbf{ETR}), a dynamic mechanism that aligns optimization constraints with signal quality. ETR constructs a signal-aware landscape through dual-level elasticity: at the micro level, it scales clipping boundaries based on advantage magnitude to accelerate learning from high-confidence paths; at the macro level, it leverages group variance to implicitly allocate larger update budgets to tasks in the optimal learning zone. Extensive experiments on AIME and MATH benchmarks demonstrate that ETR consistently outperforms GRPO, achieving superior accuracy while effectively mitigating policy entropy degradation to ensure sustained exploration.

cs.LG

Bound state in the continuum and multiple atom state transfer applications in a waveguide QED setup

Bound states in the continuum (BICs) have been extensively exploited to enhance light--matter interactions in metamaterials, yet their emergence and utility in multi-atom waveguide platforms remain far less explored. Here we study atom--waveguide-dressed BICs in a one-dimensional coupled-resonator waveguide, where two spatially separated atomic arrays couple to distinct resonators with time-dependent strengths. We show that these BICs support a standing-wave photonic mode and enable the transfer of an arbitrary unknown quantum state between the two arrays with fidelities exceeding $99\%$. The protocol remains robust against both disorder and intrinsic dissipation. Our results establish BICs as long-lived resources for high-fidelity quantum information processing in waveguide-QED architectures.

quant-ph

Quantum state preparation and transfer based on the bound state in the doublon continuum

Bound states in the continuum (BICs) have attracted intense interest, yet their many-particle counterparts remain largely unexplored in waveguide quantum electrodynamics. We identify and characterize a bound state embedded in the doublon continuum (BIDC) that emerges when four atoms couple to a coupled-resonator waveguide with strong on-site interaction. Exploiting this interaction-enabled BIDC, we show that (i) a distant, four-atom entangled state can be prepared with high fidelity, and (ii) quantum entangled states can be coherently transferred between spatially separated nodes. Our results establish a scalable mechanism for multi-particle state generation and routing in waveguide platforms, opening a route to interaction-protected quantum communication with many-particle BICs.

quant-ph

Enantiodetection in a cavity QED setup with finite chiral molecules

We investigate enantiodetection for both a single cyclic three-level chiral molecule and finite ensembles of such molecules by monitoring the steady-state intracavity photon number in a cavity-QED platform. Our scheme exploits the intrinsic global $\pi$-phase difference between opposite enantiomers to engineer destructive and/or constructive interference pathways, enabling a direct readout of enantiomeric excess with an error below $5\%$. To capture mesoscopic many-molecule effects beyond mean field while avoiding brute-force master-equation simulations, we employ a generalized discrete truncated Wigner approximation, which is well suited for systems with many yet finite molecules. These results pave the way for implementing enantiodetection in realistic quantum-optical settings.

quant-ph

Don't Just Search, Understand: Semantic Path Planning Agent for Spherical Tensegrity Robots in Unknown Environments

Endowed with inherent dynamical properties that grant them remarkable ruggedness and adaptability, spherical tensegrity robots stand as prototypical examples of hybrid softrigid designs and excellent mobile platforms. However, path planning for these robots in unknown environments presents a significant challenge, requiring a delicate balance between efficient exploration and robust planning. Traditional path planners, which treat the environment as a geometric grid, often suffer from redundant searches and are prone to failure in complex scenarios due to their lack of semantic understanding. To overcome these limitations, we reframe path planning in unknown environments as a semantic reasoning task. We introduce a Semantic Agent for Tensegrity robots (SATPlanner) driven by a Large Language Model (LLM). SATPlanner leverages high-level environmental comprehension to generate efficient and reliable planning strategies.At the core of SATPlanner is an Adaptive Observation Window mechanism, inspired by the "fast" and "slow" thinking paradigms of LLMs. This mechanism dynamically adjusts the perceptual field of the agent: it narrows for rapid traversal of open spaces and expands to reason about complex obstacle configurations. This allows the agent to construct a semantic belief of the environment, enabling the search space to grow only linearly with the path length (O(L)) while maintaining path quality. We extensively evaluate SATPlanner in 1,000 simulation trials, where it achieves a 100% success rate, outperforming other real-time planning algorithms. Critically, SATPlanner reduces the search space by 37.2% compared to the A* algorithm while achieving comparable, near-optimal path lengths. Finally, the practical feasibility of SATPlanner is validated on a physical spherical tensegrity robot prototype.

cs.RO

Engineering atomic superradiance scaling in cavity QED system with collective and individual emission channels

The coherent emission of multiple atoms gives rise to superradiance, a cornerstone phenomenon in quantum optics with wide-ranging applications in quantum information processing and precision metrology. Despite its importance, how the superradiant scaling with respect to the number of participating atoms can be effectively controlled remains largely unexplored. In this work, we investigate a cavity-QED system and demonstrate that atom-photon coupling can significantly alter the emission behavior--suppressing the collective superradiant scaling while enhancing the scaling associated with individual atomic emissions. Our study provides a pathway toward controllable collective emission in state-of-the-art experimental platforms.

quant-ph

CLPO: Curriculum Learning meets Policy Optimization for LLM Reasoning

Online reinforcement learning with verifiable rewards (RLVR) has become an effective paradigm for improving the reasoning abilities of large language models, but most methods still optimize reasoning trajectories over the static problem set, wasting rollout budget on solved or overly difficult problems. We propose \textbf{CLPO (Curriculum Learning meets Policy Optimization)}, a self-evolving curriculum framework that uses on-policy rollout accuracy to identify solved, medium-difficulty, and hard problems, then restructures selected tasks according to the model's current capability. Hard problems are simplified to become learnable, while medium-difficulty problems are diversified to provide useful training variation. This allows the learning curriculum to co-evolve with the policy rather than remaining fixed as the model's capability boundary shifts. Rather than treating these rewrites as static data augmentation, CLPO optimizes restructuring trajectories with credit assigned by the downstream accuracy gain of the rewritten problem, requiring no additional human annotations beyond the original verifiable answers. Experiments across mathematical reasoning and out-of-domain general reasoning benchmarks show that CLPO substantially outperforms GRPO and DAPO on Qwen3-8B by 10.21 and 7.75 average points, respectively. Ablation studies on math and code domains further show that both the restructuring mode and the rewriting loss contribute to the final gains, demonstrating that CLPO provides a scalable and robust pathway for eliciting stronger reasoning capabilities through a self-evolving curriculum.

cs.AI

Global Optimality in Multi-Flyby Asteroid Trajectory Optimization: Theory and Application Techniques

Designing optimal trajectories for multi-flyby asteroid missions is scientifically critical but technically challenging due to nonlinear dynamics, intermediate constraints, and numerous local optima. This paper establishes a method that approaches global optimality for multi-flyby trajectory optimization under a given sequence. The original optimal control problem with interior-point equality constraints is transformed into a multi-stage decision formulation. This reformulation enables direct application of dynamic programming in lower dimensions, and follows Bellman's principle of optimality. Moreover, the method provides a quantifiable bound on global optima errors introduced by discretization and approximation assumptions, thus ensuring a measure of confidence in the obtained solution. The method accommodates both impulsive and low-thrust maneuver schemes in rendezvous and flyby scenarios. Several computational techniques are introduced to enhance efficiency, including a specialized solution for bi-impulse cases and an adaptive step refinement strategy. The proposed method is validated through three problems: 1) an impulsive variant of the fourth Global Trajectory Optimization competition problem (GTOC4), 2) the GTOC11 problem, and 3) the original low-thrust GTOC4 problem. Each case demonstrates improvements in fuel consumption over the best-known trajectories. These results give evidence of the generality and effectiveness of the proposed method in global trajectory optimization.

math.OC

Giant emitter magnetometer

Leveraging the sensitive dependence of a giant atom's radiation rate on its frequency [A. F. Kockum, $et~al$., Phys. Rev. A 90, 013837 (2014)], we propose an effective magnetometer model based on single giant emitter. In this model, the emitter's frequency is proportional to the applied bias magnetic field. The self-interference effect causes the slope of the dissipation spectrum to vary linearly with the number of emitter-coupling points. The giant emitter magnetometer achieves a sensitivity as high as $10^{-8}-10^{-9}\,{\rm T/\sqrt{Hz}}$, demonstrating the significant advantages of the self-interference effect compared to small emitters. We hope our proposal will expand the applications of giant emitters in precision measurement and magnetometry.

quant-ph

Controllable superradiance scaling in photonic waveguide

We investigate the superradiance of two-level target atoms (TAs) coupled to a photonic waveguide, demonstrating that the scaling of the superradiance strength can be controlled on demand by an ensemble of control atoms (CAs). The scaling with respect to the number of TAs can be lower, higher, or equal to the traditional Dicke superradiance, depending on the relative positioning of the ensembles and the type of CAs (e.g., small or giant). These phenomena are attributed to unconventional atomic correlations. Furthermore, we observe chiral superradiance of the TAs, where the degree of chirality can be enhanced by giant CAs instead of small ones. The effects discussed in this work could be observed in waveguide QED experiments, offering a potential avenue for manipulating superradiance.

quant-ph

Seal: Advancing Speech Language Models to be Few-Shot Learners

Existing auto-regressive language models have demonstrated a remarkable capability to perform a new task with just a few examples in prompt, without requiring any additional training. In order to extend this capability to a multi-modal setting (i.e. speech and language), this paper introduces the Seal model, an abbreviation for speech language model. It incorporates a novel alignment method, in which Kullback-Leibler divergence loss is performed to train a projector that bridges a frozen speech encoder with a frozen language model decoder. The resulting Seal model exhibits robust performance as a few-shot learner on two speech understanding tasks. Additionally, consistency experiments are conducted to validate its robustness on different pre-trained language models.

cs.CL

3D Geometry-aware Deformable Gaussian Splatting for Dynamic View Synthesis

In this paper, we propose a 3D geometry-aware deformable Gaussian Splatting method for dynamic view synthesis. Existing neural radiance fields (NeRF) based solutions learn the deformation in an implicit manner, which cannot incorporate 3D scene geometry. Therefore, the learned deformation is not necessarily geometrically coherent, which results in unsatisfactory dynamic view synthesis and 3D dynamic reconstruction. Recently, 3D Gaussian Splatting provides a new representation of the 3D scene, building upon which the 3D geometry could be exploited in learning the complex 3D deformation. Specifically, the scenes are represented as a collection of 3D Gaussian, where each 3D Gaussian is optimized to move and rotate over time to model the deformation. To enforce the 3D scene geometry constraint during deformation, we explicitly extract 3D geometry features and integrate them in learning the 3D deformation. In this way, our solution achieves 3D geometry-aware deformation modeling, which enables improved dynamic view synthesis and 3D dynamic reconstruction. Extensive experimental results on both synthetic and real datasets prove the superiority of our solution, which achieves new state-of-the-art performance. The project is available at https://npucvr.github.io/GaGS/

cs.CV

Forward Flow for Novel View Synthesis of Dynamic Scenes

This paper proposes a neural radiance field (NeRF) approach for novel view synthesis of dynamic scenes using forward warping. Existing methods often adopt a static NeRF to represent the canonical space, and render dynamic images at other time steps by mapping the sampled 3D points back to the canonical space with the learned backward flow field. However, this backward flow field is non-smooth and discontinuous, which is difficult to be fitted by commonly used smooth motion models. To address this problem, we propose to estimate the forward flow field and directly warp the canonical radiance field to other time steps. Such forward flow field is smooth and continuous within the object region, which benefits the motion model learning. To achieve this goal, we represent the canonical radiance field with voxel grids to enable efficient forward warping, and propose a differentiable warping process, including an average splatting operation and an inpaint network, to resolve the many-to-one and one-to-many mapping issues. Thorough experiments show that our method outperforms existing methods in both novel view rendering and motion modeling, demonstrating the effectiveness of our forward flow motion modeling. Project page: https://npucvr.github.io/ForwardFlowDNeRF

cs.CV

GRASS: Unified Generation Model for Speech-to-Semantic Tasks

This paper explores the instruction fine-tuning technique for speech-to-semantic tasks by introducing a unified end-to-end (E2E) framework that generates target text conditioned on a task-related prompt for audio data. We pre-train the model using large and diverse data, where instruction-speech pairs are constructed via a text-to-speech (TTS) system. Extensive experiments demonstrate that our proposed model achieves state-of-the-art (SOTA) results on many benchmarks covering speech named entity recognition, speech sentiment analysis, speech question answering, and more, after fine-tuning. Furthermore, the proposed model achieves competitive performance in zero-shot and few-shot scenarios. To facilitate future work on instruction fine-tuning for speech-to-semantic tasks, we release our instruction dataset and code.

cs.CL

Unveiling the Impact of Cognitive Distraction on Cyclists Psycho-behavioral Responses in an Immersive Virtual Environment

The National Highway Traffic Safety Administration reported that the number of bicyclist fatalities has increased by more than 35% since 2010. One of the main reasons associated with cyclists' crashes is the adverse effect of high cognitive load due to distractions. However, very limited studies have evaluated the impact of secondary tasks on cognitive distraction during cycling. This study leverages an Immersive Virtual Environment (IVE) simulation environment to explore the effect of secondary tasks on cyclists' cognitive distraction through evaluating their behavioral and physiological responses. Specifically, by recruiting 75 participants, this study explores the effect of listening to music versus talking on the phone as a standardized secondary tasks on participants' behavior (i.e., speed, lane position, input power, head movement) as well as, physiological responses including participants' heart rate variability and skin conductance metrics. Our results show that (1) listening to high-tempo music can lead to a significantly higher speed, a lower standard deviation of speed, and higher input power. Additionally, the trend is more significant for cyclists who had a strong habit of daily music listening (> 4 hours/day). In the high cognitive workload situation (simulated hands-free phone talking), cyclists had a lower speed with less input power and less head movement variation. Our results indicate that participants' HRV (HF, pnni-50) and EDA features (numbers of SCR peaks) are sensitive to cyclists' cognitive load changes in the IVE simulator.

cs.HC