arXiv ScienceSearch

arXiv subjects

Qian Cao

Publications and source records attributed to Qian Cao.

At least 19 recordsLinked to original sources

Countervailing Curation Strategic Disclosure and the Design of Attention

Opposing advocates may be unable to fabricate evidence but can choose which true observations to present. We study how a court, editor, or platform should divide potential exposure between an advocate who prefers a higher decision and one who prefers a lower decision. Each advocate controls a separate evidence pool, exposure is committed before the state and evidence are known, and a selection-naive receiver averages displayed observations. Under baseline linear preferences and common knowledge of the realized pools, all Nash equilibria in every finite pool induce the same action and common disclosure bar: the upward advocate reveals observations above it and the downward advocate those below it. Attention therefore changes selection as well as weight. Under a common evidence law, more exposure makes an advocate speak less often and more extremely while moving the decision in its preferred direction. In large pools, the unique state-contingent exposure share that reproduces the complete-evidence decision balances the advocates' directional tail moments, not their observed speech. Because this share generally depends on the unknown state, we characterize the optimal ex ante compromise and solve a primitive two-state economy with a unique second-best policy. For finite pools, we derive the exact risk-minimizing adjustment, separating selection-induced cutoff bias, the receiver's fixed benchmark, and sampling variance. The attention-to-cutoff feedback survives partial weighting of empty slots and vanishes at full imputation.

econ.TH

Mixed-State Symmetry-Protected Topology and Strong-to-Weak Spontaneous Symmetry Breaking in a Superconducting Qubit Array

We experimentally investigate how symmetry-protected topological order in a one-dimensional cluster state is transformed by measurement and decoherence in a five-qubit superconducting array. We first characterize the state's nonlocal string order and show that controlled dephasing selectively suppresses one symmetry sector while leaving the other robust, consistent with average symmetry-protected topological order. We then measure one sublattice in a tunable basis and show that the remaining qubits are driven between a long-range-entangled GHZ state and a paramagnetic state. When the measurement record is discarded, the conventional long-range correlator vanishes while a nonlinear fidelity correlator remains finite, providing a finite-size signature of strong-to-weak spontaneous symmetry breaking. These experiments demonstrate how conditioning, averaging, and decoherence reveal distinct manifestations of order encoded in the same underlying cluster state, and establish a superconducting-circuit setting for probing mixed-state symmetry and topology.

quant-ph

StorySpark: Module-wise Evolutionary Search for Story Premise Generation

A story premise is the creative spark from which a full narrative can grow. Yet LLM-based story generation has mostly emphasized later-stage planning, controllability, coherence, and prose expansion, while premise-level ideation remains comparatively underexplored. We introduce StorySpark, a module-wise evolutionary search framework for story premise generation. StorySpark operates over interpretable narrative modules such as background, persona, event, ending, and twist, treating each active module not as a static field to fill once, but as a local search space conditioned on the partial premise built so far. For each module, it generates alternatives, evaluates them in context, refines them through feedback-driven mutation and recombination, preserves complementary strengths with Pareto-guided selection, and reallocates frontier capacity to balance branch coverage with promising directions. Multi-view automatic and human evaluations show that StorySpark produces stronger final premises than competitive baselines, with especially consistent gains in originality; when expanded with the same story writer, its premises also lead to higher-quality downstream stories while maintaining completeness, fascination, and diverse usable narrative directions.

cs.CL

Measurement and Control of the Complex Berry Phase in a Quantum System

The Berry phase is a geometric phase acquired during adiabatic evolution over a closed loop in parameter space. It plays an essential role in geometric quantum gates and other phase-based protocols. In non-Hermitian systems, the Berry phase is complex, introducing fundamentally new geometric effects, including state amplification. In this work, we report experimental measurement of both the real and imaginary components of a Berry phase in a fully quantum system using a superconducting transmon circuit with engineered dissipation. We also demonstrate the path-dependent effects of the imaginary part on the dissipation and its utility in the implementation of non-unitary quantum control. These findings establish a clear geometric distinction between the real and imaginary components of the Berry phase and experimentally confirm the unique adiabatic behavior of non-Hermitian quantum systems.

quant-ph

Qwen-Scope: Turning Sparse Features into Development Tools for Large Language Models

Large language models have achieved remarkable capabilities across diverse tasks, yet their internal decision-making processes remain largely opaque, limiting our ability to inspect, control, and systematically improve them. This opacity motivates a growing body of research in mechanistic interpretability, with sparse autoencoders (SAEs) emerging as one of the most promising tools for decomposing model activations into sparse, interpretable feature representations. We introduce Qwen-Scope, an open-source suite of SAEs built on the Qwen model family, comprising 14 groups of SAEs across 7 model variants from the Qwen3 and Qwen3.5 series, covering both dense and mixture-of-expert architectures. Built on top of these SAEs, we show that SAEs can go beyond post-hoc analysis to serve as practical interfaces for model development along four directions: (i) inference-time steering, where SAE feature directions control language, concepts, and preferences without modifying model weights; (ii) evaluation analysis, where activated SAE features provide a representation-level proxy for benchmark redundancy and capability coverage; (iii) data-centric workflows, where SAE features support multilingual toxicity classification and safety-oriented data synthesis; and (iv) post-training optimization, where SAE-derived signals are incorporated into supervised fine-tuning and reinforcement learning objectives to mitigate undesirable behaviors such as code-switching and repetition. Together, these results demonstrate that SAEs can serve not only as post-hoc analysis tools, but also as reusable representation-level interfaces for diagnosing, controlling, evaluating, and improving large language models. By open-sourcing Qwen-Scope, we aim to support mechanistic research and accelerate practical workflows that connect model internals to downstream behavior.

cs.CL

Analogical Reasoning as a Doctor: A Foundation Model for Gastrointestinal Endoscopy Diagnosis

Gastrointestinal diseases impose a growing global health burden, and endoscopy is a primary tool for early diagnosis. However, routine endoscopic image interpretation still suffers from missed lesions and limited efficiency. Although AI-assisted diagnosis has shown promise, existing models often lack generalizability, adaptability, robustness, and scalability because of limited medical data, domain shift, and heterogeneous annotations. To address these challenges, we develop RATNet, a foundation model for gastrointestinal endoscopy imaging based on analogical reasoning. RATNet acquires and transfers knowledge from heterogeneous expert annotations across five gastrointestinal endoscopy datasets through a cyclic pre-training strategy. Its architecture consists of an encoder, a relevance-knowledge acquisition and transfer (RAT) module, a projector, and a multi-task head, and supports fine-tuning, linear probing, and zero-shot transfer. Evaluations show that RATNet outperforms existing foundation models, including GastroNet and GastroVision, across six scenarios: diagnosis of common gastrointestinal diseases, few-shot learning for rare diseases, zero-shot transfer to new medical sites, robustness under long-tailed disease distributions, adaptation to novel diseases, and privacy-preserving deployment via federated learning. Its advantage comes from an analogical reasoning mechanism that matches image-derived posterior knowledge to a learned prior knowledge base and transfers relative knowledge to guide diagnosis, improving generalization and resistance to bias. RATNet is open and cost-effective, supports automatic integration of heterogeneous annotations without manual label unification, and reduces data acquisition costs, making it a practical foundation for intelligent gastrointestinal diagnosis, especially in resource-limited settings.

cs.CV

Statistical modeling of breast cancer radiomic features and hazard using image registration-aided longitudinal CT data

Patients with metastatic breast cancer (mBC) undergo repeated computed tomography (CT) imaging during treatment to monitor disease progression. Accurate longitudinal tracking of individual lesions across scans from multiple radiologists is essential for reliable radiomic analysis and clinical decision-making. We conducted a retrospective study using serial chest CT scans from the Phase III MONALEESA-3 and MONALEESA-7 trials and developed statistical models for multi-source data integration and survival analysis. First, we introduced a Registration-based Automated Matching and Correspondence (RAMAC) algorithm to establish lesion correspondence across annotations from different radiologists and imaging time points using the Hungarian algorithm. Second, using the RAMAC-processed dataset, we developed interpretable radiomic survival models for progression-free survival prediction by combining baseline radiomic features, post-treatment changes at Weeks 8, 16, and 24, and demographic variables. To address the high dimensionality of longitudinal radiomic data, feature reduction was performed using an L1-penalized additive Cox proportional hazards model and best subset selection followed by Cox modeling. Model performance was evaluated using the concordance index (C-index). Incorporating additional imaging time points improved predictive performance, increasing the mean C-index from 0.58 at baseline to 0.64. Joint modeling further showed significant associations between longitudinal radiomic features and survival outcomes over time.

stat.AP

Restoring Exploration after Post-Training: Latent Exploration Decoding for Large Reasoning Models

Large Reasoning Models (LRMs) have recently achieved strong mathematical and code reasoning performance through Reinforcement Learning (RL) post-training. However, we show that modern reasoning post-training induces an unintended exploration collapse: temperature-based sampling no longer increases pass@$n$ accuracy. Empirically, the final-layer posterior of post-trained LRMs exhibit sharply reduced entropy, while the entropy of intermediate layers remains relatively high. Motivated by this entropy asymmetry, we propose Latent Exploration Decoding (LED), a depth-conditioned decoding strategy. LED aggregates intermediate posteriors via cumulative sum and selects depth configurations with maximal entropy as exploration candidates. Without additional training or parameters, LED consistently improves pass@1 and pass@16 accuracy by 0.61 and 1.03 percentage points across multiple reasoning benchmarks and models. Furthermore, integrating LED into reinforcement learning, e.g., using GRPO as the rollout strategy, yields faster reward improvement and higher final performance, due to the efficient exploration capability of LED. Project page: https://github.com/AlbertTan404/LED.

cs.CL

DPWriter: Reinforcement Learning with Diverse Planning Branching for Creative Writing

Reinforcement learning (RL)-based enhancement of large language models (LLMs) often leads to reduced output diversity, undermining their utility in open-ended tasks like creative writing. Current methods lack explicit mechanisms for guiding diverse exploration and instead prioritize optimization efficiency and performance over diversity. This paper proposes an RL framework structured around a semi-structured long Chain-of-Thought (CoT), in which the generation process is decomposed into explicitly planned intermediate steps. We introduce a Diverse Planning Branching method that strategically introduces divergence at the planning phase based on diversity variation, alongside a group-aware diversity reward to encourage distinct trajectories. Experimental results on creative writing benchmarks demonstrate that our approach significantly improves output diversity without compromising generation quality, consistently outperforming existing baselines.

cs.CL

Polygonal Spatiotemporal Optical Vortices Wavepackets with Prescribed Vortex Structure

Optical vortices carrying orbital angular momentum offer additional degrees of freedom. According to the orientation of orbital angular momentum, optical vortices can be classified into spatial optical vortex beam carrying longitudinalorbital angular momentum and spatiotemporal optical vortices carrying transverse orbital angular momentum. As an emerging subset of optical vortices, polygonal optical vortices provide a unique platform for a wide range of frontier applications by introducing a new degree of freedom in the form of a customizable intensity structure. In the spatial domain, polygonal spatial optical vortex beam carrying longitudinal orbital angular momentum have already demonstrated great potential in optical manipulation and two-photon lithography. However, polygonal spatiotemporal optical vortex wavepackets contain multiple sub spatiotemporal optical vortices carrying transverse orbital angular momentum remains unrealized to date. In this work, we theoretically propose and experimentally demonstrate polygonal spatiotemporal optical vortices wavepackets embedded with prescribed vortex structures. Within the structure, a prescribed number of sub spatiotemporal optical vortices carrying transverse orbital angular momentum is set along a designed polygonal spatiotemporal trajectory. Using the spatiotemporal holographic shaping approach, we generate polygonal perfect spatiotemporal optical vortex wavepacket and use the combination of multiple polygonal perfect spatiotemporal optical vortex wavepacket to form polygonal spatiotemporal optical vortex wavepacket with the prescribed vortex structure. A full control over multiple key properties of the polygonal spatiotemporal optical vortex wavepackets such as the geometry, number of phase singularities, and spatiotemporal distribution of sub spatiotemporal optical vortices is also achieved.

physics.optics

Nonlinear quantum evolution of a dissipative superconducting qubit

Unitary and dissipative models of quantum dynamics are linear maps on the space of states or density matrices. This linearity encodes the superposition principle, a key feature of quantum theory. However, this principle can break down in effective non-Hermitian dynamics arising from postselected quantum evolution. We theoretically characterize and experimentally investigate this breakdown in a dissipative superconducting transmon circuit. Within the circuit's three-level manifold, no-jump postselection generates an effective non-Hermitian Hamiltonian governing the excited two-level subspace and an anti-Hermitian nonlinearity. We prepare different initial states and use quantum state tomography to track their evolution under this effective, nonlinear Hamiltonian. By comparing the evolution of a superposition-state to a superposition of individually-evolved basis states, we test linearity and observe clear violations which we quantify across the exceptional-point (EP) degeneracy of the non-Hermitian Hamiltonian. We extend the analysis to density matrices, revealing a breakdown in linearity for the two-level subspace while demonstrating that linearity is preserved in the full three-level system. These results provide direct evidence of nonlinearity in non-Hermitian quantum evolution, highlighting unique features that are absent in classical non-Hermitian systems.

quant-ph

Spatiotemporal Topological Combs for Robust High-Dimensional Information Transmission

Sculpting light across its independent degrees of freedom-from orbital angular momentum to the discrete wavelengths of optical frequency combs-has unlocked vast communication bandwidth by enabling massively parallel information channels. However, the Shannon-Hartley theorem sets a hard limit by tying channel capacity to the trade-off between SNR and rate, a central challenge in communication. Inspired by lock-in amplification in electronics, we encode data on THz optical burst carriers so the signal resides beyond the conventional noise band, yielding exceptional robustness. By leveraging a programmable all-degree-of-freedom (All-DoF) modulator, we generate a spatiotemporal topological comb (ST-Comb) that structures light into a vast, highentropy state space for high-dimensional information encoding. Crucially, we find that the associated topological winding number is preserved under diverse perturbations, ensuring stable information encoding and retrieval. This paradigm illustrates how structured light can simultaneously expand channel dimensionality and maintain robustness, charting a pathway to chip-scale, reconfigurable photonic platforms for the PHz era, while also opening previously inaccessible regimes of light-matter interaction.

physics.optics

GeoBS: Information-Theoretic Quantification of Geographic Bias in AI Models

The widespread adoption of AI models, especially foundation models (FMs), has made a profound impact on numerous domains. However, it also raises significant ethical concerns, including bias issues. Although numerous efforts have been made to quantify and mitigate social bias in AI models, geographic bias (in short, geo-bias) receives much less attention, which presents unique challenges. While previous work has explored ways to quantify geo-bias, these measures are model-specific (e.g., mean absolute deviation of LLM ratings) or spatially implicit (e.g., average fairness scores of all spatial partitions). We lack a model-agnostic, universally applicable, and spatially explicit geo-bias evaluation framework that allows researchers to fairly compare the geo-bias of different AI models and to understand what spatial factors contribute to the geo-bias. In this paper, we establish an information-theoretic framework for geo-bias evaluation, called GeoBS (Geo-Bias Scores). We demonstrate the generalizability of the proposed framework by showing how to interpret and analyze existing geo-bias measures under this framework. Then, we propose three novel geo-bias scores that explicitly take intricate spatial factors (multi-scalability, distance decay, and anisotropy) into consideration. Finally, we conduct extensive experiments on 3 tasks, 8 datasets, and 8 models to demonstrate that both task-specific GeoAI models and general-purpose foundation models may suffer from various types of geo-bias. This framework will not only advance the technical understanding of geographic bias but will also establish a foundation for integrating spatial fairness into the design, deployment, and evaluation of AI systems.

cs.AI

Inverse Weak measurement in SERF magnetometer

Weak measurement techniques have been extensively applied in the field of quantum precision measurement to detect ultra-small signals due to the amplification effect. In this work, we propose an optical detection system for a spin-exchange relaxation-free (SERF) magnetometer based on the inverse weak measurement (IWM) framework. By using the spatial pattern of a probe laser as the measurement pointer, we successfully detect ultra-weak magnetic fields. In our model, the spatial pattern of the probe laser is weakly coupled to its polarization, which is sensitive to external magnetic fields. Through post-selection on the optical polarization, the ultra-small magnetic field is significantly amplified with the amplification factor inversely proportional to the coupling strength, as reflected in the measured displacement of the final spatial pattern. By analysing the response curve of the probe laser displacement to the magnetic field, we identify the point of maximum sensitivity, achieving a magnetic field sensitivity of 182.8 fT/Hz1/2. Furthermore, in the IWM scheme, the detected signals depend only on the internal degrees of freedom of the probe laser, making the system robust against the fluctuations in laser power. To demonstrate this advantage, we compute the Allan standard deviation of the output signals for both conventional and IWM detection methods. The results indicate that the IWM-based method improves stability of detection by one to two orders of magnitude. This work presents a novel detection approach that integrates weak measurement techniques, offering a significant enhancement in the performance of SERF magnetometers.

physics.ins-det

From Heuristics to Data: Quantifying Site Planning Layout Indicators with Deep Learning and Multi-Modal Data

The spatial layout of urban sites shapes land-use efficiency and spatial organization. Traditional site planning often relies on experiential judgment and single-source data, limiting systematic quantification of multifunctional layouts. We propose a Site Planning Layout Indicator (SPLI) system, a data-driven framework integrating empirical knowledge with heterogeneous multi-source data to produce structured urban spatial information. The SPLI supports multimodal spatial data systems for analytics, inference, and retrieval by combining OpenStreetMap (OSM), Points of Interest (POI), building morphology, land use, and satellite imagery. It extends conventional metrics through five dimensions: (1) Hierarchical Building Function Classification, refining empirical systems into clear hierarchies; (2) Spatial Organization, quantifying seven layout patterns (e.g., symmetrical, concentric, axial-oriented); (3) Functional Diversity, transforming qualitative assessments into measurable indicators using Functional Ratio (FR) and Simpson Index (SI); (4) Accessibility to Essential Services, integrating facility distribution and transport networks for comprehensive accessibility metrics; and (5) Land Use Intensity, using Floor Area Ratio (FAR) and Building Coverage Ratio (BCR) to assess utilization efficiency. Data gaps are addressed through deep learning, including Relational Graph Neural Networks (RGNN) and Graph Neural Networks (GNN). Experiments show the SPLI improves functional classification accuracy and provides a standardized basis for automated, data-driven urban spatial analytics.

cs.LG

Seed-X: Building Strong Multilingual Translation LLM with 7B Parameters

Multilingual translation stands as a challenging task for large language models (LLMs) to handle intricate language patterns and stilted translations that arise in automated translations. In this paper, we introduce Seed-X, a family of open-source LLMs comprising instruct and reasoning models, pushing the limits of translation capability with 7B parameter size. The base model is pre-trained on a diverse, high-quality dataset encompassing both monolingual and bilingual content across 28 languages, harnessing the full potential of multilingual data. The instruct model is then finetuned to translate by Chain-of-Thought (CoT) reasoning and further enhanced through reinforcement learning (RL) to achieve better generalization across diverse language pairs. Seed-X achieves performance comparable to leading closed-source models, including Gemini-2.5 and GPT-4o, across 28 languages, and significantly outperforms larger open-source models in both automatic metrics and human evaluations. We share the best practices through our optimization process, and make the parameter public available for advancing translation research and applications.

cs.CL

Spatiotemporal coupled Airy-Airy wavepacket and its propagation dynamics

Airy beams, celebrated for their self-acceleration, diffraction-free propagation, and self-healing properties, have garnered significant interest in optics and photonics, with applications spanning ultrafast optics, laser processing, nonlinear optics, and optical communications. Recent research primarily aims at independent control of Airy beams in both spatial and spatiotemporal domains. In a pioneering approach, we have successfully generated and controlled a spatiotemporal coupled (STc) Airy-Airy wavepacket, achieving its rotation while preserving vertical distribution in the spatiotemporal domain. Furthermore, we have investigated the self-acceleration and self-healing properties of the STc Airy-Airy wavepacket in this domain, noting that its dynamically adjustable rotation and spatiotemporal coupling capability provide a novel strategy for managing ultrafast lasers, with potential advancements in optical micromanipulation and time-domain coding communication.

physics.optics

DiffVLA: Vision-Language Guided Diffusion Planning for Autonomous Driving

Research interest in end-to-end autonomous driving has surged owing to its fully differentiable design integrating modular tasks, i.e. perception, prediction and planing, which enables optimization in pursuit of the ultimate goal. Despite the great potential of the end-to-end paradigm, existing methods suffer from several aspects including expensive BEV (bird's eye view) computation, action diversity, and sub-optimal decision in complex real-world scenarios. To address these challenges, we propose a novel hybrid sparse-dense diffusion policy, empowered by a Vision-Language Model (VLM), called Diff-VLA. We explore the sparse diffusion representation for efficient multi-modal driving behavior. Moreover, we rethink the effectiveness of VLM driving decision and improve the trajectory generation guidance through deep interaction across agent, map instances and VLM output. Our method shows superior performance in Autonomous Grand Challenge 2025 which contains challenging real and reactive synthetic scenarios. Our methods achieves 45.0 PDMS.

cs.AI