arXiv ScienceSearch

arXiv subjects

Yujia Zhang

Publications and source records attributed to Yujia Zhang.

At least 19 recordsLinked to original sources

Scalable Bayesian Optimization of Composite Functions for Image-Based Inverse Problems in Materials Characterization

Estimating physical parameters from scientific images is a common inverse problem in materials characterization that often relies on expensive physics-based simulations. In electron microscopy, specimen thickness and crystal mistilt are critical parameters that govern how electrons scatter through the sample, and therefore the accuracy of any atomic-scale structure recovered from it. They are commonly inferred by matching experimental position-averaged convergent-beam electron diffraction (PACBED) patterns to simulated ones, but grid searches scale poorly and neural-network methods require extensive pretraining that may not transfer to new conditions. Here, we propose scalable Bayesian optimization of composite functions (SBOCF), a simulation-efficient method that exploits the known composite structure of the image-matching objective and the intermediate information contained in simulated images. By representing PACBED images with patch-level summaries and two correction terms, SBOCF preserves the original pixel-wise objective while reducing the number of modeled outputs from 24,649 to 11. Under a budget of 50 simulator evaluations, SBOCF outperformed standard Bayesian optimization with expected improvement on synthetic SrTiO3 benchmarks with thick and thin specimens, reducing the median final SSE by up to 290x in the thick-sample case. On experimental data, SBOCF produced parameter estimates consistent with previously reported values without task-specific pretraining. For a simulated mistilted specimen, using the SBOCF estimates in a downstream ptychographic reconstruction recovered sharp atoms that were otherwise blurred. These results establish SBOCF as a promising approach for inverse problems involving expensive simulators and high-dimensional structured outputs.

cs.LG

Entirely nonlocal quantum magic without entanglement

Nonstabilizerness, or magic, is an archetypal \emph{quantum} resource that is necessary for quantum computational advantage. Here we uncover a phenomenon seemingly at odds with the quantum nature of magic: entirely nonlocal magic (ENM)---magic present only in correlations and absent from each party's marginal---can live without entanglement. We systematically study this separation and show it is universal and operationally reversible: every magical state or channel can be encoded into and recovered from a separable ENM realization using only local stabilizer processing and classical communication. We leverage this mechanism to devise an activation key protocol in which a classical key controls access to non-Clifford operations. We further formulate magic secret sharing, in which computational power inaccessible to any party alone becomes accessible through cooperation. On a superconducting quantum processor, we experimentally demonstrate activation key and network computing primitives, together with separable ENM state preparation and extraction protocols. Together, our results establish that magic can be classically activated, localized, and secret-shared without entanglement, providing new resource-control primitives for distributed quantum computation.

quant-ph

Temporally Grounded Compositional Camera Motion Understanding via Geometric Knowledge Distillation

Understanding camera motion is fundamental to video perception, with applications in spatial intelligence and controllable video generation. Multimodal large language models (MLLMs) provide a natural interface for this task, but existing work typically assigns one or more labels to an entire clip. Such clip-level recognition overlooks two defining properties of real camera motion: it can change within a shot, and multiple movements can occur simultaneously. We therefore formulate camera-motion understanding as temporally grounded, compositional recognition, which requires a model to localize motion-consistent intervals and identify every movement active within each interval. We introduce CamChoreo, a benchmark of 4,229 real single-shot clips with expert-annotated temporal segments. Its annotations use a compact vocabulary of 20 direction-aware labels, and nearly half of the segments contain compound camera motion, with multiple movement primitives active simultaneously. Recognizing such fine-grained, compositional motion is hard for current MLLMs, whose visual encoders emphasize semantic content rather than the geometric evidence on which camera motion depends. Directly injecting features from a frozen 3D foundation model addresses this gap, but requires running the expensive geometry model on every input; we refer to this baseline as CamInject. We instead propose CamDistill, which distills the same geometric knowledge into lightweight camera tokens during training and removes the 3D model at inference. CamDistill matches the accuracy of direct feature injection without running the 3D teacher at inference. Together, CamChoreo and CamDistill advance camera-motion understanding from clip-level labeling to temporally grounded, compositional recognition. Project page: https://ddz16.github.io/cammotion.github.io/.

cs.CV

On the transition to large fluxes and access to second stability in gyrokinetic simulations of electromagnetic turbulence in STEP

This work investigates the nonlinear transition to large heat fluxes observed in local gyrokinetic simulations of electromagnetic turbulence in STEP. Using the stress-balance framework of Zhang et al. (arXiv:2606.04616, arXiv:2607.11789), we confirm that the onset of extreme transport correlates with a critical value of $q^{2}\beta_{e}$, where $q$ is the safety factor and $\beta_{e}$ is the ratio of electron thermal pressure to magnetic pressure, and relate this to a limit on the poloidal beta $\beta_{\mathrm{pol}}$. Crucially, this critical value lies below any relevant linear stability limit in the ($q$, $\beta_{e}$) space (e.g., the onset of ideal or kinetic ballooning modes). Using an extensive set of nonlinear gyrokinetic simulations, we demonstrate that the transition to large fluxes in STEP is governed by a balance between the electrostatic and magnetic-flutter stresses. We argue, and also show numerically, that larger-major-radius tokamaks reach the electromagnetic non-zonal regime at lower $\beta_{e}$, making this MHD-controlled saturation limit more accessible in reactor-scale devices than in small spherical tokamaks. We also demonstrate that access to a second-stable regime enables re-saturation at larger values of $\beta^{\prime}$. We further show that the ideal ballooning mode (IBM) threshold serves as a useful proxy for delineating this second-stable region and also as a qualitative guide for the onset of large fluxes. These results provide a predictive framework for identifying no-go zone predictions from local gyrokinetics and offer new insight into the electromagnetic saturation physics relevant to STEP and other high-$\beta_{e}$ devices.

physics.plasm-ph

A Multi-Modal Framework with Cross-Subject Pseudo-Labeling and Semantic Alignment for Micro-Gesture Recognition

Micro-gestures (MGs) are spontaneous and subtle body movements that frequently convey hidden human emotions. Recognizing MGs in untrimmed videos remains highly challenging due to their extremely low signal-to-noise ratio, severe long-tailed class distribution, and the inherent domain shift encountered in cross-subject evaluation scenarios. In this paper, we propose a comprehensive multi-modal framework for Track 1 of the 4th MiGA-IJCAI Challenge. To capture fine-grained representations, we design a saliency-guided multi-modal extraction pipeline integrating 68-keypoint skeleton joint coordinates, 3D heatmap volumes, and high-resolution RGB visual features. We introduce a gentle square-root smoothed weighting mechanism paired with an Orthogonal Semantic Embedding Loss to protect tail classes without compromising overall recognition capabilities. More importantly, to bridge the cross-subject generalization gap, we propose a Cross-Modal Pseudo-Labeling (CMPL) strategy for unsupervised domain adaptation, which significantly boosts single-modal robustness. A temperature-scaled soft-voting mechanism is finally utilized to alleviate overconfidence during late fusion. Extensive experiments demonstrate that our framework achieves a competitive F1-score of 68.13\%, securing the 4th place.

cs.CV

See Better, Foresee Better, Act Wiser: Physically Grounded Proactive Modeling and Decision Making

Reliable proactive agents must choose an action and judge whether current evidence is sufficient to act. We study retail service from sparse third-person video: before an explicit customer request, an agent must use limited human-object interaction evidence to intervene or remain silent. Physical grounding here means converting observations into task-relevant retail state, not modeling low-level dynamics. We introduce the Proactive Intent World Model (PIWM): See constructs the perceptual basis, Foresee models counterfactual consequences, and Act selects an action. Performance is poor when the agent must extract information from raw video and decide directly, but improves substantially with structured inputs extracted and annotated from a professional retail perspective. AIDA-stage constraints and BDI-state ablations further support role- and goal-directed selection and organization of decision-relevant cues. Counterfactual prediction performs well in standalone evaluation, yet planning methods that query these forecasts at inference time degrade sharply: locally useful consequence prediction does not reliably improve action selection. This gap may reflect incomplete process understanding, uncertainty in fine-grained single-step outcomes, and insufficient joint modeling of scenes and temporal evolution. Hold remains the hardest action in structured-state evaluation, exposing a related challenge in temporal awareness. PIWM advances static intent recognition toward intent world modeling by organizing observations under task knowledge, anticipating candidate interventions, and treating intervention and non-intervention jointly. Future work will introduce long-horizon interaction trajectories and temporal consequence supervision to improve sustained reasoning and intervention timing.

cs.CL

Transformer-based Neural Operators for 3D Wind Field Prediction over Complex Mountainous Terrain

Accurate prediction of three-dimensional (3D) wind fields over complex mountainous terrain is essential for renewable energy deployment and regional weather modeling. Traditional computational fluid dynamics (CFD) simulations face two fundamental bottlenecks: expert-intensive mesh generation around irregular topography, and iterative solvers that require hours to days even on high-performance clusters. Recent neural operator approaches accelerate inference, but typically fail to resolve the sharp, localized velocity gradients induced by complex terrain features. Here, we present a transformer-based dual-attention neural-operator framework for 3D wind field prediction over complex mountainous terrain, and validate its effectiveness through two instantiations on representative point-based (mesh-free) and graph-based neural-operator architectures, namely Patch-solver and Patch-GTO. Trained on a large CFD-generated dataset spanning diverse terrain geometries and inflow conditions, the framework enables rapid prediction of steady-state wind field while maintaining competitive accuracy. It also demonstrates robust zero-shot transfer to real-world mountainous sites across several diverse locations, outperforming existing neural operator baselines by 10% in relative error. We further verify that incorporating sparse observational data (1% spatial coverage) reduces prediction error by 16.89% relative to the corresponding model without sparse data input and by 32.75% relative to advanced neural operator baselines on unseen terrains. This framework establishes a generalizable computational paradigm across domains, promising to be a real-time tool for wind resource assessment over complex mountainous terrain and related atmosphere-surface interaction studies.

physics.flu-dyn

EvalVerse: Pipeline-Aware and Expert-Calibrated Benchmarking for Professional Cinematic Video Generation

The rapid evolution of generative video foundation models has propelled the field toward professional-grade cinematic synthesis. To achieve such demanding quality, the community transitions towards Reinforcement Learning (RL) and agentic workflows. However, reliable evaluation has emerged as a critical bottleneck. Existing benchmarks predominantly evaluate ''whether it is right'' (basic prompt-following) while fundamentally neglecting ''whether it is good'' (cinematic quality, acting, and aesthetics). Furthermore, current automated metrics lack the domain-specific rigor required to provide trustworthy signals, creating a severe credibility gap between human aesthetic perception and machine scoring. To bridge this gap, we introduce EvalVerse, a comprehensive, pipeline-aware, and expert-calibrated evaluation framework. We treat video generation assessment not merely as an engineering task, but as a core scientific problem: the systematic digitization of subjective cinematic expertise. First, we organize domain knowledge into an evaluation taxonomy aligned with the professional filmmaking workflow (pre-production, production, and post-production). Second, we distill human expert judgments into a curated dataset with large-scale human annotations. Third, we inject this knowledge into Vision-Language Models (VLMs) through an expert-calibrated fine-tuning strategy, enabling the VLM to perform explicit Chain-of-Thought reasoning. Compared to previous works, EvalVerse not only retains compatibility with foundational ''rightness'' metrics, but also significantly expands the criteria to ''goodness'' and broaden the task coverage to complex multi-shot sequencing and audio-visual integration. Consequently, by providing granular diagnostic signals, EvalVerse transcends a static leaderboard and establishes a fundamental infrastructure for future work, such as reward models and evaluator agent.

cs.CV

Non-Local and Non-Markovian Effects of a Microscopic Two-Level Defect in Superconducting Quantum Circuits

Microscopic two-level systems (TLS) -- ubiquitous atomic-scale defects in solid-state quantum devices -- are a dominant source of qubit decoherence, yet their role is often considered local and short-memoried. Here, we report the observation of a coherent TLS that couples simultaneously to two spatially distant superconducting qubits. The TLS is identified to reside within the tunable coupler linking the qubits, enabling controllability of the TLS-qubit coupling strength via coupler frequency -- a capability absent in earlier studies. This tunability allows us to systematically probe how TLS distorts qubit dynamics, revisiting the decoherence model in the presence of non-Markovian TLS dephasing noise. This is corroborated by the reconstructed $1/f$ noise spectrum of TLS frequency fluctuation spanning more than ten orders of magnitude (0.1\,mHz -- 1\,MHz) that reveals discrete fluctuator signatures. Quantum process tomography further unveils TLS-induced correlated qubit dynamics, highlighting the long-lived TLS as an effective source of non-Markovianity. Our findings expose a previously overlooked interaction mechanism in scalable quantum architectures: defects embedded in coupling elements can simultaneously affect multiple qubits with variable impact. Beyond immediate implications for system characterization and calibration, this situation provides a powerful testbed for studying defect-driven quantum dynamics, refining error suppression strategies, and advancing architecture design for scalable quantum technologies.

quant-ph

MLLMs Know When Before Speaking: Revealing and Recovering Temporal Grounding via Attention Cues

Video temporal grounding (VTG), which localizes the start and end times of a queried event in an untrimmed video, is a key test of whether multimodal large language models (MLLMs) understand not only what happens but also when it happens. Although modern MLLMs describe video content fluently, their timestamp predictions remain unreliable, while existing remedies either require costly post-training on temporal annotations or rely on coarse training-free heuristics. In this work, we probe the cross-modal attention of MLLMs and uncover a perception-generation gap. Our key finding is that MLLMs often know the target interval during prefill, but lose this signal when generating the final answer. In the prefill stage, a sparse set of attention heads, which we call \emph{Temporal Grounding Heads} (TG-Heads), concentrates query-to-video attention on the ground-truth interval. During autoregressive decoding, however, the answer tokens shift attention away from this interval toward visually salient but query-irrelevant segments. This observation motivates an inference-time read-then-regenerate framework. We first convert TG-Head prefill attention into a debiased frame-level relevance signal and extract the high-attention interval it highlights. We then re-invoke the MLLM with visual context restricted to this interval, using video cropping or attention masking to suppress distractors. Without parameter updates and architectural changes, our framework consistently improves MiMo-VL-7B, Qwen3-VL-8B, and TimeLens-8B on three VTG benchmarks, with gains of up to +3.5 mIoU. The project website can be found at https://ddz16.github.io/mllmsknowwhen.github.io/.

cs.CV

Learning Spatiotemporal Sensitivity in Video LLMs via Counterfactual Reinforcement Learning

Video large language models (Video LLMs) achieve strong benchmark accuracy, yet often answer video questions through shortcuts such as single-frame cues and language priors rather than by tracking spatiotemporal dynamics. This issue is exacerbated in RL post-training, where correctness-only rewards can further reinforce shortcut policies that obtain high reward without tracking video dynamics. We address this by asking a controlled counterfactual question: if the visual world changed while the question remained fixed, should the answer change or stay the same? Based on this view, we propose \textbf{Counterfactual Relational Policy Optimization (CRPO)}, a dual-branch RL framework for improving \emph{spatiotemporal sensitivity}. CRPO constructs counterfactual videos through horizontal flips and temporal reversals, trains on both original and counterfactual branches, and introduces a \textbf{Counterfactual Relation Reward (CRR)} between their answers. CRR encourages answers to change for dynamic questions and remain unchanged for static questions. This cross-branch constraint makes it difficult for shortcut policies to be consistently rewarded across both branches. To evaluate this property, we introduce \textbf{DyBench}, a paired counterfactual video benchmark with 3,014 videos covering reversible dynamics, moving direction, and event sequence, together with a strict pair-accuracy metric that prevents fixed-answer shortcuts from inflating scores. Experiments show that CRPO outperforms prior RL methods on spatiotemporal-sensitive evaluations while maintaining competitive general video performance. On Qwen3-VL-8B, CRPO improves DyBench P-Acc by +7.7 and TimeBlind I-Acc by +8.2 over the base model, indicating improved spatiotemporal sensitivity rather than stronger reliance on static shortcuts. The project website can be found at https://ddz16.github.io/crpo.github.io/ .

cs.CV

HotLoop Optimization of Petawatt Laser Focal Spot via a Twin-Focus Scheme

Achieving diffraction-limited focusing of high-power laser pulses to generate ultra-high intensities is crucial for developing compact laser-driven particle accelerators and exploring strong-field quantum electrodynamics. However, accurately diagnosing and optimizing the focal spots of petawatt (PW) laser pulses remains a significant challenge. In this work, we present an experimental methodology utilizing a twin-focus scheme to precisely characterize the intensity distribution and wavefront of focused PW femtosecond laser pulses, and employ it to elucidate their power-dependent evolution. Furthermore, we optimized the focal spots at full power via our in-situ wavefront correction method termed ``HotLoop', achieving a Strehl ratio of 0.80 for 1 PW laser pulses. Consequently, the cutoff proton energies in laser proton acceleration experiments were significantly enhanced. The success of this approach underscores the necessity of in-situ high-energy wavefront correction for ultra-high intensity laser-matter interactions.

physics.optics

Sheet As Token: A Graph-Enhanced Representation for Multi-Sheet Spreadsheet Understanding

Workbook-scale spreadsheet understanding is increasingly important for language-model-based data analysis agents, but remains challenging because relevant information is often distributed across multiple sheets with heterogeneous schemas, layouts, and implicit relationships. Existing retrieval-augmented approaches typically decompose spreadsheets into rows, columns, or blocks to improve scalability; however, such chunk-centric representations can fragment worksheets into isolated text spans and weaken global sheet-level semantics. We propose Sheet As Token (SAT), a graph-enhanced framework that treats each worksheet as a unified semantic unit for multi-sheet spreadsheet retrieval. SAT serializes sparse schema-aware features, including sheet name, shape, and column headers, and encodes each worksheet into a compact dense token. Given a query, SAT retrieves candidates with a BGE-initialized Sheet Encoder and refines them with a gated relational GNN. In strict full-corpus evaluation, SAT reaches 0.9173 NDCG@5 on IndustryTab-614 and 0.9222 on IndustryTab-1K, relative improvements of 44.6% and 46.7% over zero-shot BGE RAG, respectively. SAT therefore improves both retrieval accuracy and serving efficiency: on IndustryTab-1K, it exceeds a Qwen3.5-9B RAG reranker by 12.5% while reducing online latency from 2.61 s to 9.24 ms, approximately 283X faster. These results show that SAT provides accurate and latency-efficient retrieval in the evaluated fixed-corpus setting. Code and data are available at https://github.com/SHITIANYU-hue/SheetasToken .

cs.AI

A Spatial-Resolved Proton Energy Spectrometer Based on a Scintillation-Fiber Cube

Advanced particle acceleration methods have produced high-peak-current ion beams with broad energy spread and complex spatial distribution. There is an urgent need to develop online spatial-resolved energy spectrometers for high-energy pulsed ions. This paper introduces a novel spectrometer based on a scintillation-fiber cube for online diagnosis of proton beams with broadband energy spread and complex spatial distribution. We present its working principles, experimental setup, and comprehensive calibration using monoenergetic and spatially uniform proton beams generated by a synchrotron accelerator. Calibration results confirm an energy measurement range of 6-93 MeV, a relative energy uncertainty of 0.6% at 80 MeV, and a pixel size of 0.5 mm for beam profile reconstruction. By exploiting a custom-designed energy degrader, we generated a complex proton beam and measured it with the scintillation-fiber cube spectrometer (SFICS). The results demonstrate the spectrometer's potential for online measurement of the energy spectrum and spatial distribution of complex proton beams.

physics.acc-ph

Power laws in the sea ice floe size distribution: a stochastic theory

Sea ice is a complex system, and observations have shown that ice segments (i.e., floes) have a wide range of sizes, with a floe size distribution that follows a power law. However, a theory for the power law and its exponent have remained elusive. Here, floe-resolving numerical simulations are investigated with a discrete element model, in order to gain further information by gathering statistics of fracture and welding events. Then, based on the insights from the floe-resolving simulations, a stochastic fragmentation-coagulation theory is proposed. Exact solutions are found with a power law. The power-law exponent can take a variety of values, and it depends on the fracture and welding rates. Such behavior is reminiscent of seasonal changes in the power-law exponent, which have been reported in past analyses of observational data.

physics.ao-ph

Micro-Macro Retrieval: Reducing Long-Form Hallucination in Large Language Models

Large Language Models (LLMs) achieve impressive performance across many tasks but remain prone to hallucination, especially in long-form generation where redundant retrieved contexts and lengthy reasoning chains amplify factual errors. Recent studies highlight a critical phenomenon: the closer key information appears to the model outputs, the higher the factual accuracy. However, existing retrieval-augmented language models (RALMs) lack effective mechanisms to ensure this proximity - external evidence is injected into reasoning via multi-turn retrieval, but this cannot ensure key information stays close to the outputs. We propose Micro-Macro Retrieval (M2R), a novel retrieve-while-generate framework to fill this gap. At the macro level, M2R retrieves coarse-grained evidence from external sources; at the micro level, it extracts essential results from a key information repository built during reasoning and reuses them while generating answers. This design directly addresses the key-information-to-output proximity bottleneck, effectively reducing hallucination in long-form tasks. M2R is trained with a curriculum learning-based reinforcement learning strategy using customized rule-based rewards, enabling stable acquisition of retrieval and grounding skills. Extensive experiments across different benchmarks demonstrate the effectiveness of M2R, especially in lengthy-context settings.

cs.CL

Visualizing spin-polarization of an altermagnet KV$_2$Se$_2$O via spin-selective tunneling

Altermagnetism, a recently identified magnetic phase that combines vanishing net magnetization with momentum-dependent spin splitting, challenges the conventional dichotomy between ferromagnets and antiferromagnets. While several candidate materials have been proposed, direct experimental evidence linking crystal symmetry, electronic structure and d-wave spin polarization remains scarce. Here we report the visualization of a metallic d-wave altermagnet in KV2Se2O. Through spin-selective scanning tunneling microscopy powered by a topological insulator tip, we uncover symmetry-protected momentum-dependent spin splitting that follows a characteristic d-wave form factor. Our results establish KV2Se2O as a tunable platform to study the interplay between spin-valley locking, Fermi-surface instability and unconventional magnetism, and open a pathway toward symmetry-engineered spintronics without net magnetization.

cond-mat.mtrl-sci

Utonia: Toward One Encoder for All Point Clouds

We dream of a future where point clouds from all domains can come together to shape a single model that benefits them all. Toward this goal, we present Utonia, a first step toward training a single self-supervised point transformer encoder across diverse domains, spanning remote sensing, outdoor LiDAR, indoor RGB-D sequences, object-centric CAD models, and point clouds lifted from RGB-only videos. Despite their distinct sensing geometries, densities, and priors, Utonia learns a consistent representation space that transfers across domains. This unification improves perception capability while revealing intriguing emergent behaviors that arise only when domains are trained jointly. Beyond perception, we observe that Utonia representations can also benefit embodied and multimodal reasoning: conditioning vision-language-action policies on Utonia features improves robotic manipulation, and integrating them into vision-language models yields gains on spatial reasoning. We hope Utonia can serve as a step toward foundation models for sparse 3D data, and support downstream applications in AR/VR, robotics, and autonomous driving.

cs.CV