arXiv ScienceSearch

arXiv subjects

Pengfei Zhang

Publications and source records attributed to Pengfei Zhang.

At least 19 recordsLinked to original sources

The Platonic brain bridge hypothesis: human brain networks as an architectural prior for multimodal large language models

Multimodal large language models predict brain activity, but brain alignment has been a measurement, not a design tool. We propose the Platonic brain bridge hypothesis: omni models, multimodal large language models that process video, audio and text jointly, converge on brain-like representations usable in both directions. From model to brain, brain-likeness of seven omni models is stable across participants, rises with every input channel in three bases, and our encoders lead the Algonauts 2025 out-of-distribution leaderboard. From brain to model, Brain-MoE fixes the expert partition of a frozen base to the seven networks of human cortex, trains experts on network-labelled Brain-AVQA questions, raises held-out accuracy in all 15 model-benchmark pairs by 6.42 percentage points on average and exceeds capacity-matched random experts in 14. Brain-Scope localizes the correspondence to sparse features whose removal weakens brain prediction. Human brain organization is therefore a usable architectural prior for multimodal large language models.

q-bio.NC

Proving olympiad geometry theorems on a superconducting quantum processor

Automated theorem proving seeks to use computational systems to prove or disprove mathematical and logical statements [1, 2]. It underpins a wide range of applications, and enhancing theorem-proving capabilities remains a central objective in artificial intelligence [3]. Although recent neuro-symbolic systems have achieved remarkable progress [4-7], their operation is ultimately constrained by classical computational architectures. Quantum computing [8], by contrast, enables information encoding and coherent parallelism beyond classical limits [9-14], raising the possibility of accelerating structured symbolic deduction [15]. Here we report the experimental realization of automated geometry theorem proving on a fully programmable superconducting quantum processor. We develop two complementary quantum proving frameworks. The first implements Wu's algebraic elimination method using quantum pseudo-division, with multivariate polynomials represented in superposition states, enabling quantum algebraic theorem proving. The second implements the full-angle method as backward symbolic reasoning through a hybrid quantum strategy-guided architecture, demonstrating a general route toward quantum symbolic proof search. As illustrative examples, we prove two theorems on a superconducting quantum processor: the perpendicularity of the diagonals of a square and a 1978 International Mathematical Olympiad geometry problem. Our results establish, at the experimental level, automated logical reasoning as a viable task for near-term quantum processors and provide a concrete pathway toward quantum-enhanced symbolic intelligence.

quant-ph

Beyond Noise Steering: Dual-Latent Space Reinforcement Learning for Generative Robot Policy

Pretrained generative robot policies learn expressive action priors from demonstrations. However, existing reinforcement learning methods only steer the noisy space but fail to modulate intermediate action representations during the generation process, resulting in performance degradation and inefficiency. To address this limitation, we propose a novel Dual-Latent Space Reinforcement Learning (DLSRL) framework, which complements initial-noise steering with representation-level control inside the frozen generator. Specifically, our actor network predicts two distinct latent variables: an initial-noise latent variable that steers behavior generation, and an action-representation latent variable for intermediate feature modulation. Moreover, this representation latent variable is mapped to adapter features and ingeniously injected into the hidden states of intermediate action tokens via residual connections. Our dual-control design enables direct adjustment of action representations without updating the base policy. Experiments across generative policy architectures and robotic manipulation tasks show that DLSRL effectively accelerates online robot policy adaptation and achieves competitive performance. Our code is available at \href{https://github.com/xianchaoxiu/DLSRL}{https://github.com/xianchaoxiu/DLSRL}.

cs.RO

Do Video Generators Track the World Across Segments? A Benchmark and Method for World-State Reasoning in Video Continuation

Video generators build long videos by composing shorter parts, either by generating segments one after another or by autoregressively extending chunks. Each new part usually depends on memories of historical observations, such as recent frames, selected key frames, memory banks, or cached features. These memories preserve visible evidence from the past, but current generators do not reliably turn such evidence into a world-state interface: what holds in the video world after previous actions and how it should change under the next prompt. A past frame remains valid history, but it may not describe the state needed by the next segment; some states must instead be inferred from occluded or implicit changes rather than copied from a directly observed frame. This creates a simple but overlooked question for video continuation: given a previous video, its prompt, and a new prompt, can a model generate a continuation that reflects the state determined by both the historical video and the new prompt? To answer this question, we introduce Statebench, a benchmark that targets this gap by testing continuations over three state categories: past-visible states, occluded-process states, and complex-transition states. We further propose Stateagent, which explicitly maintains an entity-state representation, updates it under the new prompt, grounds the predicted post-action state as a future end frame, and renders the next video. Experiments show that our method improves controlled video continuation by raising the all-case state score (SCS-All) from 45.2 to 69.3, and also benefits story generation at the one-minute scale. Code is avaliable at https://github.com/AMAP-ML/StateAgent.

cs.CV

Reconstruction of anomalous air showers with SKA-Low

Double-bump showers are a surprising class of extensive air showers (EAS) predicted by Monte Carlo simulations, which, so far, no experiment has been able to directly detect. They occur when a high-energy secondary particle, the leading particle, travels significantly farther than the rest, creating a distinct double-peaked longitudinal profile. The unique radio footprint of double-bump showers, characterized by multiple pulses in the signals and interference patterns in the frequency spectra, enables reconstruction of longitudinal profiles from radio observations. With its dense antenna array and broad frequency range, SKA-Low will be the first observatory capable of detecting these features, offering a new opportunity to probe hadronic interactions and use the distinctive signatures of elements to provide new mass composition measurements. The goal of this analysis is to take the first steps toward using these radio signatures to reconstruct the relevant parameters of the longitudinal profile of a double-bump shower. We will start by explaining the radio signal of double-bump}showers compared to that of average showers. Then we will create a simple 2-point emission model to explain the interference patterns in the frequency spectra, which can be inverted to obtain rudimentary estimates of atmospheric depth of both peaks. Lastly, we implement a brute-force approach to reconstruct multiple parameters of the double bump.

astro-ph.HE

Beyond $X_\mathrm{max}$ : Reconstructing Air Shower Profiles with Information Field Theory with SKA-Low

While radio measurements of extensive air showers have shown to achieve a high precision of $X_\mathrm{max}$ sensitivity, it has been shown that parameters beyond $X_\mathrm{max}$ can also be reconstructed. These shape parameters contain additional sensitivity to the hadronic physics in the shower as well as its mass composition. In this work, we showcase a reconstruction framework to recover the full longitudinal profile from realistic radio measurements. The framework is based on Information Field Theory that infers the full profile with a forward-based model, which uses a Gaisser-Hillas profile with weakly informative shower priors, SMIET with a template library to synthesise pulses at any event geometry, and a realistic antenna response and noise level emulating that of SKA-Low. We verify the self-consistency of our framework with $\sim 900$ events generated with SMIET with antennas placed on the $\vec{v} \times (\vec{v} \times \vec{B})$ axis. The framework recovers the full profile within uncertainty and capture correlations between shower parameters. We yield an $X_\mathrm{max}$ resolution of $< 9$ g cm$^{-2}$ as well as resolutions of the width and asymmetry with minimal bias. The profile is also recovered with a bias of $< 4$% at all atmospheric depths $< 1200$ g cm$^{-2}$. We aim to apply this framework with pulses simulated from CoREAS with measured noise, ultimately extending the framework to realistic antenna layouts such as from LOFAR or SKA-Low.

astro-ph.IM

AudioLens: Multi-Perspective Speech Clustering with Reasoning Audio-Language Models

Audio clustering is a fundamental task for organizing rapidly growing speech collections, supporting applications such as conversational analysis and speech-driven discovery. However, existing methods rely on fixed acoustic similarity metrics or ASR-based text pipelines, limiting their ability to reorganize the same audio collection under different user-specified perspectives, especially when clustering depends on both linguistic and paralinguistic cues. We introduce audio multi-perspective clustering, where a model directly partitions speech recordings according to a natural-language perspective while inferring both the number of clusters and their assignments. To study this setting, we construct AudioLens-Bench, a benchmark spanning multiple application domains and evaluating both in-perspective and cross-perspective generalization. We further propose AudioLens-R1, an end-to-end large audio-language model trained with reasoning distillation and preference optimization. Experiments show that AudioLens-R1 consistently outperforms all baselines, improving overall ARI by 12.99 points and V-measure by 11.62 points. These results demonstrate the promise of native audio-language models for flexible, perspective-conditioned structure discovery over speech collections.

cs.SD

Multimodal Adaptive Control for Safe Robotic Craniotomy Under Partial Observability

Autonomous robotic craniotomy requires continuous regulation of tool-tissue interactions to mitigate mechanical overload and thermal damage while maintaining surgical efficiency. However, this process is inherently partially observable due to unknown, time-varying tissue properties and the inability to directly measure cutting temperatures under physical occlusion. To address these challenges, we propose RL-MACRO, a cybernetic closed-loop intelligence framework that couples multimodal perception, adaptive decision-making, and robotic execution. This framework empowers the surgical robot to autonomously perceive inaccessible states from partial sensory feedback and dynamically optimize its behaviors under uncertain environment. A CNN-LSTM observer first fuses force and sound feedback to reconstruct the hidden temperature state (R^2=0.939, MAE = 1.717 deg C). This reconstructed temperature, alongside multi-sensor features, forms the belief state for an offline Implicit Q-Learning (IQL) policy. A novel dual-head Actor dynamically coordinates the feed rate, spindle speed, and cutting depth to optimize efficiency within strict safety bounds. These decisions are seamlessly translated into spatial motions via online trajectory re-planning and velocity servoing. Experiments on bovine ribs and six ex vivo goat skulls validate the system's robust perception, adaptive recovery from force/temperature excursions, and smooth execution on irregular surfaces, establishing a data-driven cybernetic paradigm for safe and efficient autonomous bone cutting.

cs.RO

DreamX-Phi 1.0: Action-Conditioned Video World Model for Robotic Manipulation

We present \textbf{DreamX-Phi 1.0}, an action-conditioned video world model for robotic manipulation that, given an observed frame, a language instruction, and a prescribed action sequence comprising end-effector poses and gripper states, predicts the resulting future observations. Yet realism alone does not guarantee faithfulness: a convincing rollout can still move the wrong arm or lose the manipulated object. To ensure the prediction respects each arm's commanded path, we inject per-arm $\mathrm{SE}(3)$ transformations into attention via \textbf{PRoPE-style geometric encoding}, preserving arm identity and rigid-motion structure. Action control alone does not fully constrain scene geometry or the evolution of small manipulated objects. We therefore add a lightweight \textbf{depth branch} for scene-level geometry and use \textbf{SAM3 masks} with a frozen \textbf{V-JEPA teacher} to maintain object consistency throughout grasping. We further distill the multi-step generator into a few-step student via distribution-matching distillation for efficient deployment. At the time of writing, \model{} achieves first place on Track~1 and second place on Track~2 of the WorldArena~2.0 Challenge. Our model and code will be publicly available.

cs.CV

Evaluation-Verification Reward for Consistent Multi-Reference Image Editing

While recent image editing models have made rapid progress, multi-reference editing remains challenging, particularly in maintaining visual consistency across references and ensuring overall visual harmony. Reinforcement learning has proven highly effective for text-to-image generation and single-image editing, but its extension to multi-reference editing is hindered by the absence of suitable reward models that capture multi-image relational constraints. Moreover, naively using multimodal large language models(MLLMs) as zero-shot evaluators faces a key tension between hallucination-prone long-form reasoning and the limited deductive power of short-form judgments. We address these issues with a Multi-dimensional Evaluation-Verification Reward(EVR). EVR decomposes evaluation into distinct visual criteria; for each criterion, an MLLM Evaluator generates multiple candidate hypotheses, and a Verifier grounds each claim in concrete visual evidence to accept or reject it, producing reliable and fine-grained reward signals. Together with a scalable data pipeline, our method enables RL fine-tuning of off-the-shelf editors without architectural changes. Extensive experiments show substantial gains over the base Qwen-Image-Edit, improving consistency and harmony to match or surpass NanoBanana.

cs.CV

Robust Estimation of Sparse Numerical Vectors under Local Differential Privacy

Local differential privacy (LDP) protocols are vulnerable to poisoning attacks. Existing research have proposed efficient defense strategies for single-item users. However, in practice, a user may possess multiple items. The defense against poisoning attacks for multi-item users is challenging, because due to larger output spaces, the adversary can conduct more powerful attacks without being detected. In this paper, we address the robust sparse vector mean estimation problem, in which each user has a vector with $m$ nonzero coordinates. We propose Randomized Projection with Clipping (RPC). Firstly, the server sends a random binary vector to each user. The user then projects its local data on the vector, and clip the value to restrict the attacker's capability. To handle clipping bias, we propose a correction method based on a careful analysis that gives an exact expression of the bias. As a result, bias-variance tradeoff is no longer needed, thus the clipping threshold can be further reduced to shrink the output space and enhance robustness. We provide a rigorous theoretical guarantee of the estimation error under all possible attacks. Numerical experiments show that under trusted environments, our new method achieves comparable or better performance than existing methods, indicating that our method is already an efficient estimator in its own right. Under untrusted environments, our method is also significantly more robust to poisoning attacks.

stat.ML

Listen, Do Not Copy: Internalizing Audio-Grounded Scaffold Context for Robust Omni-Model Speech Understanding

Omni models transcribe clean, single-speaker speech well, but their accuracy drops sharply when speakers overlap and the scene is noisy, exactly where knowing who said what matters most. A natural fix is a short scene description. We show why this is risky: answer-bearing text lets the model copy instead of listen, so the score rises although nothing has been heard; a silent test exposes this shortcut at once. We call this failure mode perception bypass and address it with Audio-Grounded Scaffold Context (AGSC). AGSC links three steps: first, we build clues from audio to guide listening without giving the answer; second, answer-overlap and silence tests probe them for leakage and audio dependence; finally, those clues scaffold training but vanish at test time, yielding no-clue capability. Across three heterogeneous Omni models, training on AGSC lowers no-clue capped mean permutation word error rate (mpWER) on overlapping, noisy speech from 25%-71% to 9%-15%. For streaming control, we formulate a joint GDPO task in which the model learns when to use a clue and how to produce a speaker-attributed transcript from separately normalized format, gate, and transcript rewards. After internalization, AGSC adds almost no inference overhead.

cs.SD

AI Empowered Communication and Radar Modulation Recognition: A Survey

Automatic modulation recognition (AMR) is of vital importance for ensuring communication and radar reliability, efficient spectrum utilization and resistance to electronic interference. The development of artificial intelligence (AI) technology is reshaping the technological paradigm of AMR, promoting its transition from traditional modes relying on manual features to data-driven intelligent recognition. This change is not only reflected in the significant improvement of recognition accuracy, but also injects strong momentum into the intelligent evolution of both communication and radar systems through algorithm innovation, architecture optimization, and scenario expansion. In order to clarify the current development status and bottlenecks of AMR, and to find breakthrough directions, we make a comprehensive survey of recent AI-based technologies for AMR in this paper, including model-based machine learning (ML) methods and data-driven deep learning (DL) methods. We first investigate the modulation types used in current communication and radar systems. Next, we summarize the typically used features in the field of AMR, and discuss their inherent advantages and disadvantages. Then, we introduce the basic AI models for AMR and conduct a hierarchical investigation of AMR methods for communication and radar. Finally, based on existing research works, we highlight open issues and propose future research directions for AMR.

eess.SP

AG-REPA: Causal Layer Selection for Representation Alignment in Audio Flow Matching

REPresentation Alignment (REPA) improves the training of generative flow models by aligning intermediate hidden states with pretrained teacher features, but its effectiveness in token-conditioned audio Flow Matching critically depends on the choice of supervised layers, which is typically made heuristically based on the depth. In this work, we introduce Attribution-Guided REPresentation Alignment (AG-REPA), a novel causal layer selection strategy for representation alignment in audio Flow Matching. Firstly, we find that layers that best store semantic/acoustic information (high teacher-space similarity) are not necessarily the layers that contribute most to the velocity field that drives generation, and we call it Store-Contribute Dissociation (SCD). To turn this insight into an actionable training guidance, we propose a forward-only gate ablation (FoG-A) that quantifies each layer's causal contribution via the induced change in the predicted velocity field, enabling sparse layer selection and adaptive weighting for alignment. Across unified speech and general-audio training (LibriSpeech + AudioSet) under different token-conditioning topologies, AG-REPA consistently outperforms REPA baselines. Overall, our results show that alignment is most effective when applied to the causally dominant layers that drive the velocity field, rather than to layers that are representationally rich but functionally passive.

cs.SD

Human-Inspired Framework for Robotic Craniotomy: Integrating Multimodal Fusion and Adaptive Trajectory Adjustment

Manual craniotomy is a high-risk, skill-dependent procedure associated with surgeon fatigue and potential dural injury. While robotic approaches have improved safety, existing open-loop systems rely solely on preoperative images and cannot compensate for intraoperative registration errors or tissue deformation. To address this, we propose a human-inspired closed-loop robotic craniotomy framework that intelligently integrates preoperative planning with intraoperative execution. An adaptive dual-contour fusion algorithm is employed to generate trajectories that conform to complex cranial geometries while maintaining a consistent tool-bone relative pose. For intraoperative perception, a multimodal two-stage cross-modal attention block (CMA)-temporal convolutional network (TCN)-Transformer network combined with an adaptive Bayesian filter fuses force and acoustic signals to achieve robust breakthrough detection under varying bone conditions. Upon detection, an in-situ projection-based trajectory adjustment strategy dynamically compensates for depth deviations, enabling safe residual bone isolation. Experiments on bovine ribs show a breakthrough prediction accuracy of 97%, a detection latency of 0.048 +/- 0.097 s, and a maximum overshoot of 0.29 mm. All four ex vivo cranial experiments were successfully completed without dural injury. These results demonstrate that the proposed cybernetic framework enables safe and autonomous craniotomy with highly effective closed-loop control.

cs.RO

Detector-level simulation closure of radio air-shower reconstruction with native RF-chain inversion and out-of-fold endpoint interpolation

Autonomous radio arrays reconstruct air showers from trigger-selected digitized voltage traces rather than ideal electric-field footprints. We present a detector-level simulation-closure test linking RF-chain-triggered ZHAireS event packages to reconstructed direction, energy and quality diagnostics. Voltage amplitudes and station timing determine geometry through a robust joint timing--amplitude axis fit. Native-grid RF-chain inversion recovers 50--200 MHz electric-field peaks along the shower-plane v cross B axis. Symmetric four-arm iron and proton templates generated with 2 ns time bins provide endpoint estimates that are combined by a nested five-fold out-of-fold interpolation estimator. One estimator is used over the full angular range without a fitted multiplicative energy scale. Of 972 triggered event packages, 725 pass common quality selection in raw-noisy weighted, hard-gated denoised equal-weight and clean equal-weight branches. In the primary reconstructed-zenith range 60--85 degrees, 693 common events have mean energy residuals of -0.05, -0.30 and -0.66 percent with standard deviations of 11.04, 10.90 and 10.75 percent. The hard-gated denoised branch has a median angular separation of 0.052 degrees. Boundary zenith ranges are reported separately; the upper boundary has a mean energy residual of +14.2 percent and a 19.9 percent standard deviation. These results describe cross-fitted closure within a matched simulation framework, not deployed-array resolution or independent energy calibration.

astro-ph.IM

Discovery of $γ$-Ray Pulsations from the Extreme-Spin-Down Millisecond Pulsar PSR J0435+3233

Motivated by the recent discovery of PSR~J0435+3233, a millisecond pulsar with an exceptionally large period derivative of $\dot{P}=4.9\times10^{-17}\ {\rm s\,s^{-1}}$ and a high spin-down luminosity of $\dot{E}=5.89\times10^{37}\ {\rm erg\,s^{-1}}$, we analyze $\sim$17~yr of observations obtained with the \textit{Fermi} Large Area Telescope for the pulsar. We identify the cataloged source 4FGL~J0435.5+3232, located only $0.01^{\circ}$ from the radio timing position, as the $γ$-ray counterpart of PSR~J0435+3233. Using $γ$-ray events collected within the validity interval of the radio timing ephemeris, we detect its $γ$-ray pulsations at a $\sim6.8σ$ confidence level. In this interval, the phase-resolved analysis shows that the pulsed emission is concentrated predominantly within the rotational-phase of $ϕ\sim0.44$--$0.69$. We also derive its $γ$-ray luminosity of $L_γ=6.26\times10^{32}\ {\rm erg\,s^{-1}}$, assuming a distance of 1.2~kpc and isotropic emission. This luminosity corresponds to an apparent $γ$-ray efficiency of $η_γ\sim1.1\times10^{-5}$, revealing an exceptionally low $γ$-ray output despite the pulsar's young-pulsar-like rotational-energy budget. Our detection establishes PSR~J0435+3233 as a $γ$-ray MSP, and the striking combination of its high spin-down power and low apparent $γ$-ray efficiency provides a new probe of particle acceleration, radiation beaming, and viewing geometry in the magnetospheres of millisecond pulsars with extreme rotational properties.

astro-ph.HE

Emergence of the Scrooge Ensemble in the Sachdev-Ye-Kitaev Model

The probabilistic nature of quantum measurement provides a direct window into the structure and complexity of many-body wave functions. When only part of a system is measured, the remaining degrees of freedom form an ensemble of post-measurement states whose statistical structure can reveal a stronger form of thermalization, known as deep thermalization. Recent numerical evidence suggests that this phenomenon is characterized by convergence of the projected ensemble to the Scrooge ensemble, a maximally random ensemble compatible with a given density matrix. In this Letter, we use the solvable Sachdev-Ye-Kitaev (SYK) model to unveil the mechanism by which the Scrooge ensemble emerges in many-body systems. By formulating measurement probabilities and post-measurement states in terms of path integrals, we analytically characterize all moments of the projected ensemble and show that they exactly match those of the Scrooge ensemble, even at short evolution times. We further connect this result to the saddle-point structure of the measurement path integral, which naturally generates the replica permutations underlying Scrooge statistics. Our results establish the solvable SYK model as a tractable setting for exploring universal statistics of quantum measurements in chaotic many-body dynamics.

quant-ph