arXiv ScienceSearch

arXiv subjects

Jun Jiang

Publications and source records attributed to Jun Jiang.

At least 19 recordsLinked to original sources

A computable representation of the physical laboratory enables verifiable workflows

Making science computable requires representations of both scientific knowledge and the physical world in which scientific claims are tested. A computable representation of the physical laboratory is established through typed research objects, capability-bound operations and a compositional workflow algebra. It provides the physical-world counterpart to machine-readable knowledge, expressing workflows as programs over evolving laboratory states with explicit dependencies, decisions, iteration and concurrency. The representation was implemented in a modular agentic robotic laboratory by binding formal operations to executable Function Skills. For diverse scientific intents, capability-relative workflows were generated, while stateful simulation propagated object transformations and verified operation preconditions and laboratory constraints before dispatch. The proposed representation and its engineering framework jointly establish a general computational interface between agent reasoning and capability-bound physical transformations, providing a foundation for end-to-end autonomous scientific discovery.

cs.AI

Stress-testing large language model agents in a robotic chemistry laboratory

AI is evaluated through knowledge, reasoning and plan generation, yet scientific agency requires reliable physical action and adaptation to evidence. Here, we use a robotic chemistry laboratory as a physical-world testbed to make scientific agency measurable. Its 45 modular workstations exposed as machine-readable skills enabled 4,608 trials. Only 3.3% of trials produced expert-assessed executable workflows under laboratory constraints; even the best system achieved 28.1%. Long-horizon planning remained a challenge: only three executable workflows exceeded 30 operations, although the longest contained 44. Across five rounds, experimental feedback prompted local adjustments but no workflow-level replanning or analytical-method redesign. By making physical executability and evidence-driven replanning measurable, our study provides an evidence-based assessment of deployment readiness and a diagnostic framework to guide closed-loop improvements towards physically grounded autonomous research.

cs.AI

JEPA-CFM: A Joint Embedding Predictive Architecture-based Channel Foundation Model for Robust Fluid Antenna Systems

Fluid antenna systems (FAS) have emerged as a promising technology for sixth-generation (6G) wireless networks. By allowing antenna elements to move freely within a compact region, FAS can exploit rich spatial diversity without additional hardware. However, acquiring real-time channel state information (CSI), extrapolating channel values to unmeasured antenna ports, and determining accurate user positions remain major obstacles. These challenges stem mainly from strong spatial correlations within the limited aperture and the scarcity of observable data. To overcome these limitations, this paper introduces joint embedding predictive architecture (JEPA)-based channel foundation model (CFM) specifically designed for FAS. The model adopts JEPA to learn versatile representations by extracting high-level latent embeddings of masked or unobserved channel segments. Unlike conventional approaches that attempt pixel-by-pixel reconstruction of raw CSI coefficients, JEPA-CFM focuses on predicting abstract structures in a compact feature space. The pre-training objective combines three complementary loss terms: the standard masked autoencoder reconstruction loss, the JEPA latent prediction loss, and a sliced isotropic Gaussian regularization (SIGReg) term. Together, these components prevent representation collapse and significantly enhance robustness under severe spatial correlation and highly sparse observations. After pre-training, the encoder is frozen, and lightweight task-specific heads are attached: a decoder for channel extrapolation and a global average pooling layer followed by a multi-layer perceptron regression head for wireless positioning. Extensive simulations in the realistic DeepMIMO urban scenario demonstrate that JEPA-CFM substantially outperforms the conventional masked autoencoder baseline in channel extrapolation and wireless positioning.

eess.SP

Between Safe Boundaries: Exploiting Temporal Consistency for Jailbreaking Text-To-Video Generation Models

Recently, text-to-video (T2V) models have been widely deployed, sparking growing concerns over their robustness against jailbreak attacks. Existing jailbreak methods, mostly adapted from text-to-image attacks, suffer notable drawbacks when applied to T2V systems. They fail to fully leverage temporal consistency, an inherent characteristic of video generation. Besides, these methods demand heavy video query optimization, which is infeasible in practical black-box scenarios. Their adversarial prompt search is also driven by heuristic local signals, lacking principled structured exploration strategies. To tackle these limitations, we propose BSB, a structured, query-efficient jailbreak framework for T2V models. BSB harnesses temporal consistency by encoding harmful intent as the transition between two individually harmless boundary states. Under this paradigm, the attack targets boundary-state pairs whose interpolation tends to produce unsafe intermediate frames during video generation. Directly evaluating all candidate pairs within the video space incurs prohibitive computation cost. Instead, BSB conducts Monte Carlo Tree Search (MCTS) in a cheaper textual proxy space and regularly calibrates search outcomes with sparse video-level evaluations. We conduct comprehensive experiments on mainstream commercial T2V models including Veo 3.1, Sora 2, Seedance and Kling v1. Results show BSB surpasses all existing jailbreak baselines, delivering an average 18.6% relative gain in attack success rate over the strongest competitor across evaluated models. Our findings identify temporal consistency as an understudied yet vital attack surface for T2V models and verify that structured search facilitates effective vulnerability discovery under constrained query budgets.

cs.CR

CFM-Bench: A Unified Multi-Domain, Multi-Task Benchmark for Channel Foundation Models

Channel foundation models (CFMs) are commonly evaluated in model-specific pipelines that differ in data, radio configurations, partitions, adaptation procedures, task definitions, and metrics, preventing reproducible comparison across CFMs and against task-specific networks. We release CFM-Bench, a unified multi-domain, multi-task benchmark comprising 157,900 official single-frame examples from six domains spanning 3GPP statistical simulation, two ray-tracing pipelines, terrestrial and aerial measurements, and synchronized vehicular multimodal simulation. Source-specific interfaces preserve complex channel state information (CSI) and the physical metadata available in each domain while allowing documented model-specific preprocessing. To reduce spatio-temporal leakage, official partitions isolate complete trajectories, measurement sessions, flights, vehicle links, simulation realizations, or buffered spatial regions. CFM-Bench excludes all benchmark splits from foundation-model pretraining, reserves the official training split for downstream fine-tuning, and reports the data used during model development. Six task groups across PHY, RAN, and ISAC applications cover CSI feedback, frequency and temporal channel extrapolation, propagation-state classification, current- and future-beam prediction, and single-frame and temporal localization. Representative experiments on CSI feedback, channel extrapolation, current-beam prediction, and wireless positioning provide reproducible reference results for pretrained channel-prediction models and task-specific neural networks. These results show that relative model performance can vary across data domains, highlighting the importance of using common data partitions, task definitions, and evaluation metrics. CFM-Bench provides a common substrate for evaluating the transferability of channel representations across models, domains, and tasks.

cs.AI

Language models guide symbolic equation discovery by controlling search

Scientific equation discovery must combine broad domain priors with strict numerical testing. Symbolic regression supplies numerical grounding but faces a combinatorial search space, whereas many language-model systems ask the model to propose or select formulas directly. We test a different division of labour. We compare role specifications in which the language model acts as equation author, candidate decider or search controller, alongside end-to-end language-model and purely numerical baselines. In the controller setting we propose here, implemented as LLM-PySR, language models specify variables, operators, transformations and search depth; symbolic regression enumerates and fits expressions; and deterministic metrics govern retention. Across 74 AI-Feynman equations and seven complex formula-recovery tasks, search control achieved the strongest observed balance of accuracy, complexity, stability and cost. On an independent battery dataset, LLM-PySR identified a compact piecewise-linear relation between early voltage-curve displacement and cycle life. The results suggest that language models should shape hypothesis exploration rather than decide which equations survive.

cs.AI

Labimus: A Simulation and Benchmark for Humanoid Dexterous Manipulation in Chemical Laboratory

Laboratory automation has made remarkable progress through robotic platforms and AI-driven scientific reasoning. However, many laboratory operations (e.g., solid--solid transfer) remain inherently dynamic and require real-time adaptation to different materials and experimental conditions. Such precision-critical manipulations are difficult to standardize, motivating the use of humanoid robots with dexterous hands. Despite this opportunity, no existing benchmark evaluates humanoid manipulation in precision-critical laboratory environments. We present Labimus, to our knowledge, the first benchmark for humanoid dexterous manipulation in organic chemistry laboratories. Labimus reconstructs over 30 functionally faithful assets from real organic chemistry workstations through real-to-sim modeling, collectively covering the core operations of routine organic chemistry experiments. The benchmark integrates articulated laboratory instruments, particle-based powder physics, and closed-loop instrument readouts, enabling a complete manipulation-to-measurement pipeline. It further defines six atomic operations and a seven-step solid-weighing workflow derived from real laboratory standard operating procedures. We introduce a precision-aware evaluation protocol designed to jointly measure task completion, experimental precision, and long-horizon execution. We benchmark three representative policies under procedural layouts and environmental perturbations. Results reveal a precision gap: policies that successfully complete laboratory tasks can still fail to satisfy the quantitative tolerances required by experimental protocols. Our benchmark exposes a fundamental disconnect between task completion and experimental validity, providing a new testbed for developing reliable humanoid robots for scientific laboratories.

cs.RO

CSI-CLIP++: A Scalable Channel Foundation Model for Wireless Communication via CIR-CSI Consistency

Self-supervised learning can exploit large-scale unlabeled channel data to improve the transferability of wireless AI models. Existing channel foundation models are often built on single-domain representations or reconstruction-oriented objectives, which may not explicitly capture the physical correspondence between frequency- and delay-domain channel views. This paper proposes CSI-CLIP++, a scalable channel foundation model for MIMO wireless channels. CSI-CLIP++ treats frequency-domain channel state information (CSI) and delay-domain channel impulse response (CIR) as paired views of the same propagation process and learns transferable representations through CSI-CIR contrastive alignment. The pretrained CSI encoder is adapted to channel identification, beam prediction, and positioning, representing PHY, RAN, and ISAC applications. Experiments on large-scale DeepMIMO scenarios show consistent gains over supervised baselines across environments, carrier frequencies, and data scales. CSI-CLIP++ improves beam prediction Top-1 accuracy by up to 19.31 percentage points and achieves competitive positioning performance, including cross-simulator transfer on a Sionna RT dataset. Backbone scaling results further show that the proposed objective remains effective across encoder architectures and benefits from larger model capacity.

eess.SP

6G Native AI and Channel Foundation Models

The integration of artificial intelligence (AI) and wireless communications is widely regarded as a core objective of sixth-generation (6G) systems. However, both the meaning of native AI and the type of AI capability that should be embedded into future wireless systems remain open to interpretation. This paper discusses 6G native AI from a system-design perspective and argues that native AI should be co-designed, optimized, and deployed as an intrinsic component of the wireless system rather than as a removable post-deployment add-on. From this perspective, conventional task-specific supervised models are difficult to use as the main technical basis of native AI because they depend heavily on labeled data, generalize poorly across propagation conditions, and require fragmented designs for different channel-related tasks. Motivated by these limitations, we position channel foundation models (CFMs) as a channel-centric foundation-model paradigm for 6G native AI. We define the scope of CFMs, clarify their differences from task-specific wireless AI models and large language models, and summarize three pretraining families: generative, discriminative, and hybrid pretraining. We further discuss how CFMs may support physical-layer processing, radio access network intelligence, and integrated sensing and communications. Preliminary CSI-CLIP-based results are included as bounded evidence that CFM-style pretraining can improve positioning and beam prediction when task-specific labels are limited.

eess.SP

Spin correlations and quantum entanglement in $\gamma\gamma \to t\bar t$ at polarized photon colliders with NLO QCD corrections

We study spin correlations and quantum entanglement in the process $\gamma\gamma\to t\bar t$ with the photons coming from Compton backscattered laser beam. We present predictions for the cross sections and spin observables including next-to-leading order (NLO) QCD corrections under various beam-polarization configurations. The NLO QCD corrections significantly enhance the total cross section while having only a minor impact on spin observables. Using the spin density matrix of the $t\bar t$ system, we further investigate quantum entanglement and Bell nonlocality. We find that the entanglement is the strongest near the $t\bar t$ invariant-mass threshold, while both entanglement and Bell nonlocality are highly sensitive to the initial beam polarization. Our results provide a theoretical basis for future studies of quantum correlations in top-quark pair production at photon colliders.

hep-ph

Non-abelian Extensions of Lie algebras with derivations

In this paper, we investigate non-abelian extensions of Lie algebras with derivations from several different perspectives. We show that the theory of non-abelian extensions of a Lie algebra with a derivation can be characterized by means of the second non-abelian cohomology, the Deligne groupoid, the homotopy category of strict Lie $2$-algebras with strict derivations, and the notion of a $(\mathfrak{g}, D)$-kernel, respectively. Moreover, within this unified framework, we address the following existence problem: given a non-abelian extension of Lie algebras $$ 0\longrightarrow\mathfrak{h}\overset{i}\longrightarrow\hat{\mathfrak{g}}\overset{p}\longrightarrow\mathfrak{g}\longrightarrow 0, $$ let $(K,D)\in\mathrm{Der}(\mathfrak{h})\times\mathrm{Der}(\mathfrak{g})$ be a pair of derivations of $\mathfrak{h}$ and $\mathfrak{g}$ respectively. When does there exist a derivation $\hat{D}$ of $\hat{\mathfrak{g}}$ such that $\hat{D}|_{\mathfrak{h}}=K$ and $D\circ p=p\circ\hat{D}.$ We provide an obstruction class for the existence of such a lift.

math.RA

Efficient and Principled Scientific Discovery through Bayesian Optimization: A Tutorial

Traditional scientific discovery relies on an iterative hypothesise-experiment-refine cycle that has driven progress for centuries, but its intuitive, ad-hoc implementation often wastes resources, yields inefficient designs, and misses critical insights. This tutorial presents Bayesian Optimisation (BO), a principled probability-driven framework that formalises and automates this core scientific cycle. BO uses surrogate models (e.g., Gaussian processes) to model empirical observations as evolving hypotheses, and acquisition functions to guide experiment selection, balancing exploitation of known knowledge and exploration of uncharted domains to eliminate guesswork and manual trial-and-error. We first frame scientific discovery as an optimisation problem, then unpack BO's core components, end-to-end workflows, and real-world efficacy via case studies in catalysis, materials science, organic synthesis, and molecule discovery. We also cover critical technical extensions for scientific applications, including batched experimentation, heteroscedasticity, contextual optimisation, and human-in-the-loop integration. Tailored for a broad audience, this tutorial bridges AI advances in BO with practical natural science applications, offering tiered content to empower cross-disciplinary researchers to design more efficient experiments and accelerate principled scientific discovery.

cs.LG

Learning to See through Illumination Extremes with Event Streaming in Multimodal Large Language Models

Multimodal Large Language Models (MLLMs) perform strong vision-language reasoning under standard conditions but fail in extreme illumination, where RGB inputs lose irrevocable structure and semantics. We propose Event-MLLM, an event-enhanced model that performs all-light visual reasoning by dynamically fusing event streams with RGB frames. Two key components drive our approach: an Illumination Indicator - a learnable signal derived from a DINOv2 branch that represents exposure degradation and adaptively modulates event-RGB fusion - and an Illumination Correction Loss that aligns fused features with non-degraded (normal-light) semantics in the latent space, compensating for information lost in extreme lighting. We curate the first multi-illumination event-instruction corpus for MLLMs, with 2,241 event-RGB samples (around 6 QA pairs each) across diverse scenes and 17 brightness rates (0.05x - 20x), plus an instruct-following benchmark for reasoning, counting, and fine-grained recognition under extreme lighting. Experiments show that Event-MLLM markedly outperforms general-purpose, illumination-adaptive, and event-only baselines, setting a new state of the art in robust multimodal perception and reasoning under challenging illumination.

cs.CV

AI/ML for mobile networks: Current status in Rel. 19 and challenges ahead

The transformative power of artificial intelligence (AI) and machine learning (ML) is recognized as a key enabler for sixth generation (6G) mobile networks by both academia and industry. Research on AI/ML in mobile networks has been ongoing for years, and the 3rd generation partnership project (3GPP) launched standardization efforts to integrate AI into mobile networks. However, a comprehensive review of the current status and challenges of the standardization of AI/ML for mobile networks is still missing. To this end, we provided a comprehensive review of the standardization efforts by 3GPP on AI/ML for mobile networks. This includes an overview of the general AI/ML framework, representative use cases (i.e., CSI feedback, beam management and positioning), and corresponding evaluation matrices. We emphasized the key research challenges on dataset preparation, generalization evaluation and baseline AI/ML models selection. Using CSI feedback as a case study, given the test dataset 2, we demonstrated that the pre-training-fine-tuning paradigm (i.e., pre-training using dataset 1 and fine-tuning using dataset 2) outperforms training on dataset 2. Moreover, we observed the highest performance enhancements in Transformer-based models through fine-tuning, showing its great generalization potential at large floating-point operations (FLOPs). Finally, we outlined future research directions for the application of AI/ML in mobile networks.

eess.SP

Explainable Interictal Epileptiform Discharge Detection Method Based on Scalp EEG and Retrieval-Augmented Generation

The detection of interictal epileptiform discharge (IED) is crucial for the diagnosis of epilepsy, but automated methods often lack interpretability. This study proposes IED-RAG, an explainable multimodal framework for joint IED detection and report generation. Our approach employs a dual-encoder to extract electrophysiological and semantic features, aligned via contrastive learning in a shared EEG-text embedding space. During inference, clinically relevant EEG-text pairs are retrieved from a vector database as explicit evidence to condition a large language model (LLM) for the generation of evidence-based reports. Evaluated on a private dataset from Wuhan Children's Hospital and the public TUH EEG Events Corpus (TUEV), the framework achieved balanced accuracies of 89.17\% and 71.38\%, with BLEU scores of 89.61\% and 64.14\%, respectively. The results demonstrate that retrieval of explicit evidence enhances both diagnostic performance and clinical interpretability compared to standard black-box methods.

eess.SP

WMVLM: Evaluating Diffusion Model Image Watermarking via Vision-Language Models

Digital watermarking is essential for securing generated images from diffusion models. Accurate watermark evaluation is critical for algorithm development, yet existing methods have significant limitations: they lack a unified framework for both residual and semantic watermarks, provide results without interpretability, neglect comprehensive security considerations, and often use inappropriate metrics for semantic watermarks. To address these gaps, we propose WMVLM, the first unified and interpretable evaluation framework for diffusion model image watermarking via vision-language models (VLMs). We redefine quality and security metrics for each watermark type: residual watermarks are evaluated by artifact strength and erasure resistance, while semantic watermarks are assessed through latent distribution shifts. Moreover, we introduce a three-stage training strategy to progressively enable the model to achieve classification, scoring, and interpretable text generation. Experiments show WMVLM outperforms state-of-the-art VLMs with strong generalization across datasets, diffusion models, and watermarking methods.

cs.CV

ReinPath: A Multimodal Reinforcement Learning Approach for Pathology

Interpretability is significant in computational pathology, leading to the development of multimodal information integration from histopathological image and corresponding text data.However, existing multimodal methods have limited interpretability due to the lack of high-quality dataset that support explicit reasoning and inference and simple reasoning process.To address the above problems, we introduce a novel multimodal pathology large language model with strong reasoning capabilities.To improve the generation of accurate and contextually relevant textual descriptions, we design a semantic reward strategy integrated with group relative policy optimization.We construct a high-quality pathology visual question answering (VQA) dataset, specifically designed to support complex reasoning tasks.Comprehensive experiments conducted on this dataset demonstrate that our method outperforms state-of-the-art methods, even when trained with only 20% of the data.Our method also achieves comparable performance on downstream zero-shot image classification task compared with CLIP.

cs.CV

Where to find $X(17)$?

The Atomki anomaly puts forward the hypothesis of an $X(17)$ particle to explain its observation. Utilizing experimental results from the Atomki experiments, measurements of the electron's anomalous magnetic moment, beam dump experiments, the KLOE-2 experiment, the PADME experiment, and the parity-violating M{\o}ller scattering experiment, we derive constraints on the $Xee$ coupling of the $X(17)$ boson to electrons. It is found that the scalar and pseudoscalar models can be excluded by Atomki experiments due to the parity conservation, and the pure axial-vector model is excluded at 98\% C.L. Meanwhile, the analyses in both pure vector and the vector $\pm$ axial-vector models consistently show that the $Xee$ coupling is of the vector type and has an almost fixed value, $\left(6.78 \pm 0.042\right) \times 10^{-4} \lesssim |\varepsilon_e^v| \lesssim \left(6.93 \pm 1.66\right) \times 10^{-4}$ in unit of electric charge $e$.

hep-ph