arXiv ScienceSearch

arXiv subjects

Shenyi Zhang

Publications and source records attributed to Shenyi Zhang.

17 recordsLinked to original sources

MMAligner: Safeguarding Multimodal Large Language Models through Representation Calibration

Multimodal large language models (MLLMs) often refuse unsafe text prompts yet generate harmful responses to semantically equivalent multimodal inputs. Existing defenses either rely on external guardrails, which add inference overhead without repairing intrinsic flaws, or safety fine-tuning, which treats alignment as black-box optimization and may sacrifice utility or require large multimodal datasets. To identify the cause of this safety disparity, we analyze MLLM representations geometrically. We find that safety mechanisms learned from text persist across modalities: a shared safety subspace and refusal boundary remain effective, and representations inside this boundary consistently trigger refusals. However, unsafe multimodal inputs undergo a representation shift that places most of them outside the boundary, allowing them to bypass the model's intrinsic safety mechanism. This indicates that multimodal safety degradation stems from representation misalignment rather than the absence of safety capability. Based on this finding, we propose MMAligner, a safeguarding method that calibrates unsafe multimodal representations into the pre-existing refusal region. MMAligner applies a hard lower bound to ensure refusal, a soft upper bound to avoid excessive modification, and a preservation objective for benign inputs. Experiments across multiple open-source MLLMs show that MMAligner raises the average refusal rate on unsafe multimodal inputs to 99% with less than 2% utility degradation and minimal training data, substantially improving the safety-utility trade-off over existing baselines. (*Due to the notification from arXiv, "The Abstract field cannot be longer than 1,920 characters", the Abstract that appeared is shortened.)

cs.CR

VOID: Defeating Unauthorized Mimicry in Latent Diffusion Models

While Latent Diffusion Models (LDMs) have revolutionized visual synthesis, they are increasingly exploited for unauthorized mimicry of individuals. Existing defenses inject deceptive perturbations to steer the generated images toward irrelevant targets. However, this approach hinges on an ungrounded assumption: subtle perturbations can maintain their deceptive efficacy throughout an LDM's extensive generation process. In reality, the model's innate restoration mechanism will remove such perturbations and cause individual identities to re-emerge in the images generated. We propose VOID, a defense framework that overcomes this conundrum by manipulating an LDM's intrinsic stochasticity. VOID perturbs the diffusion pipeline in two novel ways: 1) amplifying the latent encoding errors to shatter an image's semantic structure, and 2) counteracting the target guidance signals to suppress the model's restoration capabilities. This results in a semantic corruption that thwarts any unauthorized mimicry. Notably, the security gain does not come at the price of visual utility, as VOID simultaneously manages to confine perturbations to human-imperceptible regions of protected images. Our comprehensive evaluation of 24 state-of-the-art defenses against 10 mimicry attacks on 5 datasets demonstrates VOID's unprecedented protection power: it increases the average Frechet Inception Distance (FID) from 113 to 365, a 223% improvement over the strongest defense to date.

cs.CV

Sparse Tokens Suffice: Jailbreaking Audio Language Models via Token-Aware Gradient Optimization

Jailbreak attacks on audio language models (ALMs) optimize audio perturbations to elicit unsafe generations, and they typically update the entire waveform densely throughout optimization. In this work, we investigate the necessity of such dense optimization by analyzing the structure of token-aligned gradients in ALMs. We find that gradient energy is highly non-uniform across audio tokens, indicating that only a small subset of token-aligned audio regions dominates the optimization signal. Motivated by this observation, we propose Token-Aware Gradient Optimization (TAGO), which enables sparse jailbreak optimization by retaining only waveform gradients aligned with audio tokens that have high gradient energy, while masking the remaining gradients at each iteration. Across three ALMs, TAGO outperforms baselines, and substantial sparsification preserves strong attack success rates (e.g. on Qwen3-Omni, $\mathrm{ASR}_{l}$ remains at 86% with a token retention ratio of 0.25, compared to 87% with full token retention). These results demonstrate that dense waveform updates are largely redundant, and we advocate that future audio jailbreak and safety alignment research should further leverage this heterogeneous token-level gradient structure.

cs.CR

Sculpting of Martian brain terrain reveals the drying of ancient Mars

The Martian brain terrain (MBT), characterized by its unique brain-like morphology, is a potential geological archive for finding hints of paleoclimatic conditions during its formation period. The morphological similarity of MBT to self-organized patterned ground on Earth suggests a shared formation mechanism. However, the lack of quantitative descriptions and robust physical modeling of self-organized stone transport jointly limits the study of the thermal and aqueous conditions governing MBT's formation. Here we established a specialized quantitative system for extracting the morphological features of MBT, taking a typical region located in the northern Arabia Terra as an example, and then employed a numerical model to investigate its formation mechanisms. Our simulation results accurately replicate the observed morphology of MBT, matching its key geometric metrics with deviations <15%. Crucially, however, we find that the self-organized transport can solely produce relief <0.5 m, insufficient to explain the formation of MBT with average relief of 3.29 \pm 0.65 m. We attribute this discrepancy to sculpting driven by late-stage sublimation, constraining cumulative subsurface ice loss in this region to ~3 meters over the past ~3 Ma. These findings demonstrate that MBT's formation is a multi-stage process: initial patterning driven by freeze-thaw cycles implying liquid water followed by vertical sculpting via sublimation requiring a dry environment. This evolution provides physical evidence for the transition of the ancient Martian climate from a wetter period to a colder hyper-arid state.

physics.geo-ph

Amulet: Fast TEE-Shielded Inference for On-Device Model Protection

On-device machine learning (ML) introduces new security concerns about model privacy. Storing valuable trained ML models on user devices exposes them to potential extraction by adversaries. The current mainstream solution for on-device model protection is storing the weights and conducting inference within Trusted Execution Environments (TEEs). However, due to limited trusted memory that cannot accommodate the whole model, most existing approaches employ a partitioning strategy, dividing a model into multiple slices that are loaded into the TEE sequentially. This frequent interaction between untrusted and trusted worlds dramatically increases inference latency, sometimes by orders of magnitude. In this paper, we propose Amulet, a fast TEE-shielded on-device inference framework for ML model protection. Amulet incorporates a suite of obfuscation methods specifically designed for common neural network architectures. After obfuscation by the TEE, the entire transformed model can be securely stored in untrusted memory, allowing the inference process to execute directly in untrusted memory with GPU acceleration. For each inference request, only two rounds of minimal-overhead interaction between untrusted and trusted memory are required to process input samples and output results. We also provide theoretical proof from an information-theoretic perspective that the obfuscated model does not leak information about the original weights. We comprehensively evaluated Amulet using diverse model architectures ranging from ResNet-18 to GPT-2. Our approach incurs inference latency only 2.8-4.8x that of unprotected models with negligible accuracy loss, achieving an 8-9x speedup over baseline methods that execute inference entirely within TEEs, and performing approximately 2.2x faster than the state-of-the-art obfuscation-based method.

cs.CR

The Solar Close Observations and Proximity Experiments (SCOPE) mission

The Solar Close Observations and Proximity Experiments (SCOPE) mission will send a spacecraft into the solar atmosphere at a low altitude of just 5 R_sun from the solar center. It aims to elucidate the mechanisms behind solar eruptions and coronal heating, and to directly measure the coronal magnetic field. The mission will perform in situ measurements of the current sheet between coronal mass ejections and their associated solar flares, and energetic particles produced by either reconnection or fast-mode shocks driven by coronal mass ejections. This will help to resolve the nature of reconnections in current sheets, and energetic particle acceleration regions. To investigate coronal heating, the mission will observe nano-flares on scales smaller than 70 km in the solar corona and regions smaller than 40 km in the photosphere, where magnetohydrodynamic waves originate. To study solar wind acceleration mechanisms, the mission will also track the process of ion charge-state freezing in the solar wind. A key achievement will be the observation of the coronal magnetic field at unprecedented proximity to the solar photosphere. The polar regions will also be observed at close range, and the inner edge of the solar system dust disk may be identified for the first time. This work presents the detailed background, science, and mission concept of SCOPE and discusses how we aim to address the questions mentioned above.

astro-ph.SR

Selective Masking Adversarial Attack on Automatic Speech Recognition Systems

Extensive research has shown that Automatic Speech Recognition (ASR) systems are vulnerable to audio adversarial attacks. Current attacks mainly focus on single-source scenarios, ignoring dual-source scenarios where two people are speaking simultaneously. To bridge the gap, we propose a Selective Masking Adversarial attack, namely SMA attack, which ensures that one audio source is selected for recognition while the other audio source is muted in dual-source scenarios. To better adapt to the dual-source scenario, our SMA attack constructs the normal dual-source audio from the muted audio and selected audio. SMA attack initializes the adversarial perturbation with a small Gaussian noise and iteratively optimizes it using a selective masking optimization algorithm. Extensive experiments demonstrate that the SMA attack can generate effective and imperceptible audio adversarial examples in the dual-source scenario, achieving an average success rate of attack of 100% and signal-to-noise ratio of 37.15dB on Conformer-CTC, outperforming the baselines.

cs.CR

Competing Accelerated Failure Time Models for Multiple Concurrent Failure Mechanisms

The rising prevalence of complex diseases characterised by multiple coexisting and interacting etiological processes poses critical challenges for survival analysis and precision medicine, particularly as population ageing renders mutually exclusive models increasingly untenable. We propose a competing accelerated failure time (cAFT) framework to understand the individual-specific temporal dynamics of disease competition and interaction based on a first-to-fail principle. Specifically, we introduce an individualised, time-varying winning probability to quantify the relative contributions of latent causes and provide an interpretable basis for patient stratification within distinct subtypes. Consistency and asymptotic normality are established for the maximum likelihood estimation of the parameters, with practical implementation via an expectation-maximisation (EM) algorithm. We illustrate the model's effectiveness and efficiency through numerical simulations and real-world applications, including biomarker discovery for 28-day survival in sepsis and overall survival in lung adenocarcinoma. Compared with standard AFT and Cox proportional hazards models, the cAFT model consistently improves predictive accuracy (C-index and iAUC gains of 5--10%) and reveals subtype-dependent gene effects within distinct biological pathways across heterogeneous patient subgroups. Conclusively, the cAFT model provides deeper insights into patient prognosis and potential personalised therapeutic strategies.

stat.ME

JBShield: Defending Large Language Models from Jailbreak Attacks through Activated Concept Analysis and Manipulation

Despite the implementation of safety alignment strategies, large language models (LLMs) remain vulnerable to jailbreak attacks, which undermine these safety guardrails and pose significant security threats. Some defenses have been proposed to detect or mitigate jailbreaks, but they are unable to withstand the test of time due to an insufficient understanding of jailbreak mechanisms. In this work, we investigate the mechanisms behind jailbreaks based on the Linear Representation Hypothesis (LRH), which states that neural networks encode high-level concepts as subspaces in their hidden representations. We define the toxic semantics in harmful and jailbreak prompts as toxic concepts and describe the semantics in jailbreak prompts that manipulate LLMs to comply with unsafe requests as jailbreak concepts. Through concept extraction and analysis, we reveal that LLMs can recognize the toxic concepts in both harmful and jailbreak prompts. However, unlike harmful prompts, jailbreak prompts activate the jailbreak concepts and alter the LLM output from rejection to compliance. Building on our analysis, we propose a comprehensive jailbreak defense framework, JBShield, consisting of two key components: jailbreak detection JBShield-D and mitigation JBShield-M. JBShield-D identifies jailbreak prompts by determining whether the input activates both toxic and jailbreak concepts. When a jailbreak prompt is detected, JBShield-M adjusts the hidden representations of the target LLM by enhancing the toxic concept and weakening the jailbreak concept, ensuring LLMs produce safe content. Extensive experiments demonstrate the superior performance of JBShield, achieving an average detection accuracy of 0.95 and reducing the average attack success rate of various jailbreak attacks to 2% from 61% across distinct LLMs.

cs.CR

Zero-Query Adversarial Attack on Black-box Automatic Speech Recognition Systems

In recent years, extensive research has been conducted on the vulnerability of ASR systems, revealing that black-box adversarial example attacks pose significant threats to real-world ASR systems. However, most existing black-box attacks rely on queries to the target ASRs, which is impractical when queries are not permitted. In this paper, we propose ZQ-Attack, a transfer-based adversarial attack on ASR systems in the zero-query black-box setting. Through a comprehensive review and categorization of modern ASR technologies, we first meticulously select surrogate ASRs of diverse types to generate adversarial examples. Following this, ZQ-Attack initializes the adversarial perturbation with a scaled target command audio, rendering it relatively imperceptible while maintaining effectiveness. Subsequently, to achieve high transferability of adversarial perturbations, we propose a sequential ensemble optimization algorithm, which iteratively optimizes the adversarial perturbation on each surrogate model, leveraging collaborative information from other models. We conduct extensive experiments to evaluate ZQ-Attack. In the over-the-line setting, ZQ-Attack achieves a 100% success rate of attack (SRoA) with an average signal-to-noise ratio (SNR) of 21.91dB on 4 online speech recognition services, and attains an average SRoA of 100% and SNR of 19.67dB on 16 open-source ASRs. For commercial intelligent voice control devices, ZQ-Attack also achieves a 100% SRoA with an average SNR of 15.77dB in the over-the-air setting.

cs.CR

Boosting Adversarial Transferability with Low-Cost Optimization via Maximin Expected Flatness

Transfer-based attacks craft adversarial examples on white-box surrogate models and directly deploy them against black-box target models, offering model-agnostic and query-free threat scenarios. While flatness-enhanced methods have recently emerged to improve transferability by enhancing the loss surface flatness of adversarial examples, their divergent flatness definitions and heuristic attack designs suffer from unexamined optimization limitations and missing theoretical foundation, thus constraining their effectiveness and efficiency. This work exposes the severely imbalanced exploitation-exploration dynamics in flatness optimization, establishing the first theoretical foundation for flatness-based transferability and proposing a principled framework to overcome these optimization pitfalls. Specifically, we systematically unify fragmented flatness definitions across existing methods, revealing their imbalanced optimization limitations in over-exploration of sensitivity peaks or over-exploitation of local plateaus. To resolve these issues, we rigorously formalize average-case flatness and transferability gaps, proving that enhancing zeroth-order average-case flatness minimizes cross-model discrepancies. Building on this theory, we design a Maximin Expected Flatness (MEF) attack that enhances zeroth-order average-case flatness while balancing flatness exploration and exploitation. Extensive evaluations across 22 models and 24 current transfer-based attacks demonstrate MEF's superiority: it surpasses the state-of-the-art PGN attack by 4% in attack success rate at half the computational cost and achieves 8% higher success rate under the same budget. When combined with input augmentation, MEF attains 15% additional gains against defense-equipped models, establishing new robustness benchmarks. Our code is available at https://github.com/SignedQiu/MEFAttack.

cs.CV

Hijacking Attacks against Neural Networks by Analyzing Training Data

Backdoors and adversarial examples are the two primary threats currently faced by deep neural networks (DNNs). Both attacks attempt to hijack the model behaviors with unintended outputs by introducing (small) perturbations to the inputs. Backdoor attacks, despite the high success rates, often require a strong assumption, which is not always easy to achieve in reality. Adversarial example attacks, which put relatively weaker assumptions on attackers, often demand high computational resources, yet do not always yield satisfactory success rates when attacking mainstream black-box models in the real world. These limitations motivate the following research question: can model hijacking be achieved more simply, with a higher attack success rate and more reasonable assumptions? In this paper, we propose CleanSheet, a new model hijacking attack that obtains the high performance of backdoor attacks without requiring the adversary to tamper with the model training process. CleanSheet exploits vulnerabilities in DNNs stemming from the training data. Specifically, our key idea is to treat part of the clean training data of the target model as "poisoned data," and capture the characteristics of these data that are more sensitive to the model (typically called robust features) to construct "triggers." These triggers can be added to any input example to mislead the target model, similar to backdoor attacks. We validate the effectiveness of CleanSheet through extensive experiments on 5 datasets, 79 normally trained models, 68 pruned models, and 39 defensive models. Results show that CleanSheet exhibits performance comparable to state-of-the-art backdoor attacks, achieving an average attack success rate (ASR) of 97.5% on CIFAR-100 and 92.4% on GTSRB, respectively. Furthermore, CleanSheet consistently maintains a high ASR, when confronted with various mainstream backdoor defenses.

cs.CR

Solar Ring Mission: Building a Panorama of the Sun and Inner-heliosphere

Solar Ring (SOR) is a proposed space science mission to monitor and study the Sun and inner heliosphere from a full 360{\deg} perspective in the ecliptic plane. It will deploy three 120{\deg}-separated spacecraft on the 1-AU orbit. The first spacecraft, S1, locates 30{\deg} upstream of the Earth, the second, S2, 90{\deg} downstream, and the third, S3, completes the configuration. This design with necessary science instruments, e.g., the Doppler-velocity and vector magnetic field imager, wide-angle coronagraph, and in-situ instruments, will allow us to establish many unprecedented capabilities: (1) provide simultaneous Doppler-velocity observations of the whole solar surface to understand the deep interior, (2) provide vector magnetograms of the whole photosphere - the inner boundary of the solar atmosphere and heliosphere, (3) provide the information of the whole lifetime evolution of solar featured structures, and (4) provide the whole view of solar transients and space weather in the inner heliosphere. With these capabilities, Solar Ring mission aims to address outstanding questions about the origin of solar cycle, the origin of solar eruptions and the origin of extreme space weather events. The successful accomplishment of the mission will construct a panorama of the Sun and inner-heliosphere, and therefore advance our understanding of the star and the space environment that holds our life.

astro-ph.SR

Primary and albedo protons detected by the Lunar Lander Neutron and Dosimetry (LND) experiment on the lunar farside

The Lunar Lander Neutron and Dosimetry (LND) Experiment aboard the Chang$'$E-4 Lander on the lunar-far side measures energetic charged and neutral particles and monitors the corresponding radiation levels. During solar quiet times, galactic cosmic rays (GCRs) are the dominating component of charged particles on the lunar surface. Moreover, the interaction of GCRs with the lunar regolith also results in upward directed albedo protons which are measured by the LND. In this work, we used calibrated LND data to study the GCR primary and albedo protons. We calculate the averaged GCR proton spectrum in the range of 9 368 MeV and the averaged albedo proton flux between 64.7 and 76.7 MeV from June 2019 (the 7th lunar day after Chang$'$E-4$'$s landing) to July 2020 (the 20th lunar day). We compare the primary proton measurements of LND with the Electron Proton Helium INstrument (EPHIN) on SOHO. The comparison shows a reasonable agreement of the GCR proton spectra among different instruments and illustrates the capability of LND. Likewise, the albedo proton measurements of LND are also comparable with measurements by the Cosmic Ray Telescope for the Effects of Radiation (CRaTER) during solar minimum. Our measurements confirm predictions from the Radiation Environment and Dose at the Moon (REDMoon) model. Finally, we provide the ratio of albedo protons to primary protons for measurements in the energy range of 64.7-76.7 MeV which confirms simulations over a broader energy range.

physics.space-ph

Black-box Adversarial Attacks on Commercial Speech Platforms with Minimal Information

Adversarial attacks against commercial black-box speech platforms, including cloud speech APIs and voice control devices, have received little attention until recent years. The current "black-box" attacks all heavily rely on the knowledge of prediction/confidence scores to craft effective adversarial examples, which can be intuitively defended by service providers without returning these messages. In this paper, we propose two novel adversarial attacks in more practical and rigorous scenarios. For commercial cloud speech APIs, we propose Occam, a decision-only black-box adversarial attack, where only final decisions are available to the adversary. In Occam, we formulate the decision-only AE generation as a discontinuous large-scale global optimization problem, and solve it by adaptively decomposing this complicated problem into a set of sub-problems and cooperatively optimizing each one. Our Occam is a one-size-fits-all approach, which achieves 100% success rates of attacks with an average SNR of 14.23dB, on a wide range of popular speech and speaker recognition APIs, including Google, Alibaba, Microsoft, Tencent, iFlytek, and Jingdong, outperforming the state-of-the-art black-box attacks. For commercial voice control devices, we propose NI-Occam, the first non-interactive physical adversarial attack, where the adversary does not need to query the oracle and has no access to its internal information and training data. We combine adversarial attacks with model inversion attacks, and thus generate the physically-effective audio AEs with high transferability without any interaction with target devices. Our experimental results show that NI-Occam can successfully fool Apple Siri, Microsoft Cortana, Google Assistant, iFlytek and Amazon Echo with an average SRoA of 52% and SNR of 9.65dB, shedding light on non-interactive physical attacks against voice control devices.

cs.CR

First Solar energetic particles measured on the Lunar far-side

On 2019 May 6, the Lunar Lander Neutron & Dosimetry (LND) Experiment on board the Chang'E-4 on the far-side of the Moon detected its first small solar energetic particle (SEP) event with proton energies up to 21MeV. Combined proton energy spectra are studied based on the LND, SOHO/EPHIN and ACE/EPAM measurements which show that LND could provide a complementary dataset from a special location on the Moon, contributing to our existing observations and understanding of space environment. Velocity dispersion analysis (VDA) has been applied to the impulsive electron event and weak proton enhancement and the results demonstrate that electrons are released only 22 minutes after the flare onset and $\sim$15 minutes after type II radio burst, while protons are released more than one hour after the electron release. The impulsive enhancement of the in-situ electrons and the derived early release time indicate a good magnetic connection between the source and Earth. However, stereoscopic remote-sensing observations from Earth and STA suggest that the SEPs are associated with an active region nearly 100$^\circ$ away from the magnetic footpoint of Earth. This suggests that the propagation of these SEPs could not follow a nominal Parker spiral under the ballistic mapping model and the release and propagation mechanism of electrons and protons are likely to differ significantly during this event.

astro-ph.SR

The Lunar Lander Neutron and Dosimetry (LND) Experiment on Chang'E 4

Chang'E 4 is the first mission to the far side of the Moon and consists of a lander, a rover, and a relay spacecraft. Lander and rover were launched at 18:23 UTC on December 7, 2018 and landed in the von K\'arm\'an crater at 02:26 UTC on January 3, 2019. Here we describe the Lunar Lander Neutron \& Dosimetry experiment (LND) which is part of the Chang'E 4 Lander scientific payload. Its chief scientific goal is to obtain first active dosimetric measurements on the surface of the Moon. LND also provides observations of fast neutrons which are a result of the interaction of high-energy particle radiation with the lunar regolith and of their thermalized counterpart, thermal neutrons, which are a sensitive indicator of subsurface water content.

astro-ph.IM