arXiv ScienceSearch

arXiv subjects

Yuan Gan

Publications and source records attributed to Yuan Gan.

At least 19 recordsLinked to original sources

FedOT: Ownership Verification and Leakage Tracing via Watermarks for Federated LDMs

Training Latent Diffusion Models (LDMs) within Federated Learning (FL) has attracted increasing attention due to its ability to combine the powerful generative capacity of LDMs with the privacy-preserving properties of FL. However, FL requires sharing the global model with multiple participants, which risks unauthorized model distribution or resale by malicious clients. While an intuitive approach is to adopt existing VAE-based watermarking techniques for LDMs in FL, this strategy falls short in addressing such threats due to two fundamental challenges: (1) Existing methods support ownership verification but lack the ability to trace model leakage to a specific malicious client; (2) VAE-based watermarks are vulnerable, as they can be removed simply by replacing the decoder with a clean counterpart. In this paper, we propose FedOT, the first framework for ownership verification and leakage tracing in federated LDMs. Specifically, to address the first challenge, we design a chunked watermark, where the first part is for ownership verification, and the second part is used for client identification. Furthermore, to overcome the second challenge and secure the model against VAE replacement attack, we introduce Latent Vector Transformation (LVT), which strengthens the connection between the VAE and U-Net latent spaces by modifying the original latent distribution of the VAE. Consequently, any attempt to replace the VAE for watermark removal leads to significant image quality degradation, making the LDM model unusable. Extensive experiments demonstrate that FedOT achieves superior performance in both ownership verification and traceability. Project page: https://spyzixuan.github.io/FedOT/.

cs.CV

CapTalk: Text-Guided Stylization and Speech-Driven 3D Head Animation

Audio-driven 3D facial animation aims to generate synchronized lip movements and vivid facial expressions from arbitrary audio clips. While existing methods can produce synchronized lip motions, they often rely on predefined identity or style latent features, which limits users' ability to freely control speaking styles. Moreover, applying a fixed style or identity to an entire audio segment typically results in facial animation styles that do not adapt to the emotional content of the audio. To address these challenges, we revisit the entanglement between style and emotion, construct a large-scale dataset with textual descriptions of both style and emotion, and propose a novel talking head generation framework that enables separate control over style and emotion. Our model takes as input both textual descriptions of speaking style and character emotion, as well as the driving audio stream, enabling real-time generation of highly synchronized lip movements and facial expressions that match the provided descriptions. Furthermore, our model supports dynamic emotion control during inference, allowing it to handle scenarios where the target emotion changes throughout the speech.

cs.CV

NTIRE 2026 3D Restoration and Reconstruction in Real-world Adverse Conditions: RealX3D Challenge Results

This paper presents a comprehensive review of the NTIRE 2026 3D Restoration and Reconstruction (3DRR) Challenge, detailing the proposed methods and results. The challenge seeks to identify robust reconstruction pipelines that are robust under real-world adverse conditions, specifically extreme low-light and smoke-degraded environments, as captured by our RealX3D benchmark. A total of 279 participants registered for the competition, of whom 33 teams submitted valid results. We thoroughly evaluate the submitted approaches against state-of-the-art baselines, revealing significant progress in 3D reconstruction under adverse conditions. Our analysis highlights shared design principles among top-performing methods and provides insights into effective strategies for handling 3D scene degradation.

cs.CV

RealX3D: A Physically-Degraded 3D Benchmark for Multi-view Visual Restoration and Reconstruction

We introduce RealX3D, a real-capture benchmark for multi-view visual restoration and 3D reconstruction under diverse physical degradations. RealX3D groups corruptions into four families, including illumination, scattering, occlusion, and blurring, and captures each at multiple severity levels using a unified acquisition protocol that yields pixel-aligned LQ/GT views. Each scene includes high-resolution capture, RAW images, and dense laser scans, from which we derive world-scale meshes and metric depth. Benchmarking a broad range of optimization-based and feed-forward methods shows substantial degradation in reconstruction quality under physical corruptions, underscoring the fragility of current multi-view pipelines in real-world challenging environments.

cs.CV

Silence is Golden: Leveraging Adversarial Examples to Nullify Audio Control in LDM-based Talking-Head Generation

Advances in talking-head animation based on Latent Diffusion Models (LDM) enable the creation of highly realistic, synchronized videos. These fabricated videos are indistinguishable from real ones, increasing the risk of potential misuse for scams, political manipulation, and misinformation. Hence, addressing these ethical concerns has become a pressing issue in AI security. Recent proactive defense studies focused on countering LDM-based models by adding perturbations to portraits. However, these methods are ineffective at protecting reference portraits from advanced image-to-video animation. The limitations are twofold: 1) they fail to prevent images from being manipulated by audio signals, and 2) diffusion-based purification techniques can effectively eliminate protective perturbations. To address these challenges, we propose Silencer, a two-stage method designed to proactively protect the privacy of portraits. First, a nullifying loss is proposed to ignore audio control in talking-head generation. Second, we apply anti-purification loss in LDM to optimize the inverted latent feature to generate robust perturbations. Extensive experiments demonstrate the effectiveness of Silencer in proactively protecting portrait privacy. We hope this work will raise awareness among the AI security community regarding critical ethical issues related to talking-head generation techniques. Code: https://github.com/yuangan/Silencer.

cs.GR

Efficient Emotional Adaptation for Audio-Driven Talking-Head Generation

Audio-driven talking-head synthesis is a popular research topic for virtual human-related applications. However, the inflexibility and inefficiency of existing methods, which necessitate expensive end-to-end training to transfer emotions from guidance videos to talking-head predictions, are significant limitations. In this work, we propose the Emotional Adaptation for Audio-driven Talking-head (EAT) method, which transforms emotion-agnostic talking-head models into emotion-controllable ones in a cost-effective and efficient manner through parameter-efficient adaptations. Our approach utilizes a pretrained emotion-agnostic talking-head transformer and introduces three lightweight adaptations (the Deep Emotional Prompts, Emotional Deformation Network, and Emotional Adaptation Module) from different perspectives to enable precise and realistic emotion controls. Our experiments demonstrate that our approach achieves state-of-the-art performance on widely-used benchmarks, including LRW and MEAD. Additionally, our parameter-efficient adaptations exhibit remarkable generalization ability, even in scenarios where emotional training videos are scarce or nonexistent. Project website: https://yuangan.github.io/eat/

cs.SD

Gate-tuned ambipolar superconductivity with strong pairing interaction in intrinsic gapped monolayer 1T'-MoTe2

Gate tunable two-dimensional (2D) superconductors offer significant advantages when studying superconducting phase transitions. Here, we address superconductivity in exfoliated 1T'-MoTe2 monolayers with an intrinsic band gap of ~7.3 meV using electrostatic doping. Despite large differences in the dispersion of the conduction and the valence bands, superconductivity can be achieved easily for both electrons and holes. The onset of superconductivity occurs near 7-8K for both charge carrier types. This temperature is much higher than in bulk samples. Also the in-plane upper critical field is strongly enhanced and exceeds the BCS Pauli limit in both cases. Gap information is extracted using point-contact spectroscopy. The gap ratio exceeds multiple times the value expected for BCS weak-coupling. All these observations suggest a strong enhancement of the pairing interaction.

cond-mat.supr-con

Enhanced low-energy magnetic excitations evidencing the Cu-induced localization in an Fe-based superconductor Fe$_{0.98}$Te$_{0.5}$Se$_{0.5}$

We have performed inelastic neutron scattering measurements on optimally-doped Fe$_{0.98}$Te$_{0.5}$Se$_{0.5}$ and 10% Cu-doped Fe$_{0.88}$Cu$_{0.1}$Te$_{0.5}$Se$_{0.5}$ to investigate the substitution effects on the spin excitations in the whole energy range up to 300 meV. It is found that substitution of Cu for Fe enhances the low-energy spin excitations ($\le$ 100 meV), especially around the (0.5, 0.5) point, and leaves the high-energy magnetic excitations intact. In contrast to the expectation that Cu with spin 1/2 will dilute the magnetic moments contributed by Fe with a larger spin, we find that the 10% Cu doping enlarges the effective fluctuating moment from 2.85 to 3.13 $\mu_{\rm B}$/Fe, although there is no long- or short-range magnetic order around (0.5, 0.5) and (0.5, 0). The presence of enhanced magnetic excitations in the 10% Cu doped sample which is in the insulating state indicates that the magnetic excitations must have some contributions from the local moments, reflecting the dual nature of the magnetism in iron-based superconductors. We attribute the substitution effects to the localization of the itinerant electrons induced by Cu dopants. These results also indicate that the Cu doping does not act as electron donor as in a rigid-band shift model, but more as scattering centers that localize the system.

cond-mat.supr-con

VidFace: A Full-Transformer Solver for Video FaceHallucination with Unaligned Tiny Snapshots

In this paper, we investigate the task of hallucinating an authentic high-resolution (HR) human face from multiple low-resolution (LR) video snapshots. We propose a pure transformer-based model, dubbed VidFace, to fully exploit the full-range spatio-temporal information and facial structure cues among multiple thumbnails. Specifically, VidFace handles multiple snapshots all at once and harnesses the spatial and temporal information integrally to explore face alignments across all the frames, thus avoiding accumulating alignment errors. Moreover, we design a recurrent position embedding module to equip our transformer with facial priors, which not only effectively regularises the alignment mechanism but also supplants notorious pre-training. Finally, we curate a new large-scale video face hallucination dataset from the public Voxceleb2 benchmark, which challenges prior arts on tackling unaligned and tiny face snapshots. To the best of our knowledge, we are the first attempt to develop a unified transformer-based solver tailored for video-based face hallucination. Extensive experiments on public video face benchmarks show that the proposed method significantly outperforms the state of the arts.

cs.CV

A New Automatic Tool for CME Detection and Tracking with Machine Learning Techniques

With the accumulation of big data of CME observations by coronagraphs, automatic detection and tracking of CMEs has proven to be crucial. The excellent performance of convolutional neural network in image classification, object detection and other computer vision tasks motivates us to apply it to CME detection and tracking as well. We have developed a new tool for CME Automatic detection and tracking with MachinE Learning (CAMEL) techniques. The system is a three-module pipeline. It is first a supervised image classification problem. We solve it by training a neural network LeNet with training labels obtained from an existing CME catalog. Those images containing CME structures are flagged as CME images. Next, to identify the CME region in each CME-flagged image, we use deep descriptor transforming to localize the common object in an image set. A following step is to apply the graph cut technique to finely tune the detected CME region. To track the CME in an image sequence, the binary images with detected CME pixels are converted from cartesian to polar coordinate. A CME event is labeled if it can move in at least two frames and reach the edge of coronagraph field of view. For each event, a few fundamental parameters are derived. The results of four representative CMEs with various characteristics are presented and compared with those from four existing automatic and manual catalogs. We find that CAMEL can detect more complete and weaker structures, and has better performance to catch a CME as early as possible.

astro-ph.SR

Raman evidence for dimerization and Mott collapse in $\alpha$-RuCl$_3$ under pressures

We perform Raman spectroscopy studies on $\alpha$-RuCl$_3$ at room temperature to explore its phase transitions of magnetism and chemical bonding under pressures. The Raman measurements resolve two critical pressures, about $p_1=1.1$~GPa and $p_2=1.7$~GPa, involving very different intertwining behaviors between the structural and magnetic excitations. With increasing pressures, a stacking order phase transition of $\alpha$-RuCl$_3$ layers develops at $p_1=1.1$~GPa, indicated by the new Raman phonon modes and the modest Raman magnetic susceptibility adjustment. The abnormal softening and splitting of the Ru in-plane Raman mode provide direct evidence of the in-plane dimerization of the Ru-Ru bonds at $p_2=1.7$~GPa. The Raman susceptibility is greatly enhanced with pressure increasing and sharply suppressed after the dimerization. We propose that the system undergoes Mott collapse at $p_2=1.7$~GPa and turns into a dimerized correlated band insulator. Our studies demonstrate competitions between Kitaev physics, magnetism, and chemical bondings in Kitaev compounds.

cond-mat.str-el

Unusual phonon density of states and response to the superconducting transition in the In-doped topological crystalline insulator Pb$_{0.5}$Sn$_{0.5}$Te

We present inelastic neutron scattering results of phonons in (Pb$_{0.5}$Sn$_{0.5}$)$_{1-x}$In$_x$Te powders, with $x=0$ and 0.3. The $x=0$ sample is a topological crystalline insulator, and the $x=0.3$ sample is a superconductor with a bulk superconducting transition temperature $T_c$ of 4.7 K. In both samples, we observe unexpected van Hove singularities in the phonon density of states at energies of 1--2.5 meV, suggestive of local modes. On cooling the superconducting sample through $T_c$, there is an enhancement of these features for energies below twice the superconducting-gap energy. We further note that the superconductivity in (Pb$_{0.5}$Sn$_{0.5}$)$_{1-x}$In$_x$Te occurs in samples with normal-state resistivities of order 10 m$\Omega$~cm, indicative of bad-metal behavior. Calculations based on density functional theory suggest that the superconductivity is easily explainable in terms of electron-phonon coupling; however, they completely miss the low-frequency modes and do not explain the large resistivity. While the bulk superconducting state of (Pb$_{0.5}$Sn$_{0.5}$)$_{0.7}$In$_{0.3}$Te appears to be driven by phonons, a proper understanding will require ideas beyond simple BCS theory.

cond-mat.supr-con

Multi-carrier transport in ZrTe5 film

The single layer of Zirconium pentatelluride (ZrTe5) has been predicted to be a large-gap two-dimensional (2D) topological insulator, which has attracted particular attention in the topological phase transitions and potential device application. Here we investigated the transport properties in ZrTe5 films with the dependence of thickness from a few nm to several hundred nm. We find that the temperature of the resistivity anomaly's peak (Tp) is inclining to increase as the thickness decreases, and around a critical thickness of ~40 nm, the dominating carriers in the films change from n-type to p-type. With comprehensive studying of the Shubnikov-de Hass (SdH) oscillations and Hall resistance at variable temperatures, we demonstrate the multi-carrier transport instinct in the thin films. We extract the carrier densities and mobilities of two majority carriers using the simplified two-carrier model. The electron carriers can be attributed to the Dirac band with a non-trivial Berry's phase {\pi}, while the hole carriers may originate from the surface chemical reaction or unintentional doping during the microfabrication process. It is necessary to encapsulate ZrTe5 film in the inert or vacuum environment to make a substantial improvement in the device quality.

cond-mat.mes-hall

Large-Scale 3D Shape Reconstruction and Segmentation from ShapeNet Core55

We introduce a large-scale 3D shape understanding benchmark using data and annotation from ShapeNet 3D object database. The benchmark consists of two tasks: part-level segmentation of 3D shapes and 3D reconstruction from single view images. Ten teams have participated in the challenge and the best performing teams have outperformed state-of-the-art approaches on both tasks. A few novel deep learning architectures have been proposed on various 3D representations on both tasks. We report the techniques used by each team and the corresponding performances. In addition, we summarize the major discoveries from the reported results and possible trends for the future work in the field.

cs.CV

Suppression of the antiferromagnetic order when approaching the superconducting state in a phase-separated crystal of K$_x$Fe$_{2-y}$Se$_2$

We have combined elastic and inelastic neutron scattering techniques, magnetic susceptibility and resistivity measurements to study single-crystal samples of K$_x$Fe$_{2-y}$Se$_2$, which contain the superconducting phase that has a transition temperature of $\sim$31 K. In the inelastic neutron scattering measurements, we observe both the spin-wave excitations resulting from the block antiferromagnetic ordered phase and the resonance that is associated with the superconductivity in the superconducting phase, demonstrating the coexistence of these two orders. From the temperature dependence of the intensity of the magnetic Bragg peaks, we find that well before entering the superconducting state, the development of the magnetic order is interrupted, at $\sim$42 K. We consider this result to be evidence for the physical separation of the antiferromagnetic and superconducting phases; the suppression is possibly due to the proximity effect of the superconducting fluctuations on the antiferromagnetic order.

cond-mat.supr-con

3D Shape Segmentation via Shape Fully Convolutional Networks

We desgin a novel fully convolutional network architecture for shapes, denoted by Shape Fully Convolutional Networks (SFCN). 3D shapes are represented as graph structures in the SFCN architecture, based on novel graph convolution and pooling operations, which are similar to convolution and pooling operations used on images. Meanwhile, to build our SFCN architecture in the original image segmentation fully convolutional network (FCN) architecture, we also design and implement a generating operation} with bridging function. This ensures that the convolution and pooling operation we have designed can be successfully applied in the original FCN architecture. In this paper, we also present a new shape segmentation approach based on SFCN. Furthermore, we allow more general and challenging input, such as mixed datasets of different categories of shapes} which can prove the ability of our generalisation. In our approach, SFCNs are trained triangles-to-triangles by using three low-level geometric features as input. Finally, the feature voting-based multi-label graph cuts is adopted to optimise the segmentation results obtained by SFCN prediction. The experiment results show that our method can effectively learn and predict mixed shape datasets of either similar or different characteristics, and achieve excellent segmentation results.

cs.CV

Spin-wave excitations evidencing the Kitaev interaction in single crystalline $\alpha$-RuCl$_3$

Kitaev interactions underlying a quantum spin liquid have been long sought, but experimental data from which their strengths can be determined directly is still lacking. Here, by carrying out inelastic neutron scattering measurements on high-quality single crystals of $\alpha$-RuCl$_3$, we observe spin-wave spectra with a gap of $\sim$2 meV around the M point of the two-dimensional Brillouin zone. We derive an effective-spin model in the strong-coupling limit based on energy bands obtained from first-principle calculations, and find that the anisotropic Kitaev interaction $K$ term and the isotropic antiferromagentic off-diagonal exchange interaction $\Gamma$ term are significantly larger than the Heisenberg exchange coupling $J$ term. Our experimental data can be well fit using an effective-spin model with $K=-6.8$ meV and $\Gamma=9.5$ meV. These results demonstrate explicitly that Kitaev physics is realized in real materials.

cond-mat.str-el

Real-space characterization of reactivity towards water at Bi2Te3(111) surface

Surface reactivity is important in modifying the physical and chemical properties of surface sensitive materials, such as the topological insulators (TIs). Even though many studies addressing the reactivity of TIs towards external gases have been reported, it is still under heavy debate whether and how the topological insulators react with H$_2$O. Here, we employ scanning tunneling microscopy (STM) to directly probe the surface reaction of Bi$_2$Te$_3$ towards H$_2$O. Surprisingly, it is found that only the top quintuple layer is reactive to H$_2$O, resulting in a hydrated Bi bilayer as well as some Bi islands, which passivate the surface and prevent from the subsequent reaction. A reaction mechanism is proposed with H$_2$Te and hydrated Bi as the products. Unexpectedly, our study indicates the reaction with water is intrinsic and not dependent on any surface defects. Since water inevitably exists, these findings provide key information when considering the reactions of Bi$_2$Te$_3$ with residual gases or atmosphere.

cond-mat.mtrl-sci