arXiv ScienceSearch

arXiv subjects

Zijian Jiang

Publications and source records attributed to Zijian Jiang.

18 recordsLinked to original sources

CRIP: Channel Level Representation Injection for Personalized One-Shot Federated Learning

One-shot federated learning (OSFL) has emerged as a promising collaborative model learning framework with only a single round of communication, offering significant advantages in communication efficiency and privacy preservation. However, OSFL often faces inherent limitations under severe domain heterogeneity across clients due to the lack of iterative knowledge exchange. Most existing OSFL methods require an auxiliary public dataset for knowledge distillation or leverage statistical information for parameter-level aggregation, overlooking feature shift caused by domain heterogeneity. To address these challenges, we propose CRIP, a personalized OSFL framework that operates in the representation space via channel-level feature alignment. To achieve this, each client uploads its feature extractor to the server, which broadcasts all extractors back to every client. Since not all source clients share compatible feature distributions with the target client, indiscriminate fusion of cross-client features would introduce domain-specific noise. Therefore, CRIP effectively measures the channel-wise representational similarity between the target client and each source client on a small local mini-batch, and selectively fuses only the most compatible features. Extensive experiments on domain-heterogeneous benchmarks such as DomainNet, PACS, and Office-Home demonstrate that CRIP consistently outperforms local models and state-of-the-art baselines, validating the effectiveness of representation-space personalization under extreme domain heterogeneity.

cs.LG

Instruction-based Image Editing: A Survey on Data, Models, Evaluation, and Applications

Instruction-based Image Editing (IIE) aims to transform a given image into a new one based on textual instructions. Advances in Large Language Models (LLMs) and Vision-Language Models (VLMs) have accelerated progress toward practical ``one-sentence image editing" systems. This survey presents a systematic taxonomy and comprehensive review of IIE research, structured around five core dimensions: (1) task definition and hierarchical categorization of editing operations, (2) methodologies for training data construction, (3) architectural evolution from GAN-based to diffusion and autoregressive paradigms, (4) standardized evaluation metrics and benchmark development, and (5) introduction of commercial solutions. Our analysis shows critical technological milestones across model generations. We further propose a Comprehensive, in-Depth, and Diagnostic benchmark for IIE task (CDD-IIE Bench), which can rigorously assess the multiple aspects of model performance. Through empirical comparisons of open-source solutions, we highlight their respective capabilities and limitations. Finally, we discuss future research directions to advance the field.

cs.CV

W2T: LoRA Weights Already Know What They Can Do

Each LoRA checkpoint compactly stores task-specific updates in low-rank weight matrices, offering an efficient way to adapt large language models to new tasks and domains. In principle, these weights already encode what the adapter does and how well it performs. In this paper, we ask whether this information can be read directly from the weights, without running the base model or accessing training data. A key obstacle is that a single LoRA update can be factorized in infinitely many ways. Without resolving this ambiguity, models trained on the factors may fit the particular factorization rather than the underlying update. To this end, we propose \methodfull, which maps each LoRA update to a provably canonical form via QR decomposition followed by SVD, so that all equivalent factorizations share the same representation. The resulting components are then tokenized and processed by a Transformer to produce a weight-space embedding. Across language and vision LoRA collections, W2T achieves strong results on attribute classification, performance prediction, and adapter retrieval, demonstrating that LoRA weights reliably indicate model behavior once factorization ambiguity is removed. Code is available at https://github.com/xiaolonghan2000/Weight2Token.

cs.LG

Proteus-ID: ID-Consistent and Motion-Coherent Video Customization

Video identity customization seeks to synthesize realistic, temporally coherent videos of a specific subject, given a single reference image and a text prompt. This task presents two core challenges: (1) maintaining identity consistency while aligning with the described appearance and actions, and (2) generating natural, fluid motion without unrealistic stiffness. To address these challenges, we introduce Proteus-ID, a novel diffusion-based framework for identity-consistent and motion-coherent video customization. First, we propose a Multimodal Identity Fusion (MIF) module that unifies visual and textual cues into a joint identity representation using a Q-Former, providing coherent guidance to the diffusion model and eliminating modality imbalance. Second, we present a Time-Aware Identity Injection (TAII) mechanism that dynamically modulates identity conditioning across denoising steps, improving fine-detail reconstruction. Third, we propose Adaptive Motion Learning (AML), a self-supervised strategy that reweights the training loss based on optical-flow-derived motion heatmaps, enhancing motion realism without requiring additional inputs. To support this task, we construct Proteus-Bench, a high-quality dataset comprising 200K curated clips for training and 150 individuals from diverse professions and ethnicities for evaluation. Extensive experiments demonstrate that Proteus-ID outperforms prior methods in identity preservation, text alignment, and motion quality, establishing a new benchmark for video identity customization. Codes and data are publicly available at https://grenoble-zhang.github.io/Proteus-ID/.

cs.CV

PUGS: Zero-shot Physical Understanding with Gaussian Splatting

Current robotic systems can understand the categories and poses of objects well. But understanding physical properties like mass, friction, and hardness, in the wild, remains challenging. We propose a new method that reconstructs 3D objects using the Gaussian splatting representation and predicts various physical properties in a zero-shot manner. We propose two techniques during the reconstruction phase: a geometry-aware regularization loss function to improve the shape quality and a region-aware feature contrastive loss function to promote region affinity. Two other new techniques are designed during inference: a feature-based property propagation module and a volume integration module tailored for the Gaussian representation. Our framework is named as zero-shot physical understanding with Gaussian splatting, or PUGS. PUGS achieves new state-of-the-art results on the standard benchmark of ABO-500 mass prediction. We provide extensive quantitative ablations and qualitative visualization to demonstrate the mechanism of our designs. We show the proposed methodology can help address challenging real-world grasping tasks. Our codes, data, and models are available at https://github.com/EverNorif/PUGS

cs.CV

Ctrl-U: Robust Conditional Image Generation via Uncertainty-aware Reward Modeling

In this paper, we focus on the task of conditional image generation, where an image is synthesized according to user instructions. The critical challenge underpinning this task is ensuring both the fidelity of the generated images and their semantic alignment with the provided conditions. To tackle this issue, previous studies have employed supervised perceptual losses derived from pre-trained models, i.e., reward models, to enforce alignment between the condition and the generated result. However, we observe one inherent shortcoming: considering the diversity of synthesized images, the reward model usually provides inaccurate feedback when encountering newly generated data, which can undermine the training process. To address this limitation, we propose an uncertainty-aware reward modeling, called Ctrl-U, including uncertainty estimation and uncertainty-aware regularization, designed to reduce the adverse effects of imprecise feedback from the reward model. Given the inherent cognitive uncertainty within reward models, even images generated under identical conditions often result in a relatively large discrepancy in reward loss. Inspired by the observation, we explicitly leverage such prediction variance as an uncertainty indicator. Based on the uncertainty estimation, we regularize the model training by adaptively rectifying the reward. In particular, rewards with lower uncertainty receive higher loss weights, while those with higher uncertainty are given reduced weights to allow for larger variability. The proposed uncertainty regularization facilitates reward fine-tuning through consistency construction. Extensive experiments validate the effectiveness of our methodology in improving the controllability and generation quality, as well as its scalability across diverse conditional scenarios. Codes are publicly available at https://grenoble-zhang.github.io/Ctrl-U-Page/.

cs.CV

Suppression of Spin Pumping at Metal Interfaces

An electrically conductive metal typically transmits or absorbs a spin current. Here, we report on evidence that interfacing two metal thin films can suppress spin transmission and absorption. We examine spin pumping in ferromagnet/spacer/ferromagnet heterostructures, in which the spacer -- consisting of metallic Cu and Cr thin films -- separates the ferromagnetic spin-source and spin-sink layers. The Cu/Cr spacer largely suppresses spin pumping -- i.e., neither transmitting nor absorbing a significant amount of spin current -- even though Cu or Cr alone transmits a sizable spin current. The antiferromagnetism of Cr is not essential for the suppression of spin pumping, as we observe similar suppression with Cu/V spacers where V is a nonmagnetic analogue of Cr. We speculate that diverse combinations of spin-transparent metals may form interfaces that suppress spin pumping, although the underlying mechanism remains unclear. Our work may stimulate a new perspective on understanding and engineering spin transport in metallic multilayers.

cond-mat.mtrl-sci

Spectrum of non-Hermitian deep-Hebbian neural networks

Neural networks with recurrent asymmetric couplings are important to understand how episodic memories are encoded in the brain. Here, we integrate the experimental observation of wide synaptic integration window into our model of sequence retrieval in the continuous time dynamics. The model with non-normal neuron-interactions is theoretically studied by deriving a random matrix theory of the Jacobian matrix in neural dynamics. The spectra bears several distinct features, such as breaking rotational symmetry about the origin, and the emergence of nested voids within the spectrum boundary. The spectral density is thus highly non-uniformly distributed in the complex plane. The random matrix theory also predicts a transition to chaos. In particular, the edge of chaos provides computational benefits for the sequential retrieval of memories. Our work provides a systematic study of time-lagged correlations with arbitrary time delays, and thus can inspire future studies of a broad class of memory models, and even big data analysis of biological time series.

q-bio.NC

Associative memory model with arbitrary Hebbian length

Conversion of temporal to spatial correlations in the cortex is one of the most intriguing functions in the brain. The learning at synapses triggering the correlation conversion can take place in a wide integration window, whose influence on the correlation conversion remains elusive. Here, we propose a generalized associative memory model with arbitrary Hebbian length. The model can be analytically solved, and predicts that a small Hebbian length can already significantly enhance the correlation conversion, i.e., the stimulus-induced attractor can be highly correlated with a significant number of patterns in the stored sequence, thereby facilitating state transitions in the neural representation space. Moreover, an anti-Hebbian component is able to reshape the energy landscape of memories, akin to the function of sleep. Our work thus establishes the fundamental connection between associative memory, Hebbian length, and correlation conversion in the brain.

cond-mat.dis-nn

Eigenvalue spectrum of neural networks with arbitrary Hebbian length

Associative memory is a fundamental function in the brain. Here, we generalize the standard associative memory model to include long-range Hebbian interactions at the learning stage, corresponding to a large synaptic integration window. In our model, the Hebbian length can be arbitrarily large. The spectral density of the coupling matrix is derived using the replica method, which is also shown to be consistent with the results obtained by applying the free probability method. The maximal eigenvalue is then obtained by an iterative equation, related to the paramagnetic to spin glass transition in the model. Altogether, this work establishes the connection between the associative memory with arbitrary Hebbian length and the asymptotic eigen-spectrum of the neural-coupling matrix.

cond-mat.dis-nn

The impact of data volume on performance of deep learning based building rooftop extraction using very high spatial resolution aerial images

Building rooftop data are of importance in several urban applications and in natural disaster management. In contrast to traditional surveying and mapping, by using high spatial resolution aerial images, deep learning-based building rooftops extraction methods are efficient and accurate. Although more training data is preferred in deep learning-based tasks, the effect of data volume on building extraction models is underexplored. Therefore, the paper explores the impact of data volume on the performance of building rooftop extraction from very-high-spatial-resolution (VHSR) images using deep learning-based methods. To do so, we manually labelled 0.12m spatial resolution aerial images and perform a comparative analysis of models trained on datasets of different sizes using popular deep learning architectures for segmentation tasks, including Fully Convolutional Networks (FCN)-8s, U-Net and DeepLabv3+. The experiments showed that with more training data, algorithms converged faster and achieved higher accuracy, while better algorithms were able to better mitigate the lack of training data.

cs.CV

Investigating and Recommending Co-Changed Entities for JavaScript Programs

JavaScript (JS) is one of the most popular programming languages due to its flexibility and versatility, but maintaining JS code is tedious and error-prone. In our research, we conducted an empirical study to characterize the relationship between co-changed software entities (e.g., functions and variables), and built a machine learning (ML)-based approach to recommend additional entity to edit given developers' code changes. Specifically, we first crawled 14,747 commits in 10 open-source projects; for each commit, we created one or more change dependency graphs (CDGs) to model the referencer-referencee relationship between co-changed entities. Next, we extracted the common subgraphs between CDGs to locate recurring co-change patterns between entities. Finally, based on those patterns, we extracted code features from co-changed entities and trained an ML model that recommends entities-to-change given a program commit. According to our empirical investigation, (1) three recurring patterns commonly exist in all projects; (2) 80%--90% of co-changed function pairs either invoke the same function(s), access the same variable(s), or contain similar statement(s); (3) our ML-based approach CoRec recommended entity changes with high accuracy (73%--78%). CoRec complements prior work because it suggests changes based on program syntax, textual similarity, as well as software history; it achieved higher accuracy than two existing tools in our evaluation.

cs.SE

Relationship between manifold smoothness and adversarial vulnerability in deep learning with local errors

Artificial neural networks can achieve impressive performances, and even outperform humans in some specific tasks. Nevertheless, unlike biological brains, the artificial neural networks suffer from tiny perturbations in sensory input, under various kinds of adversarial attacks. It is therefore necessary to study the origin of the adversarial vulnerability. Here, we establish a fundamental relationship between geometry of hidden representations (manifold perspective) and the generalization capability of the deep networks. For this purpose, we choose a deep neural network trained by local errors, and then analyze emergent properties of trained networks through the manifold dimensionality, manifold smoothness, and the generalization capability. To explore effects of adversarial examples, we consider independent Gaussian noise attacks and fast-gradient-sign-method (FGSM) attacks. Our study reveals that a high generalization accuracy requires a relatively fast power-law decay of the eigen-spectrum of hidden representations. Under Gaussian attacks, the relationship between generalization accuracy and power-law exponent is monotonic, while a non-monotonic behavior is observed for FGSM attacks. Our empirical study provides a route towards a final mechanistic interpretation of adversarial vulnerability under adversarial attacks.

cs.LG

Sub-Nanosecond Spin-Transfer Torque in an Ensemble of Superparamagnetic-Like Nanomagnets

Spin currents can exert spin-transfer torques on magnetic systems even in the limit of vanishingly small net magnetization, as is the case for antiferromagnets. Here, we experimentally show that a spin-transfer torque is operative in a material with weak, short-range magnetic order -- namely, a macroscopic ensemble of superparamagnetic-like Co nanomagnets. We employ element- and time-resolved X-ray ferromagnetic resonance (XFMR) spectroscopy to directly detect sub-ns dynamics of the Co nanomagnets, excited into precession with cone angle $\geq$0.003$^{\circ}$ by an oscillating spin current. XFMR measurements reveal that as the net moment of the ensemble decreases, the strength of the spin-transfer torque increases relative to those of magnetic field torques. Our findings point to spin-transfer torque as an effective way to manipulate the state of nanomagnet ensembles at sub-ns timescales.

cond-mat.mes-hall

Magnetic Damping in Epitaxial Fe Alloyed with Vanadium and Aluminum

To develop low-moment, low-damping metallic ferromagnets for power-efficient spintronic devices, it is crucial to understand how magnetic relaxation is impacted by the addition of nonmagnetic elements. Here, we compare magnetic relaxation in epitaxial Fe films alloyed with light nonmagnetic elements of V and Al. FeV alloys exhibit lower intrinsic damping compared to pure Fe, reduced by nearly a factor of 2, whereas damping in FeAl alloys increases with Al content. Our experimental and computational results indicate that reducing the density of states at the Fermi level, rather than the average atomic number, has a more significant impact in lowering damping in Fe alloyed with light elements. Moreover, FeV is confirmed to exhibit an intrinsic Gilbert damping parameter of $\simeq$0.001, among the lowest ever reported for ferromagnetic metals.

cond-mat.mtrl-sci

Current-induced spin-orbit field in permalloy interfaced with ultrathin Ti and Cu

How spin-orbit torques emerge from materials with weak spin-orbit coupling (e.g., light metals) is an open question in spintronics. Here, we report on a field-like spin-orbit torque (i.e., in-plane spin-orbit field transverse to the current axis) in SiO$_2$-sandwiched permalloy (Py), with the top Py-SiO$_2$ interface incorporating ultrathin Ti or Cu. In both SiO$_2$/Py/Ti/SiO$_2$ and SiO$_2$/Py/Cu/SiO$_2$, this spin-orbit field opposes the classical Oersted field. While the magnitude of the spin-orbit field is at least a factor of 3 greater than the Oersted field, we do not observe evidence for a significant damping-like torque in SiO$_2$/Py/Ti/SiO$_2$ or SiO$_2$/Py/Cu/SiO$_2$. Our findings point to contributions from a Rashba-Edelstein effect or spin-orbit precession at the (Ti, Cu)-inserted interface.

cond-mat.mtrl-sci

Conductivity-Like Gilbert Damping due to Intraband Scattering in Epitaxial Iron

Confirming the origin of Gilbert damping by experiment has remained a challenge for many decades, even for simple ferromagnetic metals. In this Letter, we experimentally identify Gilbert damping that increases with decreasing electronic scattering in epitaxial thin films of pure Fe. This observation of conductivity-like damping, which cannot be accounted for by classical eddy current loss, is in excellent quantitative agreement with theoretical predictions of Gilbert damping due to intraband scattering. Our results resolve the longstanding question about a fundamental damping mechanism and offer hints for engineering low-loss magnetic metals for cryogenic spintronics and quantum devices.

cond-mat.mtrl-sci

Dynamic nuclear spin polarization induced by Edelstein effect at Bi(111) surfaces

Nuclear spin polarization induced by hyperfine interaction and the Edelstein effect due to strong spin-orbit interaction is investigated by quantum transport in Bi(111) thin film samples. The Bi(111) films are deposited on mica by van der Waals epitaxial growth. The Bi(111) films show micrometer-sized triangular islands with 0.39 nm step height, corresponding to the Bi(111) bilayer height. At low temperatures a high current density is applied to generate a non-equilibrium carrier spin polarization by the Edelstein effect at the Bi(111) surface, which then induces dynamic nuclear polarization by hyperfine interaction. Comparative quantum magnetotransport antilocalization measurements indicate a suppression of antilocalization by the in-plane Overhauser field from the nuclear polarization and allow a quantification of the Overhauser field. Hence nuclear polarization was both achieved and quantified by a purely electronic transport-based approach.

cond-mat.mes-hall