arXiv ScienceSearch

arXiv subjects

Sheng Chang

Publications and source records attributed to Sheng Chang.

18 recordsLinked to original sources

Changing the Game: The Bounce-Bind Ising Machine

The Ising model, originally proposed a century ago, has become a cornerstone of combinatorial optimization in recent decades. However, Ising machines remain constrained by a fundamental hardware-speed trade-off. We introduce the Bounce-Bind Ising Machine (BBIM), a mechanism with a single tunable parameter that modulates spin dynamics without altering the energy landscape, building upon the classic golf-ball analogy but replacing it with a dynamic tennis ball/shot put system. The Bounce mode (accelerating escapes from local minima) and Bind mode (enabling rapid convergence) dynamically balance speed and quality. Benchmarked on dense MAX-CUT (edge density=0.5), BBIM achieves a peak speedup of 6.15 times at n=200. For sparse 3-Regular 3-XORSAT (second-order), the peak speedup reaches 27.3 times at n=160. Both results incur negligible additional hardware resource consumption. This work demonstrates a critical pathway to circumventing the hardware-speed bottleneck and its practical applicability to large-scale optimization hardware, validated on structurally distinct benchmarks.

cs.AR

Conversational Learning Diagnosis via Reasoning Multi-Turn Interactive Learning

Learning diagnosis is a critical task that monitors students' cognitive state during educational activities, with the goal of enhancing learning outcomes. With advancements in language models (LMs), many AI-driven educational studies have shifted towards conversational learning scenarios, where students engage in multi-turn interactive dialogues with tutors. However, conversational learning diagnosis remains underdeveloped, and most existing techniques acquire students' cognitive state through intuitive instructional prompts on LMs to analyze the dialogue text. This direct prompting approach lacks a solid psychological foundation and fails to ensure the reliability of the generated analytical text. In this study, we introduce ParLD, a preview-analyze-reason framework for conversational learning diagnosis, which leverages multi-agent collaboration to diagnose students' cognitive state over multiple dialogue turns. Specifically, ParLD comprises three main components: (1) Behavior Previewer, which generates a student behavior schema based on previous states and learning content; (2) State Analyzer, which diagnoses the tutor-student dialogue and behavior schema to update the cognitive state; and (3) Performance Reasoner, which predicts the student's future responses and provides verifiable feedback to support ParLD's self-reflection with the Chain Reflector. They operate sequentially and iteratively during each interaction turn to diagnose the student's cognitive state. We conduct experiments to evaluate both performance prediction and tutoring support, emphasizing the effectiveness of ParLD in providing reliable and insightful learning diagnosis.

cs.CY

S2M2ECG: Spatio-temporal bi-directional State Space Model Enabled Multi-branch Mamba for ECG

As one of the most effective methods for cardiovascular disease (CVD) diagnosis, multi-lead Electrocardiogram (ECG) signals present a characteristic multi-sensor information fusion challenge that has been continuously researched in deep learning domains. Despite the numerous algorithms proposed with different DL architectures, maintaining a balance among performance, computational complexity, and multi-source ECG feature fusion remains challenging. Recently, state space models (SSMs), particularly Mamba, have demonstrated remarkable effectiveness across various fields. Their inherent design for high-efficiency computation and linear complexity makes them particularly suitable for low-dimensional data like ECGs. This work proposes S2M2ECG, an SSM architecture featuring three-level fusion mechanisms: (1) Spatio-temporal bi-directional SSMs with segment tokenization for low-level signal fusion, (2) Intra-lead temporal information fusion with bi-directional scanning to enhance recognition accuracy in both forward and backward directions, (3) Cross-lead feature interaction modules for spatial information fusion. To fully leverage the ECG-specific multi-lead mechanisms inherent in ECG signals, a multi-branch design and lead fusion modules are incorporated, enabling individual analysis of each lead while ensuring seamless integration with others. Experimental results reveal that S2M2ECG achieves superior performance in the rhythmic, morphological, and clinical scenarios. Moreover, its lightweight architecture ensures it has nearly the fewest parameters among existing models, making it highly suitable for efficient inference and convenient deployment. Collectively, S2M2ECG offers a promising alternative that strikes an excellent balance among performance, computational complexity, and ECG-specific characteristics, paving the way for high-performance, lightweight computations in CVD diagnosis.

eess.SP

Amber Pruner: Leveraging N:M Activation Sparsity for Efficient Prefill in Large Language Models

In the era of large language models (LLMs), N:M sparsity has emerged as a structured compression technique critical for accelerating inference. While prior work has primarily focused on weight sparsity, it often suffers from significant accuracy degradation. Activation sparsity, though promising, is typically training-dependent and faces challenges in generalization. To address these limitations, we introduce Amber Pruner, a training-free N:M activation sparsity method designed specifically for the prefill stage, targeting the acceleration of linear projection layers in LLMs. Extensive experiments across multiple models and sparsity ratios (2:4, 4:8, and 8:16) demonstrate that Amber Pruner can effectively sparsify and accelerate more than 55% of linear computations without requiring model retraining. To further enhance generality and efficiency, we propose Outstanding-sparse, a unified framework that integrates Amber Pruner with post-training W8A8 quantization. Our approach preserves strong performance across a range of downstream tasks, with notable advantages in generative tasks. This work pioneers a new frontier in activation sparsity, providing foundational insights that are poised to guide the co-evolution of algorithms and architectures in the design of next-generation AI systems.

cs.LG

Efficient Oriented Object Detection with Enhanced Small Object Recognition in Aerial Images

Achieving a balance between computational efficiency and detection accuracy in the realm of rotated bounding box object detection within aerial imagery is a significant challenge. While prior research has aimed at creating lightweight models that enhance computational performance and feature extraction, there remains a gap in the performance of these networks when it comes to the detection of small and multi-scale objects in remote sensing (RS) imagery. To address these challenges, we present a novel enhancement to the YOLOv8 model, tailored for oriented object detection tasks and optimized for environments with limited computational resources. Our model features a wavelet transform-based C2f module for capturing associative features and an Adaptive Scale Feature Pyramid (ASFP) module that leverages P2 layer details. Additionally, the incorporation of GhostDynamicConv significantly contributes to the model's lightweight nature, ensuring high efficiency in aerial imagery analysis. Featuring a parameter count of 21.6M, our approach provides a more efficient architectural design than DecoupleNet, which has 23.3M parameters, all while maintaining detection accuracy. On the DOTAv1.0 dataset, our model demonstrates a mean Average Precision (mAP) that is competitive with leading methods such as DecoupleNet. The model's efficiency, combined with its reduced parameter count, makes it a strong candidate for aerial object detection, particularly in resource-constrained environments.

cs.CV

Heisenberg-limit spin squeezing with spin Bogoliubov Hamiltonian

It is well established that the optimal spin squeezing under a one-axis-twisting Hamiltonian follows a scaling law of $J^{-2/3}$ for $J$ interacting atoms after a quench dynamics. Here we prove analytically and numerically that the spin squeezing of the ground state of the one-axis-twisting Hamiltonian actually reaches the Heisenberg limit $J^{-1}$. By constructing a bilinear Bogoliubov Hamiltonian with the raising and lowering spin operators, we exactly diagonalize the spin Bogoliubov Hamiltonian, which includes the one-axis twisting Hamiltonian as a limiting case. The ground state of the spin Bogoliubov Hamiltonian exhibits wonderful spin squeezing, which approaches to the Heisenberg limit in the case of the one-axis twisting Hamiltonian. It is possible to realize experimentally the spin squeezed ground state of the one-axis-twisting Hamiltonian in dipolar spinor condensates, ultracold atoms in optical lattices, spins in a cavity, or alkali atoms in a vapor cell.

cond-mat.quant-gas

A Novel Field-Free SOT Magnetic Tunnel Junction With Local VCMA-Induced Switching

By integrating the local voltage-controlled magnetic anisotropy (VCMA) effect, Dzyaloshinskii-Moriya interaction (DMI) effect, and spin-orbit torque (SOT) effect, we propose a novel device structure for field-free magnetic tunnel junction (MTJ). Micromagnetic simulation shows that the device utilizes the chiral symmetry breaking caused by the DMI effect to induce a non-collinear spin texture under the influence of SOT current. This, combined with the perpendicular magnetic anisotropy (PMA) gradient generated by the local VCMA effect, enables deterministic switching of the MTJ state without an external field. The impact of variations in DMI strength and PMA gradient on the magnetization dynamics is analyzed.

eess.SP

A Novel Interface Database of Graphene Nanoribbon from Density Functional Theory

Interfaces play a crucial role in determining the overall performance and functionality of electronic devices and systems. Driven by the data science, machine learning (ML) reveals excellent guidance for material selection and device design, in which an advanced database is crucial for training models with state-of-the-art (SOTA) precision. However, a systematic database of interfaces is still in its infancy due to the difficulties in collecting raw data in experiment and the expensive first-principles computational cost in density functional theory (DFT). In this paper, we construct ample interface structures of graphene nanoribbons (GNR), whose interfacial morphology can be precisely fabricated based on specific molecular precursors. The GNR interfaces serve as promising candidates since their bandgaps can be modulated. Their physical properties including energy bands and density of states (DOS) maps are obtained under reasonable calculation parameters. This database can provide theoretical guidance for the design of electronic devices and accelerate the ML study of various physical quantities.

cond-mat.mtrl-sci

First-principles Prediction of Potential Candidate Materials MCu$_3$X$_4$ (M = V, Nb, Ta; X = S, Se, Te) for Neuromorphic Computing

Inspired by the neuro-synaptic frameworks in the human brain, neuromorphic computing is expected to overcome the bottleneck of traditional von-Neumann architecture and be used in artificial intelligence. Here, we predict a class of potential candidate materials, MCu$_3$X$_4$ (M = V, Nb, Ta; X = S, Se, Te), for neuromorphic computing applications through first-principles calculations based on density functional theory. We find that when MCu$_3$X$_4$ are inserted with Li atom, the systems would transform from semiconductors to metals due to the considerable electron filling [~0.8 electrons per formula unit (f.u.)] and still maintain well structural stability. Meanwhile, the inserted Li atom also has a low diffusion barrier (~0.6 eV/f.u.), which ensures the feasibility to control the insertion/extraction of Li by gate voltage. These results establish that the system can achieve the reversible switching between two stable memory states, i.e., high/low resistance state, indicating that it could potentially be used to design synaptic transistor to enable neuromorphic computing. Our work provides inspiration for advancing the search of candidate materials related to neuromorphic computing from the perspective of theoretical calculations.

cond-mat.mtrl-sci

A Latent Encoder Coupled Generative Adversarial Network (LE-GAN) for Efficient Hyperspectral Image Super-resolution

Realistic hyperspectral image (HSI) super-resolution (SR) techniques aim to generate a high-resolution (HR) HSI with higher spectral and spatial fidelity from its low-resolution (LR) counterpart. The generative adversarial network (GAN) has proven to be an effective deep learning framework for image super-resolution. However, the optimisation process of existing GAN-based models frequently suffers from the problem of mode collapse, leading to the limited capacity of spectral-spatial invariant reconstruction. This may cause the spectral-spatial distortion on the generated HSI, especially with a large upscaling factor. To alleviate the problem of mode collapse, this work has proposed a novel GAN model coupled with a latent encoder (LE-GAN), which can map the generated spectral-spatial features from the image space to the latent space and produce a coupling component to regularise the generated samples. Essentially, we treat an HSI as a high-dimensional manifold embedded in a latent space. Thus, the optimisation of GAN models is converted to the problem of learning the distributions of high-resolution HSI samples in the latent space, making the distributions of the generated super-resolution HSIs closer to those of their original high-resolution counterparts. We have conducted experimental evaluations on the model performance of super-resolution and its capability in alleviating mode collapse. The proposed approach has been tested and validated based on two real HSI datasets with different sensors (i.e. AVIRIS and UHD-185) for various upscaling factors and added noise levels, and compared with the state-of-the-art super-resolution models (i.e. HyCoNet, LTTR, BAGAN, SR- GAN, WGAN).

eess.IV

A Novel CropdocNet for Automated Potato Late Blight Disease Detection from the Unmanned Aerial Vehicle-based Hyperspectral Imagery

Late blight disease is one of the most destructive diseases in potato crop, leading to serious yield losses globally. Accurate diagnosis of the disease at early stage is critical for precision disease control and management. Current farm practices in crop disease diagnosis are based on manual visual inspection, which is costly, time consuming, subject to individual bias. Recent advances in imaging sensors (e.g. RGB, multiple spectral and hyperspectral cameras), remote sensing and machine learning offer the opportunity to address this challenge. Particularly, hyperspectral imagery (HSI) combining with machine learning/deep learning approaches is preferable for accurately identifying specific plant diseases because the HSI consists of a wide range of high-quality reflectance information beyond human vision, capable of capturing both spectral-spatial information. The proposed method considers the potential disease specific reflectance radiation variance caused by the canopy structural diversity, introduces the multiple capsule layers to model the hierarchical structure of the spectral-spatial disease attributes with the encapsulated features to represent the various classes and the rotation invariance of the disease attributes in the feature space. We have evaluated the proposed method with the real UAV-based HSI data under the controlled field conditions. The effectiveness of the hierarchical features has been quantitatively assessed and compared with the existing representative machine learning/deep learning methods. The experiment results show that the proposed model significantly improves the accuracy performance when considering hierarchical-structure of spectral-spatial features, comparing to the existing methods only using spectral, or spatial or spectral-spatial features without consider hierarchical-structure of spectral-spatial features.

cs.CV

A Multilayer Neural Network Merging Image Preprocessing and Pattern Recognition by Integrating Diffusion and Drift Memristors

With the development of research on novel memristor model and device, neural networks by integrating various memristor models have become a hot research topic recently. However, state-of-the-art works still build such neural networks using drift memristor only. Furthermore, some other related works are only applied to a few individual applications including pattern recognition and edge detection. In this paper, a novel kind of multilayer neural network is proposed, in which diffusion and drift memristor models are applied to construct a system merging image preprocessing and pattern recognition. Specifically, the entire network consists of two diffusion memristive cellular layers for image preprocessing and one drift memristive feedforward layer for pattern recognition. Experimental results show that good recognition accuracy of noisy MNIST is obtained due to the fusion of image preprocessing and pattern recognition. Moreover, owing to high-efficiency in-memory computing and brief spiking encoding methods, high processing speed, high throughput, and few hardware resources of the entire network are achieved.

cs.ET

Fully Memristive Spiking-Neuron Learning Framework and its Applications on Pattern Recognition and Edge Detection

Fully memristive spiking-neuron learning framework, which uses drift and diffusion memristor models as axon and dendrite respectively, becomes a hot topic recently with the development of memristor devices. Normally, some other devices like resistor or capacitor are still necessary on recent works of fully memristive learning framework. However, theoretically, one neuron needs axon and dendrite only, which makes technique process simpler and learning framework more similar to biologic brain. In this paper, a fully memristive spiking-neuron learning framework is introduced, in which a neuron structure is just built of one drift and one diffusion memristive models. To verify it merits, a feedforward neural network for pattern recognition and a cellular neural network for edge detection are designed. Experiment results show that compared to other memristive neural networks, our framework's the processing speed is much faster and the hardware resource is saved in pattern recognition due to its simple structure. Further due to the dynamic filtering function of diffusion memristor model in our learning framework, its peak signal noise ratio (PSNR) is much higher than traditional algorithms in edge detection.

cs.ET

A Biologically Interpretable Two-stage Deep Neural Network (BIT-DNN) For Vegetation Recognition From Hyperspectral Imagery

Spectral-spatial based deep learning models have recently proven to be effective in hyperspectral image (HSI) classification for various earth monitoring applications such as land cover classification and agricultural monitoring. However, due to the nature of "black-box" model representation, how to explain and interpret the learning process and the model decision, especially for vegetation classification, remains an open challenge. This study proposes a novel interpretable deep learning model -- a biologically interpretable two-stage deep neural network (BIT-DNN), by incorporating the prior-knowledge (i.e. biophysical and biochemical attributes and their hierarchical structures of target entities) based spectral-spatial feature transformation into the proposed framework, capable of achieving both high accuracy and interpretability on HSI based classification tasks. The proposed model introduces a two-stage feature learning process: in the first stage, an enhanced interpretable feature block extracts the low-level spectral features associated with the biophysical and biochemical attributes of target entities; and in the second stage, an interpretable capsule block extracts and encapsulates the high-level joint spectral-spatial features representing the hierarchical structure of biophysical and biochemical attributes of these target entities, which provides the model an improved performance on classification and intrinsic interpretability with reduced computational complexity. We have tested and evaluated the model using four real HSI datasets for four separate tasks (i.e. plant species classification, land cover classification, urban scene recognition, and crop disease recognition tasks). The proposed model has been compared with five state-of-the-art deep learning models.

cs.CV

The MBPEP: a deep ensemble pruning algorithm providing high quality uncertainty prediction

Machine learning algorithms have been effectively applied into various real world tasks. However, it is difficult to provide high-quality machine learning solutions to accommodate an unknown distribution of input datasets; this difficulty is called the uncertainty prediction problems. In this paper, a margin-based Pareto deep ensemble pruning (MBPEP) model is proposed. It achieves the high-quality uncertainty estimation with a small value of the prediction interval width (MPIW) and a high confidence of prediction interval coverage probability (PICP) by using deep ensemble networks. In addition to these networks, unique loss functions are proposed, and these functions make the sub-learners available for standard gradient descent learning. Furthermore, the margin criterion fine-tuning-based Pareto pruning method is introduced to optimize the ensembles. Several experiments including predicting uncertainties of classification and regression are conducted to analyze the performance of MBPEP. The experimental results show that MBPEP achieves a small interval width and a low learning error with an optimal number of ensembles. For the real-world problems, MBPEP performs well on input datasets with unknown distributions datasets incomings and improves learning performance on a multi task problem when compared to that of each single model.

cs.LG

A Hardware Friendly Unsupervised Memristive Neural Network with Weight Sharing Mechanism

Memristive neural networks (MNNs), which use memristors as neurons or synapses, have become a hot research topic recently. However, most memristors are not compatible with mainstream integrated circuit technology and their stabilities in large-scale are not very well so far. In this paper, a hardware friendly MNN circuit is introduced, in which the memristive characteristics are implemented by digital integrated circuit. Through this method, spike timing dependent plasticity (STDP) and unsupervised learning are realized. A weight sharing mechanism is proposed to bridge the gap of network scale and hardware resource. Experiment results show the hardware resource is significantly saved with it, maintaining good recognition accuracy and high speed. Moreover, the tendency of resource increase is slower than the expansion of network scale, which infers our method's potential on large scale neuromorphic network's realization.

cs.ET

Progress on a Miniature Cold-Atom Frequency Standard

Atomic clocks play a crucial role in timekeeping, communications, and navigation systems. Recent efforts enabled by heterogeneous MEMS integration have led to the commercial introduction of Chip-Scale Atomic Clocks (CSAC) with a volume of 16 cm3, power consumption of 120 mW, and instability (Allan Deviation) of σ(τ = 1 sec) < 2e-10. In order to reduce the temperature sensitivity of next-generation CSACs for timing applications, the interaction of atoms with the environment must be minimized, which can be accomplished in an architecture based on trapped, laser-cooled atoms. In this paper, we present results describing the development of a miniature cold-atom apparatus for operation as a frequency standard. Our architecture is based on laser-cooling a sample of neutral atoms in a Magneto-Optical Trap (MOT) using a conical retro-reflector in a miniature vacuum chamber. Trapping the atoms in vacuum and performing microwave interrogation in the dark reduces the temperature sensitivity compared to vapor-cell CSACs. We present details of the component development associated with the laser systems, opto-electronics, and vacuum package for miniature cold-atom technology. Finally, we conclude by characterizing the optimum alkali background pressure for such a cold-atom frequency standard.

physics.atom-ph

Nucleon Mass Splitting at Finite Isospin Chemical Potential

We investigate nucleon mass splitting at finite isospin chemical potential in the frame of two flavor Nambu--Jona-Lasinio model. It is analytically proved that, in the phase with explicit isospin symmetry breaking the proton mass decreases and the neutron mass increases linearly in the isospin chemical potential.

nucl-th